Operations
Operate the AI portfolio through the portfolio console. This reference covers the underlying services and Kubernetes/GitOps deployment.
Prerequisites
- Rust (stable) —
rustup default stable - Node 22+ and
pnpm(fornexus-ui) - MongoDB 6+ reachable on the network
- NATS with JetStream enabled
- A Kubernetes cluster +
kubectlfor agent runs - A Telegram bot token (optional, for the gateway)
- A Taiga project + service-account credentials (optional, for board sync)
Secrets never reach a log line
nexus-observability does not merely print events: it captures every one,
publishes it to NATS and persists it to MongoDB for the Admin UI's platform
log page. A credential logged after the bus is up is therefore written to a
database and rendered in a web page, not just left in kubectl logs.
Found 2026-09-06: nexus-accum-host-0 logged its full NATS URL at INFO on
every boot, password included. Measured blast radius — smaller than it first
looked: the connect line is emitted inside nexus_events::connect(), i.e.
before the NATS transport that ships logs exists, so it never reached MongoDB.
A scan of the 671k-row logs collection found 0 rows carrying an
unredacted nats://user:pass@ URL, and the cluster runs no node-level log
collector (no promtail/loki/fluent/vector). Actual exposure was pod stdout via
kubectl logs only. The rule below still holds for any credential logged
after startup — that one really would be persisted and rendered.
Fixed in nexus-platform e340b58 with
nexus_observability::redact_url, applied at both sites that logged a
credential-bearing URL — the JetStream connect line and the agent runner's
clone line and failure message (repo.url can already carry a token, see
authed_url). The boot line now reads
nats://nexus:***@nexus-nats.nexus-nats.svc.cluster.local:4222.
The rule: never log a connection string raw — pass it through
redact_url first. It keeps the username (useful, not secret), replaces the
password with ***, redacts a bare token that is the whole userinfo, and
leaves a credential-free or unparseable URL untouched. An @ in a path or
query is not mistaken for userinfo.
Both NATS credentials were rotated on 2026-09-06 17:19Z (operator
instruction), nexus and runner: new 40-char URL-safe secrets written to
nexus-nats-auth (ns nexus-nats) and into NEXUS_NATS_URL /
NEXUS_AGENT_NATS_URL in nexus-app (ns nexus), NATS restarted, then all
seven consumers. The workspace .secrets was updated to match, and both
passwords satisfy bootstrap-secrets.sh's URL-safe charset guard so a future
bootstrap reproduces the same URLs. Verified: every pod reconnected, the only
authorization violation lines are from outgoing pods inside the ~4 s cutover
window — which is itself the proof the old credential is now rejected — and
the job pipeline resumed. The old secrets are kept nowhere but the operator's
local scratch backup.
Run locally
MongoDB + NATS
docker run -d --name nexus-mongo -p 27017:27017 mongo:6
docker run -d --name nexus-nats -p 4222:4222 nats:latest -js
Nexus Core
cd apps/nexus-core
cp .env.example .env
cargo run --release
Minimum settings:
MONGODB_URL=mongodb://localhost:27017
MONGODB_DB=nexus
NATS_URL=nats://localhost:4222
KUBE_NAMESPACE=nexus-agents
JWT_SIGNING_KEY=<random string>
Verify:
curl http://localhost:8080/healthz
open http://localhost:8080/swagger-ui/
Worker
cd apps/nexus-worker
cargo run --release
UI
The admin UI is a separate repo (nexus-ui):
cd ../nexus-ui # the standalone Next.js repo
cp .env.example .env.local
pnpm install
pnpm dev # http://localhost:3000
NEXUS_CORE_URL=http://localhost:8080
NEXUS_ADMIN_TOKEN=dev-token
AUTH_SECRET=at-least-16-chars-long-secret
Deploy on Kubernetes (GitOps)
Deployment lives in nexus-gitops: Helm charts, per-environment values, Argo CD apps / Flux manifests, and the MongoDB
- NATS manifests.
# App-of-apps: applies every Application under argocd/apps (project, infra,
# nexus chart with prod values, redis, langfuse)
kubectl apply -f argocd/root.yaml
Image tags are written into environments/prod/values.yaml by each repo's
CI on push to main; Argo CD deploys within its ~3-minute poll.
Namespaces
Run agent jobs in a dedicated namespace (e.g. nexus-agents), separate
from the control plane (nexus). Core's service account can create Jobs only in
the agent namespace.
Scoped RBAC
verbs: [create, get, list, watch, delete]
resources: [jobs, pods, pods/log]
Observability
Every service built on nexus-observability serves Prometheus metrics on
:9100 (NEXUS_METRICS_ADDR; off disables — agent runner Jobs set that).
This became true with nexus-platform f183437 (2026-09-05); before it the
names below were constants nothing recorded and the port refused connections.
nexus_agent_runs_total{agent,outcome} recorded (worker run_job)
nexus_agent_run_failures_total{agent,reason} recorded (reason = timeout | error)
nexus_task_duration_seconds{job} recorded
process_cpu_seconds_total / process_resident_memory_bytes / process_open_fds
nexus_board_sync_errors_total registered, no call sites yet
nexus_llm_tokens_total / nexus_llm_cost_usd_total registered, no call sites yet
Since nexus-platform ac38aa1 (2026-09-05 17:12 UTC) the worker also
exports the treasury contract (nexus_nav_*, nexus_benchmark_*,
nexus_mark_*, nexus_exec_*, nexus_sleeve_*, nexus_outbox_*,
nexus_ladder_* — table on
Measuring success)
and, since 7e1adcb, the budget gate's nexus_portfolio_* gauges and
nexus_accum_actions_gated_total{class,reason}; since fdd7912
(21:54 UTC) the lending-hygiene gauges listed under
Lending hygiene; from ee0cd80 (deployed
2026-09-05, live-verified 23:4x UTC) the budget ledger's nexus_portfolio_ledger_version (the CAS
token — should only ever go up), nexus_portfolio_reserve_conflicts_total
(lost CAS writes, retried) and nexus_portfolio_fence_rejections_total
(settlements refused for a stale / missing fence — see
Budget ledger).
P&L-by-source gauges (deployed 2026-09-05, first rows 21:12 UTC)
nexus-platform 750c5ce (pushed ≈ 20:1x UTC with 3a91236 … 7d6ed59)
adds the owner's "where are we winning money" gauges to the same worker
exporter, refreshed from snapshot.meta.pnl after every treasury_nav
run. Semantics and the source vocabulary:
Measuring success → P&L by source.
| Metric | Labels | Meaning |
|---|---|---|
nexus_pnl_source_usdt | source, window = 1d / 7d / 30d / inception | P&L over the window by source (jito_staking, lending:<venue>, lp_fees:<pool>, sleeve_realized:<pair>, sleeve_unrealized:<pair>, market:SOL|BTC|STABLE|LP, costs:tx|venue|infra, unexplained) |
nexus_pnl_unexplained_usdt | window | ΔNAV − external flows − Σ explained sources over the window — the reconciliation residual |
nexus_sleeve_round_trips_total | — | closed swing round trips in the ledger (complete history) |
nexus_sleeve_realized_usdt_total | — | sleeve realized P&L over the complete ledger, net of known fees |
nexus_sleeve_unrealized_usdt | — | open sleeve inventory qty × (mark − vwap) |
nexus_staking_apr_realized | — | JitoSOL pool-rate change since inception, annualised |
nexus_lending_apr_realized | venue | lending interest since inception over the average balance, annualised |
nexus_lending_interest_usdt | venue, window | lending interest over the window |
| nexus_pnl_cost_anomalies_total | field | since 2553748 (2026-09-06): on-chain cost figures the sanity rules refused — ≤ 0, non-finite, above PNL_COST_SANITY_USD (25), or a rent_* field on a native-SOL pair. Never bucketed, never subtracted |
Label names follow the rule below (source,
window, venue, field — none of them registry-owned). Nothing alerts
on these yet: nexus_pnl_unexplained_usdt should get a rule later (a
residual that grows period after period means a movement the flow ledger
did not see) — not written until a few days of live rows show what
"normal" is. Read nexus_pnl_cost_anomalies_total the same way: a flat
counter is the rules idle, a rising one means the exec is reporting a cost
figure the platform will not book, and it wants a look rather than an
alert. Note that unexplained on any window reaching back before
2026-09-05 20:12 UTC is mostly the boundary of the ledger — check
series_starts_at first (the residual).
Metric labels: names the registry already owns
Never use component, job, instance, pod or namespace as a
metric label in a Nexus service. nexus-observability adds component
to every metric globally, and Prometheus adds job / instance (with
pod / namespace from the ServiceMonitor relabeling). A metric that
declares one of those names itself produces a duplicate label and
Prometheus rejects the whole scrape, not just that series.
This happened on 2026-09-05: nexus_nav_component_usdt{component,asset}
shipped in ac38aa1 and every worker scrape failed from 17:12 to ≈ 17:4x
UTC (label name "component" is not unique) — the NAV panels stayed
empty and NexusNavStale kept firing although snapshots existed. The
label is nav_component since nexus-platform 513a173 / d128e17 and
dashboard nexus-gitops a5912c7. Symptom to recognise: the target shows
DOWN in Prometheus with that error text while :9100/metrics answers
200 in-cluster; the fix is a rename, never a relabel.
Scraping: the nexus-metrics ServiceMonitor in the chart → kube-prometheus-stack
(monitoring). Dashboards (Grafana folder Nexus, sidecar ConfigMap
nexus-grafana-dashboard): Nexus Platform (runs, failures, durations,
process) and, since nexus-gitops 05c3ffd (2026-09-05), Nexus Treasury
(uid nexus-treasury: NAV vs benchmarks, excess return, components,
exposure, attribution, flows & costs, marks, execution safety; since
nexus-gitops b5b6079, 2026-09-05 ≈ 20:10 UTC, the row "Where the money
comes from" — P&L by source 30 d / inception bar gauges, unexplained
30 d, swing realized / unrealized / round trips, JitoSOL and Kamino realised
APR, lending interest by venue and P&L groups over time; populated since
the 21:12 UTC snapshot on worker 7d6ed59). Review 2026-09-07 (gitops e4df941): checked against Prometheus, six panels were empty because they queried window="30d" while the exporter publishes 1d + inception only until the series covers a month — excess return, attribution, P&L by source, unexplained, lending interest and P&L groups now read the inception window (P&L by source also 24 h); "budget gate verdicts" queried a counter that has never incremented and now shows the gate's judged exposure per asset (nexus_portfolio_exposure_after_frac); BTC dust is out of the way — exposure draws assets above 0.5 % of NAV only, the cbBTC mark panel is gone (a $4 holding; the mark stays in Prometheus and on the Treasury page), cap headroom shows the caps in play, and the bar gauges list sources with |amount| > 0.05 USDT; benchmark legends read "hold BTC/SOL" / "DCA BTC/SOL"; USDT panels carry a USDT suffix and the snapshot age is a duration. Every remaining panel query was run against Prometheus after the reload. Alerts:
Alertmanager → Telegram group NexusHome (see the
operator guide). Traces go to
Langfuse (NEXUS_LANGFUSE_*), not OTLP — there is no OpenTelemetry exporter
in the tree.
Alerting (PrometheusRule nexus-platform-alerts)
One PrometheusRule, templates/alerts.yaml in the chart (value
monitoring.alerts: true), picked up by the kps Prometheus (empty
ruleSelector). Groups and what they watch:
| Group | Since | Rules |
|---|---|---|
nexus.backups | nexus-gitops 9f6aa48 (2026-09-05 15:5x UTC) | NexusMongoBackupStale, NexusMongoBackupJobFailed |
nexus.signer | 9f6aa48 | NexusSignerRestarted, NexusSignerDown |
nexus.auth | 9f6aa48 | NexusCoreAuthWouldDeny, NexusCoreAuthDenied (Core nexus_core_auth_decisions_total) |
nexus.agents | 9f6aa48 | NexusAgentRunsFailing |
nexus.financial | 05c3ffd (2026-09-05 ≈ 16:27 UTC) | NexusNavStale (snapshot > 3 h or absent, 15 m), NexusNavIncomplete (2 h), NexusNavDrawdown 10 % / NexusNavDrawdownCritical 20 % (30 m), NexusNavDailyLoss 5 %, NexusStablecoinDepeg (|mark − 1| > 1 %, critical), NexusMarkStale 30 min, NexusSignerUnreachable (critical), NexusSignerPaused, NexusSignerAuthDenied, NexusSignerPolicyDenial (critical), NexusDailyCapNearlySpent 80 %, NexusSleeveUnresolved 30 min, NexusOutboxStuck 30 min (critical), NexusLadderRungFiringStuck 30 min |
nexus.budget | a33055d (deployed 2026-09-05 23:01 UTC) | NexusBudgetUnavailable (nexus_portfolio_budget_available == 0 for 30 m, warning — stale NAV / ladder unreadable; in enforce this denies every risk-increasing action), NexusBudgetFenceRejection (increase(nexus_portfolio_fence_rejections_total[15m]) > 0, critical — a stale executor tried to settle a reservation after a takeover: check for a duplicated money movement, runbook), NexusSolCapBreachedEnforce (nexus_portfolio_cap_headroom_usdt{cap="sol_exposure"} < 0 for 6 h while the worker Deployment exists, info — the F6 owner decision: trim rungs, raise the cap, or accept the sleeve pausing). The same commit adds the Treasury dashboard row "Portfolio budget (caps, reservations, ledger)": budget available, ledger version, reserve CAS conflicts, fence rejections, reserved (ladder), SOL cap / liquid reserve headroom, gate verdicts (24 h increase), cap headroom (USDT) |
Thresholds for the drawdown / daily-loss rules are chart values
monitoring.financial.{drawdownWarnFrac,drawdownCritFrac,dailyLossWarnFrac}
(prod 0.10 / 0.20 / 0.05). The nexus.financial group reads the treasury
metric contract documented under
Measuring success → As built;
the worker exporter that emits it is deployed since 2026-09-05 17:12 UTC
(nexus-platform ac38aa1; scrapeable since ≈ 17:4x UTC after the
component label fix). NexusOutboxStuck reads
nexus_outbox_rows{state} and stays silent while EXEC_OUTBOX=0; no rule
reads the signer's /v1/status.costs_today or the
P&L gauges (nexus_pnl_unexplained_usdt is the candidate)
yet. The four earlier groups are
live on metrics that already existed.
Apalis queue: orphaned in-flight jobs
The worker's job queue is Apalis on Redis (nexus_jobs:*). A job a pod
fetched but never finished sits in nexus_jobs:inflight:<consumer> and is
reclaimed as an orphan only once that consumer's heartbeat in
nexus_jobs:consumers has expired. With a shared consumer name that
never happens during a rollout: the new pod's heartbeat keeps the name
alive, so the outgoing pod's job stays in-flight forever. This lost the
first treasury-nav run on 2026-09-05 (17:12 UTC rollout of ac38aa1;
the old binary fetched a job kind it could not decode).
Fixed on both sides, both rolled with the same day's images: the worker
Deployment uses strategy: Recreate (nexus-gitops 441291a, one consumer
version at a time) and the consumer name is per pod,
nexus-jobs-<POD_NAME> (nexus-platform 9063b8d, POD_NAME from the
downward API, gitops 078055c) — a dead pod's name stops heartbeating and
its in-flight jobs are reclaimed.
Runbook, if it recurs (a scheduled job "ran" but wrote nothing):
-
Symptom — the job id is in
nexus_jobs:inflight:<consumer>and there is no completion log line for it on the running pod (kubectl logs deploy/nexus-worker | grep <job>). -
Confirm —
HGETALL nexus_jobs:consumers(or the equivalent Redis read): if the consumer holding the job is still heartbeating and is not the running pod, the job is orphaned behind a live name. -
Re-enqueue by hand (the worker's own Redis,
nexus-redisinnexus):SPOP nexus_jobs:inflight:<consumer> # returns the job idRPUSH nexus_jobs:active <id>RPUSH nexus_jobs:signal 1Then watch the job's log line / its collection (for
treasury-nav, a newaccum_nav_snapshotsrow).SPOPon a set with several members pops an arbitrary one —SMEMBERSfirst andSREMthe exact id if the set is not a singleton. -
Never
DELthe inflight set: the job is gone, not re-queued.
Portfolio budget gate: log → enforce
The strategy-wide caps (audit 2026-09-05 F6 / H5,
HARVEST §5) run in
log mode in prod since 2026-09-05 (nexus-gitops d848adf,
worker.budget.mode: log): every risk-increasing movement is evaluated
against the caps and reserved in accum_reservations, but a movement the
caps would deny is allowed and counted. Since nexus-platform fda55f4
(live with fdd7912, 21:54 UTC) log mode never denies — not even on a
stale NAV: the 18:36 / 19:06 UTC incident, where budget_unavailable
bypassed log mode and refused hygiene for two sweeps, is fixed
(e6ce1af fresh valuation on a stale stored snapshot, 84f29a5 oracle
re-read on stale marks).
Before flipping:
-
Watch, over at least a day of sweeps (hygiene runs every 30 min):
sum by (reason) (increase(nexus_accum_actions_gated_total{class="budget"}[24h]))Every
would_block:<reason>series is a movement enforce would have stopped. On 2026-09-05 the dry run predictedwould_block:sol_exposure_capandwould_block:liquid_reserveon every swing buy / LP add while the ladder is fully armed (projected SOL exposure 69 % vs the 60 % cap) — so the flip is preceded by an owner decision: trim the deep rungs, raiseworker.budget.maxSolExposureFrac(≈0.75clears the ladder as configured today), or accept the sleeve pausing whenever the ladder is armed. -
Check
GET /v1/desk/budget(available: true,headroom[]per cap,reservations[]only what you expect) andnexus_portfolio_budget_available == 1— in enforce, a stale NAV (nav_max_age_secs7 200) denies every risk-increasing action withbudget_unavailable, so the hourlytreasury-navjob must be healthy. -
Flip:
worker.budget.mode: enforceinenvironments/prod/values.yaml(optionallymaxSolExposureFrac,maxBtcExposureFrac,minLiquidReserveUsdc; any other limit is anACCUM_LIMIT_*env on the worker). Argo rolls the worker (Recreate, one short gap in sweeps). -
After: the same query must show the plain reasons (
sol_exposure_cap, …) instead ofwould_block:; a denial also writes⛔ budget: …into the🧹 ACCUM HYGIENETelegram report. Rollback ismode: logagain.
The gate's defaults and the reservation model are on HARVEST §5; the API shape on the API reference.
What enforce does not give you (owner re-verification 2026-09-05,
register F6), and what changed since: on the running worker (fdd7912)
the gate is a worker-only projection — exec-plane rung fires, direct
execute-role calls to the signer and Telegram approvals happen outside it
— and two independent sweeps can read the same headroom and both reserve
(no aggregate CAS across reservation ids). nexus-platform ee0cd80 and nexus-solana a1eb8f3 (both deployed and
live-verified 2026-09-05 23:33–23:4x UTC)
close the first two of those: every reservation is projected against one
CAS document (accum_budget_ledger), Telegram-approved buys reserve on
it, and the exec reads Core's verdict before a rung fires — the
Budget ledger runbook below. Direct execute-role calls
to the signer stay outside it, so a flip to enforce is still a policy
step, not a global financial lock; the signer's per-action / daily caps
stay the last line. Before flipping after that roll, add to step 2:
nexus_portfolio_ledger_version advancing every sweep,
nexus_portfolio_fence_rejections_total at 0, and the exec's
/v1/status.budget_check.last_fetch.ok == true.
Budget ledger: inspecting, fence rejections, stuck reservations
Applies from nexus-platform ee0cd80 (deployed 2026-09-05, ledger
live-verified 23:4x UTC). Model and field names:
HARVEST §5 → Budget ledger;
routes: API reference → Desk budget mutations.
Inspect. The authority is one document — nexus.accum_budget_ledger,
_id: "portfolio". Read it through Core (open read, no key needed):
curl -s http://nexus-core/v1/desk/budget | jq '.verdict, .ledger.version, (.ledger.entries[] | {id, state, fence, lease, amount_usdc, expires_at_ms})'
or directly in Mongo (nexus_rw / nexus_admin credentials from the
nexus-app Secret):
db.accum_budget_ledger.findOne({_id: "portfolio"}, {version: 1, nav_snapshot_id: 1, limits_version: 1, totals_usdc: 1, "reservations.id": 1, "reservations.state": 1, "reservations.fence": 1, "reservations.lease": 1, "reservations.expires_at_ms": 1, updated_at: 1})
What normal looks like: version goes up on every sweep that reserves,
consumes or expires something (the gauge nexus_portfolio_ledger_version
mirrors it — flat for hours means the worker is not sweeping or the
budget is unavailable); the standing ladder_rung:ladder:SOL /
ladder_rung:ladder:BTC entries are always present while the ladder is
armed; executing entries carry a lease whose owner is a live pod
(worker:<pod>) or the approving gateway (telegram:<user>) and an
until_ms in the near future; held entries have no lease and expire at
expires_at_ms (TTL 21 600 s). accum_reservations rows are the
write-through audit trail of the same ids — useful for history, never
for headroom. If the document is missing, the next sweep seeds it from
the live held rows (rebuilt_from_rows_at_ms records that); nothing
to run by hand. nexus_portfolio_reserve_conflicts_total ticking is
expected when jobs overlap (the loser re-reads and re-evaluates); eight
losses in a row on one write is a denial reservation_error and worth a
look at Mongo latency.
A fence rejection (nexus_portfolio_fence_rejections_total + 1,
alert NexusBudgetFenceRejection, critical) means an executor tried to
consume or release an entry with a fence that is no longer current
(StaleFence), or to consume an entry nobody had claimed
(NotClaimed). The first happens when an executor's lease
(claim_lease_secs, default 900 s) ran out, another executor took the
claim over (the fence moved on) and the first one later came back to
settle — i.e. two executors may both have acted on one reservation.
It is a signal, not a loss by itself: find the entry
(.ledger.entries[] | select(.id == …) — last_owner is who holds it
now, the worker's error line accum: reservation settlement REFUSED — execution claim was taken over names the presented and current fences), then check the exec for two sends under
the same reference (GET /v1/sleeve for a swing buy, /v1/positions for
a supply, /v1/status.costs_today.sends) and reconcile the one that
should not have happened through the sleeve or lending routes. In log
mode the movement itself was never blocked, so the rejection is the only
trace of the overlap.
Releasing a stuck reservation. A held entry is never stuck: it
expires at expires_at_ms and the next sweep releases it (released_reason: "expired"). An executing entry whose owner died stays executing until
its lease ends; after that it is taken over by the next executor that
needs it (the worker re-claims its own ids each sweep) — the takeover is
the recovery, no operator action needed. An executing entry counts
against headroom only while its TTL (expires_at_ms) or its lease is
live, so even an abandoned claim stops weighing on the caps after
21 600 s. Only when an entry's lease is dead, nobody will claim it again
(a Telegram approval that was abandoned, a pod that will not come back)
and you need the headroom back before the TTL, release it by hand with a
service key (the read key gets 403), claiming first so the fence is
yours:
# 1. take the claim over (the lease must be dead — 409 lease_held otherwise)
curl -s -X POST -H "X-Nexus-Api-Key: $NEXUS_API_KEY" -H 'content-type: application/json' \
http://nexus-core/v1/desk/budget/reservations/swing_buy:approval:abc:0/claim \
-d '{"owner":"operator:<name>"}' # → {"id","fence":N,"lease_until_ms","took_over_from"}
# 2. release with that fence (never consume — consume says the money moved)
curl -s -X POST -H "X-Nexus-Api-Key: $NEXUS_API_KEY" -H 'content-type: application/json' \
http://nexus-core/v1/desk/budget/reservations/swing_buy:approval:abc:0/release \
-d '{"fence":N,"reason":"operator: approval abandoned, no send on the exec"}'
Never release an entry whose owner might still be executing (a live
until_ms): the call answers 409 lease_held for exactly that reason,
and forcing it by editing the document would let the running executor's
settlement be refused as stale while its money moved. A standing
ladder_rung entry is not released by hand — cancel rungs on the exec
(POST /v1/ladder/cancel) and the next sweep shrinks the standing
reservation.
Residual acceptance checks carried from the audit (the L-rows on the register crosswalk), for the operator to keep in view next to these runbooks:
- L6 — swing action cooldown, deep-bottom freeze, yield-tier / hot-wallet limits and the SOL↔JitoSOL clip guard are still not code behind the signer gate; effective reserve floors (code 0.2 SOL / 5 000 USDC vs the desk's 0.5 SOL) have to be checked across a worker restart, not read once.
- L7 — the LP fractions above are observational; allowed-pool validation, fee return net of impermanent loss, range-health alarms and deterministic unwind are separate controls (no open LP today).
- L9 — the exec's second RPC (
SOLANA_RPC_URLS, since 2026-09-05 15:5x UTC) has never been exercised by an outage; account-owner, discriminator, feed-id, full-verification, confidence and slot-lag checks on the oracle remain open. There is no failover runbook yet.
Lending hygiene: knobs, gauges, vetoing a decision
The standing lending hygiene runs under the owner lending policy of 2026-09-05 — yield-first, rare moves; the USDC stack stays in the JLP reserve, 85–99 % utilisation is the accepted band (85–95 % until 2026-09-24; utilisation alone no longer recalls above 20× liquidity), the tripwire is the liquidity floor (HARVEST §5). Every rule is confirmed across sweeps; one read never moves money.
Knobs — chart worker.hygiene.* (nexus-gitops d5318ad, live since
≈ 21:15 UTC; prod inherits the chart values) → worker env. Empty = code
default, which is the same number. Every knob is read by the running
worker since nexus-platform fdd7912 (21:54 UTC); between ≈ 21:15 and
21:54 UTC the previous worker (7d6ed59) read only the three lines
(ACCUM_UTIL_RECALL, ACCUM_UTIL_REPARK, ACCUM_ROTATE_MIN_SPREAD,
single read).
| Env | Values key | Prod | Meaning |
|---|---|---|---|
ACCUM_UTIL_RECALL | utilRecall | 0.99 (prod since 2026-09-24; chart default 0.95) | withdraw-all line, confirmed over …_CONFIRM_READS sweeps — exempt while liquidity ≥ …_LIQ_EXEMPT_X × our position |
ACCUM_UTIL_RECALL_HARD | utilRecallHard | 0.995 (prod since 2026-09-24; chart default 0.985) | single-read emergency recall — same exemption |
ACCUM_UTIL_RECALL_LIQ_EXEMPT_X | utilRecallLiqExemptX | 20 | 2026-09-24: utilisation alone never recalls while available liquidity covers this × our position; the liquidity floor still recalls on its own. Kamino main sat at 93–99.9 % every night and the old lines pulled the whole stack out and back eleven nights running with $1–8 M free against $60 k. 0 disables |
ACCUM_UTIL_RECALL_CONFIRM_READS | utilRecallConfirmReads | 2 | consecutive sweeps a recall condition (utilisation or liquidity floor) must hold; floored at 1 |
ACCUM_LENDING_LIQ_FLOOR_X | liqFloorX | 10 | available liquidity must cover this × our position (recall) / × the amount moved (destination) |
ACCUM_UTIL_REPARK | utilRepark | 0.88 | a destination must sit under this |
ACCUM_ROTATE_MIN_SPREAD | rotateMinSpread | 0.02 | APY advantage a destination needs (2 pp) |
ACCUM_ROTATE_CONFIRM_SWEEPS | rotateConfirmSweeps | 48 | consecutive sweeps the spread must hold (≈ 24 h at the 30-min cadence); floored at 1 |
ACCUM_ROTATE_COOLDOWN_SECS | rotateCooldownSecs | 604800 | 7 d between rotations of the same money |
ACCUM_LENDING_MIN_HOLD_SECS | minHoldSecs | 172800 | 48 h before a supplied position may rotate |
Unparsable or negative values keep the default. ACCUM_REPARK_VENUES (the
allowlisted supply venues, kamino) is unchanged.
Gauges (worker exporter, since fdd7912, 21:54 UTC):
| Metric | Labels | Meaning |
|---|---|---|
nexus_accum_lending_moves_total | kind = recall / rotate / repark / held | movements executed, or wanted and withheld with a reason; held also counts an officer round trip the sweep refused |
nexus_lending_reserve_utilization | venue, label | last read per reserve |
nexus_lending_reserve_liquidity_x_position | venue, label | available liquidity as a multiple of our position — the floor is 10 |
nexus_lending_rotation_cooldown_seconds_left | venue | time before the money there may rotate again |
Reading them: liquidity_x_position falling toward 10 on the reserve that
holds the position is the tripwire; utilisation inside 85–95 % is the
accepted band, not an alarm; a rising held count with no rotate is the
policy working. The held reasons (spread 2.1 pp held 5/48 sweeps,
cooldown 6d 3h left, dest liquidity 8× < 10×) are in the
🧹 ACCUM HYGIENE Telegram report. No alert rule reads these yet.
Vetoing a decision before the sweep executes it (the executed-marker
veto). The executor runs the newest accum_decision_json digest once,
then writes a marker digest accum_exec_done whose text is that
decision's ts (the RFC 3339 ts field of the accum_decision_json
row); a decision whose ts matches one of the last 10 markers is skipped
as "decision already executed". The 20:29 UTC round-trip decision of
2026-09-05 is stopped exactly that way — its marker (acted=0) was written
by the executor after the withdraw leg failed, and the owner's veto keeps
it there. To stop a decision that has not run yet, insert the marker
by hand into nexus.operator_digests before the next accum-execute
sweep (every 30 min):
// mongosh, database `nexus`
const d = db.operator_digests.find({ kind: "accum_decision_json" })
.sort({ _id: -1 }).limit(1).next();
db.operator_digests.insertOne({ kind: "accum_exec_done", text: d.ts,
ts: new Date().toISOString() });
To skip only the money movements and let the order pass run, insert
kind: "accum_exec_done_treasury" instead (treasury actions carry their
own once-per-decision marker); since 2026-09-07 LP actions carry a third,
accum_exec_done_lp, so an LP open the signer refused before signing can
be retried without replaying the funding legs: delete the decision's
accum_exec_done and accum_exec_done_lp markers, keep
accum_exec_done_treasury, and trigger accum_execute — the exec re-arms
the same lp_add intent (rule 5) while the unstake / withdraw intents
would only have been replayed (and their flows recorded twice). The veto is per decision — the next cycle's
decision executes normally — which is why the desk prompt carries the
policy too (nexus-gitops → worker.hygiene).
The desk's session store: why a cycle loses a step
The accum mob-host keeps every member session in one SQLite file on its realm
PVC, /realms/nexus-accum/sessions.sqlite3 (meerkat's store: WAL,
synchronous=FULL, busy_timeout 5 s). Nothing prunes it, and a member
session is not small: metadata_json.session_transcript_history_state_v1
accumulates the member's whole transcript — about 1.8 MB per member per
cycle — and the row is rewritten on every turn. A host that has been up a
day carries 60 MB rows; the file itself reached 3.97 GB on a 5 GiB PVC by
2026-09-08 (652 sessions, oldest 08-28).
That is what a lost specialist read looks like from the outside:
accum_health {"missing":["structure"],"steps_ok":4,"degraded":true}
host log session-store projection update failed after committed runtime
checkpoint; quarantining rejected runtime snapshot
… SQLite error: database is locked
flow terminal accum-cycle Failed output_keys ["venue","flows","narrative","challenge"]
The analyst ran (its turn is in Langfuse with a normal latency); rewriting its 60 MB row while another flow held the write lock took longer than the 5 s busy timeout, so the runtime quarantined the snapshot and unregistered the session and the output never reached the panel. The Thesis Officer then writes the brief headed "PANEL INCOMPLETE — 4 of 5". The next cycle can fail harder: no terminal at all, no digest, the store untouched and zero open fds on it — a wedge.
Checking it (the host has python3, no sqlite3 CLI; read-only is safe):
kubectl -n nexus exec nexus-accum-host-0 -- sh -c 'ls -la /realms/nexus-accum/; df -h /realms'
kubectl -n nexus exec nexus-accum-host-0 -- python3 -c "
import sqlite3
c=sqlite3.connect('file:/realms/nexus-accum/sessions.sqlite3?mode=ro', uri=True, timeout=5)
q='select session_id, message_count, length(cast(session_json as blob))+length(cast(metadata_json as blob)) sz from sessions order by sz desc limit 5'
[print(r) for r in c.execute(q)]"
Anything over ~1 GB total, or single sessions over ~20 MB, is the warning.
Resetting it. Sessions are working memory — the desk's durable memory is in Mongo (digests, theses, knowledge) and Langfuse — so a reset costs one rollover, nothing more. The host recreates the schema on boot, so the reset is a file move plus a pod delete, done while the host is idle between cycles (it holds no fd on the store then):
kubectl -n nexus exec nexus-accum-host-0 -- sh -c \
'ls -l /proc/1/fd | grep -c sessions.sqlite3' # must be 0
kubectl -n nexus exec nexus-accum-host-0 -- mv \
/realms/nexus-accum/sessions.sqlite3 /realms/nexus-accum/sessions.sqlite3.bak-$(date +%Y%m%d)
kubectl -n nexus delete pod nexus-accum-host-0
Argo CD runs selfHeal: true on this application: kubectl scale statefulset nexus-accum-host --replicas=0 is reverted within seconds and the pod comes
back. Swap the file in place instead.
The two jobs that keep it from happening (both were missing from
scheduled_jobs until 2026-09-08, which is why the 07:57 wedge went unwatched):
| Job | Cadence | What it does |
|---|---|---|
mob_watchdog | 10 min | Langfuse silence > 2× the mob's cadence, host idle ⇒ restart the pod. The only net that catches a wedge — accum_cycle_health's degraded streak cannot, because a cycle that never terminates never records a degraded verdict. |
session_rollover | 6 h | Restarts the host when idle so sessions never grow past ~11 MB. Defers on the mid-cycle busy gate. |
Confirm they are armed with GET /v1/schedules (kinds mob_watchdog /
session_rollover); create one with POST /v1/schedules and a body like
{"name":…,"kind":"mob_watchdog","schedule":{"kind":"interval","minutes":10}, "args":{},"deliver":"local","enabled":true,"state":"scheduled"}.
Triggering a cycle by hand, and why it may do nothing
Every accum-* job is a row in nexus.scheduled_jobs. To make one due
immediately:
curl -X POST "$NEXUS_CORE/v1/schedules/<id>/run" \
-H "x-nexus-api-key: $NEXUS_API_KEY"
POST /v1/schedules/{id}/run does not enqueue the job: it sets
next_run_at to the epoch and clears last_fire_key, so the scheduler
picks it up on its next tick. It answers {"ok": true}, or 404 when no
document has that _id. It is a mutation, so under
NEXUS_AUTH_MODE=enforce it needs a service key in
x-nexus-api-key — the read-only key gets 403 read_only_key, and the UI
proxy principal additionally needs its X-Nexus-Proxy-Origin attestation.
Ids: job_treasury_nav is a code constant
(apps/nexus-worker/src/treasury/mod.rs) and is seeded by the worker. The
accumulation ids — the operator uses job_accum_execute — are not
defined anywhere in nexus-platform; they are _ids of documents created
out of band. List them before scripting:
// mongosh, database `nexus`
db.scheduled_jobs.find({}, { _id: 1, name: 1, enabled: 1, next_run_at: 1 })
Three reasons a triggered accum_execute correctly moves no money:
- The decision pass short-circuits on its marker. The executor runs
the newest
accum_decision_jsondigest once and then writes anaccum_exec_donedigest whosetextis that decision'sts; a decision whosetsmatches one of the last 10 markers is skipped as "decision already executed". (accum_exec_done_treasuryis the separate marker for the treasury half andaccum_exec_done_lpfor the LP half — the veto procedure above uses all three.) - The hygiene planner will not move less than
MIN_MOVE_USD— a hard-coded 1 000 USD constant inapps/nexus-worker/src/mob_jobs/lending_hygiene.rs, not an env var. It gates both the planner's decisions and the chunked rotation / repark loops. - A
lending_supplyis capped atexec USDC − float floorand skipped below 1 USDC (HARVEST §5).
That combination is exactly what happened on 2026-09-06 ≈ 09:0x UTC: the
operator-triggered sweep moved nothing (moves_money: false) because the
995.55 USDC of idle float was under MIN_MOVE_USD and the decision pass
had its done-marker. The path was exercised directly instead — see below.
A verified live lending supply through the fixed path
2026-09-06 ≈ 09:0x UTC, the first exercise of the F2 idempotency path and
the proof that the -32002 blockhash defect is gone:
POST /v1/lending/supply { amount: 995.55 USDC → kamino JLP, intent_id: <fresh> }
→ sent: true, signature 4pv2ZsbL…, confirmed
What to check on such a supply, and what it read that day:
| Check | Observed |
|---|---|
| exec USDC before → after | 5,995.55 → 5,000.00 |
| venue position | JLP 56,448.93 → 57,445.10 @ 11.02 % |
costs.tx_fee_lamports | 5 000 lamports ($0.0005), rent_lamports: 0 |
| S1 receipt bound (pre-sign) | min 821,596,315 cTokens, 1 % tolerance — fired before signing |
GET /v1/intents/{id} | confirmed, attempts: 1 |
GET /v1/positions after | the new figure, not the pre-write one (the write invalidated the venue) |
The last row is the point of the position cache:
before ea96739 the aggregate served 56,448.93 while
/v1/lending/position already read 57,445.10 — and NAV and the budget
gate read the aggregate.
Reading exec positions: the cache and ?fresh
Since nexus-solana ea96739 (2026-09-06) GET /v1/positions and
GET /v1/lp/positions are served from a snapshot refreshed every
EXEC_POSITIONS_REFRESH_SECS (default 30 s, floor 5). Every answer carries
as_of, stale and venues_unreadable[].
stale: truemeans older than 3 refresh intervals, or a half has never been read, or a venue is unreadable. Treat it as "do not use this figure to size a movement".venues_unreadablenames the venue whose scan failed. A failed scan keeps the last good value — it never becomes an empty list, because an LP position the sleeve reconciler cannot see is indistinguishable from coin missing from the wallet.?freshforces a synchronous full scan (≈ 10 s of RPC). It accepts1,true,yesor a bare?fresh;?fresh=0and absence do not. Use it for an operator read that must not be a cache hit — never in a loop or a dashboard.- A confirmed lending supply / withdraw / withdraw-all or LP open /
remove / collect / close marks its venue dirty, so the next read
refreshes that venue before answering. You do not need
?freshafter your own write. LENDING_CACHE_SECSnow governs onlyGET /v1/lending/rates.
The LP position entering its range
2026-09-06, the behaviour the liquidity page describes as an automated swing. The Orca SOL/USDC position opened at 08:29 UTC with 10 SOL into the aligned range [105.96, 112.02] while spot was 105.12 — i.e. out of range, 100 % SOL, earning nothing, which is a laddered sell from 106 to 112, not a broken position. When the price crossed back in, the position became 9.26 SOL + 78.45 USDC and took its first fees: 0.001584 SOL + 0.1963 USDC.
Two consequences for an operator reading the pages: an out-of-range
position showing "earning nothing" is the range doing its job, not an
incident; and the LP legs are counted in the exposure families since
nexus-platform 72967bd, so opening one moves exposure.SOL and the F6
sol_exposure headroom (LP as legs).
Deploying the exec: the price-plane window
Every exec / oracle deploy produces ≈ 10–15 s of price plane unreadable, and during it the H1 depeg guard fails closed: signing
is denied, and a rung that triggers in that window is not fired.
Why: both binaries ship in the same image, so an exec rollout restarts the oracle pod too, and the exec's 5 s price fetch has no retry. The direction is safe — an unreadable oracle denies USDC actions rather than guessing — and it is self-inflicted only on deploys.
What to do: nothing during the window. Confirm afterwards with
GET /v1/status.policy (the depeg band readable again) and the exec log,
and do not interpret a denial timestamped within a minute of a rollout as
a policy incident. Queued improvement (not built): keep the last good
USDC cross for EXEC_DEPEG_GRACE_SECS (≈ 60 s) with its age already
exposed on /v1/status.policy.depeg.age_s, and only then go Unreadable.
Canonical host: three ways a redirect can fail
www.nexusapp.dev could never hold a session — the session cookie is set
without a domain, so it is host-only — and every gated page there looped
through /sign-in. Getting the 308 in place took three attempts, each a
failure mode worth recognising:
- An ingress
permanent-redirectwith the wrong backend port. The redirect Ingress named port 3000 while thenexus-uiService listens on 80; the backend was unresolvable and nginx answered a plain 404 — not an error mentioning the port (gitops329bab4→d41d779). - ingress-nginx rejected the annotation value because it contained
$request_uri. Variables are not allowed in that annotation; the Ingress is refused at admission, so the symptom is a failed apply, not a bad redirect. - Build-time inlining in Next.js.
redirects()innext.configis baked into the routes manifest at build time, andprocess.envinside Edge middleware is inlined at build time too — so a runtimeCANONICAL_HOSTfrom the Secret was invisible to both.
As built (nexus-ui ca1d75e, built 2026-09-06): the rule lives in
src/middleware.ts, derives the target from the request's own Host
header by stripping a leading www., and honours CANONICAL_HOST when
the runtime does expose it. Aliases route normally to nexus-ui:80 again
(gitops 91232c7) and bootstrap-secrets.sh derives CANONICAL_HOST from
APP_BASE_URL. General lesson: in a Next.js app, anything that must be
configurable at runtime cannot live in next.config or in Edge
process.env — derive it from the request, or accept a rebuild.
The UI's exec proxy budget
nexus-ui proxies exec reads through /api/solana/[...path] with a
per-method timeout: GET_TIMEOUT_MS 25 000 since 4dca5c0
(2026-09-06, was 8 000) and POST_TIMEOUT_MS 600 000 for catalog
refreshes. The 8 s budget aborted the 9.5 s LP position scan, so
/trading/liquidity rendered "positions unavailable" over a live Orca
position. Both halves were fixed: the budget here, and the scan itself on
the exec side (9.5 s → ≈ 2 ms, above). If a page
reports a venue "unavailable", check the exec's own latency before
assuming the venue is down.
Platform logs (Admin UI)
Every component installs a tracing subscriber that (1) writes JSON to
stdout (for kubectl logs) and (2) ships structured records over NATS
(plain subject logs.<component>, not JetStream). The worker runs a single
log sink that tails logs.> and writes into a capped MongoDB collection
(logs). The Admin UI Logs page (GET /v1/logs) reads that collection —
filter by component / level / text, with live polling. This is the fastest way
to confirm agents connect, NATS/Mongo are reachable, and LLM calls aren't
erroring. See Logging.
Backups
As built (audit H9, partially addressed):
- MongoDB (
nexus-mongodb): single replica on amicrok8s-hostpathvolume. A dailymongodumpCronJob (03:15, 14-day retention) covers it and the HARVEST desk database — into an on-node 40 Gi PVC. No off-site copy, no encryption, no restore drill has been run. - NATS JetStream, Redis, Qdrant: single replicas, hostPath volumes, no backup.
This is a local backup, not disaster recovery (audit 2026-09-05 O1):
the database and its backup share the same node and disk, so it protects
against logical data loss only. Owner decision, 2026-09-05: the off-node
copy is deferred ("not for now") — the local hostpath dump stays the only
backup. The candidate (encrypted nightly copy to the .98 host, integrity
checks, a restore drill into isolated infrastructure) is recorded, not
scheduled. The target (3-node replica set, encrypted off-site copies with
object lock, quarterly restore drills) is not in place. NexusMongoBackupStale
/ NexusMongoBackupJobFailed (group nexus.backups) page when the local
dump stops.
Release flow
Every push to main of nexus-platform, nexus-ui, nexus-solana and
nexus-docs deploys: CI writes the image tag into
environments/prod/values.yaml and Argo rolls it. A PR flow with branch
protection (audit O3) was declined by the owner on 2026-09-05;
direct-to-main stays the deploy flow. Rollback is a values-file edit
(e.g. solanaExec.tag back to the previous sha-…).
Secrets
.envfiles never committed.- Target: production secrets from a secrets manager (Vault / cloud secret manager / sealed-secrets in GitOps) with scheduled rotation.
- As built (2026-09-05): Kubernetes Secrets are created by
nexus-gitops/bootstrap-secrets.shfrom the git-ignored.secretsfile; no secrets manager, no scheduled rotation — rotations are ad hoc. - Agents can run on a subscription CLI login (Claude/Codex) instead of an API key — see LLM backends.
Production checklist
- MongoDB replica set + backups tested via restore — daily on-node dump only; off-node copy deferred by the owner (2026-09-05)
- NATS authenticated — 2026-09-05:
nexus/runnerusers, runner subject ACL (audit H6); still a single node - Agent namespace + scoped RBAC applied — enforced since 2026-09-05: authorizer
Node,RBAC(audit C1) - Metrics + alerting deployed and verified — 2026-09-05: /metrics live, Alertmanager → Telegram,
nexus-platform-alerts(5 groups); the NAV exporter is live since 17:12 UTC (scrapeable since ≈ 17:4x) - Secrets in a manager, rotation scheduled
- Approval gates configured for sensitive actions — Telegram approvals run actions in order; a failed action leaves the approval
failedand pending (2026-09-05) - Quarantine policy thresholds set
- Taiga integration (if used) rate-limited + webhook HMAC-SHA1 verified
- Telegram allowlist locked down — IDs set; empty list denies everyone since
445e107(audit C7)