Skip to main content

Operations

Operate the AI portfolio through the portfolio console. This reference covers the underlying services and Kubernetes/GitOps deployment.

Prerequisites​

  • Rust (stable) — rustup default stable
  • Node 22+ and pnpm (for nexus-ui)
  • MongoDB 6+ reachable on the network
  • NATS with JetStream enabled
  • A Kubernetes cluster + kubectl for agent runs
  • A Telegram bot token (optional, for the gateway)
  • A Taiga project + service-account credentials (optional, for board sync)

Secrets never reach a log line​

nexus-observability does not merely print events: it captures every one, publishes it to NATS and persists it to MongoDB for the Admin UI's platform log page. A credential logged after the bus is up is therefore written to a database and rendered in a web page, not just left in kubectl logs.

Found 2026-09-06: nexus-accum-host-0 logged its full NATS URL at INFO on every boot, password included. Measured blast radius — smaller than it first looked: the connect line is emitted inside nexus_events::connect(), i.e. before the NATS transport that ships logs exists, so it never reached MongoDB. A scan of the 671k-row logs collection found 0 rows carrying an unredacted nats://user:pass@ URL, and the cluster runs no node-level log collector (no promtail/loki/fluent/vector). Actual exposure was pod stdout via kubectl logs only. The rule below still holds for any credential logged after startup — that one really would be persisted and rendered.

Fixed in nexus-platform e340b58 with nexus_observability::redact_url, applied at both sites that logged a credential-bearing URL — the JetStream connect line and the agent runner's clone line and failure message (repo.url can already carry a token, see authed_url). The boot line now reads nats://nexus:***@nexus-nats.nexus-nats.svc.cluster.local:4222.

The rule: never log a connection string raw — pass it through redact_url first. It keeps the username (useful, not secret), replaces the password with ***, redacts a bare token that is the whole userinfo, and leaves a credential-free or unparseable URL untouched. An @ in a path or query is not mistaken for userinfo.

Both NATS credentials were rotated on 2026-09-06 17:19Z (operator instruction), nexus and runner: new 40-char URL-safe secrets written to nexus-nats-auth (ns nexus-nats) and into NEXUS_NATS_URL / NEXUS_AGENT_NATS_URL in nexus-app (ns nexus), NATS restarted, then all seven consumers. The workspace .secrets was updated to match, and both passwords satisfy bootstrap-secrets.sh's URL-safe charset guard so a future bootstrap reproduces the same URLs. Verified: every pod reconnected, the only authorization violation lines are from outgoing pods inside the ~4 s cutover window — which is itself the proof the old credential is now rejected — and the job pipeline resumed. The old secrets are kept nowhere but the operator's local scratch backup.

Run locally​

MongoDB + NATS​

docker run -d --name nexus-mongo -p 27017:27017 mongo:6
docker run -d --name nexus-nats -p 4222:4222 nats:latest -js

Nexus Core​

cd apps/nexus-core
cp .env.example .env
cargo run --release

Minimum settings:

MONGODB_URL=mongodb://localhost:27017
MONGODB_DB=nexus
NATS_URL=nats://localhost:4222
KUBE_NAMESPACE=nexus-agents
JWT_SIGNING_KEY=<random string>

Verify:

curl http://localhost:8080/healthz
open http://localhost:8080/swagger-ui/

Worker​

cd apps/nexus-worker
cargo run --release

UI​

The admin UI is a separate repo (nexus-ui):

cd ../nexus-ui # the standalone Next.js repo
cp .env.example .env.local
pnpm install
pnpm dev # http://localhost:3000
NEXUS_CORE_URL=http://localhost:8080
NEXUS_ADMIN_TOKEN=dev-token
AUTH_SECRET=at-least-16-chars-long-secret

Deploy on Kubernetes (GitOps)​

Deployment lives in nexus-gitops: Helm charts, per-environment values, Argo CD apps / Flux manifests, and the MongoDB

  • NATS manifests.
# App-of-apps: applies every Application under argocd/apps (project, infra,
# nexus chart with prod values, redis, langfuse)
kubectl apply -f argocd/root.yaml

Image tags are written into environments/prod/values.yaml by each repo's CI on push to main; Argo CD deploys within its ~3-minute poll.

Namespaces​

Run agent jobs in a dedicated namespace (e.g. nexus-agents), separate from the control plane (nexus). Core's service account can create Jobs only in the agent namespace.

Scoped RBAC​

verbs: [create, get, list, watch, delete]
resources: [jobs, pods, pods/log]

See Security & permissions.

Observability​

Every service built on nexus-observability serves Prometheus metrics on :9100 (NEXUS_METRICS_ADDR; off disables — agent runner Jobs set that). This became true with nexus-platform f183437 (2026-09-05); before it the names below were constants nothing recorded and the port refused connections.

nexus_agent_runs_total{agent,outcome} recorded (worker run_job)
nexus_agent_run_failures_total{agent,reason} recorded (reason = timeout | error)
nexus_task_duration_seconds{job} recorded
process_cpu_seconds_total / process_resident_memory_bytes / process_open_fds
nexus_board_sync_errors_total registered, no call sites yet
nexus_llm_tokens_total / nexus_llm_cost_usd_total registered, no call sites yet

Since nexus-platform ac38aa1 (2026-09-05 17:12 UTC) the worker also exports the treasury contract (nexus_nav_*, nexus_benchmark_*, nexus_mark_*, nexus_exec_*, nexus_sleeve_*, nexus_outbox_*, nexus_ladder_* — table on Measuring success) and, since 7e1adcb, the budget gate's nexus_portfolio_* gauges and nexus_accum_actions_gated_total{class,reason}; since fdd7912 (21:54 UTC) the lending-hygiene gauges listed under Lending hygiene; from ee0cd80 (deployed 2026-09-05, live-verified 23:4x UTC) the budget ledger's nexus_portfolio_ledger_version (the CAS token — should only ever go up), nexus_portfolio_reserve_conflicts_total (lost CAS writes, retried) and nexus_portfolio_fence_rejections_total (settlements refused for a stale / missing fence — see Budget ledger).

P&L-by-source gauges (deployed 2026-09-05, first rows 21:12 UTC)​

nexus-platform 750c5ce (pushed ≈ 20:1x UTC with 3a91236 … 7d6ed59) adds the owner's "where are we winning money" gauges to the same worker exporter, refreshed from snapshot.meta.pnl after every treasury_nav run. Semantics and the source vocabulary: Measuring success → P&L by source.

MetricLabelsMeaning
nexus_pnl_source_usdtsource, window = 1d / 7d / 30d / inceptionP&L over the window by source (jito_staking, lending:<venue>, lp_fees:<pool>, sleeve_realized:<pair>, sleeve_unrealized:<pair>, market:SOL|BTC|STABLE|LP, costs:tx|venue|infra, unexplained)
nexus_pnl_unexplained_usdtwindowΔNAV − external flows − Σ explained sources over the window — the reconciliation residual
nexus_sleeve_round_trips_total—closed swing round trips in the ledger (complete history)
nexus_sleeve_realized_usdt_total—sleeve realized P&L over the complete ledger, net of known fees
nexus_sleeve_unrealized_usdt—open sleeve inventory qty × (mark − vwap)
nexus_staking_apr_realized—JitoSOL pool-rate change since inception, annualised
nexus_lending_apr_realizedvenuelending interest since inception over the average balance, annualised
nexus_lending_interest_usdtvenue, windowlending interest over the window

| nexus_pnl_cost_anomalies_total | field | since 2553748 (2026-09-06): on-chain cost figures the sanity rules refused — ≤ 0, non-finite, above PNL_COST_SANITY_USD (25), or a rent_* field on a native-SOL pair. Never bucketed, never subtracted |

Label names follow the rule below (source, window, venue, field — none of them registry-owned). Nothing alerts on these yet: nexus_pnl_unexplained_usdt should get a rule later (a residual that grows period after period means a movement the flow ledger did not see) — not written until a few days of live rows show what "normal" is. Read nexus_pnl_cost_anomalies_total the same way: a flat counter is the rules idle, a rising one means the exec is reporting a cost figure the platform will not book, and it wants a look rather than an alert. Note that unexplained on any window reaching back before 2026-09-05 20:12 UTC is mostly the boundary of the ledger — check series_starts_at first (the residual).

Metric labels: names the registry already owns​

Never use component, job, instance, pod or namespace as a metric label in a Nexus service. nexus-observability adds component to every metric globally, and Prometheus adds job / instance (with pod / namespace from the ServiceMonitor relabeling). A metric that declares one of those names itself produces a duplicate label and Prometheus rejects the whole scrape, not just that series.

This happened on 2026-09-05: nexus_nav_component_usdt{component,asset} shipped in ac38aa1 and every worker scrape failed from 17:12 to ≈ 17:4x UTC (label name "component" is not unique) — the NAV panels stayed empty and NexusNavStale kept firing although snapshots existed. The label is nav_component since nexus-platform 513a173 / d128e17 and dashboard nexus-gitops a5912c7. Symptom to recognise: the target shows DOWN in Prometheus with that error text while :9100/metrics answers 200 in-cluster; the fix is a rename, never a relabel.

Scraping: the nexus-metrics ServiceMonitor in the chart → kube-prometheus-stack (monitoring). Dashboards (Grafana folder Nexus, sidecar ConfigMap nexus-grafana-dashboard): Nexus Platform (runs, failures, durations, process) and, since nexus-gitops 05c3ffd (2026-09-05), Nexus Treasury (uid nexus-treasury: NAV vs benchmarks, excess return, components, exposure, attribution, flows & costs, marks, execution safety; since nexus-gitops b5b6079, 2026-09-05 ≈ 20:10 UTC, the row "Where the money comes from" — P&L by source 30 d / inception bar gauges, unexplained 30 d, swing realized / unrealized / round trips, JitoSOL and Kamino realised APR, lending interest by venue and P&L groups over time; populated since the 21:12 UTC snapshot on worker 7d6ed59). Review 2026-09-07 (gitops e4df941): checked against Prometheus, six panels were empty because they queried window="30d" while the exporter publishes 1d + inception only until the series covers a month — excess return, attribution, P&L by source, unexplained, lending interest and P&L groups now read the inception window (P&L by source also 24 h); "budget gate verdicts" queried a counter that has never incremented and now shows the gate's judged exposure per asset (nexus_portfolio_exposure_after_frac); BTC dust is out of the way — exposure draws assets above 0.5 % of NAV only, the cbBTC mark panel is gone (a $4 holding; the mark stays in Prometheus and on the Treasury page), cap headroom shows the caps in play, and the bar gauges list sources with |amount| > 0.05 USDT; benchmark legends read "hold BTC/SOL" / "DCA BTC/SOL"; USDT panels carry a USDT suffix and the snapshot age is a duration. Every remaining panel query was run against Prometheus after the reload. Alerts: Alertmanager → Telegram group NexusHome (see the operator guide). Traces go to Langfuse (NEXUS_LANGFUSE_*), not OTLP — there is no OpenTelemetry exporter in the tree.

Alerting (PrometheusRule nexus-platform-alerts)​

One PrometheusRule, templates/alerts.yaml in the chart (value monitoring.alerts: true), picked up by the kps Prometheus (empty ruleSelector). Groups and what they watch:

GroupSinceRules
nexus.backupsnexus-gitops 9f6aa48 (2026-09-05 15:5x UTC)NexusMongoBackupStale, NexusMongoBackupJobFailed
nexus.signer9f6aa48NexusSignerRestarted, NexusSignerDown
nexus.auth9f6aa48NexusCoreAuthWouldDeny, NexusCoreAuthDenied (Core nexus_core_auth_decisions_total)
nexus.agents9f6aa48NexusAgentRunsFailing
nexus.financial05c3ffd (2026-09-05 ≈ 16:27 UTC)NexusNavStale (snapshot > 3 h or absent, 15 m), NexusNavIncomplete (2 h), NexusNavDrawdown 10 % / NexusNavDrawdownCritical 20 % (30 m), NexusNavDailyLoss 5 %, NexusStablecoinDepeg (|mark − 1| > 1 %, critical), NexusMarkStale 30 min, NexusSignerUnreachable (critical), NexusSignerPaused, NexusSignerAuthDenied, NexusSignerPolicyDenial (critical), NexusDailyCapNearlySpent 80 %, NexusSleeveUnresolved 30 min, NexusOutboxStuck 30 min (critical), NexusLadderRungFiringStuck 30 min
nexus.budgeta33055d (deployed 2026-09-05 23:01 UTC)NexusBudgetUnavailable (nexus_portfolio_budget_available == 0 for 30 m, warning — stale NAV / ladder unreadable; in enforce this denies every risk-increasing action), NexusBudgetFenceRejection (increase(nexus_portfolio_fence_rejections_total[15m]) > 0, critical — a stale executor tried to settle a reservation after a takeover: check for a duplicated money movement, runbook), NexusSolCapBreachedEnforce (nexus_portfolio_cap_headroom_usdt{cap="sol_exposure"} < 0 for 6 h while the worker Deployment exists, info — the F6 owner decision: trim rungs, raise the cap, or accept the sleeve pausing). The same commit adds the Treasury dashboard row "Portfolio budget (caps, reservations, ledger)": budget available, ledger version, reserve CAS conflicts, fence rejections, reserved (ladder), SOL cap / liquid reserve headroom, gate verdicts (24 h increase), cap headroom (USDT)

Thresholds for the drawdown / daily-loss rules are chart values monitoring.financial.{drawdownWarnFrac,drawdownCritFrac,dailyLossWarnFrac} (prod 0.10 / 0.20 / 0.05). The nexus.financial group reads the treasury metric contract documented under Measuring success → As built; the worker exporter that emits it is deployed since 2026-09-05 17:12 UTC (nexus-platform ac38aa1; scrapeable since ≈ 17:4x UTC after the component label fix). NexusOutboxStuck reads nexus_outbox_rows{state} and stays silent while EXEC_OUTBOX=0; no rule reads the signer's /v1/status.costs_today or the P&L gauges (nexus_pnl_unexplained_usdt is the candidate) yet. The four earlier groups are live on metrics that already existed.

Apalis queue: orphaned in-flight jobs​

The worker's job queue is Apalis on Redis (nexus_jobs:*). A job a pod fetched but never finished sits in nexus_jobs:inflight:<consumer> and is reclaimed as an orphan only once that consumer's heartbeat in nexus_jobs:consumers has expired. With a shared consumer name that never happens during a rollout: the new pod's heartbeat keeps the name alive, so the outgoing pod's job stays in-flight forever. This lost the first treasury-nav run on 2026-09-05 (17:12 UTC rollout of ac38aa1; the old binary fetched a job kind it could not decode).

Fixed on both sides, both rolled with the same day's images: the worker Deployment uses strategy: Recreate (nexus-gitops 441291a, one consumer version at a time) and the consumer name is per pod, nexus-jobs-<POD_NAME> (nexus-platform 9063b8d, POD_NAME from the downward API, gitops 078055c) — a dead pod's name stops heartbeating and its in-flight jobs are reclaimed.

Runbook, if it recurs (a scheduled job "ran" but wrote nothing):

  1. Symptom — the job id is in nexus_jobs:inflight:<consumer> and there is no completion log line for it on the running pod (kubectl logs deploy/nexus-worker | grep <job>).

  2. Confirm — HGETALL nexus_jobs:consumers (or the equivalent Redis read): if the consumer holding the job is still heartbeating and is not the running pod, the job is orphaned behind a live name.

  3. Re-enqueue by hand (the worker's own Redis, nexus-redis in nexus):

    SPOP nexus_jobs:inflight:<consumer> # returns the job id
    RPUSH nexus_jobs:active <id>
    RPUSH nexus_jobs:signal 1

    Then watch the job's log line / its collection (for treasury-nav, a new accum_nav_snapshots row). SPOP on a set with several members pops an arbitrary one — SMEMBERS first and SREM the exact id if the set is not a singleton.

  4. Never DEL the inflight set: the job is gone, not re-queued.

Portfolio budget gate: log → enforce​

The strategy-wide caps (audit 2026-09-05 F6 / H5, HARVEST §5) run in log mode in prod since 2026-09-05 (nexus-gitops d848adf, worker.budget.mode: log): every risk-increasing movement is evaluated against the caps and reserved in accum_reservations, but a movement the caps would deny is allowed and counted. Since nexus-platform fda55f4 (live with fdd7912, 21:54 UTC) log mode never denies — not even on a stale NAV: the 18:36 / 19:06 UTC incident, where budget_unavailable bypassed log mode and refused hygiene for two sweeps, is fixed (e6ce1af fresh valuation on a stale stored snapshot, 84f29a5 oracle re-read on stale marks).

Before flipping:

  1. Watch, over at least a day of sweeps (hygiene runs every 30 min):

    sum by (reason) (increase(nexus_accum_actions_gated_total{class="budget"}[24h]))

    Every would_block:<reason> series is a movement enforce would have stopped. On 2026-09-05 the dry run predicted would_block:sol_exposure_cap and would_block:liquid_reserve on every swing buy / LP add while the ladder is fully armed (projected SOL exposure 69 % vs the 60 % cap) — so the flip is preceded by an owner decision: trim the deep rungs, raise worker.budget.maxSolExposureFrac (≈ 0.75 clears the ladder as configured today), or accept the sleeve pausing whenever the ladder is armed.

  2. Check GET /v1/desk/budget (available: true, headroom[] per cap, reservations[] only what you expect) and nexus_portfolio_budget_available == 1 — in enforce, a stale NAV (nav_max_age_secs 7 200) denies every risk-increasing action with budget_unavailable, so the hourly treasury-nav job must be healthy.

  3. Flip: worker.budget.mode: enforce in environments/prod/values.yaml (optionally maxSolExposureFrac, maxBtcExposureFrac, minLiquidReserveUsdc; any other limit is an ACCUM_LIMIT_* env on the worker). Argo rolls the worker (Recreate, one short gap in sweeps).

  4. After: the same query must show the plain reasons (sol_exposure_cap, …) instead of would_block:; a denial also writes ⛔ budget: … into the 🧹 ACCUM HYGIENE Telegram report. Rollback is mode: log again.

The gate's defaults and the reservation model are on HARVEST §5; the API shape on the API reference.

What enforce does not give you (owner re-verification 2026-09-05, register F6), and what changed since: on the running worker (fdd7912) the gate is a worker-only projection — exec-plane rung fires, direct execute-role calls to the signer and Telegram approvals happen outside it — and two independent sweeps can read the same headroom and both reserve (no aggregate CAS across reservation ids). nexus-platform ee0cd80 and nexus-solana a1eb8f3 (both deployed and live-verified 2026-09-05 23:33–23:4x UTC) close the first two of those: every reservation is projected against one CAS document (accum_budget_ledger), Telegram-approved buys reserve on it, and the exec reads Core's verdict before a rung fires — the Budget ledger runbook below. Direct execute-role calls to the signer stay outside it, so a flip to enforce is still a policy step, not a global financial lock; the signer's per-action / daily caps stay the last line. Before flipping after that roll, add to step 2: nexus_portfolio_ledger_version advancing every sweep, nexus_portfolio_fence_rejections_total at 0, and the exec's /v1/status.budget_check.last_fetch.ok == true.

Budget ledger: inspecting, fence rejections, stuck reservations​

Applies from nexus-platform ee0cd80 (deployed 2026-09-05, ledger live-verified 23:4x UTC). Model and field names: HARVEST §5 → Budget ledger; routes: API reference → Desk budget mutations.

Inspect. The authority is one document — nexus.accum_budget_ledger, _id: "portfolio". Read it through Core (open read, no key needed):

curl -s http://nexus-core/v1/desk/budget | jq '.verdict, .ledger.version, (.ledger.entries[] | {id, state, fence, lease, amount_usdc, expires_at_ms})'

or directly in Mongo (nexus_rw / nexus_admin credentials from the nexus-app Secret):

db.accum_budget_ledger.findOne({_id: "portfolio"}, {version: 1, nav_snapshot_id: 1, limits_version: 1, totals_usdc: 1, "reservations.id": 1, "reservations.state": 1, "reservations.fence": 1, "reservations.lease": 1, "reservations.expires_at_ms": 1, updated_at: 1})

What normal looks like: version goes up on every sweep that reserves, consumes or expires something (the gauge nexus_portfolio_ledger_version mirrors it — flat for hours means the worker is not sweeping or the budget is unavailable); the standing ladder_rung:ladder:SOL / ladder_rung:ladder:BTC entries are always present while the ladder is armed; executing entries carry a lease whose owner is a live pod (worker:<pod>) or the approving gateway (telegram:<user>) and an until_ms in the near future; held entries have no lease and expire at expires_at_ms (TTL 21 600 s). accum_reservations rows are the write-through audit trail of the same ids — useful for history, never for headroom. If the document is missing, the next sweep seeds it from the live held rows (rebuilt_from_rows_at_ms records that); nothing to run by hand. nexus_portfolio_reserve_conflicts_total ticking is expected when jobs overlap (the loser re-reads and re-evaluates); eight losses in a row on one write is a denial reservation_error and worth a look at Mongo latency.

A fence rejection (nexus_portfolio_fence_rejections_total + 1, alert NexusBudgetFenceRejection, critical) means an executor tried to consume or release an entry with a fence that is no longer current (StaleFence), or to consume an entry nobody had claimed (NotClaimed). The first happens when an executor's lease (claim_lease_secs, default 900 s) ran out, another executor took the claim over (the fence moved on) and the first one later came back to settle — i.e. two executors may both have acted on one reservation. It is a signal, not a loss by itself: find the entry (.ledger.entries[] | select(.id == …) — last_owner is who holds it now, the worker's error line accum: reservation settlement REFUSED — execution claim was taken over names the presented and current fences), then check the exec for two sends under the same reference (GET /v1/sleeve for a swing buy, /v1/positions for a supply, /v1/status.costs_today.sends) and reconcile the one that should not have happened through the sleeve or lending routes. In log mode the movement itself was never blocked, so the rejection is the only trace of the overlap.

Releasing a stuck reservation. A held entry is never stuck: it expires at expires_at_ms and the next sweep releases it (released_reason: "expired"). An executing entry whose owner died stays executing until its lease ends; after that it is taken over by the next executor that needs it (the worker re-claims its own ids each sweep) — the takeover is the recovery, no operator action needed. An executing entry counts against headroom only while its TTL (expires_at_ms) or its lease is live, so even an abandoned claim stops weighing on the caps after 21 600 s. Only when an entry's lease is dead, nobody will claim it again (a Telegram approval that was abandoned, a pod that will not come back) and you need the headroom back before the TTL, release it by hand with a service key (the read key gets 403), claiming first so the fence is yours:

# 1. take the claim over (the lease must be dead — 409 lease_held otherwise)
curl -s -X POST -H "X-Nexus-Api-Key: $NEXUS_API_KEY" -H 'content-type: application/json' \
http://nexus-core/v1/desk/budget/reservations/swing_buy:approval:abc:0/claim \
-d '{"owner":"operator:<name>"}' # → {"id","fence":N,"lease_until_ms","took_over_from"}
# 2. release with that fence (never consume — consume says the money moved)
curl -s -X POST -H "X-Nexus-Api-Key: $NEXUS_API_KEY" -H 'content-type: application/json' \
http://nexus-core/v1/desk/budget/reservations/swing_buy:approval:abc:0/release \
-d '{"fence":N,"reason":"operator: approval abandoned, no send on the exec"}'

Never release an entry whose owner might still be executing (a live until_ms): the call answers 409 lease_held for exactly that reason, and forcing it by editing the document would let the running executor's settlement be refused as stale while its money moved. A standing ladder_rung entry is not released by hand — cancel rungs on the exec (POST /v1/ladder/cancel) and the next sweep shrinks the standing reservation.

Residual acceptance checks carried from the audit (the L-rows on the register crosswalk), for the operator to keep in view next to these runbooks:

  • L6 — swing action cooldown, deep-bottom freeze, yield-tier / hot-wallet limits and the SOL↔JitoSOL clip guard are still not code behind the signer gate; effective reserve floors (code 0.2 SOL / 5 000 USDC vs the desk's 0.5 SOL) have to be checked across a worker restart, not read once.
  • L7 — the LP fractions above are observational; allowed-pool validation, fee return net of impermanent loss, range-health alarms and deterministic unwind are separate controls (no open LP today).
  • L9 — the exec's second RPC (SOLANA_RPC_URLS, since 2026-09-05 15:5x UTC) has never been exercised by an outage; account-owner, discriminator, feed-id, full-verification, confidence and slot-lag checks on the oracle remain open. There is no failover runbook yet.

Lending hygiene: knobs, gauges, vetoing a decision​

The standing lending hygiene runs under the owner lending policy of 2026-09-05 — yield-first, rare moves; the USDC stack stays in the JLP reserve, 85–99 % utilisation is the accepted band (85–95 % until 2026-09-24; utilisation alone no longer recalls above 20× liquidity), the tripwire is the liquidity floor (HARVEST §5). Every rule is confirmed across sweeps; one read never moves money.

Knobs — chart worker.hygiene.* (nexus-gitops d5318ad, live since ≈ 21:15 UTC; prod inherits the chart values) → worker env. Empty = code default, which is the same number. Every knob is read by the running worker since nexus-platform fdd7912 (21:54 UTC); between ≈ 21:15 and 21:54 UTC the previous worker (7d6ed59) read only the three lines (ACCUM_UTIL_RECALL, ACCUM_UTIL_REPARK, ACCUM_ROTATE_MIN_SPREAD, single read).

EnvValues keyProdMeaning
ACCUM_UTIL_RECALLutilRecall0.99 (prod since 2026-09-24; chart default 0.95)withdraw-all line, confirmed over …_CONFIRM_READS sweeps — exempt while liquidity ≥ …_LIQ_EXEMPT_X × our position
ACCUM_UTIL_RECALL_HARDutilRecallHard0.995 (prod since 2026-09-24; chart default 0.985)single-read emergency recall — same exemption
ACCUM_UTIL_RECALL_LIQ_EXEMPT_XutilRecallLiqExemptX202026-09-24: utilisation alone never recalls while available liquidity covers this × our position; the liquidity floor still recalls on its own. Kamino main sat at 93–99.9 % every night and the old lines pulled the whole stack out and back eleven nights running with $1–8 M free against $60 k. 0 disables
ACCUM_UTIL_RECALL_CONFIRM_READSutilRecallConfirmReads2consecutive sweeps a recall condition (utilisation or liquidity floor) must hold; floored at 1
ACCUM_LENDING_LIQ_FLOOR_XliqFloorX10available liquidity must cover this × our position (recall) / × the amount moved (destination)
ACCUM_UTIL_REPARKutilRepark0.88a destination must sit under this
ACCUM_ROTATE_MIN_SPREADrotateMinSpread0.02APY advantage a destination needs (2 pp)
ACCUM_ROTATE_CONFIRM_SWEEPSrotateConfirmSweeps48consecutive sweeps the spread must hold (≈ 24 h at the 30-min cadence); floored at 1
ACCUM_ROTATE_COOLDOWN_SECSrotateCooldownSecs6048007 d between rotations of the same money
ACCUM_LENDING_MIN_HOLD_SECSminHoldSecs17280048 h before a supplied position may rotate

Unparsable or negative values keep the default. ACCUM_REPARK_VENUES (the allowlisted supply venues, kamino) is unchanged.

Gauges (worker exporter, since fdd7912, 21:54 UTC):

MetricLabelsMeaning
nexus_accum_lending_moves_totalkind = recall / rotate / repark / heldmovements executed, or wanted and withheld with a reason; held also counts an officer round trip the sweep refused
nexus_lending_reserve_utilizationvenue, labellast read per reserve
nexus_lending_reserve_liquidity_x_positionvenue, labelavailable liquidity as a multiple of our position — the floor is 10
nexus_lending_rotation_cooldown_seconds_leftvenuetime before the money there may rotate again

Reading them: liquidity_x_position falling toward 10 on the reserve that holds the position is the tripwire; utilisation inside 85–95 % is the accepted band, not an alarm; a rising held count with no rotate is the policy working. The held reasons (spread 2.1 pp held 5/48 sweeps, cooldown 6d 3h left, dest liquidity 8× < 10×) are in the 🧹 ACCUM HYGIENE Telegram report. No alert rule reads these yet.

Vetoing a decision before the sweep executes it (the executed-marker veto). The executor runs the newest accum_decision_json digest once, then writes a marker digest accum_exec_done whose text is that decision's ts (the RFC 3339 ts field of the accum_decision_json row); a decision whose ts matches one of the last 10 markers is skipped as "decision already executed". The 20:29 UTC round-trip decision of 2026-09-05 is stopped exactly that way — its marker (acted=0) was written by the executor after the withdraw leg failed, and the owner's veto keeps it there. To stop a decision that has not run yet, insert the marker by hand into nexus.operator_digests before the next accum-execute sweep (every 30 min):

// mongosh, database `nexus`
const d = db.operator_digests.find({ kind: "accum_decision_json" })
.sort({ _id: -1 }).limit(1).next();
db.operator_digests.insertOne({ kind: "accum_exec_done", text: d.ts,
ts: new Date().toISOString() });

To skip only the money movements and let the order pass run, insert kind: "accum_exec_done_treasury" instead (treasury actions carry their own once-per-decision marker); since 2026-09-07 LP actions carry a third, accum_exec_done_lp, so an LP open the signer refused before signing can be retried without replaying the funding legs: delete the decision's accum_exec_done and accum_exec_done_lp markers, keep accum_exec_done_treasury, and trigger accum_execute — the exec re-arms the same lp_add intent (rule 5) while the unstake / withdraw intents would only have been replayed (and their flows recorded twice). The veto is per decision — the next cycle's decision executes normally — which is why the desk prompt carries the policy too (nexus-gitops → worker.hygiene).

The desk's session store: why a cycle loses a step​

The accum mob-host keeps every member session in one SQLite file on its realm PVC, /realms/nexus-accum/sessions.sqlite3 (meerkat's store: WAL, synchronous=FULL, busy_timeout 5 s). Nothing prunes it, and a member session is not small: metadata_json.session_transcript_history_state_v1 accumulates the member's whole transcript — about 1.8 MB per member per cycle — and the row is rewritten on every turn. A host that has been up a day carries 60 MB rows; the file itself reached 3.97 GB on a 5 GiB PVC by 2026-09-08 (652 sessions, oldest 08-28).

That is what a lost specialist read looks like from the outside:

accum_health {"missing":["structure"],"steps_ok":4,"degraded":true}
host log session-store projection update failed after committed runtime
checkpoint; quarantining rejected runtime snapshot
… SQLite error: database is locked
flow terminal accum-cycle Failed output_keys ["venue","flows","narrative","challenge"]

The analyst ran (its turn is in Langfuse with a normal latency); rewriting its 60 MB row while another flow held the write lock took longer than the 5 s busy timeout, so the runtime quarantined the snapshot and unregistered the session and the output never reached the panel. The Thesis Officer then writes the brief headed "PANEL INCOMPLETE — 4 of 5". The next cycle can fail harder: no terminal at all, no digest, the store untouched and zero open fds on it — a wedge.

Checking it (the host has python3, no sqlite3 CLI; read-only is safe):

kubectl -n nexus exec nexus-accum-host-0 -- sh -c 'ls -la /realms/nexus-accum/; df -h /realms'
kubectl -n nexus exec nexus-accum-host-0 -- python3 -c "
import sqlite3
c=sqlite3.connect('file:/realms/nexus-accum/sessions.sqlite3?mode=ro', uri=True, timeout=5)
q='select session_id, message_count, length(cast(session_json as blob))+length(cast(metadata_json as blob)) sz from sessions order by sz desc limit 5'
[print(r) for r in c.execute(q)]"

Anything over ~1 GB total, or single sessions over ~20 MB, is the warning.

Resetting it. Sessions are working memory — the desk's durable memory is in Mongo (digests, theses, knowledge) and Langfuse — so a reset costs one rollover, nothing more. The host recreates the schema on boot, so the reset is a file move plus a pod delete, done while the host is idle between cycles (it holds no fd on the store then):

kubectl -n nexus exec nexus-accum-host-0 -- sh -c \
'ls -l /proc/1/fd | grep -c sessions.sqlite3' # must be 0
kubectl -n nexus exec nexus-accum-host-0 -- mv \
/realms/nexus-accum/sessions.sqlite3 /realms/nexus-accum/sessions.sqlite3.bak-$(date +%Y%m%d)
kubectl -n nexus delete pod nexus-accum-host-0
Do not scale the StatefulSet to zero

Argo CD runs selfHeal: true on this application: kubectl scale statefulset nexus-accum-host --replicas=0 is reverted within seconds and the pod comes back. Swap the file in place instead.

The two jobs that keep it from happening (both were missing from scheduled_jobs until 2026-09-08, which is why the 07:57 wedge went unwatched):

JobCadenceWhat it does
mob_watchdog10 minLangfuse silence > 2× the mob's cadence, host idle ⇒ restart the pod. The only net that catches a wedge — accum_cycle_health's degraded streak cannot, because a cycle that never terminates never records a degraded verdict.
session_rollover6 hRestarts the host when idle so sessions never grow past ~11 MB. Defers on the mid-cycle busy gate.

Confirm they are armed with GET /v1/schedules (kinds mob_watchdog / session_rollover); create one with POST /v1/schedules and a body like {"name":…,"kind":"mob_watchdog","schedule":{"kind":"interval","minutes":10}, "args":{},"deliver":"local","enabled":true,"state":"scheduled"}.

Triggering a cycle by hand, and why it may do nothing​

Every accum-* job is a row in nexus.scheduled_jobs. To make one due immediately:

curl -X POST "$NEXUS_CORE/v1/schedules/<id>/run" \
-H "x-nexus-api-key: $NEXUS_API_KEY"

POST /v1/schedules/{id}/run does not enqueue the job: it sets next_run_at to the epoch and clears last_fire_key, so the scheduler picks it up on its next tick. It answers {"ok": true}, or 404 when no document has that _id. It is a mutation, so under NEXUS_AUTH_MODE=enforce it needs a service key in x-nexus-api-key — the read-only key gets 403 read_only_key, and the UI proxy principal additionally needs its X-Nexus-Proxy-Origin attestation.

Ids: job_treasury_nav is a code constant (apps/nexus-worker/src/treasury/mod.rs) and is seeded by the worker. The accumulation ids — the operator uses job_accum_execute — are not defined anywhere in nexus-platform; they are _ids of documents created out of band. List them before scripting:

// mongosh, database `nexus`
db.scheduled_jobs.find({}, { _id: 1, name: 1, enabled: 1, next_run_at: 1 })

Three reasons a triggered accum_execute correctly moves no money:

  1. The decision pass short-circuits on its marker. The executor runs the newest accum_decision_json digest once and then writes an accum_exec_done digest whose text is that decision's ts; a decision whose ts matches one of the last 10 markers is skipped as "decision already executed". (accum_exec_done_treasury is the separate marker for the treasury half and accum_exec_done_lp for the LP half — the veto procedure above uses all three.)
  2. The hygiene planner will not move less than MIN_MOVE_USD — a hard-coded 1 000 USD constant in apps/nexus-worker/src/mob_jobs/lending_hygiene.rs, not an env var. It gates both the planner's decisions and the chunked rotation / repark loops.
  3. A lending_supply is capped at exec USDC − float floor and skipped below 1 USDC (HARVEST §5).

That combination is exactly what happened on 2026-09-06 ≈ 09:0x UTC: the operator-triggered sweep moved nothing (moves_money: false) because the 995.55 USDC of idle float was under MIN_MOVE_USD and the decision pass had its done-marker. The path was exercised directly instead — see below.

A verified live lending supply through the fixed path​

2026-09-06 ≈ 09:0x UTC, the first exercise of the F2 idempotency path and the proof that the -32002 blockhash defect is gone:

POST /v1/lending/supply { amount: 995.55 USDC → kamino JLP, intent_id: <fresh> }
→ sent: true, signature 4pv2ZsbL…, confirmed

What to check on such a supply, and what it read that day:

CheckObserved
exec USDC before → after5,995.55 → 5,000.00
venue positionJLP 56,448.93 → 57,445.10 @ 11.02 %
costs.tx_fee_lamports5 000 lamports ($0.0005), rent_lamports: 0
S1 receipt bound (pre-sign)min 821,596,315 cTokens, 1 % tolerance — fired before signing
GET /v1/intents/{id}confirmed, attempts: 1
GET /v1/positions afterthe new figure, not the pre-write one (the write invalidated the venue)

The last row is the point of the position cache: before ea96739 the aggregate served 56,448.93 while /v1/lending/position already read 57,445.10 — and NAV and the budget gate read the aggregate.

Reading exec positions: the cache and ?fresh​

Since nexus-solana ea96739 (2026-09-06) GET /v1/positions and GET /v1/lp/positions are served from a snapshot refreshed every EXEC_POSITIONS_REFRESH_SECS (default 30 s, floor 5). Every answer carries as_of, stale and venues_unreadable[].

  • stale: true means older than 3 refresh intervals, or a half has never been read, or a venue is unreadable. Treat it as "do not use this figure to size a movement".
  • venues_unreadable names the venue whose scan failed. A failed scan keeps the last good value — it never becomes an empty list, because an LP position the sleeve reconciler cannot see is indistinguishable from coin missing from the wallet.
  • ?fresh forces a synchronous full scan (≈ 10 s of RPC). It accepts 1, true, yes or a bare ?fresh; ?fresh=0 and absence do not. Use it for an operator read that must not be a cache hit — never in a loop or a dashboard.
  • A confirmed lending supply / withdraw / withdraw-all or LP open / remove / collect / close marks its venue dirty, so the next read refreshes that venue before answering. You do not need ?fresh after your own write.
  • LENDING_CACHE_SECS now governs only GET /v1/lending/rates.

The LP position entering its range​

2026-09-06, the behaviour the liquidity page describes as an automated swing. The Orca SOL/USDC position opened at 08:29 UTC with 10 SOL into the aligned range [105.96, 112.02] while spot was 105.12 — i.e. out of range, 100 % SOL, earning nothing, which is a laddered sell from 106 to 112, not a broken position. When the price crossed back in, the position became 9.26 SOL + 78.45 USDC and took its first fees: 0.001584 SOL + 0.1963 USDC.

Two consequences for an operator reading the pages: an out-of-range position showing "earning nothing" is the range doing its job, not an incident; and the LP legs are counted in the exposure families since nexus-platform 72967bd, so opening one moves exposure.SOL and the F6 sol_exposure headroom (LP as legs).

Deploying the exec: the price-plane window​

Every exec / oracle deploy produces ≈ 10–15 s of price plane unreadable, and during it the H1 depeg guard fails closed: signing is denied, and a rung that triggers in that window is not fired.

Why: both binaries ship in the same image, so an exec rollout restarts the oracle pod too, and the exec's 5 s price fetch has no retry. The direction is safe — an unreadable oracle denies USDC actions rather than guessing — and it is self-inflicted only on deploys.

What to do: nothing during the window. Confirm afterwards with GET /v1/status.policy (the depeg band readable again) and the exec log, and do not interpret a denial timestamped within a minute of a rollout as a policy incident. Queued improvement (not built): keep the last good USDC cross for EXEC_DEPEG_GRACE_SECS (≈ 60 s) with its age already exposed on /v1/status.policy.depeg.age_s, and only then go Unreadable.

Canonical host: three ways a redirect can fail​

www.nexusapp.dev could never hold a session — the session cookie is set without a domain, so it is host-only — and every gated page there looped through /sign-in. Getting the 308 in place took three attempts, each a failure mode worth recognising:

  1. An ingress permanent-redirect with the wrong backend port. The redirect Ingress named port 3000 while the nexus-ui Service listens on 80; the backend was unresolvable and nginx answered a plain 404 — not an error mentioning the port (gitops 329bab4 → d41d779).
  2. ingress-nginx rejected the annotation value because it contained $request_uri. Variables are not allowed in that annotation; the Ingress is refused at admission, so the symptom is a failed apply, not a bad redirect.
  3. Build-time inlining in Next.js. redirects() in next.config is baked into the routes manifest at build time, and process.env inside Edge middleware is inlined at build time too — so a runtime CANONICAL_HOST from the Secret was invisible to both.

As built (nexus-ui ca1d75e, built 2026-09-06): the rule lives in src/middleware.ts, derives the target from the request's own Host header by stripping a leading www., and honours CANONICAL_HOST when the runtime does expose it. Aliases route normally to nexus-ui:80 again (gitops 91232c7) and bootstrap-secrets.sh derives CANONICAL_HOST from APP_BASE_URL. General lesson: in a Next.js app, anything that must be configurable at runtime cannot live in next.config or in Edge process.env — derive it from the request, or accept a rebuild.

The UI's exec proxy budget​

nexus-ui proxies exec reads through /api/solana/[...path] with a per-method timeout: GET_TIMEOUT_MS 25 000 since 4dca5c0 (2026-09-06, was 8 000) and POST_TIMEOUT_MS 600 000 for catalog refreshes. The 8 s budget aborted the 9.5 s LP position scan, so /trading/liquidity rendered "positions unavailable" over a live Orca position. Both halves were fixed: the budget here, and the scan itself on the exec side (9.5 s → ≈ 2 ms, above). If a page reports a venue "unavailable", check the exec's own latency before assuming the venue is down.

Platform logs (Admin UI)​

Every component installs a tracing subscriber that (1) writes JSON to stdout (for kubectl logs) and (2) ships structured records over NATS (plain subject logs.<component>, not JetStream). The worker runs a single log sink that tails logs.> and writes into a capped MongoDB collection (logs). The Admin UI Logs page (GET /v1/logs) reads that collection — filter by component / level / text, with live polling. This is the fastest way to confirm agents connect, NATS/Mongo are reachable, and LLM calls aren't erroring. See Logging.

Backups​

As built (audit H9, partially addressed):

  • MongoDB (nexus-mongodb): single replica on a microk8s-hostpath volume. A daily mongodump CronJob (03:15, 14-day retention) covers it and the HARVEST desk database — into an on-node 40 Gi PVC. No off-site copy, no encryption, no restore drill has been run.
  • NATS JetStream, Redis, Qdrant: single replicas, hostPath volumes, no backup.

This is a local backup, not disaster recovery (audit 2026-09-05 O1): the database and its backup share the same node and disk, so it protects against logical data loss only. Owner decision, 2026-09-05: the off-node copy is deferred ("not for now") — the local hostpath dump stays the only backup. The candidate (encrypted nightly copy to the .98 host, integrity checks, a restore drill into isolated infrastructure) is recorded, not scheduled. The target (3-node replica set, encrypted off-site copies with object lock, quarterly restore drills) is not in place. NexusMongoBackupStale / NexusMongoBackupJobFailed (group nexus.backups) page when the local dump stops.

Release flow​

Every push to main of nexus-platform, nexus-ui, nexus-solana and nexus-docs deploys: CI writes the image tag into environments/prod/values.yaml and Argo rolls it. A PR flow with branch protection (audit O3) was declined by the owner on 2026-09-05; direct-to-main stays the deploy flow. Rollback is a values-file edit (e.g. solanaExec.tag back to the previous sha-…).

Secrets​

  • .env files never committed.
  • Target: production secrets from a secrets manager (Vault / cloud secret manager / sealed-secrets in GitOps) with scheduled rotation.
  • As built (2026-09-05): Kubernetes Secrets are created by nexus-gitops/bootstrap-secrets.sh from the git-ignored .secrets file; no secrets manager, no scheduled rotation — rotations are ad hoc.
  • Agents can run on a subscription CLI login (Claude/Codex) instead of an API key — see LLM backends.

Production checklist​

  • MongoDB replica set + backups tested via restore — daily on-node dump only; off-node copy deferred by the owner (2026-09-05)
  • NATS authenticated — 2026-09-05: nexus/runner users, runner subject ACL (audit H6); still a single node
  • Agent namespace + scoped RBAC applied — enforced since 2026-09-05: authorizer Node,RBAC (audit C1)
  • Metrics + alerting deployed and verified — 2026-09-05: /metrics live, Alertmanager → Telegram, nexus-platform-alerts (5 groups); the NAV exporter is live since 17:12 UTC (scrapeable since ≈ 17:4x)
  • Secrets in a manager, rotation scheduled
  • Approval gates configured for sensitive actions — Telegram approvals run actions in order; a failed action leaves the approval failed and pending (2026-09-05)
  • Quarantine policy thresholds set
  • Taiga integration (if used) rate-limited + webhook HMAC-SHA1 verified
  • Telegram allowlist locked down — IDs set; empty list denies everyone since 445e107 (audit C7)