Monitoring & alerting — fleet health UI
A redesign of the admin UI around one question: is everything healthy, and
if not, did Telegram already tell me? Scope is the whole fleet, not just
trading. As built (2026-09-04): the accumulation cycle-health stack
exists — per-cycle health records on the mob host, a /health that fails
after two degraded cycles so Kubernetes restarts it, worker-side staleness
pages once per outage and once on recovery, a Langfuse silence watchdog,
GET /v1/desk/accum-health, the /trading/cycles page and a header pill.
The fleet-wide collector, component_health store, rule engine, fleet grid
and alert history are not built. Details in
HARVEST — as built §8.
Update 2026-09-05: the generic layer now exists outside the UI —
/metrics is real on every platform service, a ServiceMonitor scrapes it,
the Nexus Platform Grafana dashboard exists, and Alertmanager delivers
kube-prometheus-stack's rules to Telegram (it previously discarded all of
them). See Operations → Observability.
1. Why
The fleet now spans two clusters' worth of moving parts — platform services, two executors, mob hosts, the Solana venue services, the CEX manager, a self-hosted Jupiter Metis server, a dedicated Solana node, and four infra stores. Failures today surface as symptoms (pages down, silent mob wedges, stale oracles) discovered by humans. With HARVEST placing real treasury transactions unattended, "discovered by humans" stops being acceptable.
2. Components covered
| Group | Components | Key signals |
|---|---|---|
| Platform | nexus-core, nexus-worker, nexus-telegram | heartbeat, API error rate, job-queue lag |
| Trading | nexus-oracle (nexus-executor / nexus-executor-live retired 2026-08-28) | ingest freshness, restart counts |
| Mobs | accum-host (trading-host and learning-host were retired 2026-08-28) | flow execution liveness (catches the silent-wedge failure mode — built: degraded-cycle streak → liveness restart), brief/decision output freshness (built) |
| Solana venue | nexus-solana-oracle, nexus-solana-exec | feed freshness, divergence-guard state, tranche queue, reconciliation status |
| Owned infra | Jupiter Metis server, Solana node | Metis liveness + quote latency; node slot lag vs an external reference (a node silently serving stale slots is the worst failure — it poisons every read) |
| External venue | cex-manager | reachability, WS stream health |
| Stores | MongoDB, NATS, Redis, Qdrant | reachability, replication/queue depth |
3. Architecture
services ──► /health + heartbeats + Prometheus metrics
│
▼
collector (nexus-worker job or sidecar)
│ writes component_health store
▼
rule engine: missed heartbeats, error rates,
slot lag, staleness, utilization, breaker trips
│
┌──────────┴──────────┐
▼ ▼
nexus-ui notify.events ──► Telegram
monitoring section (severity + once-per-transition dedup,
same pattern as the executor's
drawdown bands)
- nexus-solana services reuse the
prochain-solana-monitoringPrometheus crate andprochain-thread-monitorstall detection from the fleet; platform services expose health vianexus-observability. - Alert rules live in config (Mongo-backed, runtime-editable), not code.
- Every alert has a severity (info / high / critical) and dedups on state transition — no repeat spam while a condition persists, one message when it clears.
4. UI redesign
The admin UI gains a monitoring section as its operational home page:
- Fleet grid — every component as a status tile (healthy / degraded / down / unknown), restart counts, last heartbeat.
- Per-component drill-down — key metrics, recent alerts, links to the relevant operational page (e.g. oracle-health, decisions).
- Alert history — what fired, when, severity, when it cleared, whether Telegram delivered.
- HARVEST panel — ladder state, % deployed, yield accrued, tranche queue, divergence-guard status, allocation per venue with % idle, and the desk's thesis-score sparkline with forecast-calibration record (see the HARVEST and accumulation desk specs).
5. Out of scope
- Full Grafana/Prometheus stack replacement — Langfuse and existing dashboards stay; this is the operator's single pane, not a metrics warehouse.
- Auto-remediation (restart-on-wedge etc.) — alert first, automate later. As built: restart-on-wedge shipped for the accumulation host in two independent forms (host liveness on degraded cycles; worker restart at 3× cadence of silence), after a 20-hour silent wedge on 2026-09-03.
Related
- HARVEST — as built — the cycle-health stack that exists
- HARVEST (Solana accumulation)
- Operations: logging