Skip to main content

Monitoring & alerting — fleet health UI

Design document — partially built

A redesign of the admin UI around one question: is everything healthy, and if not, did Telegram already tell me? Scope is the whole fleet, not just trading. As built (2026-09-04): the accumulation cycle-health stack exists — per-cycle health records on the mob host, a /health that fails after two degraded cycles so Kubernetes restarts it, worker-side staleness pages once per outage and once on recovery, a Langfuse silence watchdog, GET /v1/desk/accum-health, the /trading/cycles page and a header pill. The fleet-wide collector, component_health store, rule engine, fleet grid and alert history are not built. Details in HARVEST — as built §8.

Update 2026-09-05: the generic layer now exists outside the UI — /metrics is real on every platform service, a ServiceMonitor scrapes it, the Nexus Platform Grafana dashboard exists, and Alertmanager delivers kube-prometheus-stack's rules to Telegram (it previously discarded all of them). See Operations → Observability.

1. Why​

The fleet now spans two clusters' worth of moving parts — platform services, two executors, mob hosts, the Solana venue services, the CEX manager, a self-hosted Jupiter Metis server, a dedicated Solana node, and four infra stores. Failures today surface as symptoms (pages down, silent mob wedges, stale oracles) discovered by humans. With HARVEST placing real treasury transactions unattended, "discovered by humans" stops being acceptable.

2. Components covered​

GroupComponentsKey signals
Platformnexus-core, nexus-worker, nexus-telegramheartbeat, API error rate, job-queue lag
Tradingnexus-oracle (nexus-executor / nexus-executor-live retired 2026-08-28)ingest freshness, restart counts
Mobsaccum-host (trading-host and learning-host were retired 2026-08-28)flow execution liveness (catches the silent-wedge failure mode — built: degraded-cycle streak → liveness restart), brief/decision output freshness (built)
Solana venuenexus-solana-oracle, nexus-solana-execfeed freshness, divergence-guard state, tranche queue, reconciliation status
Owned infraJupiter Metis server, Solana nodeMetis liveness + quote latency; node slot lag vs an external reference (a node silently serving stale slots is the worst failure — it poisons every read)
External venuecex-managerreachability, WS stream health
StoresMongoDB, NATS, Redis, Qdrantreachability, replication/queue depth

3. Architecture​

services ──► /health + heartbeats + Prometheus metrics
│
▼
collector (nexus-worker job or sidecar)
│ writes component_health store
▼
rule engine: missed heartbeats, error rates,
slot lag, staleness, utilization, breaker trips
│
┌──────────┴──────────┐
▼ ▼
nexus-ui notify.events ──► Telegram
monitoring section (severity + once-per-transition dedup,
same pattern as the executor's
drawdown bands)
  • nexus-solana services reuse the prochain-solana-monitoring Prometheus crate and prochain-thread-monitor stall detection from the fleet; platform services expose health via nexus-observability.
  • Alert rules live in config (Mongo-backed, runtime-editable), not code.
  • Every alert has a severity (info / high / critical) and dedups on state transition — no repeat spam while a condition persists, one message when it clears.

4. UI redesign​

The admin UI gains a monitoring section as its operational home page:

  1. Fleet grid — every component as a status tile (healthy / degraded / down / unknown), restart counts, last heartbeat.
  2. Per-component drill-down — key metrics, recent alerts, links to the relevant operational page (e.g. oracle-health, decisions).
  3. Alert history — what fired, when, severity, when it cleared, whether Telegram delivered.
  4. HARVEST panel — ladder state, % deployed, yield accrued, tranche queue, divergence-guard status, allocation per venue with % idle, and the desk's thesis-score sparkline with forecast-calibration record (see the HARVEST and accumulation desk specs).

5. Out of scope​

  • Full Grafana/Prometheus stack replacement — Langfuse and existing dashboards stay; this is the operator's single pane, not a metrics warehouse.
  • Auto-remediation (restart-on-wedge etc.) — alert first, automate later. As built: restart-on-wedge shipped for the accumulation host in two independent forms (host liveness on degraded cycles; worker restart at 3× cadence of silence), after a 20-hour silent wedge on 2026-09-03.