Skip to main content

Operator guide

Get Nexus running and drive your first goal end-to-end.

1. Stand up the platform​

Follow Operations to run MongoDB, NATS, Nexus Core, the worker, and the UI — locally or on Kubernetes via nexus-gitops.

2. Create an agent​

In the UI → Agents → Create agent. Fill in name, model, runtime image and resources, attach a skill or two, choose memory scopes, set permissions. No deploy needed — it's a document. See Concepts: Agents.

3. Create a goal​

From the UI or Telegram:

/nexus create goal "Build authentication system for my app"

Nexus drafts tasks and waits for your approval.

4. Approve the plan​

In the UI → Board, review the drafted tasks, adjust, and approve. Approved tasks become ready_for_agent and (if Taiga is wired) appear as Taiga user stories.

5. Watch a run​

The scheduler leases a ready task and launches a Kubernetes Job. In the UI → Runs, follow live logs, files changed, and commands run. Approve any gated action (e.g. database_migration) in Approvals.

6. Tune​

If an agent underperforms, adjust its skills/permissions or clone it. Repeated failures trigger quarantine.

7. Alerting (Telegram)​

kube-prometheus-stack's Alertmanager delivers to the Telegram group NexusHome (chat -5137428184) using the HARVEST bot token, stored in the Secret monitoring/alertmanager-telegram (never in Helm values). Routing:

  • every alert goes to Telegram — severity: info is intentionally not filtered, because CPUThrottlingHigh ships at info and is exactly the alert that would have caught 76 days of MinIO restarts;
  • only Watchdog (dead-man's switch) and InfoInhibitor (a routing helper) are dropped;
  • group_by: [namespace, alertname], repeat every 12 h, resolved messages on.

Before 2026-09-05 the root route was receiver: "null" and all 133 rules were discarded — if an alert seems missing, check the route before the rule. The Alertmanager DM to a user needs that user to message the bot first (Telegram returns 403 otherwise); the group works because the bot is a member.

8. Dashboards (Grafana)​

grafana.bitview.club → folder Nexus → Nexus Platform (delivered from nexus-gitops/charts/nexus/files/dashboards/nexus-platform.json):

RowPanels
Service health/metrics targets up, pods not ready, restarts 24 h, job failure rate, jobs/min, OOMKills
Jobsthroughput by kind (top 8), failures by kind × reason, p50 / p95 duration, full table
Resources per serviceprocess CPU / RSS / FDs, CPU throttling by container, memory used vs limit, restarts by pod
LLM usagecollapsed until nexus_llm_* is recorded

Colors are assigned per service in a fixed order; a series that disappears does not repaint the others.

9. Secrets​

Every platform Secret is created by nexus-gitops/infra/bootstrap-secrets.sh from variables in the git-ignored .secrets at the workspace root. To add or rotate a key: put it in .secrets, then

set -a; . ./.secrets; set +a
./infra/bootstrap-secrets.sh

The script refuses to run if a key that exists in the cluster is missing from its input (that is how a re-run once nearly deleted GITHUB_TOKEN), and it never writes an empty optional value. The Solana EXEC key is only ever written from EXEC_KEY_FILE and is left untouched otherwise. Full table in nexus-gitops → Secrets.

10. Resource sizing — the lesson​

Every Bitnami sub-chart in the cluster (MinIO, ZooKeeper, PostgreSQL, Valkey, MongoDB) shipped with a nano / micro resourcesPreset — laptop-sized ceilings on a 128-core / 528 GiB node. A CPU-throttled process cannot answer its liveness probe inside the timeout, so the kubelet SIGKILLs a healthy service (exit 137). That is what restarted Langfuse's MinIO 122 times over 76 days. Rule: on any new chart set resourcesPreset: "none" and explicit resources, then watch CPU throttling by container and memory used vs limit on the dashboard — sustained throttling above ~25 % or memory above ~85 % of the limit is a limit problem, not an application problem.

11. Langfuse is GitOps now​

langfuse.bitview.club is the Argo CD application langfuse (nexus-gitops/argocd/apps/langfuse.yaml) — chart 1.5.35, every credential an existingSecret / secretKeyRef into namespace langfuse. Never run helm upgrade langfuse … by hand again; Argo will fight it. To change a value, edit the Application and push.

12. Service credentials (2026-09-05)​

Everything below is created by infra/bootstrap-secrets.sh from .secrets; nothing lives in git.

  • NATS — the broker requires credentials. Control-plane services use NEXUS_NATS_URL (nexus user, unrestricted); agent Jobs receive NEXUS_AGENT_NATS_URL (runner user: may publish only nexus.agent.run.*, logs.nexus-agent-runner, tools.invoke and read stream info; subscribe only to _INBOX.>). Runbook: nexus-gitops/infra/nats/AUTH-CUTOVER.md. Lesson from the cutover: async_nats::connect(url) ignores user:pass@ in the URL — the services had only ever been "authenticated" by the broker's no_auth_user mapping, and removing it took the bus down for ~8 minutes. nexus-events now passes the credentials explicitly; a no_auth_user change needs a broker restart (config reload is refused for that setting).
  • HARVEST desk Mongo (192.168.100.97:27017) — authentication enforced. Users: nexus_rw (desk services), hermes_rw, cryptomanager (CEX manager), backup_ro (nightly dump), nexus_admin (operator). The port only accepts the pod CIDR, the host, the operator workstation and the Solana box (systemd unit mongod-firewall); the log rotates daily.
  • CEX manager (192.168.100.97:8232) — every call needs X-Api-Key (CEX_MANAGER_API_KEY, same value in nexus-app and the host's ~/.profile). Keycloak client-credentials on top is on the TODO list.
  • Core API — NEXUS_API_KEY in nexus-app is the service key Core checks (log-only today: watch nexus_core_auth_decisions_total{decision="would_deny"} in Grafana, then set core.auth.mode: enforce in environments/prod/values.yaml). Your operator identity is the email in core.auth.operatorAllowlist.
  • Keycloak also answers on https://identity.nexusapp.dev; the realm issuer and the UI's KEYCLOAK_ISSUER still point at identity.bitview.club until the coordinated switch.