Operator guide
Get Nexus running and drive your first goal end-to-end.
1. Stand up the platform
Follow Operations to run MongoDB, NATS, Nexus Core, the worker, and the UI — locally or on Kubernetes via nexus-gitops.
2. Create an agent
In the UI → Agents → Create agent. Fill in name, model, runtime image and resources, attach a skill or two, choose memory scopes, set permissions. No deploy needed — it's a document. See Concepts: Agents.
3. Create a goal
From the UI or Telegram:
/nexus create goal "Build authentication system for my app"
Nexus drafts tasks and waits for your approval.
4. Approve the plan
In the UI → Board, review the drafted tasks, adjust, and approve. Approved
tasks become ready_for_agent and (if Taiga is wired) appear as Taiga user
stories.
5. Watch a run
The scheduler leases a ready task and launches a Kubernetes Job. In the UI →
Runs, follow live logs, files changed, and commands run. Approve any gated
action (e.g. database_migration) in Approvals.
6. Tune
If an agent underperforms, adjust its skills/permissions or clone it. Repeated failures trigger quarantine.
7. Alerting (Telegram)
kube-prometheus-stack's Alertmanager delivers to the Telegram group
NexusHome (chat -5137428184) using the HARVEST bot token, stored in the
Secret monitoring/alertmanager-telegram (never in Helm values). Routing:
- every alert goes to Telegram —
severity: infois intentionally not filtered, becauseCPUThrottlingHighships at info and is exactly the alert that would have caught 76 days of MinIO restarts; - only
Watchdog(dead-man's switch) andInfoInhibitor(a routing helper) are dropped; group_by: [namespace, alertname], repeat every 12 h, resolved messages on.
Before 2026-09-05 the root route was receiver: "null" and all 133 rules
were discarded — if an alert seems missing, check the route before the rule.
The Alertmanager DM to a user needs that user to message the bot first
(Telegram returns 403 otherwise); the group works because the bot is a member.
8. Dashboards (Grafana)
grafana.bitview.club → folder Nexus → Nexus Platform (delivered
from nexus-gitops/charts/nexus/files/dashboards/nexus-platform.json):
| Row | Panels |
|---|---|
| Service health | /metrics targets up, pods not ready, restarts 24 h, job failure rate, jobs/min, OOMKills |
| Jobs | throughput by kind (top 8), failures by kind × reason, p50 / p95 duration, full table |
| Resources per service | process CPU / RSS / FDs, CPU throttling by container, memory used vs limit, restarts by pod |
| LLM usage | collapsed until nexus_llm_* is recorded |
Colors are assigned per service in a fixed order; a series that disappears does not repaint the others.
9. Secrets
Every platform Secret is created by nexus-gitops/infra/bootstrap-secrets.sh
from variables in the git-ignored .secrets at the workspace root. To add
or rotate a key: put it in .secrets, then
set -a; . ./.secrets; set +a
./infra/bootstrap-secrets.sh
The script refuses to run if a key that exists in the cluster is missing from
its input (that is how a re-run once nearly deleted GITHUB_TOKEN), and it
never writes an empty optional value. The Solana EXEC key is only ever written
from EXEC_KEY_FILE and is left untouched otherwise. Full table in
nexus-gitops → Secrets.
10. Resource sizing — the lesson
Every Bitnami sub-chart in the cluster (MinIO, ZooKeeper, PostgreSQL, Valkey,
MongoDB) shipped with a nano / micro resourcesPreset — laptop-sized
ceilings on a 128-core / 528 GiB node. A CPU-throttled process cannot answer
its liveness probe inside the timeout, so the kubelet SIGKILLs a healthy
service (exit 137). That is what restarted Langfuse's MinIO 122 times over 76
days. Rule: on any new chart set resourcesPreset: "none" and explicit
resources, then watch CPU throttling by container and memory used vs
limit on the dashboard — sustained throttling above ~25 % or memory above
~85 % of the limit is a limit problem, not an application problem.
11. Langfuse is GitOps now
langfuse.bitview.club is the Argo CD application langfuse
(nexus-gitops/argocd/apps/langfuse.yaml) — chart 1.5.35, every credential an
existingSecret / secretKeyRef into namespace langfuse. Never run
helm upgrade langfuse … by hand again; Argo will fight it. To change a
value, edit the Application and push.
12. Service credentials (2026-09-05)
Everything below is created by infra/bootstrap-secrets.sh from .secrets;
nothing lives in git.
- NATS — the broker requires credentials. Control-plane services use
NEXUS_NATS_URL(nexususer, unrestricted); agent Jobs receiveNEXUS_AGENT_NATS_URL(runneruser: may publish onlynexus.agent.run.*,logs.nexus-agent-runner,tools.invokeand read stream info; subscribe only to_INBOX.>). Runbook:nexus-gitops/infra/nats/AUTH-CUTOVER.md. Lesson from the cutover:async_nats::connect(url)ignoresuser:pass@in the URL — the services had only ever been "authenticated" by the broker'sno_auth_usermapping, and removing it took the bus down for ~8 minutes.nexus-eventsnow passes the credentials explicitly; ano_auth_userchange needs a broker restart (config reload is refused for that setting). - HARVEST desk Mongo (
192.168.100.97:27017) — authentication enforced. Users:nexus_rw(desk services),hermes_rw,cryptomanager(CEX manager),backup_ro(nightly dump),nexus_admin(operator). The port only accepts the pod CIDR, the host, the operator workstation and the Solana box (systemd unitmongod-firewall); the log rotates daily. - CEX manager (
192.168.100.97:8232) — every call needsX-Api-Key(CEX_MANAGER_API_KEY, same value innexus-appand the host's~/.profile). Keycloak client-credentials on top is on the TODO list. - Core API —
NEXUS_API_KEYinnexus-appis the service key Core checks (log-only today: watchnexus_core_auth_decisions_total{decision="would_deny"}in Grafana, then setcore.auth.mode: enforceinenvironments/prod/values.yaml). Your operator identity is the email incore.auth.operatorAllowlist. - Keycloak also answers on
https://identity.nexusapp.dev; the realm issuer and the UI'sKEYCLOAK_ISSUERstill point atidentity.bitview.clubuntil the coordinated switch.