Skip to main content

Architecture

Nexus is an Agent Operating System: a control plane that turns runtime-defined agent documents into isolated, observable executions. There are five substrates, each with exactly one job.

Nexus UI (React)
create agents · edit skills · manage
memory · view board · approve · logs
│ REST + SSE
▼
┌───────────────────────────────────────────────────────────┐
│ Nexus Core (Rust, axum) │
│ │
│ orchestration engine · scheduler · agent registry · │
│ skill registry · memory service · board sync · auth · │
│ permissions · approvals · Kubernetes job launcher │
└───────┬───────────────┬───────────────┬───────────────┬────┘
│ │ │ │
(state) │ (events)│ (board) │ (isolation)│
▼ ▼ ▼ ▼
┌────────────┐ ┌────────────┐ ┌────────────┐ ┌──────────────────┐
│ MongoDB │ │ NATS │ │ Taiga │ │ Kubernetes │
│ │ │ JetStream │ │ (adapter) │ │ │
│ source of │ │ durable │ │ external │ │ one Job per run │
│ truth for │ │ event bus │ │ projection │ │ generic runner │
│ agents, │ │ replay, │ │ of the │ │ image becomes │
│ skills, │ │ crash │ │ internal │ │ any agent from a │
│ memory, │ │ recovery │ │ board │ │ config payload │
│ tasks,runs │ └────────────┘ └────────────┘ └────────┬─────────┘
└────────────┘ │
▼
┌──────────────────────────────┐
│ nexus-agent-runner (pod) │
│ loads skills + memory + │
│ tools, clones repo, calls │
│ the LLM, runs tools, commits,│
│ reports events back to Core │
└──────────────────────────────┘

Substrates and responsibilities​

SubstrateOwnsNotes
Nexus CoreOrchestration authorityThe only component that mutates the source of truth. Everything routes through it.
MongoDBStateAgents, skills, memory, tasks, runs, approvals, board links. Dynamic configuration store.
NATS JetStreamEventsDurable, replayable stream messaging. Crash recovery for autonomous agents.
KubernetesIsolationEach agent run is an isolated Job. Resource limits, mounted secrets, scoped RBAC.
TaigaExternal projectionThe internal board is authoritative; Taiga is a synced view. Jira later.

Components​

ComponentTechResponsibility
nexus-coreRust, axum, tokio, tower, MongoDB driverAdmin/agent/board/skill/memory/webhook APIs, auth, approvals, orchestration.
nexus-workerRust, tokio, async-nats, MongoDBScheduling, agent dispatch, retries, Taiga sync, memory indexing, reconciliation.
nexus-agent-runnerRust, reqwest, tokio, ductGeneric container: load config, run skills/tools, report events. One image, many agents.
nexus-telegramRust, teloxideHigh-level command gateway: create goals, approve tasks, status, alerts.
nexus-uiNext.js 16, React 19, Tailwind v4, TanStackAdmin dashboard (separate repo): agents, skills, memory, board, runs, approvals.
nexus-docsDocusaurusThis site (docs.nexusapp.dev), built and deployed from its own repo.
nexus-oracleRustUnified perception service: reads cex-manager /Analysis/*, writes the oracle_snapshots feature feed.
nexus-mob-hostRust (meerkat)The desk host. Runs as nexus-accum-host (accumulation analysis desk). The perp nexus-trading-host and nexus-learning-host (dream mob) were retired 2026-08-28.
nexus-solana-oracleRust (nexus-solana)Read-only Solana price plane (Pyth / Metis / Coinbase quorum). No keys.
nexus-solana-execRust (nexus-solana)The only pod that mounts the EXEC signing key; builds, simulates and signs HARVEST transactions. Armed since 2026-08-29.
nexus-redisRedis (CloudPirates chart, own Argo app)Apalis job-queue backend for nexus-worker.
Retired—nexus-executor / nexus-executor-live (perp desk, retired 2026-08-28), nexus-trader (replaced by nexus-executor before that). Templates remain in the chart, disabled.
Shared cratesRustnexus-domain, nexus-db, nexus-events, nexus-board(-taiga), nexus-agent-runtime, nexus-ai-runtime, nexus-k8s, nexus-skills, nexus-memory, nexus-llm, nexus-git, nexus-auth, nexus-observability.

See Repositories for the full per-project library list and Technical overview for the per-component deep dives.

Why agents are data​

A traditional design encodes each agent role as a compiled type. Nexus rejects that. An agent is:

agent = MongoDB document (config) + attached skills + memory policy + permissions

When a task is assigned to an agent, Core assembles a run config and hands it to the runtime. The pod does not branch on role. This means:

  • You add a new agent from the UI, with no deploy.
  • The same nexus-agent-runner image serves every role.
  • Skills, memory, and permissions are composed at dispatch time.
{
"run_id": "run_123",
"agent_id": "agent_backend_implementer",
"task_id": "task_456",
"system_prompt": "...",
"skills": ["...skill markdown..."],
"memory": ["...relevant memory..."],
"tools": ["git", "shell", "cargo"],
"permissions": { "can_commit": true, "can_merge": false }
}

Trait seams​

Nexus uses traits for dynamic providers, not for fixed agent roles:

#[async_trait::async_trait]
pub trait AgentRuntime {
async fn start_run(&self, run: AgentRunRequest) -> anyhow::Result<AgentRunHandle>;
async fn cancel_run(&self, run_id: &str) -> anyhow::Result<()>;
async fn get_logs(&self, run_id: &str) -> anyhow::Result<Vec<String>>;
}

#[async_trait::async_trait]
pub trait BoardProvider {
async fn create_item(&self, item: BoardItemCreate) -> anyhow::Result<BoardItem>;
async fn update_status(&self, item_id: &str, status: BoardStatus) -> anyhow::Result<()>;
async fn add_comment(&self, item_id: &str, comment: &str) -> anyhow::Result<()>;
}

Implementations: KubernetesAgentRuntime; TaigaBoardProvider, JiraBoardProvider, NexusInternalBoardProvider. Start with Taiga — but Taiga is just an adapter.

Deployment shape​

ProcessPortsHosting
nexus-core8080 (REST + SSE, Service :80), 9100 (/metrics)Deployment (2 replicas) behind ingress nexusapp.dev.
nexus-worker9100 (/metrics)Deployment (1 replica); consumes NATS, runs the Apalis job queue on Redis.
nexus-telegram9100 (/metrics)Deployment; long-polls getUpdates.
nexus-oracle9100 (/metrics)Deployment.
nexus-accum-host8090 (host API), 9100 (/metrics)StatefulSet (nexus-mob-host image).
nexus-solana-oracle / -exec8090 / 8091Deployments; exec's :8091 is meant to be NetworkPolicy-restricted to worker/core/telegram/ui — but policies are additive and nexus-allow-monitoring also admitted the monitoring namespace on every port until it was narrowed to 9100 on 2026-09-05 (audit S2; see Security). No caller authentication on the API (audit S1, in progress). No /metrics (separate repo, not on nexus-observability).
nexus-ui3000 (Service :80)Next.js server behind the same ingress; proxies REST to nexus-core.
nexus-docs8080 (Service :80)Deployment behind docs.nexusapp.dev.
agent runs—Ephemeral Jobs in nexus-agents; /metrics deliberately off (NEXUS_METRICS_ADDR=off).
MongoDB27017Single replica, microk8s-hostpath PVC, namespace nexus-mongodb. Not a 3-node replica set (audit H9).
NATS JetStream4222, 8222Single node, namespace nexus-nats. Authenticated since 2026-09-05: control plane connects as nexus, agent Jobs as runner under a subject ACL (audit H6 closed).
HARVEST desk mongod (host)27017 on 192.168.100.97bitview_desk / adaptive_learning / nexus for the accumulation desk. Authentication enforced since 2026-09-05 with per-service users; URIs come from Secrets; the port is allow-listed by a host firewall unit (audit C5 closed).
Qdrant6333 (REST), 6334 (gRPC — what the platform uses)Single replica, API key enforced, namespace nexus-qdrant.
Redis6379Single node, AOF, nexus-redis in namespace nexus.

/metrics (as of nexus-platform f183437, 2026-09-05)​

Every service built on nexus-observability serves Prometheus text on NEXUS_METRICS_ADDR (default 0.0.0.0:9100), started from init(). Before this commit the six metric names existed only as constants and :9100 refused connections — the earlier "9100 (Prometheus)" rows on this page were false.

  • Recorded today: nexus_agent_runs_total{agent=<job kind>,outcome}, nexus_agent_run_failures_total{agent,reason=timeout|error} and nexus_task_duration_seconds{kind=<job kind>} (all from the worker's run_job), plus process_* (CPU, RSS, FDs) per service. Every series carries a component label.
  • Registered but not yet recorded (no call sites): nexus_llm_tokens_total, nexus_llm_cost_usd_total, nexus_board_sync_errors_total.
  • Scraped by the nexus-metrics ServiceMonitor (kube-prometheus-stack in monitoring); visualised in Grafana → folder Nexus → Nexus Platform.

Observability & alerting stack (shared, namespace monitoring / langfuse)​

PieceWhereNotes
Prometheus / Alertmanager / Grafanakube-prometheus-stack (kps), manual Helm releaseGrafana at grafana.bitview.club. Alertmanager routes every alert (info included) to the Telegram group NexusHome; until 2026-09-05 the root route was receiver: "null" and all 133 rules were discarded.
Langfuselangfuse.bitview.club, Argo CD app langfuseLLM traces from core/worker (NEXUS_LANGFUSE_* in the nexus-app Secret).
Platform logsNATS logs.> → Mongo capped collection → UI Logs pageSee Logging.

Nexus Core's service account can create Jobs only in the agent namespace, with a tightly scoped RBAC role. See Security & permissions.

Cluster authorizer

Since 2026-09-05 11:27Z the microk8s API server runs Node,RBAC (audit C1 closed). The roles above are therefore real: the worker's service account can create Jobs in nexus-agents, no service account can read Secrets, and token automount is off on every workload except nexus-core and nexus-worker.

Per-component deep dives​

The Technical overview is the entry point.

Where Nexus is headed​

For a deep comparison against Hermes — pluggable run backends, agent-initiated subagent delegation with shared workspaces, an in-run tool-RPC surface, MCP ingestion, a multi-surface gateway, and a self-improving learning loop (agent-authored skills, run search, trajectory export) — see the Capability parity (Hermes gap analysis).