Skip to main content

Security & permissions

Nexus runs autonomous agents that touch real repos and clusters. Security is defense in depth across four layers.

Layers​

LayerEnforces
Nexus CoreWhether an agent may be dispatched with a given tool/permission set; raises approvals for gated actions.
RunnerRefuses any action the run config didn't grant.
Kubernetes RBACThe run's service account can only touch what its role allows — enforced since 2026-09-05 (Node,RBAC).
nexus-auth crateIntended: admin users, API keys, agent tokens, tool/project permissions. As built (2026-09-05): the pure auth policy — service key (X-Nexus-Api-Key vs NEXUS_API_KEY), operator identity (X-Nexus-Operator-Sub/Email, honoured only behind a valid key, checked against NEXUS_OPERATOR_ALLOWLIST), modes off/log/enforce. nexus-core runs it as middleware on every /v1 route; production is in log-only mode (audit H11 partial). Admin JWT / agent tokens remain unimplemented.

A permission must be granted at every layer for an action to succeed — in design. Where that stands (2026-09-05): Kubernetes RBAC is enforced (the authorizer moved from AlwaysAllow to Node,RBAC; audit C1 closed). Core evaluates the service key and the operator allowlist on every /v1 request and enforces since 2026-09-05 22:37 UTC (NEXUS_AUTH_MODE=enforce, NEXUS_AUTH_READ=open — nexus-gitops c95071e, audit S3 closed after 24 h of log mode with 0 would_deny): an anonymous mutation is 401 no_key, reads stay open, decisions are counted in nexus_core_auth_decisions_total. Rollback is core.auth.mode: log.

The flip was not merely a values change (audit 2026-09-05 S3/S4): it waited for the identity binding under Identity propagation below.

Identity propagation (UI → Core)​

Status 2026-09-05: specified; implementation in progress in nexus-ui (proxy side) and nexus-platform (Core side); not deployed. The deployed UI (63a1083) forwards browser headers, injects X-Nexus-Api-Key and neither strips inbound X-Nexus-Operator-Sub/Email nor derives them from the verified session — so under enforce a signed-in browser request without identity headers would look like a privileged service request, and a browser could supply its own operator assertion (audit S4).

The contract that closes it:

Credential / headerPrincipalRule
NEXUS_UI_API_KEYui-proxyMust carry an operator identity taken from the verified session (X-Nexus-Operator-Sub/Email set by the proxy, every inbound X-Nexus-* header stripped first). Never promoted to a service principal.
NEXUS_API_KEYservice (worker, telegram, …)Full service rights; no operator assertion honoured.
NEXUS_API_KEY_READ (optional)service, read-onlySplit read/write service permissions.
X-Nexus-Proxy-OriginattestationRequired on UI-originated mutations; Core rejects a mutation under the ui-proxy key without it.

NEXUS_AUTH_MODE=enforce was gated on the UI change being deployed and on would_deny staying at zero afterwards — not on the values flip alone; both held (UI a2561d9 15:14 UTC, 383 decisions / 0 would_deny) before the 22:37 UTC flip.

Identity & tokens​

  • Admin users — JWT (jsonwebtoken), passwords hashed with argon2.
  • API keys — for machine clients.
  • Agent tokens — short-lived tokens minted per run so a pod can report back to Core, scoped to that run only.
  • Telegram allowlist — TELEGRAM_ALLOWED_USERS (in the nexus-app Secret) restricts /approve, /reject, /basis and the rest. It is set in prod and, since nexus-platform 445e107 (deployed 2026-09-05), an empty list denies everyone — the gateway fails closed (audit C7 closed).

Permission model​

agent.can_create_branch · agent.can_commit · agent.can_open_pr · agent.can_merge
agent.can_delete_files · agent.can_use_kubectl · agent.can_apply_k8s
agent.can_write_memory · agent.can_create_board_task

Approval gates​

Agents declare actions that always need a human:

requires_human_approval_for:
- production_deploy
- dependency_upgrade
- database_migration

Gated attempts raise an approval request.

Kubernetes isolation​

  • Each run is an isolated Job in a dedicated agent namespace (nexus-agents).
  • Core/worker service accounts can manage Jobs only there:
verbs: [create, get, list, watch, delete]
resources: [jobs, pods, pods/log]

Runner pod hardening​

Every agent Job is created with a locked-down pod (see nexus-k8s):

ControlSetting
Run as non-rootrunAsNonRoot: true, runAsUser/Group: 1000, fsGroup: 1000
No privilege escalationallowPrivilegeEscalation: false, privileged: false
Drop capabilitiescapabilities.drop: ["ALL"]
Immutable root FSreadOnlyRootFilesystem: true (+ writable emptyDir at /tmp and /workspace)
SeccompseccompProfile: RuntimeDefault
No cluster accessautomountServiceAccountToken: false
Auto-cleanupttlSecondsAfterFinished: 3600

Secret injection​

LLM keys are not baked into images. They live in the nexus-agent Secret in the agent namespace and are injected via envFrom (referenced as optional, so a missing secret fails the run fast rather than blocking scheduling):

GIT_TOKEN · GITHUB_TOKEN · GIT_USERNAME # what the prod Secret holds
OPENAI_API_KEY · ANTHROPIC_API_KEY # optional; absent in prod — agents run on the CLI logins below

GIT_TOKEN (falling back to GITHUB_TOKEN / GITLAB_TOKEN) authenticates clones of private https repositories in the agent's repository list. The runner injects it into the clone URL at run time and never logs it; clone errors are credential-redacted before they reach logs or run results. Repos marked read_only are a convention the prompt conveys to the agent — enforce non-push hard limits via permissions and review gates. Prefer https clones: the runner NetworkPolicy allows egress on 443, so SSH (git@…, port 22) needs the egress rule widened.

Agents using a subscription CLI backend (claude-code-cli / codex-cli) need no API key. Instead a credentials PVC is mounted read-write at /creds (CLAUDE_CONFIG_DIR=/creds/claude, CODEX_HOME=/creds/codex), populated once via an interactive login pod. See LLM backends.

NetworkPolicy​

Runner pods are confined by a NetworkPolicy (app=nexus-agent-runner):

  • Ingress: denied entirely — nothing reaches a runner.
  • Egress: DNS, the NATS event bus (port 4222 — as the runner user, which may only publish its run/log/tool subjects), and outbound TCP 443 to any IPv4 destination (0.0.0.0/0:443). That is a port restriction, not an LLM-provider allowlist (audit 2026-09-05 S7): a runner can reach any HTTPS endpoint, and DNS is not destination-restricted. An application-layer egress/SSRF policy remains open (see Planned hardening).

Enforcement requires a NetworkPolicy-capable CNI (Calico/Cilium); otherwise it is a safe no-op.

Signer ingress (nexus-solana-exec)​

The chart's nexus-solana-exec-callers policy admits :8091 only from the pods named in networkPolicy.signerCallers (worker, core, telegram, ui). That alone was not the effective boundary until 2026-09-05: Kubernetes NetworkPolicies are additive, and the nexus-allow-monitoring policy (selecting every Nexus pod) admitted the monitoring namespace on every port — the audit reached exec:8091/health from the Grafana container (audit S2, live-verified). Earlier docs saying the signer "accepts calls only from four workloads" were wrong for that reason.

Narrowed on 2026-09-05 (nexus-gitops 34f731d): nexus-allow-monitoring is limited to networkPolicy.monitoringPorts ([9100]), so the effective signer ingress is {callers:8091, monitoring:9100}. Live-verified the same day: a Grafana → exec:8091 GET times out while Grafana → worker:9100 and worker → exec:8091 still answer.

Also in 34f731d, live-verified 2026-09-05: an explicit securityContext on the signer pod (runAsNonRoot 1000, RuntimeDefault seccomp, no privilege escalation, drop ALL, read-only root filesystem with an emptyDir at /tmp) — audit S6; the live pod previously had an empty securityContext.

Signer caller authentication (EXEC_AUTH_MODE=enforce)​

Caller authentication on the signer API and semantic validation of the transactions it signs (audit S1, critical — closed) are live: code in nexus-solana 2c0be5a (deployed 2026-09-05 15:55 UTC), enforce mode since 16:27 UTC (nexus-gitops 29c983a, solanaExec.auth.mode: enforce), live-verified the same minute.

HeaderX-Nexus-Exec-Key on every request to nexus-solana-exec
Rolesworker, telegram = execute (mutating routes); core, ui = read; self = the exec's own background loops (fill-watcher, ladder tick)
Missing / unknown key401 auth:missing (or auth:unknown) on every route, /health excepted
Read-role key on a mutation403 auth:role
Counters/v1/status.auth → {allowed, would_deny, denied}; Prometheus nexus_exec_auth_decisions{decision} once the treasury exporter ships; alert NexusSignerAuthDenied
Where the keys liveSecret nexus-exec-auth (EXEC_API_KEYS, the whole caller map) mounted on the exec; each caller gets its own SOLANA_EXEC_API_KEY from the nexus-app (worker, telegram, core) or nexus-ui Secret
Rotationedit .secrets → nexus-gitops/infra/bootstrap-secrets.sh (rewrites both sides) → roll the exec and the callers. No scheduled rotation.

Verified after the flip: {allowed: 32, would_deny: 0, denied: 0}, key-less GET → 401, core (read) POST → 403, worker and telegram logs clean, 0 restarts. The flip followed 90 min of log mode with 165 allowed calls and 0 caller denials (the only would_deny entries were key-less operator probes). Operator probes must now send a key; the Core read key is the right one.

Semantic validation, enforced on all three signing paths since 2c0be5a: address-lookup-table resolution, a single-signer check, a program allowlist (loaders/ALT/stake/vote hard-denied), delegate / authority / close / transfer decoding and simulated token-delta bounds against a 90 s TxIntent. Not enforced by design (published on /v1/status not_enforced, as running on a74cc87): inner CPI programs, Token-2022 balances, LP batch sums, the Ledger-signed submit_signed relay.

Receipt bounds (nexus-solana 83bb7d8, in a1eb8f3, built 2026-09-05 23:05 UTC, live-verified 23:33 UTC — the re-verification's "a lending-supply intent bounds the outgoing mint but not a minimum receipt-token position"): a Kamino supply's cToken leg, a JitoSOL deposit / withdraw and an LP remove / close that fits one transaction are bounded pre-sign as receives legs (fungible min = expected × (1 − EXEC_RECEIPT_TOLERANCE_FRAC 0.01); an unreadable cToken rate refuses the supply); a MarginFi supply and an LP open are verified after the send from the venue read (receipt.verified) — reported, not enforceable. The not_enforced list becomes seven entries: inner_programs, token2022_balances, lp_batch_sum, lp_open_receipt_presign, lp_collect_receipt, lending_receipt_untokenized, submit_signed — nexus-solana → Receipt bounds.

This is the signer's auth. Core's own nexus-auth middleware is a separate thing, in NEXUS_AUTH_MODE=enforce since 22:37 UTC (S3 above).

Execution controls: intent idempotency and fencing​

Authentication says who may ask the signer to move money; the semantic gate says what a transaction may do; neither says how many times. Two controls, both deployed 2026-09-05 and live-verified that night, close that:

  • Exec-side idempotency + fencing (nexus-solana 23c083b in a1eb8f3, live-verified 2026-09-05 23:33 UTC; audit F2). Every mutating route accepts intent_id (the caller's durable decision id) and fence (monotonic per intent). One accum_intents record per id is driven by the three signing choke points, with the prepared → submitted CAS taken after authorize and before the broadcast, so a lost race never puts a second signature on the wire. A repeat under the same id is answered with 409 intent:mismatch (different economic parameters), 409 intent:stale_fence (an old executor after a takeover), the stored outcome with replayed: true (confirmed, or failed with a signature — nothing signed again), or 409 intent:in_flight (signature recorded, outcome unknown — reconcile on-chain); only an intent that never touched the key is re-armed. GET /v1/intents/{id} is the operator's view. Requests without intent_id are untouched; since nexus-platform 2d8eef1 (2026-09-05 23:5xZ) every Telegram step and worker exec call sends intent_id + fence — a live replay has not been observed yet (nexus-solana → Idempotency).
  • Fenced budget claims (nexus-platform 1e6a080 / b900de9 / cf219fa in ee0cd80, 23:00 UTC; audit F6 / F2). A reservation on the aggregate accum_budget_ledger must be claimed (held → executing, lease + fence) before money moves; consume / release need the fence; a dead lease can be taken over and the previous holder's settlement is then refused and counted (nexus_portfolio_fence_rejections_total, alert NexusBudgetFenceRejection, critical). Worker sweeps and Telegram-approved buys claim on it; the exec reads Core's verdict before a rung fires. Under NEXUS_AUTH_MODE=enforce the Core budget mutations (POST /v1/desk/budget/reserve, …/claim|consume|release) accept service keys only — the read key and the UI proxy get 403 (API reference, Operations → Budget ledger).

A signature that was never broadcast is not a spend (nexus-solana 0641ea5, deployed 2026-09-06 in ea96739). Record-before-send means the signature is written into the intent before the broadcast, which is correct — but it also meant that a send the node rejected at preflight settled failed with a signature, and the exactly-once rule ("failed with a signature ⇒ replay the stored outcome, never sign again") then locked that intent out for ever. It happened twice on 2026-09-06 (-32002 … Blockhash not found, a commitment mismatch fixed in 7046ba8), and both lending supplies were unrecoverable under their own ids.

send::classify now returns NeverBroadcast for exactly one shape — the JSON-RPC answer the node returns from sendTransaction instead of forwarding the transaction (SendTransactionPreflightFailure, or the bare code -32002). Only then does the record move its signature into unsent_signatures (kept across the re-arm as the audit trail), settle failed without a landed signature, and become re-armable. Every other failure — transport error, timeout, an unresolved confirm — stays Ambiguous, and the intent stays submitted: capital may be in flight and only reconciliation against the chain settles it. The safe direction is unchanged; what changed is that a proven non-spend is no longer treated as a possible one. LP batch sets keep the old behaviour from batch 1 on (an earlier batch has already landed, so replay is the safe answer), and the legacy ladder journal writes failed rather than submitted_unknown for the same verdict — nexus-solana → Send path.

What this still does not give: automatic reconciliation of an in_flight intent (the caller or an operator settles it from the signature on GET /v1/intents/{id}), and any live crash / replay rehearsal. Callers are wired: since nexus-platform 2d8eef1 every Telegram approval step and every worker exec call sends intent_id + fence, and the first live exercise of the path — a 995.55 USDC Kamino supply, confirmed, attempts: 1 — was observed on 2026-09-06 ≈ 09:0x UTC; no replay or refusal has been observed yet. Register F2.

Memory safety​

Agents cannot write permanent shared memory unattended at first — proposals go to a review queue and a human approves them. See Memory.

Secrets​

  • .env never committed.
  • As built (2026-09-05): Kubernetes Secrets created by nexus-gitops/bootstrap-secrets.sh from the operator's git-ignored .secrets file. There is no secrets manager and no scheduled rotation; rotations happen ad hoc (e.g. hermes_rw after a log leak on 2026-09-05). A recovery drill for the EXEC signing key has not been run (audit O1).

Planned hardening​

Nexus's container isolation is already at or beyond parity with comparable agents, but several runner-side safety layers are planned (see the Hermes gap analysis): a command firewall (pattern + smart-classifier + always-on hardline blocklist), application-layer egress/SSRF policy, input-trust scanning of SOUL/skill/memory/context before prompting, and per-skill credential scoping with secret redaction of tool output.