Security & permissions
Nexus runs autonomous agents that touch real repos and clusters. Security is defense in depth across four layers.
Layers
| Layer | Enforces |
|---|---|
| Nexus Core | Whether an agent may be dispatched with a given tool/permission set; raises approvals for gated actions. |
| Runner | Refuses any action the run config didn't grant. |
| Kubernetes RBAC | The run's service account can only touch what its role allows — enforced since 2026-09-05 (Node,RBAC). |
nexus-auth crate | Intended: admin users, API keys, agent tokens, tool/project permissions. As built (2026-09-05): the pure auth policy — service key (X-Nexus-Api-Key vs NEXUS_API_KEY), operator identity (X-Nexus-Operator-Sub/Email, honoured only behind a valid key, checked against NEXUS_OPERATOR_ALLOWLIST), modes off/log/enforce. nexus-core runs it as middleware on every /v1 route; production is in log-only mode (audit H11 partial). Admin JWT / agent tokens remain unimplemented. |
A permission must be granted at every layer for an action to succeed — in
design. Where that stands (2026-09-05): Kubernetes RBAC is enforced (the
authorizer moved from AlwaysAllow to Node,RBAC; audit C1 closed). Core
evaluates the service key and the operator allowlist on every /v1 request
and enforces since 2026-09-05 22:37 UTC (NEXUS_AUTH_MODE=enforce,
NEXUS_AUTH_READ=open — nexus-gitops c95071e, audit S3 closed after
24 h of log mode with 0 would_deny): an anonymous mutation is 401
no_key, reads stay open, decisions are counted in
nexus_core_auth_decisions_total. Rollback is core.auth.mode: log.
The flip was not merely a values change (audit 2026-09-05 S3/S4): it waited for the identity binding under Identity propagation below.
Identity propagation (UI → Core)
Status 2026-09-05: specified; implementation in progress in nexus-ui
(proxy side) and nexus-platform (Core side); not deployed. The deployed UI
(63a1083) forwards browser headers, injects X-Nexus-Api-Key and neither
strips inbound X-Nexus-Operator-Sub/Email nor derives them from the verified
session — so under enforce a signed-in browser request without identity
headers would look like a privileged service request, and a browser could
supply its own operator assertion (audit S4).
The contract that closes it:
| Credential / header | Principal | Rule |
|---|---|---|
NEXUS_UI_API_KEY | ui-proxy | Must carry an operator identity taken from the verified session (X-Nexus-Operator-Sub/Email set by the proxy, every inbound X-Nexus-* header stripped first). Never promoted to a service principal. |
NEXUS_API_KEY | service (worker, telegram, …) | Full service rights; no operator assertion honoured. |
NEXUS_API_KEY_READ (optional) | service, read-only | Split read/write service permissions. |
X-Nexus-Proxy-Origin | attestation | Required on UI-originated mutations; Core rejects a mutation under the ui-proxy key without it. |
NEXUS_AUTH_MODE=enforce was gated on the UI change being deployed and on
would_deny staying at zero afterwards — not on the values flip alone; both
held (UI a2561d9 15:14 UTC, 383 decisions / 0 would_deny) before the
22:37 UTC flip.
Identity & tokens
- Admin users — JWT (
jsonwebtoken), passwords hashed withargon2. - API keys — for machine clients.
- Agent tokens — short-lived tokens minted per run so a pod can report back to Core, scoped to that run only.
- Telegram allowlist —
TELEGRAM_ALLOWED_USERS(in thenexus-appSecret) restricts/approve,/reject,/basisand the rest. It is set in prod and, since nexus-platform445e107(deployed 2026-09-05), an empty list denies everyone — the gateway fails closed (audit C7 closed).
Permission model
agent.can_create_branch · agent.can_commit · agent.can_open_pr · agent.can_merge
agent.can_delete_files · agent.can_use_kubectl · agent.can_apply_k8s
agent.can_write_memory · agent.can_create_board_task
Approval gates
Agents declare actions that always need a human:
requires_human_approval_for:
- production_deploy
- dependency_upgrade
- database_migration
Gated attempts raise an approval request.
Kubernetes isolation
- Each run is an isolated Job in a dedicated agent namespace (
nexus-agents). - Core/worker service accounts can manage Jobs only there:
verbs: [create, get, list, watch, delete]
resources: [jobs, pods, pods/log]
Runner pod hardening
Every agent Job is created with a locked-down pod (see nexus-k8s):
| Control | Setting |
|---|---|
| Run as non-root | runAsNonRoot: true, runAsUser/Group: 1000, fsGroup: 1000 |
| No privilege escalation | allowPrivilegeEscalation: false, privileged: false |
| Drop capabilities | capabilities.drop: ["ALL"] |
| Immutable root FS | readOnlyRootFilesystem: true (+ writable emptyDir at /tmp and /workspace) |
| Seccomp | seccompProfile: RuntimeDefault |
| No cluster access | automountServiceAccountToken: false |
| Auto-cleanup | ttlSecondsAfterFinished: 3600 |
Secret injection
LLM keys are not baked into images. They live in the nexus-agent Secret in
the agent namespace and are injected via envFrom (referenced as optional, so a
missing secret fails the run fast rather than blocking scheduling):
GIT_TOKEN · GITHUB_TOKEN · GIT_USERNAME # what the prod Secret holds
OPENAI_API_KEY · ANTHROPIC_API_KEY # optional; absent in prod — agents run on the CLI logins below
GIT_TOKEN (falling back to GITHUB_TOKEN / GITLAB_TOKEN) authenticates clones
of private https repositories in the agent's repository list.
The runner injects it into the clone URL at run time and never logs it; clone
errors are credential-redacted before they reach logs or run results. Repos marked
read_only are a convention the prompt conveys to the agent — enforce non-push
hard limits via permissions and review gates. Prefer https clones: the runner
NetworkPolicy allows egress on 443, so SSH (git@…, port 22) needs the egress
rule widened.
Agents using a subscription CLI backend (claude-code-cli / codex-cli) need
no API key. Instead a credentials PVC is mounted read-write at /creds
(CLAUDE_CONFIG_DIR=/creds/claude, CODEX_HOME=/creds/codex), populated once via
an interactive login pod. See LLM backends.
NetworkPolicy
Runner pods are confined by a NetworkPolicy (app=nexus-agent-runner):
- Ingress: denied entirely — nothing reaches a runner.
- Egress: DNS, the NATS event bus (port 4222 — as the
runneruser, which may only publish its run/log/tool subjects), and outbound TCP 443 to any IPv4 destination (0.0.0.0/0:443). That is a port restriction, not an LLM-provider allowlist (audit 2026-09-05 S7): a runner can reach any HTTPS endpoint, and DNS is not destination-restricted. An application-layer egress/SSRF policy remains open (see Planned hardening).
Enforcement requires a NetworkPolicy-capable CNI (Calico/Cilium); otherwise it is a safe no-op.
Signer ingress (nexus-solana-exec)
The chart's nexus-solana-exec-callers policy admits :8091 only from the
pods named in networkPolicy.signerCallers (worker, core, telegram, ui).
That alone was not the effective boundary until 2026-09-05: Kubernetes
NetworkPolicies are additive, and the nexus-allow-monitoring policy
(selecting every Nexus pod) admitted the monitoring namespace on every
port — the audit reached exec:8091/health from the Grafana container
(audit S2, live-verified). Earlier docs saying the signer "accepts calls
only from four workloads" were wrong for that reason.
Narrowed on 2026-09-05 (nexus-gitops 34f731d): nexus-allow-monitoring is
limited to networkPolicy.monitoringPorts ([9100]), so the effective signer
ingress is {callers:8091, monitoring:9100}. Live-verified the same day: a
Grafana → exec:8091 GET times out while Grafana → worker:9100 and
worker → exec:8091 still answer.
Also in 34f731d, live-verified 2026-09-05: an explicit securityContext on
the signer pod (runAsNonRoot 1000, RuntimeDefault seccomp, no privilege
escalation, drop ALL, read-only root filesystem with an emptyDir at
/tmp) — audit S6; the live pod previously had an empty securityContext.
Signer caller authentication (EXEC_AUTH_MODE=enforce)
Caller authentication on the signer API and semantic validation of the
transactions it signs (audit S1, critical — closed) are live:
code in nexus-solana 2c0be5a (deployed 2026-09-05 15:55 UTC),
enforce mode since 16:27 UTC (nexus-gitops 29c983a,
solanaExec.auth.mode: enforce), live-verified the same minute.
| Header | X-Nexus-Exec-Key on every request to nexus-solana-exec |
| Roles | worker, telegram = execute (mutating routes); core, ui = read; self = the exec's own background loops (fill-watcher, ladder tick) |
| Missing / unknown key | 401 auth:missing (or auth:unknown) on every route, /health excepted |
| Read-role key on a mutation | 403 auth:role |
| Counters | /v1/status.auth → {allowed, would_deny, denied}; Prometheus nexus_exec_auth_decisions{decision} once the treasury exporter ships; alert NexusSignerAuthDenied |
| Where the keys live | Secret nexus-exec-auth (EXEC_API_KEYS, the whole caller map) mounted on the exec; each caller gets its own SOLANA_EXEC_API_KEY from the nexus-app (worker, telegram, core) or nexus-ui Secret |
| Rotation | edit .secrets → nexus-gitops/infra/bootstrap-secrets.sh (rewrites both sides) → roll the exec and the callers. No scheduled rotation. |
Verified after the flip: {allowed: 32, would_deny: 0, denied: 0}, key-less
GET → 401, core (read) POST → 403, worker and telegram logs clean, 0
restarts. The flip followed 90 min of log mode with 165 allowed calls and 0
caller denials (the only would_deny entries were key-less operator
probes). Operator probes must now send a key; the Core read key is the
right one.
Semantic validation, enforced on all three signing paths since
2c0be5a: address-lookup-table resolution, a single-signer check, a program
allowlist (loaders/ALT/stake/vote hard-denied), delegate / authority /
close / transfer decoding and simulated token-delta bounds against a 90 s
TxIntent. Not enforced by design (published on /v1/status
not_enforced, as running on a74cc87): inner CPI programs, Token-2022
balances, LP batch sums, the Ledger-signed submit_signed relay.
Receipt bounds (nexus-solana 83bb7d8, in a1eb8f3, built 2026-09-05
23:05 UTC, live-verified 23:33 UTC — the re-verification's "a lending-supply intent bounds
the outgoing mint but not a minimum receipt-token position"): a Kamino
supply's cToken leg, a JitoSOL deposit / withdraw and an LP remove / close
that fits one transaction are bounded pre-sign as receives legs
(fungible min = expected × (1 − EXEC_RECEIPT_TOLERANCE_FRAC 0.01); an
unreadable cToken rate refuses the supply); a MarginFi supply and an LP
open are verified after the send from the venue read (receipt.verified)
— reported, not enforceable. The not_enforced list becomes seven
entries: inner_programs, token2022_balances, lp_batch_sum,
lp_open_receipt_presign, lp_collect_receipt,
lending_receipt_untokenized, submit_signed —
nexus-solana → Receipt bounds.
This is the signer's auth. Core's own nexus-auth middleware is a
separate thing, in NEXUS_AUTH_MODE=enforce since 22:37 UTC (S3 above).
Execution controls: intent idempotency and fencing
Authentication says who may ask the signer to move money; the semantic gate says what a transaction may do; neither says how many times. Two controls, both deployed 2026-09-05 and live-verified that night, close that:
- Exec-side idempotency + fencing (nexus-solana
23c083bina1eb8f3, live-verified 2026-09-05 23:33 UTC; audit F2). Every mutating route acceptsintent_id(the caller's durable decision id) andfence(monotonic per intent). Oneaccum_intentsrecord per id is driven by the three signing choke points, with theprepared → submittedCAS taken afterauthorizeand before the broadcast, so a lost race never puts a second signature on the wire. A repeat under the same id is answered with 409intent:mismatch(different economic parameters), 409intent:stale_fence(an old executor after a takeover), the stored outcome withreplayed: true(confirmed, or failed with a signature — nothing signed again), or 409intent:in_flight(signature recorded, outcome unknown — reconcile on-chain); only an intent that never touched the key is re-armed.GET /v1/intents/{id}is the operator's view. Requests withoutintent_idare untouched; since nexus-platform2d8eef1(2026-09-05 23:5xZ) every Telegram step and worker exec call sendsintent_id+fence— a live replay has not been observed yet (nexus-solana → Idempotency). - Fenced budget claims (nexus-platform
1e6a080/b900de9/cf219fainee0cd80, 23:00 UTC; audit F6 / F2). A reservation on the aggregateaccum_budget_ledgermust be claimed (held → executing, lease + fence) before money moves; consume / release need the fence; a dead lease can be taken over and the previous holder's settlement is then refused and counted (nexus_portfolio_fence_rejections_total, alertNexusBudgetFenceRejection, critical). Worker sweeps and Telegram-approved buys claim on it; the exec reads Core'sverdictbefore a rung fires. UnderNEXUS_AUTH_MODE=enforcethe Core budget mutations (POST /v1/desk/budget/reserve,…/claim|consume|release) accept service keys only — the read key and the UI proxy get 403 (API reference, Operations → Budget ledger).
A signature that was never broadcast is not a spend (nexus-solana
0641ea5, deployed 2026-09-06 in ea96739). Record-before-send means the
signature is written into the intent before the broadcast, which is
correct — but it also meant that a send the node rejected at preflight
settled failed with a signature, and the exactly-once rule
("failed with a signature ⇒ replay the stored outcome, never sign
again") then locked that intent out for ever. It happened twice on
2026-09-06 (-32002 … Blockhash not found, a commitment mismatch fixed in
7046ba8), and both lending supplies were unrecoverable under their own
ids.
send::classify now returns NeverBroadcast for exactly one shape — the
JSON-RPC answer the node returns from sendTransaction instead of
forwarding the transaction (SendTransactionPreflightFailure, or the bare
code -32002). Only then does the record move its signature into
unsent_signatures (kept across the re-arm as the audit trail), settle
failed without a landed signature, and become re-armable. Every other
failure — transport error, timeout, an unresolved confirm — stays
Ambiguous, and the intent stays submitted: capital may be in flight and
only reconciliation against the chain settles it. The safe direction is
unchanged; what changed is that a proven non-spend is no longer treated
as a possible one. LP batch sets keep the old behaviour from batch 1 on
(an earlier batch has already landed, so replay is the safe answer), and
the legacy ladder journal writes failed rather than submitted_unknown
for the same verdict —
nexus-solana → Send path.
What this still does not give: automatic reconciliation of an in_flight
intent (the caller or an operator settles it from the signature on
GET /v1/intents/{id}), and any live crash / replay rehearsal. Callers
are wired: since nexus-platform 2d8eef1 every Telegram approval step
and every worker exec call sends intent_id + fence, and the first live
exercise of the path — a 995.55 USDC Kamino supply, confirmed,
attempts: 1 — was observed on 2026-09-06 ≈ 09:0x UTC; no replay or
refusal has been observed yet. Register
F2.
Memory safety
Agents cannot write permanent shared memory unattended at first — proposals go to a review queue and a human approves them. See Memory.
Secrets
.envnever committed.- As built (2026-09-05): Kubernetes Secrets created by
nexus-gitops/bootstrap-secrets.shfrom the operator's git-ignored.secretsfile. There is no secrets manager and no scheduled rotation; rotations happen ad hoc (e.g.hermes_rwafter a log leak on 2026-09-05). A recovery drill for the EXEC signing key has not been run (audit O1).
Planned hardening
Nexus's container isolation is already at or beyond parity with comparable agents, but several runner-side safety layers are planned (see the Hermes gap analysis): a command firewall (pattern + smart-classifier + always-on hardline blocklist), application-layer egress/SSRF policy, input-trust scanning of SOUL/skill/memory/context before prompting, and per-skill credential scoping with secret redaction of tool output.