Skip to main content

Goal lifecycle engine

A target spec (not a schedule) for evolving Nexus's single decompose step into a full goal lifecycle engine: a pipeline of composable stages that turns a vague human command into reviewable, executable board tasks — and keeps steering the goal through execution, re-planning, and closure.

Why this exists

When this spec was written the worker had exactly one goal-level operation: it found approved goals and decomposed them into tasks (decompose_goals in nexus-worker/src/dispatch.rs). That one-shot step answered "what tasks are needed?" and nothing else. This spec adds the missing stages around the decompose we already have, without rebuilding the executor.

Where it stands (2026-09-05, deployed at nexus-platform e3d7c4c): clarify_goals and close_goals now exist in the worker beside decompose_goals — more of this engine is built than the sections below say (they are kept as the target). Read the state table in §4 against that.

Approval semantics differ from this spec (audit 2026-09-05 F8)

As deployed, clarify_goals moves a draft goal straight to approved as soon as assess_specificity judges it specific enough — or once it has exhausted its clarify rounds — without a human approving it. A sufficiently detailed description is therefore enough to create executable work for agents that operate on repositories and infrastructure. The planned → approved human gate described in §4 and §7 is the documented contract, not the running behaviour. Whether to add a ready_for_review state and require a recorded approving identity, or to explicitly scope an automatic pre-authorisation policy, is an open product decision — not decided as of 2026-09-05.

1. The one correction that shapes everything​

A goal lifecycle is tempting to model as ~20 peer "engine modules" (clarify, classify, scope, assign, execute, review, test, deploy…). Don't. Those verbs live on three different planes, and Nexus already has the right backbone — Goal → Task → Run. Flattening them couples the planner to the executor and throws away the DAG.

PlaneEntityExecuted byOperations
Goal planeGoalOrchestrator / Planner LLM in the worker (cheap aux model)intake · clarify · classify · scope · decompose · plan · prioritize · replan · summarize · close · learn
Task planeTaskDeterministic worker logic (already exists)assign (pick_agent) · schedule/lease · spawn (K8s Job) · promote children · sync_board (reconcile)
Run planeAgentRunThe agent, inside the podprepare_workspace · execute · handoff · merge · deploy
review/test/deploy are roles, not stages

In Nexus, "review this", "test this", "deploy this" are already just tasks with a role (reviewer, tester, devops) emitted by the decomposer and ordered by depends_on. That is strictly better than a hardcoded review→test→deploy pipeline, because the DAG decides ordering per-goal. This engine emits those tasks; it does not execute them. Keep AgentRole::{Reviewer, Tester, DevOps} as the mechanism.

So the real work is goal-plane stages before and after decompose — everything on the Task and Run planes is built.

2. Where Nexus is today (grounded)​

  • GoalStatus (nexus-db): Draft → Approved → InProgress → Done (+ Cancelled).
  • Flow: POST /v1/goals creates a Draft; the worker's clarify_goals either asks a clarifying question or — for a specific draft — sets it approved itself (no human gate today, audit F8); the worker tick picks up approved goals, calls plan_tasks(db, goal) (a one-shot Planner LLM emitting n | title | role | deps), and create_task_for(...) writes Task rows (ReadyForAgent) + BoardItems, wiring depends_on. close_goals closes goals whose tasks are all done.
  • Task fields carry status, role, assigned_agent_id, lease, depends_on — but no priority, no acceptance_criteria, no estimate.
  • Goal fields carry title, description, status — but no origin surface, no labels/classification, no scope, no acceptance criteria, no clarification thread.
  • No re-entry: once decomposed, a goal is never re-planned. A blocked task just stalls.

3. Operation taxonomy (your list, grounded)​

✅ exists · ⚠️ partial · ❌ missing. Type: det deterministic · llm (use the aux model) · human gate.

A. Goal-shaping — before decompose (highest leverage; mostly missing)​

OpQuestionTodayLands inType
intakecapture raw command✅ POST /v1/goalsnexus-coredet
clarifyis this specific enough? what's missing?❌new stage → asks origin surface, waitsllm + human
classifybackend / frontend / research / bugfix?❌ (role guessed per-task)new stage → goal labelsllm
scopewhat's in / out?❌folded into decompose promptllm

B. Planning​

OpTodayGap
decompose✅ plan_tasksthin prompt — fold in classify/scope/acceptance-criteria
plan / order⚠️ deps → depends_on DAGno milestones; no plan object to review pre-approval
prioritize❌ no fieldadd Task.priority + sort
estimate❌optional complexity/risk → feeds priority + model routing
acceptance_criteria❌Vec<String> on Task — needed for goal-mode & close

C. Execution & coordination — mostly built​

OpToday
assign✅ pick_agent(role) (least-loaded, role-matched)
spawn / prepare / execute✅ lease → K8s Job → runner clones repos → agent runs
handoff⚠️ board.comment exists; add structured result.metadata (changed_files, verification, blocked_reason)
sync_board✅ reconcile → project_to_taiga
review / test / deploy✅ as role tasks (not stages)

D. Control & closure (the other big gap)​

OpQuestionTodayType
detect_blockerwhy is it stuck?⚠️ runs block, no diagnosisllm
replanplan changed — what new tasks?❌ (decompose runs once)llm
summarizewhat's the goal status?❌ (UI shows counts)llm
closeacceptance criteria met?⚠️ manual goals/{id}/completellm judge + human
learnsave what worked/failed⚠️ memory.propose/skill.propose per-runwire to close

4. Proposed goal state machine​

Don't model 20 verbs — model the GoalStatus enum that drives them. Extend the existing enum (new variants in bold):

draft
→ needs_clarification ⇄ (human answers on origin surface) [clarify]
→ classified [classify + scope]
→ planned (tasks drafted, awaiting human approval) [decompose + prioritize]
→ approved (existing human gate — keep it)
→ in_progress ⇄ replanning [lease/execute ⇄ replan]
→ review [acceptance check]
→ done | cancelled
└→ (learn fires on terminal transition)
pub enum GoalStatus {
Draft,
NeedsClarification, // new — blocked on a human answer
Classified, // new — labelled + scoped, ready to plan
Planned, // new — tasks drafted, awaiting approval (replaces implicit gap)
Approved, // existing human gate
InProgress, // existing
Replanning, // new — sub-state: a blocker triggered re-decompose
Review, // new — all tasks done, checking acceptance criteria
Done,
Cancelled,
}

Transition table (each row = one stage firing on the worker tick or an event):

FromStageToTrigger
draftClarifyneeds_clarification or classifiedLLM judges specificity; vague ⇒ ask, clear ⇒ pass
needs_clarification(ingest answer)classifiedhuman replies on origin surface / UI
classifiedClassify+Scope, then Decompose+Prioritizeplannedaux model labels; Planner emits tasks
planned(human)approved / cancelledoperator approves the drafted plan
approved(existing leasing)in_progressfirst task leased
in_progressreconcile detects blocked/failed runreplanninga run blocks
replanningReplanin_progressremediation tasks emitted
in_progressall tasks donereviewlast task completes
reviewClose (judge + human)doneacceptance criteria pass
terminalLearn—side-effect: memory/skill write-back

5. Architecture — a new nexus-orchestrator crate​

A stage pipeline, not 20 services. One new crate, consumed by nexus-worker.

nexus-platform/crates/nexus-orchestrator/
├── src/
│ ├── lib.rs // GoalStage trait + StageOutcome + the pipeline runner
│ ├── ctx.rs // Ctx: db, event bus, aux-LLM handle, board provider
│ └── stages/
│ ├── clarify.rs
│ ├── classify.rs // classify + scope
│ ├── decompose.rs // wraps today's plan_tasks; richer prompt
│ ├── prioritize.rs
│ ├── replan.rs
│ ├── summarize.rs
│ └── close.rs
/// One step of the goal lifecycle. Stages are pure w.r.t. the Goal they receive
/// and report their effects via StageOutcome — they never mutate other goals.
#[async_trait]
pub trait GoalStage: Send + Sync {
/// Which status this stage handles. The runner dispatches by goal status.
fn applies_to(&self, status: GoalStatus) -> bool;

/// Run the stage. Returns the next status + side effects (tasks to create,
/// a question to ask, events to emit). Idempotent: safe to re-run on the
/// same goal if the worker crashes mid-tick.
async fn run(&self, goal: &Goal, ctx: &Ctx) -> anyhow::Result<StageOutcome>;
}

pub enum StageOutcome {
Advance { to: GoalStatus },
AdvanceWithTasks { to: GoalStatus, tasks: Vec<PlannedTask> },
AskHuman { question: Clarification }, // → NeedsClarification, then wait
Hold, // nothing to do this tick
}

Integration with the existing worker: the tick already loops goals by status. Replace the single decompose_goals call with a pipeline.advance(goal) that selects the matching stage:

for goal in db.goals_in_active_states().await? {
match pipeline.advance(&goal, &ctx).await? {
StageOutcome::AdvanceWithTasks { to, tasks } => {
for t in tasks { create_task_for(&db, &goal, t).await?; }
db.set_goal_status(&goal.id, to).await?;
}
StageOutcome::AskHuman { question } => deliver_clarification(&goal, question, &ctx).await?,
StageOutcome::Advance { to } => db.set_goal_status(&goal.id, to).await?,
StageOutcome::Hold => {}
}
}

Aux-model lane. Cheap stages (clarify, classify, summarize, detect_blocker) must not burn the primary model. They use the auxiliary model lane (roadmap item 18) configured alongside planner/summarizer in Settings.llm. Decompose/replan may use the planner model.

Events. Each transition publishes a durable NATS subject so the UI and gateways react live: nexus.goal.clarification_requested, nexus.goal.clarified, nexus.goal.planned, nexus.goal.replanned, nexus.goal.summarized, nexus.goal.closed. (Mirrors the existing nexus.task.* / nexus.agent.run.*.)

6. Data-model changes​

Goal gains:

pub struct Goal {
// … existing fields …
pub origin: Option<Origin>, // where to ask clarifications back
pub labels: Vec<String>, // classify output (backend, research, …)
pub scope: Option<String>, // in/out-of-scope, LLM-authored
pub acceptance_criteria: Vec<String>, // drives Close
pub clarifications: Vec<Clarification>, // Q&A thread (ask-then-wait)
}

pub struct Origin { // so clarify knows which surface to reply on
pub surface: String, // "telegram" | "ui" | "api"
pub chat_id: Option<String>,
pub thread_id: Option<String>,
}

pub struct Clarification {
pub question: String,
pub asked_at: DateTime<Utc>,
pub answer: Option<String>,
pub answered_at: Option<DateTime<Utc>>,
}

Task gains: priority: i32 (default 0) and acceptance_criteria: Vec<String>. priority is also a board-UX win (badge + column sort — see the Kanban gap analysis).

7. The clarify gate (decided: ask, then wait)​

The headline new capability — what turns "build a login system" into a good plan.

  1. Stage runs on draft. The clarify stage prompts the aux model: "Given this goal and the project profile, is it specific enough to plan? If not, list the 1–3 highest-value questions." The project profile (Project.profile / profile_refs) is in scope so it won't re-ask known facts.
  2. If specific → StageOutcome::Advance { to: Classified }.
  3. If vague → StageOutcome::AskHuman { question }:
    • Goal moves to needs_clarification (hard stop — no tasks drafted).
    • The question is delivered to goal.origin: a Telegram reply on the original chat/thread, and surfaced on the Goals page with an inline answer box.
    • Publishes nexus.goal.clarification_requested.
  4. Answer ingestion.
    • UI: POST /v1/goals/{id}/clarify { answer }.
    • Telegram: a reply in the goal's thread maps back via origin.thread_id.
    • The answer is appended to goal.clarifications, the goal returns to draft, and clarify re-runs (it can ask once more, but cap at 2 rounds then pass through with stated assumptions to avoid loops).
  5. Idempotency: a goal in needs_clarification is skipped by the tick until an answer arrives, so re-ticks are free and safe.
Keep the human approval gate

Clarify does not replace the existing planned → approved human gate. Vague goals get clarified then planned then still approved. Two gates, two purposes: clarify fixes under-specification; approval authorizes spend + dispatch.

8. The replan loop​

What makes this an engine rather than a one-shot planner.

  • Trigger: reconcile.rs already handles AGENT_RUN_FAILED / a task entering Blocked. On that event, move the parent goal to replanning and emit nexus.goal.replanned intent.
  • detect_blocker (aux model): summarize why (failed build, missing dependency, auth error) from the run's error + trajectory.
  • Replan stage: re-enter decompose with the original plan + the blocker diagnosis as context, emitting only new/remediation tasks (e.g. "Investigate Plane API token", "Add retry handling") wired depends_on the blocked task — never duplicating completed work. Goal returns to in_progress.
  • Guard: cap replans per goal (e.g. 3) → otherwise needs_clarification or a human-review approval, so a pathological goal can't fan out forever.

9. V1 scope (smallest useful engine)​

Everything on the Task/Run planes already works, so V1 is small and additive:

  1. Fields: Task.priority + acceptance_criteria (Task & Goal) end-to-end (domain → db → core API → UI badge/sort). Unblocks prioritize, close, goal-mode.
  2. nexus-orchestrator crate skeleton: GoalStage trait + pipeline runner + Decompose stage wrapping today's plan_tasks (behavior-preserving refactor).
  3. Clarify stage (ask-then-wait): new statuses, Origin, Clarification, POST /v1/goals/{id}/clarify, Telegram reply mapping, Goals-page answer box.
  4. Replan on block: reconcile → Replanning → remediation tasks.
  5. Summarize stage: aux-model goal status report on the Goals page + Telegram.

That delivers the full loop — clarify → (classify/scope folded in) → decompose → execute → review(role) → replan → summarize → close — without touching leasing, K8s, or Taiga sync.

10. Out of scope / keep as-is​

  • The executor. pick_agent, leasing, K8s Jobs, runner, reconcile, project_to_taiga are done and good. The engine emits tasks; the dispatcher runs them.
  • review/test/deploy as engine stages — they stay role tasks.
  • 20 separate crates/services — one nexus-orchestrator crate with stages.
  • Bypassing the approval gate — clarify adds a gate, it doesn't remove one.

11. Open questions​

  • Classify granularity: distinct classified status, or fold classify+scope silently into the decompose prompt? (Spec keeps the status for observability; it's cheap to collapse later.)
  • Goal-mode vs review-role: does per-goal acceptance checking belong in a Close stage (judge) or as a final role: reviewer task? (Spec: Close stage judges acceptance_criteria; reviewer role still reviews code.)
  • Replan budget default (3?) and the escalation target when exhausted.
  • Clarification timeout: auto-proceed-with-assumptions after N hours, or wait indefinitely? (Spec: wait; revisit with a timeout if goals pile up.)