Skip to content

Commit f5862ed

Browse files
committed
Retry a failed fleet worker with a changed brief
After 0.3.15 the parent idled on fail and incomplete-report as if the operator had cancelled. Fail-path salvage now invites one successor with a changed brief; operator-cancel still waits; identical briefs stay refused at the prompt layer.
1 parent 6ea5969 commit f5862ed

15 files changed

Lines changed: 271 additions & 32 deletions

‎CHANGELOG.md‎

Lines changed: 9 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -11,6 +11,15 @@ matching `## [X.Y.Z]` section (plus install instructions). Do not maintain
1111
parallel copies under `docs/` or `scripts/notes/`. At cut time: rename
1212
`## [Unreleased]` to `## [X.Y.Z] - YYYY-MM-DD`, then run the release script.
1313

14+
## [Unreleased]
15+
16+
### Changed
17+
18+
- Skywalker may spawn one successor with a changed brief after a failed,
19+
incomplete-report, or interrupted-incomplete fleet worker. Operator-cancelled
20+
salvage still waits for the operator. Identical briefs stay refused. This
21+
reverses the 0.3.15 fail-path idle, not operator-cancel.
22+
1423
## [0.3.18] - 2026-09-08
1524

1625
### Added

‎docs/ARCHITECTURE.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -114,7 +114,7 @@ Two directors, selected by role:
114114

115115
- **ChatDirector** (interactive, `src/agent/director.ts`) — Extends `DefaultDirector` with task list tracking, workflow nudges, LSP auto-activation, and multi-turn chat semantics. It never terminates the session: operator declines are surfaced as replies and the reactor stays alive for the next message. Auto mode is toggled by CLI flags (`--auto` / `--no-auto`); there is currently no in-session key to toggle it (default on; constrained envelope — workspace writes and unconstrained shell auto-allow; installs, recursive rm, force/uncontained worktree changes, sensitive-path and opaque-wrapper shell still ask; contained non-force `git worktree add`/`remove`/`prune` and `list` auto-allow; shell file-mutation denied). It is not a separate edit/plan mode.
116116
- **SubAgentDirector** (delegated work, `src/subagent/index.ts`) — Drives a dispatched worker until a turn arrives with no tool calls, then replies with the final assistant text and ends the run. A tool-less turn **after tools** completes only with the four-heading envelope (Summary, Findings, Blockers, Paths); a missing envelope nudges once (**incomplete-report**) and a second tool-less turn still without the envelope salvages as **incomplete-report-stop**. Explore/read-only workers that used tools then replied with findings remain normal completes; `requireEvidence` (off by default, set per director) additionally requires at least one read before a tool-less spawn-only reply can complete. Reads done through `run_shell` count as evidence too — `src/subagent/shell-evidence.ts` classifies shell reads (`cat`, `grep`, `sed` without `-i`, …) over the same subject expansion the auto-shell policy uses — but there is no corresponding shell-write evidence or file-write requirement: a run that never touches a file still completes normally once it replies with the envelope. There is no turn budget. Operator/parent cancel after any progress returns a **cancelled** salvage report (partial findings + tool activity) instead of a bare cancel string; cancel before progress still surfaces as cancelled-by-operator. There is no repetition/no-progress/never-acted/never-edited hard stop and no fingerprint-based re-dispatch block — a genuinely stuck worker runs until it completes, stalls, hits an opt-in wall-clock deadline, or is cancelled.
117-
`spawn_agent` starts each worker and records it in the caller's fleet mailbox; `wait_agents` collects terminal reports from that mailbox. Deadline salvage prepends an advisory parent hint suggesting continuation plus a longer deadline if more wall-clock time is warranted. Cancelled salvage asks the parent to synthesize Findings and Paths and wait for the operator instead of auto-starting another specialist. Deadline hints are advisory only — an identical re-dispatch is still admitted.
117+
`spawn_agent` starts each worker and records it in the caller's fleet mailbox; `wait_agents` collects terminal reports from that mailbox. Wait JSON includes `stop_reason` from the session when present so a salvage that is wait-`done` is not mistaken for a clean complete. Deadline salvage prepends an advisory parent hint suggesting continuation plus a longer deadline if more wall-clock time is warranted. Failed, incomplete-report, and interrupted-incomplete salvage tell the parent to diagnose from the report or error and MAY spawn one successor with a changed brief. Operator-cancelled salvage asks the parent to synthesize Findings and Paths and wait for the operator instead of auto-starting another specialist. Identical re-dispatch of the same brief stays refused at the prompt / spawn-handoff layer; there is no fingerprint-based re-dispatch hard-block. Deadline hints are advisory only — an identical re-dispatch is still admitted at runtime. Parent hints are prepended on salvage reports returned to the parent.
118118

119119
#### Model-family policy (`src/agent/model-family-policy.ts`)
120120

‎src/agent/directors/skywalker/package.test.ts‎

Lines changed: 14 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -146,6 +146,20 @@ describe("skywalkerPackage", () => {
146146
expect(p).not.toContain("if the job still needs doing");
147147
});
148148

149+
test("systemPrompt fail-then-successor is distinct from operator-cancel wait", () => {
150+
const p = skywalkerPackage.systemPrompt;
151+
expect(p).toContain("incomplete-report");
152+
expect(p).toContain("interrupted-incomplete");
153+
expect(p).toContain("MAY `spawn_agent` **one** successor");
154+
expect(p).toContain("changed** brief");
155+
expect(p).toContain("wait for the operator");
156+
expect(p).toContain("Do not auto-retry");
157+
expect(p).toContain("Identical re-dispatch of the same brief stays refused");
158+
expect(p).toContain("Operator-cancel is not a re-dispatch");
159+
expect(p).not.toContain("Then start the next worker");
160+
expect(p).not.toContain("if the job still needs doing");
161+
});
162+
149163
test("systemPrompt simple path skips explorer+critic for tiny work", () => {
150164
const p = skywalkerPackage.systemPrompt;
151165
expect(p).toContain("DIY on the parent");

‎src/agent/directors/skywalker/package.ts‎

Lines changed: 5 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -76,7 +76,9 @@ Do **not** turn a "why is this stalled / why no thinking / spawn looks broken" d
7676
- Classify digs, screenshots of worker rows, and "why/how does X work" as COMMUNICATION first.
7777
- Answer from mounted tools + known architecture; at most **one** explorer worker if a single unknown path blocks the answer.
7878
- Never spawn parallel "parent UI / child UI / stream events / prompt guardrail / session dig" waves for the same question.
79-
- When workers stall, loop, or come back unfinished: synthesize what returned, report Blockers, and change approach — do **not** re-fan-out another diagnostic wave on the same topic.
79+
- When workers stall or loop: synthesize what returned, report Blockers, and change approach — do **not** re-fan-out another diagnostic wave on the same topic.
80+
- Failed wait (\`status: failed\` plus \`error\`), salvage \`incomplete-report\`, or interrupted-incomplete (\`stop_reason: interrupted\`, not operator-cancel): diagnose from the wait report or error; MAY \`spawn_agent\` **one** successor with a **changed** brief (new \`success_criteria\` / \`do_not\` / continuation from Findings). Cap is one successor for that stall. Spawn the successor — do not search the repo as a substitute.
81+
- Operator-cancel (\`stop_reason\` cancelled, or Blockers that say wait for the operator): synthesize Findings and Paths, report Blockers, and **wait for the operator**. Do not auto-retry. Do not spawn a successor because the worker was cancelled.
8082
- Do **not** search the repo yourself after a worker stops without finishing.
8183
- Permission asks and long run_shell clocks on worker rows are not a signal to spawn more diggers.
8284
@@ -85,6 +87,8 @@ Do **not** turn a "why is this stalled / why no thinking / spawn looks broken" d
8587
Child starts blank. Parent writes a complete packet: Goal, contracts copied verbatim, Scope/do_not, Done-when/success_criteria, What to report.
8688
Runtime requires success_criteria for implement/review and their default directors; recommended otherwise.
8789
Re-dispatch after a blocker is a new handoff (new criteria / new do_not), not a retry of the old one-liner.
90+
Identical re-dispatch of the same brief stays refused.
91+
Operator-cancel is not a re-dispatch — wait for the operator.
8892
When the operator brief states a function signature or return shape, put that **verbatim** into implement success_criteria (including sync vs Promise if stated or implied by existing code/tests).
8993
9094
# Verify after ship

‎src/agent/prompts.ts‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -170,7 +170,7 @@ export function buildGuidelines(
170170
"Orchestration:",
171171
"- Break multi-step or parallel work into focused worker dispatches with distinct lenses; prefer `spawn_agent` (fire several in one turn when jobs are independent), then reply with who is running and end the turn — workers keep running while you are idle, and `wait_agents` / `list_agents` on a later turn collect their reports without holding this conversation blocked.",
172172
"- Pass the typed spawn contract: `intent`, `success_criteria` (done-when; required for implement/review and their default directors), `do_not` (scope fence), and `report_focus`. Free-form `prompt` without `success_criteria` fail-closes for implement/review and their default directors.",
173-
"- After workers return, merge their Summary/Findings into a coherent answer for the operator; do not paste raw fleet-agent dumps.",
173+
"- After workers return, classify fail / incomplete-report / interrupted-incomplete vs operator-cancel vs clean complete. Fail-path: diagnose from the report or error and MAY spawn one successor with a changed brief. Operator-cancel: wait for the operator; do not auto-retry. Identical brief: refuse. Merge Summary/Findings into a coherent answer for the operator; do not paste raw fleet-agent dumps.",
174174
"- Use manage_tasks for your own coordination checklist; spawning workers is `spawn_agent` / `wait_agents`, not manage_tasks.",
175175
"- If context is compacted automatically, do not stop tasks early due to token fear; persist progress via manage_tasks and worker reports.",
176176
]),

‎src/prompts.test.ts‎

Lines changed: 8 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -129,11 +129,17 @@ test("orchestrator guidelines teach the typed task spawn contract", () => {
129129
expect(guidelines).not.toContain("weaker");
130130
});
131131

132-
test("primary chat prompt does not invite auto-starting the next worker after unfinished specialists", () => {
132+
test("primary chat prompt classifies fail-path successor vs operator-cancel wait", () => {
133133
const prompt = buildChatSystemPrompt();
134134
const guidelines = buildGuidelines({ sessionMode: "orchestrator" });
135-
expect(guidelines).not.toContain("change the brief rather than repeating it");
135+
expect(guidelines).toContain("MAY spawn one successor with a changed brief");
136+
expect(guidelines).toContain("wait for the operator");
137+
expect(guidelines).toContain("do not auto-retry");
138+
expect(guidelines).toContain("Identical brief: refuse");
136139
expect(guidelines).not.toContain("start the next worker");
140+
expect(prompt).toContain("MAY `spawn_agent` **one** successor");
141+
expect(prompt).toContain("wait for the operator");
142+
expect(prompt).toContain("Identical re-dispatch of the same brief stays refused");
137143
expect(prompt).not.toContain("Then start the next worker");
138144
expect(prompt).not.toContain("if the job still needs doing");
139145
});

‎src/subagent/agent-fleet.test.ts‎

Lines changed: 91 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -391,12 +391,14 @@ describe("spawn_agent + wait_agents", () => {
391391
agent_id: string;
392392
status: string;
393393
report?: string;
394+
stop_reason?: string;
394395
}[];
395396
expect(results).toHaveLength(1);
396397
expect(results[0]!.status).toBe("interrupted");
397398
expect(results[0]!.report).toContain("## Summary");
398399
expect(results[0]!.report).toContain("## Findings");
399400
expect(results[0]!.report).toContain("gate.ts");
401+
expect(results[0]!.stop_reason).toBe("cancelled");
400402
// Strip stays cancelled — salvage is for wait_agents, not a resurrection.
401403
expect(deps.sessions.get(id)?.status).toBe("cancelled");
402404
expect(deps.sessions.get(id)?.lifecycle.state).toBe("cancelled");
@@ -433,6 +435,87 @@ describe("spawn_agent + wait_agents", () => {
433435
expect(results[0]!.error).toBeUndefined();
434436
expect(deps.sessions.get(id)?.status).toBe("cancelled");
435437
});
438+
439+
test("incomplete-report complete is wait done with stop_reason", async () => {
440+
const deps = makeDeps(async () => ({
441+
report: forcedStopReport("incomplete-report", "Still narrating"),
442+
stopReason: "incomplete-report",
443+
}));
444+
const spawn = createSpawnAgentTool(deps);
445+
const wait = createWaitAgentsTool({ sessions: deps.sessions, fleetRecords: deps.fleetRecords });
446+
447+
const spawned = await callTool(spawn, {
448+
description: "incomplete salvage",
449+
prompt: "probe",
450+
intent: "explore",
451+
});
452+
const id = spawned.agent_id as string;
453+
const waited = await callTool(wait, { targets: [id], timeout_ms: 5000 });
454+
const results = waited.results as {
455+
status: string;
456+
report?: string;
457+
error?: string;
458+
stop_reason?: string;
459+
}[];
460+
expect(results[0]!.status).toBe("done");
461+
expect(results[0]!.stop_reason).toBe("incomplete-report");
462+
expect(results[0]!.report).toContain("narrated instead of writing a report envelope");
463+
expect(results[0]!.error).toBeUndefined();
464+
});
465+
466+
test("failed spawn_agent wait_agents returns error not report", async () => {
467+
const deps = makeDeps(async () => {
468+
throw new Error("provider blew up");
469+
});
470+
const spawn = createSpawnAgentTool(deps);
471+
const wait = createWaitAgentsTool({ sessions: deps.sessions, fleetRecords: deps.fleetRecords });
472+
473+
const spawned = await callTool(spawn, {
474+
description: "failed run",
475+
prompt: "probe",
476+
intent: "explore",
477+
});
478+
const id = spawned.agent_id as string;
479+
const waited = await callTool(wait, { targets: [id], timeout_ms: 5000 });
480+
const results = waited.results as {
481+
status: string;
482+
report?: string;
483+
error?: string;
484+
stop_reason?: string;
485+
}[];
486+
expect(results[0]!.status).toBe("failed");
487+
expect(results[0]!.error).toContain("provider blew up");
488+
expect(results[0]!.report).toBeUndefined();
489+
expect(results[0]!.stop_reason).toBeUndefined();
490+
});
491+
492+
test("interrupt salvage wait_agents includes stop_reason interrupted", async () => {
493+
const deps = makeDeps(async () => ({
494+
report: forcedStopReport("interrupted", "partial"),
495+
stopReason: "interrupted",
496+
interrupted: true,
497+
}));
498+
const spawn = createSpawnAgentTool(deps);
499+
const wait = createWaitAgentsTool({ sessions: deps.sessions, fleetRecords: deps.fleetRecords });
500+
501+
const spawned = await callTool(spawn, {
502+
description: "interrupt salvage",
503+
prompt: "probe",
504+
intent: "explore",
505+
});
506+
const id = spawned.agent_id as string;
507+
const waited = await callTool(wait, { targets: [id], timeout_ms: 5000 });
508+
const results = waited.results as {
509+
status: string;
510+
report?: string;
511+
error?: string;
512+
stop_reason?: string;
513+
}[];
514+
expect(results[0]!.status).toBe("interrupted");
515+
expect(results[0]!.stop_reason).toBe("interrupted");
516+
expect(results[0]!.report).toContain("interrupted before finishing");
517+
expect(results[0]!.error).toBeUndefined();
518+
});
436519
});
437520

438521
describe("spawn_agent same-cwd concurrency", () => {
@@ -2079,8 +2162,14 @@ describe("admission queue", () => {
20792162
const elapsed = Date.now() - startedAt;
20802163
expect(elapsed).toBeLessThan(200);
20812164
expect(waited.timed_out).toBe(false);
2082-
const results = waited.results as { agent_id: string; status: string }[];
2083-
expect(results).toEqual([{ agent_id: queuedId, status: "interrupted" }]);
2165+
const results = waited.results as {
2166+
agent_id: string;
2167+
status: string;
2168+
stop_reason?: string;
2169+
}[];
2170+
expect(results).toEqual([
2171+
{ agent_id: queuedId, status: "interrupted", stop_reason: "cancelled" },
2172+
]);
20842173
expect(started).toBe(1);
20852174

20862175
gate.resolve({ report: "ok" });

‎src/subagent/agent-fleet.ts‎

Lines changed: 9 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -102,6 +102,7 @@ interface FleetRecord {
102102
status: WaitJSONStatus;
103103
report?: string;
104104
error?: string;
105+
stopReason?: string;
105106
providerFailure?: true;
106107
/** Set once a wait_agents caller has been handed this result. */
107108
collected?: boolean;
@@ -344,6 +345,7 @@ class FleetMailbox {
344345
...(overlay.hint !== undefined ? { hint: overlay.hint } : {}),
345346
...(payload?.report !== undefined ? { report: payload.report } : {}),
346347
...(payload?.error !== undefined && status === "failed" ? { error: payload.error } : {}),
348+
...(payload?.stopReason !== undefined ? { stopReason: payload.stopReason } : {}),
347349
...(overlay.providerFailure === true ? { providerFailure: true } : {}),
348350
...(ask !== undefined ? { question: ask.question, questionId: ask.questionId } : {}),
349351
...(status === "awaiting_director" && session !== undefined
@@ -1197,13 +1199,17 @@ export function createSpawnAgentTool(deps: AgentFleetDeps): AgentTool {
11971199
const followupLive =
11981200
now?.lifecycle.state === "running" && overlay?.status === "interrupted";
11991201
if (!followupLive) {
1200-
deps.sessions.attachReport(session.id, result.report);
1202+
deps.sessions.attachReport(session.id, result.report, {
1203+
...(result.stopReason !== undefined ? { stopReason: result.stopReason } : {}),
1204+
});
12011205
}
12021206
return;
12031207
}
12041208
const alreadyCancelled = deps.sessions.get(session.id)?.status === "cancelled";
12051209
if (alreadyCancelled) {
1206-
deps.sessions.attachReport(session.id, result.report);
1210+
deps.sessions.attachReport(session.id, result.report, {
1211+
...(result.stopReason !== undefined ? { stopReason: result.stopReason } : {}),
1212+
});
12071213
return;
12081214
}
12091215
const agentRetained = result.agentRetained === true;
@@ -1397,6 +1403,7 @@ export function createWaitAgentsTool(deps: WaitAgentsDeps): AgentTool {
13971403
? { report: taken.report }
13981404
: {}),
13991405
...(taken.error !== undefined ? { error: taken.error } : {}),
1406+
...(taken.stopReason !== undefined ? { stop_reason: taken.stopReason } : {}),
14001407
...(taken.providerFailure === true ? { provider_failure: true } : {}),
14011408
...(taken.hint !== undefined ? { hint: taken.hint } : {}),
14021409
};

0 commit comments

Comments
 (0)