Skip to content

Commit 22d53b4

Browse files
Nudge once when assistants print tool-call markup as text (#885)
* Nudge once when assistants print tool-call markup as text Models sometimes emit <tool_call><function=...> wrappers as assistant text instead of real tool_call blocks. Catch that narrow shape, give one corrective nudge per no-real-tool epoch without counting it as incomplete-report narration, then fall through to the existing report policy. Thinking blocks and arbitrary XML stay out of scope; the epoch resets only on genuine tool activity or a parent follow-up. * Document verbatim tool markup recovery in fleet stop policy Document recovery of verbatim tool markup in fleet stop policy. * Let a complete report envelope win over quoted tool-call markup
1 parent cea3b8b commit 22d53b4

4 files changed

Lines changed: 200 additions & 3 deletions

File tree

‎docs/ARCHITECTURE.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -116,7 +116,7 @@ Two directors, selected by role:
116116

117117
Auto mode is toggled by CLI flags (`--auto` / `--no-auto`); there is currently no in-session key to toggle it (default on; constrained envelope — workspace writes and unconstrained shell auto-allow; installs, recursive rm, force/uncontained worktree changes, sensitive-path and opaque-wrapper shell still ask; contained non-force `git worktree add`/`remove`/`prune` and `list` auto-allow; shell file-mutation denied). It is not a separate edit/plan mode.
118118

119-
- **SubAgentDirector** (delegated work, `src/subagent/index.ts`) — Drives a dispatched worker until a turn arrives with no tool calls, then replies with the final assistant text and ends the run. A tool-less turn **after tools** completes only with the four-heading envelope (Summary, Findings, Blockers, Paths); a missing envelope nudges once (**incomplete-report**) and a second tool-less turn still without the envelope salvages as **incomplete-report-stop**. Explore/read-only workers that used tools then replied with findings remain normal completes; `requireEvidence` (off by default, set per director) additionally requires at least one read before a tool-less spawn-only reply can complete. Reads done through `run_shell` count as evidence too — `src/subagent/shell-evidence.ts` classifies shell reads (`cat`, `grep`, `sed` without `-i`, …) over the same subject expansion the auto-shell policy uses — but there is no corresponding shell-write evidence or file-write requirement: a run that never touches a file still completes normally once it replies with the envelope. There is no turn budget. Operator/parent cancel after any progress returns a **cancelled** salvage report (partial findings + tool activity) instead of a bare cancel string; cancel before progress still surfaces as cancelled-by-operator. There is no repetition/no-progress/never-acted/never-edited hard stop and no fingerprint-based re-dispatch block — a genuinely stuck worker runs until it completes, stalls, hits an opt-in wall-clock deadline, or is cancelled.
119+
- **SubAgentDirector** (delegated work, `src/subagent/index.ts`) — Drives a dispatched worker until a turn arrives with no tool calls, then replies with the final assistant text and ends the run. A tool-less turn **after tools** completes only with the four-heading envelope (Summary, Findings, Blockers, Paths). Assistant text that prints explicit `<tool_call>` markup is treated as attempted tool use, not narration: one **verbatim-tool-call** nudge asks the worker to re-issue a real `tool_call` and does not count toward the tool-less spiral. A missing envelope otherwise nudges once (**incomplete-report**) and a second tool-less turn still without the envelope salvages as **incomplete-report-stop**. Explore/read-only workers that used tools then replied with findings remain normal completes; `requireEvidence` (off by default, set per director) additionally requires at least one read before a tool-less spawn-only reply can complete. Reads done through `run_shell` count as evidence too — `src/subagent/shell-evidence.ts` classifies shell reads (`cat`, `grep`, `sed` without `-i`, …) over the same subject expansion the auto-shell policy uses — but there is no corresponding shell-write evidence or file-write requirement: a run that never touches a file still completes normally once it replies with the envelope. There is no turn budget. Operator/parent cancel after any progress returns a **cancelled** salvage report (partial findings + tool activity) instead of a bare cancel string; cancel before progress still surfaces as cancelled-by-operator. There is no repetition/no-progress/never-acted/never-edited hard stop and no fingerprint-based re-dispatch block — a genuinely stuck worker runs until it completes, stalls, hits an opt-in wall-clock deadline, or is cancelled.
120120
`spawn_agent` starts each worker and records it in the caller's fleet mailbox. On the TUI primary, mailbox mail is the collect path: occupancy takes uncollected terminals and re-enters the parent as system inbound. Nested orchestrators still collect with `wait_agents`. TUI-primary `wait_agents` may yield as a timeout (workers untouched, no take) so occupancy can deliver mail or a queued Enter steer. Already-collected waits return status without a second report or error body. Wait JSON includes `stop_reason` from the session when present so a salvage that is wait-`done` is not mistaken for a clean complete, and so parent-initiated interrupt (`interrupted`) is not mistaken for operator-cancel (`cancelled`). Deadline salvage prepends an advisory parent hint suggesting continuation plus a longer deadline if more wall-clock time is warranted. Failed and incomplete-report salvage tell the parent to diagnose from the report or error and MAY spawn one successor with a changed brief. A parent-initiated interrupt is a resumable pause: wait unblocks with `stop_reason: interrupted` (often while the session is still running and has no report); the parent should `resume_agent` or re-wait, and must not spawn a successor against a still-live worker. Successor only if that session is no longer resumable. Operator-cancelled salvage asks the parent to synthesize Findings and Paths and wait for the operator instead of auto-starting another specialist. Identical re-dispatch of the same brief stays refused at the prompt / spawn-handoff layer; there is no fingerprint-based re-dispatch hard-block. Deadline hints are advisory only — an identical re-dispatch is still admitted at runtime. Parent hints are prepended on salvage reports returned to the parent. The runtime does not auto-spawn successors.
121121

122122
#### Model-family policy (`src/agent/model-family-policy.ts`)

‎docs/PRODUCT.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -170,7 +170,7 @@ Corbits Code fans work out to short-lived **fleet agents** — workers with thei
170170
- **Tasks** are checklist items owned by one agent via `manage_tasks`.
171171
- **Fleet agents** are spawned with `spawn_agent`. On the TUI primary, mailbox mail arrives as inbound when a worker finishes or fails — do not poll `wait_agents`. Nested orchestrators still collect with `wait_agents`. Workers ask the parent with `ask_director`. That parks a question while the worker stays `running`. Nested `wait_agents` returns `awaiting_director` with a question payload — that is not terminal. The parent answers with `send_input` (`target` = the worker's session id). When the parent TUI is not blocked in `wait_agents`, a parked question arrives as a synthetic idle-send wake. Escalate to the human only with `ask_operator`.
172172

173-
Dispatch uses a structured brief (context / goal / optional goals seed) and returns a structured report. The TUI Agents strip and fleet board show who is running; live tool progress updates the status bar without dumping the child transcript into the parent chat. There is no turn budget. A tool-less final turn completes only with the four-heading report envelope; without it, one nudge is given and a second tool-less turn without the envelope salvages as `incomplete-report-stop`. A silent worker (no activity for `stallTimeoutMs`, opt-in) gets one continuation nudge, then salvages as `stalled` if a second consecutive check finds no activity. An opt-in `deadlineMs`, or an operator cancel, can also end a run early. Each of these returns a salvage report so a runaway or idle child cannot quietly burn a large token budget or look done after prose alone.
173+
Dispatch uses a structured brief (context / goal / optional goals seed) and returns a structured report. The TUI Agents strip and fleet board show who is running; live tool progress updates the status bar without dumping the child transcript into the parent chat. There is no turn budget. A tool-less final turn completes only with the four-heading report envelope. Printed `<tool_call>` markup in assistant text gets one corrective nudge to issue a real tool call and does not count as the wrap-up; without the envelope, one incomplete-report nudge is given and a second tool-less turn without the envelope salvages as `incomplete-report-stop`. A silent worker (no activity for `stallTimeoutMs`, opt-in) gets one continuation nudge, then salvages as `stalled` if a second consecutive check finds no activity. An opt-in `deadlineMs`, or an operator cancel, can also end a run early. Each of these returns a salvage report so a runaway or idle child cannot quietly burn a large token budget or look done after prose alone.
174174

175175
## Roadmap (planned, not yet shipped)
176176

‎src/subagent/nudge-director.test.ts‎

Lines changed: 157 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -84,13 +84,20 @@ function inferenceDone(
8484
}
8585

8686
function inferenceDoneText(text: string, inputTokens = 0): ReactorInboundEvent {
87+
return inferenceDoneContent([{ type: "text", text }], inputTokens);
88+
}
89+
90+
function inferenceDoneContent(
91+
content: readonly Record<string, unknown>[],
92+
inputTokens = 0,
93+
): ReactorInboundEvent {
8794
return {
8895
type: "inference.done",
8996
turn: {
9097
role: "assistant",
9198
model: "test",
9299
timestamp: 0,
93-
content: [{ type: "text", text }],
100+
content,
94101
},
95102
usage: {
96103
input: inputTokens,
@@ -419,6 +426,155 @@ const REPORT_ENVELOPE = [
419426
"src/gate.ts",
420427
].join("\n");
421428

429+
describe("SubAgentDirector verbatim tool markup recovery", () => {
430+
const verbatimToolCall =
431+
'<tool_call><function=read_file>{"path":"src/index.ts"}</function></tool_call>';
432+
433+
test("nudges once for explicit tool-call wrapper text before report policy", async () => {
434+
const director = new SubAgentDirector("system", [], undefined, 30);
435+
const caps = capabilities();
436+
437+
const correction = actions(
438+
await director.decide(inferenceDoneText(verbatimToolCall), state, caps),
439+
);
440+
expect(correction).toContainEqual({
441+
type: "checkpoint",
442+
message: "subagent-verbatim-tool-call-nudge",
443+
});
444+
expect(ephemeralTexts(inferAction(correction))?.[0]).toContain(
445+
"real tool call",
446+
);
447+
448+
const reportNudge = actions(
449+
await director.decide(inferenceDoneText(verbatimToolCall), state, caps),
450+
);
451+
expect(reportNudge).toContainEqual({
452+
type: "checkpoint",
453+
message: "subagent-incomplete-report-nudge",
454+
});
455+
456+
const stopped = actions(
457+
await director.decide(inferenceDoneText(verbatimToolCall), state, caps),
458+
);
459+
expect(stopped).toContainEqual({
460+
type: "checkpoint",
461+
message: "subagent-incomplete-report",
462+
});
463+
});
464+
465+
test("does not treat arbitrary XML or thinking as verbatim tool calls", async () => {
466+
const caps = capabilities();
467+
const arbitraryXML = new SubAgentDirector("system", [], undefined, 30);
468+
const arbitraryResult = actions(
469+
await arbitraryXML.decide(
470+
inferenceDoneText("<read_file>src/index.ts</read_file>"),
471+
state,
472+
caps,
473+
),
474+
);
475+
expect(arbitraryResult).toContainEqual({
476+
type: "checkpoint",
477+
message: "subagent-incomplete-report-nudge",
478+
});
479+
480+
const thinkingOnly = new SubAgentDirector("system", [], undefined, 30);
481+
const thinkingResult = actions(
482+
await thinkingOnly.decide(
483+
inferenceDoneContent([
484+
{ type: "thinking", thinking: verbatimToolCall },
485+
]),
486+
state,
487+
caps,
488+
),
489+
);
490+
expect(thinkingResult).toContainEqual({
491+
type: "checkpoint",
492+
message: "subagent-incomplete-report-nudge",
493+
});
494+
});
495+
496+
test("resets correction only after genuine tool activity or parent follow-up", async () => {
497+
const director = new SubAgentDirector("system", [], undefined, 30);
498+
const caps = capabilities();
499+
500+
await director.decide(inferenceDoneText(verbatimToolCall), state, caps);
501+
const narration = actions(
502+
await director.decide(inferenceDoneText("Still working"), state, caps),
503+
);
504+
expect(narration).toContainEqual({
505+
type: "checkpoint",
506+
message: "subagent-incomplete-report-nudge",
507+
});
508+
509+
await director.decide(inferenceDone(["read-1"]), state, caps);
510+
await director.decide(toolDone("read-1"), state, caps);
511+
const afterTool = actions(
512+
await director.decide(inferenceDoneText(verbatimToolCall), state, caps),
513+
);
514+
expect(afterTool).toContainEqual({
515+
type: "checkpoint",
516+
message: "subagent-verbatim-tool-call-nudge",
517+
});
518+
519+
await director.decide(messageReceived("Try again"), state, caps);
520+
const afterFollowup = actions(
521+
await director.decide(inferenceDoneText(verbatimToolCall), state, caps),
522+
);
523+
expect(afterFollowup).toContainEqual({
524+
type: "checkpoint",
525+
message: "subagent-verbatim-tool-call-nudge",
526+
});
527+
});
528+
529+
test("after the verbatim nudge a real tool call executes", async () => {
530+
const director = new SubAgentDirector("system", [], undefined, 30);
531+
const caps = capabilities();
532+
533+
await director.decide(inferenceDoneText(verbatimToolCall), state, caps);
534+
const result = actions(
535+
await director.decide(inferenceDone(["read-1"]), state, caps),
536+
);
537+
expect(result.some((action) => action.type === "execute_tools")).toBe(true);
538+
expect(result.some((action) => action.type === "reply")).toBe(false);
539+
});
540+
541+
test("after the verbatim nudge a four-heading envelope completes", async () => {
542+
const director = new SubAgentDirector("system", [], undefined, 30);
543+
const caps = capabilities();
544+
545+
await director.decide(inferenceDoneText(verbatimToolCall), state, caps);
546+
const result = actions(
547+
await director.decide(inferenceDoneText(REPORT_ENVELOPE), state, caps),
548+
);
549+
expect(result).toContainEqual({
550+
type: "checkpoint",
551+
message: "subagent-complete",
552+
});
553+
});
554+
555+
test("a complete envelope that quotes tool-call markup still completes", async () => {
556+
const director = new SubAgentDirector("system", [], undefined, 30);
557+
const caps = capabilities();
558+
559+
const reportQuotingMarkup = `${REPORT_ENVELOPE}\n\nThe model emitted ${verbatimToolCall} as text.`;
560+
const result = actions(
561+
await director.decide(
562+
inferenceDoneText(reportQuotingMarkup),
563+
state,
564+
caps,
565+
),
566+
);
567+
expect(result).toContainEqual({
568+
type: "checkpoint",
569+
message: "subagent-complete",
570+
});
571+
expect(result).not.toContainEqual({
572+
type: "checkpoint",
573+
message: "subagent-verbatim-tool-call-nudge",
574+
});
575+
});
576+
});
577+
422578
describe("SubAgentDirector incomplete-report wiring", () => {
423579
test("tool-less narration after tools gets one wrap-up nudge, not a complete", async () => {
424580
const director = new SubAgentDirector("system", [], undefined, 30);

‎src/subagent/nudge-director.ts‎

Lines changed: 41 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -47,6 +47,20 @@ const TOOL_FAILURE_RECOVERY_NUDGE =
4747
const INCOMPLETE_REPORT_NUDGE =
4848
"Write your final report now using ## Summary, ## Findings, ## Blockers, and ## Paths. Do not narrate status. No more tools unless one lookup is required to cite a line.";
4949

50+
const VERBATIM_TOOL_CALL_NUDGE =
51+
"You wrote tool-call markup as assistant text. Invoke the real tool call instead of printing its markup, or write your final report if no tool is needed.";
52+
53+
function hasVerbatimToolCallMarkup(
54+
content: readonly { type: string; text?: string }[],
55+
): boolean {
56+
return content.some(
57+
(block) =>
58+
block.type === "text" &&
59+
typeof block.text === "string" &&
60+
/<tool_call>\s*<(?:function|tool)=[A-Za-z_][\w.-]*>/.test(block.text),
61+
);
62+
}
63+
5064
function ephemeralNudgeTurn(text: string): ConversationTurn {
5165
return {
5266
role: "user",
@@ -128,6 +142,10 @@ export class SubAgentDirector extends DefaultDirector {
128142
// narration without the envelope salvages as incomplete-report
129143
// (MAX_TOOLLESS_NARRATION_CYCLES = 2).
130144
private toolLessNarrationCycles = 0;
145+
// One corrective nudge per no-real-tool epoch when the assistant prints
146+
// explicit tool-call markup as text instead of issuing a real tool_call.
147+
// Cleared only by genuine tool activity or a non-empty parent follow-up.
148+
private verbatimToolCallNudgeFired = false;
131149

132150
// Once this leaf has replied with a terminal report (complete envelope or
133151
// salvage), empty continuations from idle-compact / stall must not fall
@@ -227,6 +245,7 @@ export class SubAgentDirector extends DefaultDirector {
227245
// A real parent follow-up re-opens the brief; empty continuations do not.
228246
if (isNonEmptyParentMessage(event)) {
229247
this.reportReplied = false;
248+
this.verbatimToolCallNudgeFired = false;
230249
}
231250

232251
const afterCompact = this.compaction.resumeAfterCompact(event);
@@ -288,6 +307,7 @@ export class SubAgentDirector extends DefaultDirector {
288307
this.lastAssistantText = lastText(content);
289308
const hasToolCalls = content.some((block) => block.type === "tool_call");
290309
if (hasToolCalls) {
310+
this.verbatimToolCallNudgeFired = false;
291311
this.thrashState = nextThrashState(this.thrashState, content);
292312
}
293313

@@ -314,6 +334,27 @@ export class SubAgentDirector extends DefaultDirector {
314334
if (compacted !== null) return compacted;
315335
return terminal;
316336
}
337+
338+
// Below the stop policy so a finished report that merely quotes
339+
// tool-call markup still completes; a markup turn with no envelope
340+
// gets the corrective nudge instead of the generic wrap-up one.
341+
if (
342+
!hasToolCalls &&
343+
!this.verbatimToolCallNudgeFired &&
344+
hasVerbatimToolCallMarkup(content)
345+
) {
346+
this.verbatimToolCallNudgeFired = true;
347+
this.interventions({
348+
id: "verbatim-tool-call",
349+
class: "nudge",
350+
state: this.interventionState(),
351+
detail: "assistant emitted explicit tool-call markup as text",
352+
});
353+
return [
354+
capabilities.checkpoint("subagent-verbatim-tool-call-nudge"),
355+
inferWithSubAgentNudge(capabilities, VERBATIM_TOOL_CALL_NUDGE),
356+
];
357+
}
317358
if (stop === "incomplete-report") {
318359
// Tool-less turn after tools, no report envelope. Must not fall through
319360
// to super.decide — DefaultDirector completes any tool-less turn.

0 commit comments

Comments
 (0)