Skip to content

Commit 0c63e12

Browse files
fix(subagent): wait on idle stall pings instead of inferring (#1160)
* fix(subagent): wait on idle stall pings instead of inferring ## Summary An empty stall or compaction continuation with recent activity waits instead of starting a model turn. Compact resume, idle threshold folds, and cache-TTL folds still run before that wait, and the silence clock is not restarted. ## Verification - bun test ./src/subagent/nudge-director.test.ts (48 pass, exit 0) - bun run typecheck Fixes CL-8912 * test(subagent): fold a due cache ttl on an in-window stall ping ## Summary An in-window empty stall ping whose provider cache TTL is due returns a cache-ttl-recompress compact and a continuation. It does not infer or restart the silence clock. The governor reads the director clock so due-ness and the stall window share one timeline. ## Verification - bun test ./src/subagent/nudge-director.test.ts (49 pass, exit 0) Refs CL-8912
1 parent 1ae2496 commit 0c63e12

3 files changed

Lines changed: 220 additions & 7 deletions

File tree

‎docs/ARCHITECTURE.md‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -162,7 +162,7 @@ The ChatDirector counts consecutive assistant turns that contain tool calls and
162162

163163
#### Sub-agent stall management
164164

165-
`SubAgentDirector` tracks `lastActivityAt`, updated on every real `inference.done` and `tool.done`. Directors are pure `decide(event, ...)` functions with no timer of their own and the reactor has no proactive "idle" event, so a genuinely silent worker (e.g. parked on a long-running background command with nothing else to do) produces no event for the director to react to. `runSubAgent` (`src/subagent/index.ts`) arms an external interval, at `subAgentStallTimeoutMs`, that pings the same content-less continuation channel the compaction governor uses to re-enter an idle reactor (`requestContinuation`). The director only acts on a ping if the elapsed time since `lastActivityAt` has crossed the timeout. A ping can still be delivered while `execute_tools` is in flight; outstanding call ids from the last `inference.done` reset the silence clock and wait rather than recording `stall-nudge` with a huge `silenceMs`. The first stall past the timeout records `stallNudgeAt` and issues one continuation nudge (asking the worker to check on the background work or report status). Later empty pings inside `subAgentStallTimeoutMs` of that instant wait without stopping or treating the ping as activity; salvage fires only once a ping arrives after that grace with still no `tool.done` / turn-boundary reset. That grace is what keeps two queued interval ticks from salvaging hundreds of milliseconds after the nudge. Any real activity clears `stallNudgeAt`, so a worker that is genuinely working through a slow single turn is never penalized. After the worker has already replied with a terminal report (complete envelope or salvage), further empty continuations — idle-compact meter sync or stall pings — return `wait` instead of falling through to `DefaultDirector.infer`; only a non-empty parent message (`resume_agent` / `send_input`) re-opens the brief. A parked `ask_director` is the same class of wait: empty stall pings are dropped (not deferred) for the park duration, compact-continue hops skipped during the park flush after unpark, and `SubAgentDirector` returns `wait` on any empty continuation that still arrives while the ask is pending so a long park cannot start a billable infer.
165+
`SubAgentDirector` tracks `lastActivityAt`, updated on every real `inference.done` and `tool.done`. Directors are pure `decide(event, ...)` functions with no timer of their own and the reactor has no proactive "idle" event, so a genuinely silent worker (e.g. parked on a long-running background command with nothing else to do) produces no event for the director to react to. `runSubAgent` (`src/subagent/index.ts`) arms an external interval, at `subAgentStallTimeoutMs`, that pings the same content-less continuation channel the compaction governor uses to re-enter an idle reactor (`requestContinuation`). The director nudges only once elapsed time since `lastActivityAt` has crossed the timeout. A ping inside that window, or when stall timing is unset, returns `wait` and does not start a model turn or stamp `lastActivityAt`. A ping can still be delivered while `execute_tools` is in flight; outstanding call ids from the last `inference.done` reset the silence clock and wait rather than recording `stall-nudge` with a huge `silenceMs`. The first stall past the timeout records `stallNudgeAt` and issues one continuation nudge (asking the worker to check on the background work or report status). Later empty pings inside `subAgentStallTimeoutMs` of that instant wait without stopping or treating the ping as activity; salvage fires only once a ping arrives after that grace with still no `tool.done` / turn-boundary reset. That grace is what keeps two queued interval ticks from salvaging hundreds of milliseconds after the nudge. Any real activity clears `stallNudgeAt`, so a worker that is genuinely working through a slow single turn is never penalized. After the worker has already replied with a terminal report (complete envelope or salvage), further empty continuations — idle-compact meter sync or stall pings — return `wait` instead of falling through to `DefaultDirector.infer`; only a non-empty parent message (`resume_agent` / `send_input`) re-opens the brief. A parked `ask_director` is the same class of wait: empty stall pings are dropped (not deferred) for the park duration, compact-continue hops skipped during the park flush after unpark, and `SubAgentDirector` returns `wait` on any empty continuation that still arrives while the ask is pending so a long park cannot start a billable infer.
166166

167167
**Intervention log**: every stop and nudge is appended as one JSONL record to `interventions.jsonl` in the firing worker's trace dir (`src/subagent/intervention-log.ts`), carrying the trigger's measured value beside the threshold it crossed, the provider/model/family it fired on, and the run state at that moment (turns used vs budget, tool calls, read/edit counts). A refused parent re-dispatch is recorded on the parent side, where no worker run exists to record it. The parent also appends one `outcome` record per completed dispatch — the salvage kind `classifyBriefSalvage` assigned, or a clean-complete marker, plus the dispatch count — so the log carries dispatch outcomes as well as interventions, and a stop record can later be read alongside what the dispatch it touched actually produced. Writes are fire-and-forget and swallow their own errors — a diagnostic must not be able to fail a run. `scripts/intervention-forensics.ts` aggregates these across local sessions: per-intervention counts by model family, the measured-value distribution against the threshold, two context columns (stops that fired on runs which had already edited files; stops that fired before half the turn budget was spent — neither is a measured false-positive rate, since either is equally consistent with a correct stop or a wrong one), and outcome counts by kind. This exists because every threshold in this tree was set by judgment and four of those judgments were later reverted — a threshold change is expected to cite this data (CL-6938).
168168

‎src/subagent/nudge-director.test.ts‎

Lines changed: 211 additions & 4 deletions
Original file line numberDiff line numberDiff line change
@@ -335,10 +335,11 @@ describe("SubAgentDirector tool failure recovery", () => {
335335
expect(resumedTexts).toHaveLength(1);
336336
expect(resumedTexts?.[0]).toContain("A tool call failed");
337337

338-
const later = inferAction(
338+
const later = actions(
339339
await director.decide(messageReceived(""), longState, caps),
340340
);
341-
expect(ephemeralTexts(later)).toBeUndefined();
341+
expect(later).toEqual([{ type: "wait" }]);
342+
expect(later.some((action) => action.type === "infer")).toBe(false);
342343
});
343344

344345
test("recovery nudge appends to ephemeral turns already on the infer", async () => {
@@ -462,10 +463,11 @@ describe("SubAgentDirector tool failure recovery", () => {
462463
expect(resumedTexts).toHaveLength(1);
463464
expect(resumedTexts?.[0]).toContain("A tool call failed");
464465

465-
const later = inferAction(
466+
const later = actions(
466467
await director.decide(messageReceived(""), state, caps),
467468
);
468-
expect(ephemeralTexts(later)).toBeUndefined();
469+
expect(later).toEqual([{ type: "wait" }]);
470+
expect(later.some((action) => action.type === "infer")).toBe(false);
469471
});
470472

471473
test("successful nudged infer then later overflow does not resurrect recovery", async () => {
@@ -1530,6 +1532,211 @@ describe("SubAgentDirector ask_director park wait-guard", () => {
15301532
});
15311533
});
15321534

1535+
describe("SubAgentDirector idle stall ping", () => {
1536+
const STALL_NUDGE_TEXT =
1537+
"No activity has been observed for a while. If you are waiting on a " +
1538+
"background command, check its status now; otherwise continue working or " +
1539+
"write your report.";
1540+
1541+
test("empty ping inside the stall window waits and does not infer", async () => {
1542+
let now = 8_000_000;
1543+
const director = new SubAgentDirector(
1544+
"system",
1545+
[],
1546+
undefined,
1547+
1_000,
1548+
() => now,
1549+
);
1550+
const caps = createTestCapabilities();
1551+
1552+
await director.decide(inferenceDoneText("working"), state, caps);
1553+
1554+
now += 200;
1555+
const early = actions(
1556+
await director.decide(messageReceived(""), state, caps),
1557+
);
1558+
expect(early).toEqual([{ type: "wait" }]);
1559+
expect(early.some((action) => action.type === "infer")).toBe(false);
1560+
expect(early.some((action) => action.type === "checkpoint")).toBe(false);
1561+
1562+
// The in-window wait must not restart the silence clock. One stall
1563+
// timeout from the original activity still nudges, once.
1564+
now = 8_000_000 + 1_000;
1565+
const nudge = actions(
1566+
await director.decide(messageReceived(""), state, caps),
1567+
);
1568+
expect(nudge).toContainEqual({
1569+
type: "checkpoint",
1570+
message: "subagent-stall-nudge",
1571+
});
1572+
expect(ephemeralTexts(inferAction(nudge))).toEqual([STALL_NUDGE_TEXT]);
1573+
1574+
now += 200;
1575+
const grace = actions(
1576+
await director.decide(messageReceived(""), state, caps),
1577+
);
1578+
expect(grace).toEqual([{ type: "wait" }]);
1579+
1580+
now += 800;
1581+
const stopped = actions(
1582+
await director.decide(messageReceived(""), state, caps),
1583+
);
1584+
expect(stopped).toContainEqual({
1585+
type: "checkpoint",
1586+
message: "subagent-stalled",
1587+
});
1588+
expect(stopped.some((action) => action.type === "infer")).toBe(false);
1589+
expect(stopped.some((action) => action.type === "reply")).toBe(true);
1590+
});
1591+
1592+
test("outstanding post-compact infer still infers on an in-window empty ping", async () => {
1593+
let now = 9_000_000;
1594+
let continuations = 0;
1595+
const director = new SubAgentDirector(
1596+
"system",
1597+
[],
1598+
() => {
1599+
continuations++;
1600+
},
1601+
60_000,
1602+
() => now,
1603+
);
1604+
const caps = createTestCapabilities();
1605+
1606+
const compact = actions(
1607+
await director.decide(overflowError(), state, caps),
1608+
);
1609+
expect(compact).toEqual([
1610+
{
1611+
type: "compact",
1612+
compactor: "pruning-compactor",
1613+
reason: "context-overflow",
1614+
},
1615+
]);
1616+
expect(continuations).toBe(1);
1617+
1618+
now += 200;
1619+
const resumed = actions(
1620+
await director.decide(messageReceived(""), state, caps),
1621+
);
1622+
expect(resumed.some((action) => action.type === "infer")).toBe(true);
1623+
expect(resumed.some((action) => action.type === "wait")).toBe(false);
1624+
});
1625+
1626+
test("idle-threshold fold still compacts on an in-window empty ping", async () => {
1627+
let now = 10_000_000;
1628+
let continuations = 0;
1629+
const director = new SubAgentDirector(
1630+
"system",
1631+
[],
1632+
() => {
1633+
continuations++;
1634+
},
1635+
60_000,
1636+
() => now,
1637+
);
1638+
const caps = createTestCapabilities();
1639+
1640+
await director.decide(inferenceDone(["read-1"]), longState, caps);
1641+
await director.decide(toolDone("read-1"), longState, caps);
1642+
const complete = actions(
1643+
await director.decide(
1644+
inferenceDoneText(REPORT_ENVELOPE, 999_999),
1645+
longState,
1646+
caps,
1647+
),
1648+
);
1649+
expect(complete.some((action) => action.type === "reply")).toBe(true);
1650+
expect(continuations).toBe(1);
1651+
1652+
now += 200;
1653+
const folded = actions(
1654+
await director.decide(messageReceived(""), longState, caps),
1655+
);
1656+
expect(folded).toEqual([
1657+
{
1658+
type: "compact",
1659+
compactor: "pruning-compactor",
1660+
reason: "context-threshold",
1661+
},
1662+
]);
1663+
expect(continuations).toBe(2);
1664+
expect(folded.some((action) => action.type === "infer")).toBe(false);
1665+
expect(folded.some((action) => action.type === "wait")).toBe(false);
1666+
});
1667+
1668+
test("cache-ttl recompress still folds on an in-window empty ping", async () => {
1669+
// test-model takes the 10-minute default TTL. The stall window is longer
1670+
// so the due fold is still an in-window ping, not a stall nudge.
1671+
const activityAt = 11_000_000;
1672+
const cacheTtlMs = 10 * 60_000;
1673+
const stallTimeoutMs = 15 * 60_000;
1674+
let now = activityAt;
1675+
let continuations = 0;
1676+
const director = new SubAgentDirector(
1677+
"system",
1678+
[],
1679+
() => {
1680+
continuations++;
1681+
},
1682+
stallTimeoutMs,
1683+
() => now,
1684+
);
1685+
const caps = createTestCapabilities();
1686+
1687+
await director.decide(inferenceDoneText("working"), longState, caps);
1688+
1689+
now += cacheTtlMs + 1;
1690+
const folded = actions(
1691+
await director.decide(messageReceived(""), longState, caps),
1692+
);
1693+
expect(folded).toEqual([
1694+
{
1695+
type: "compact",
1696+
compactor: "pruning-compactor",
1697+
reason: "cache-ttl-recompress",
1698+
},
1699+
]);
1700+
expect(continuations).toBe(1);
1701+
expect(folded.some((action) => action.type === "infer")).toBe(false);
1702+
expect(folded.some((action) => action.type === "wait")).toBe(false);
1703+
1704+
// Meter-only resume of the empty fold. Same clock: still inside the
1705+
// stall window, and this wait must not count as activity either.
1706+
const resumed = actions(
1707+
await director.decide(messageReceived(""), longState, caps),
1708+
);
1709+
expect(resumed).toEqual([{ type: "wait" }]);
1710+
expect(resumed.some((action) => action.type === "infer")).toBe(false);
1711+
1712+
// The fold must not stamp lastActivityAt or clear stallNudgeAt. One
1713+
// stall timeout from the original activity still nudges, once.
1714+
now = activityAt + stallTimeoutMs;
1715+
const nudge = actions(
1716+
await director.decide(messageReceived(""), longState, caps),
1717+
);
1718+
expect(nudge).toContainEqual({
1719+
type: "checkpoint",
1720+
message: "subagent-stall-nudge",
1721+
});
1722+
expect(ephemeralTexts(inferAction(nudge))).toEqual([STALL_NUDGE_TEXT]);
1723+
expect(nudge.some((action) => action.type === "wait")).toBe(false);
1724+
});
1725+
1726+
test("no stall timeout waits on an unsolicited empty continuation", async () => {
1727+
const director = new SubAgentDirector("system", [], undefined);
1728+
const caps = createTestCapabilities();
1729+
1730+
await director.decide(inferenceDoneText("working"), state, caps);
1731+
const ping = actions(
1732+
await director.decide(messageReceived(""), state, caps),
1733+
);
1734+
expect(ping).toEqual([{ type: "wait" }]);
1735+
expect(ping.some((action) => action.type === "infer")).toBe(false);
1736+
expect(ping.some((action) => action.type === "checkpoint")).toBe(false);
1737+
});
1738+
});
1739+
15331740
function stubAdmission(
15341741
notes: { provider: string; until: number }[],
15351742
): AdmissionQueue {

‎src/subagent/nudge-director.ts‎

Lines changed: 8 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -265,6 +265,7 @@ export class SubAgentDirector extends DefaultDirector {
265265
requestContinuation,
266266
composedPrompt,
267267
toolDefinitions,
268+
now,
268269
);
269270
this.stallTimeoutMs = stallTimeoutMs;
270271
this.now = now;
@@ -343,6 +344,10 @@ export class SubAgentDirector extends DefaultDirector {
343344

344345
const stallOutcome = this.checkStallPing(event, capabilities);
345346
if (stallOutcome !== null) return stallOutcome;
347+
// Inside the stall window, or with no stall timeout, an empty
348+
// continuation is not a model turn. Leave lastActivityAt and
349+
// stallNudgeAt alone so the next real silence can still nudge.
350+
if (isEmptyContinuation(event)) return capabilities.wait();
346351

347352
// Keep the running local estimate current on every cycle (tool results and
348353
// rewrites included). Arming still happens inside noteInferenceDone, which
@@ -522,8 +527,9 @@ export class SubAgentDirector extends DefaultDirector {
522527
* stallNudgeAt. Queued pings that arrive inside the stallTimeoutMs grace
523528
* after that nudge neither stop nor restart the grace (and do not count as
524529
* activity). Stop only when a ping arrives after the grace with still no
525-
* activity. Returns null when this event is not a stall check the director
526-
* should act on (let it fall through as an ordinary continuation).
530+
* activity. Returns null when this ping is not yet silence, or when stall
531+
* timing is unconfigured. decide then waits on an empty continuation
532+
* without stamping the silence clock, instead of inferring.
527533
*/
528534
private checkStallPing(
529535
event: ReactorInboundEvent,

0 commit comments

Comments
 (0)