Skip to content

Commit 76d8b81

Browse files
committed
fix(compaction): align ttl nits with hysteresis-gap behavior
1 parent 4487f8f commit 76d8b81

4 files changed

Lines changed: 84 additions & 30 deletions

File tree

‎docs/ARCHITECTURE.md‎

Lines changed: 2 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -187,11 +187,11 @@ The agent maintains an optional **`manage_tasks`** list (create/update via the h
187187

188188
#### Context compaction (the compaction governor)
189189

190-
When a cycle's input tokens cross a threshold, the director compacts the inference-facing history (the full run is always retained in the context store). The threshold is **model-aware** — roughly 60% of the active model's real context window — so small-window models compact early enough to avoid provider context-overflow while large-window models do not compact prematurely. The compacted prefix is **append-only across passes**: the existing compacted user turn stays byte-identical; new folds become later summary turns with a harness-inserted assistant spacer (identified by reserved `model: "harness"`, plus a visible sentinel; persisted `[compaction]` tokens without a producer still freeze) between them so the prompt head can remain in the provider KV cache. Model-emitted copies of the spacer are not frozen. Spacer-only model replies are incomplete: ChatDirector nudges, then falls through loop-protection, workflow-idle, and open-task rails rather than empty-settling with work still open. The governor covers three cases:
190+
When a cycle's input tokens cross a threshold, the director compacts the inference-facing history (the full run is always retained in the context store). The threshold is **model-aware** — roughly 60% of the active model's real context window — so small-window models compact early enough to avoid provider context-overflow while large-window models do not compact prematurely. The compacted prefix is **append-only across passes**: the existing compacted user turn stays byte-identical; new folds become later summary turns with a harness-inserted assistant spacer (identified by reserved `model: "harness"`, plus a visible sentinel; persisted `[compaction]` tokens without a producer still freeze) between them so the prompt head can remain in the provider KV cache. Model-emitted copies of the spacer are not frozen. Spacer-only model replies are incomplete: ChatDirector nudges, then falls through loop-protection, workflow-idle, and open-task rails rather than empty-settling with work still open. The governor covers four cases:
191191

192192
- **Threshold at a tool pause** — Once over threshold, the follow-up `infer` after a tool batch is swapped for a `compact` cycle, and inference resumes via a host continuation message. After a compact that remains over the high watermark, the governor uses **growth hysteresis** (wait for usage to grow by ~10% of the window) instead of re-arming on every cycle; dropping under 60% is not required.
193193
- **Idle (end-of-turn)** — An interactive turn can end with a reply and then sit idle with no tool batch to intercept; the governor requests a continuation at that pause and compacts when it arrives. An operator message that races the continuation still compacts first, then re-enters inference to answer it.
194-
- **Idle recompress past the provider TTL** — The fold is a re-compress, not a cache play: provider KV caches expire on their own schedule, so an under-threshold session compresses _after_ that expiry (the next turn is a cheaper write and later reads compound on the shrunk context) instead of only at 60% tokens. Any live re-entry past the window fires it — an empty stall ping re-enters meter-only, a raced operator message re-infers — under reason `cache-ttl-recompress` through the same compactor, so the fresh tail stays raw exactly as in the threshold path. The interval follows provider cache economics (`src/provider/cache-ttl.ts`), not a global N minutes:
194+
- **Idle recompress past the provider TTL** — The fold is a re-compress, not a cache play: provider KV caches expire on their own schedule, so a session compresses _after_ that expiry (the next turn is a cheaper write and later reads compound on the shrunk context) instead of only at 60% tokens. It is not limited to under-threshold sessions: once growth hysteresis has cleared `pending`, an over-threshold session in that gap can fire the same fold. Any live re-entry past the window fires it — an empty stall ping re-enters meter-only, a raced operator message re-infers — under reason `cache-ttl-recompress` through the same compactor, so the fresh tail stays raw exactly as in the threshold path. The interval follows provider cache economics (`src/provider/cache-ttl.ts`), not a global N minutes.
195195
- **Overflow recovery** — A `context_overflow` inference error would otherwise become a terminal error reply; the governor compacts and retries instead, bounded so a history the compactor cannot shrink does not loop forever. Overflow ignores hysteresis for the compact itself.
196196

197197
| Provider segment | Idle recompress allowed after | Why |

‎src/agent/compaction.test.ts‎

Lines changed: 44 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -950,6 +950,50 @@ describe("provider-aware idle recompress (CL-8745)", () => {
950950
).toEqual(ttlCompact);
951951
});
952952

953+
test("after threshold compact and gap TTL, growth-armed compact stays blocked until a tool_call", () => {
954+
// Threshold compact (consecutive=1) plus TTL fire in the hysteresis gap
955+
// (consecutive=2) fills the shared cap. Later growth that would re-arm
956+
// the threshold path stays blocked until a tool_call occupancy resets it.
957+
let nowMs = 80_000_000;
958+
const governor = createCompactionGovernor(
959+
() => undefined,
960+
"",
961+
[],
962+
() => nowMs,
963+
);
964+
governor.noteInferenceDone(inferenceDone(overThreshold), tenTurns);
965+
expect(
966+
governor.interceptActions(toolDone(), inferAction, capabilities),
967+
).not.toBeNull();
968+
expect(governor.resumeAfterCompact(emptyMessage())).toBe("infer");
969+
970+
governor.noteInferenceDone(inferenceDone(overThreshold), tenTurns);
971+
nowMs += 11 * MINUTE_MS;
972+
expect(
973+
governor.interceptIdleContinuation(emptyMessage(), capabilities),
974+
).toEqual(ttlCompact);
975+
expect(governor.resumeAfterCompact(emptyMessage())).toBe("meter");
976+
977+
// Post-TTL snapshot, then growth past resumeDelta — pending re-arms,
978+
// but the cap is full so interceptActions stays null.
979+
governor.noteInferenceDone(inferenceDone(overThreshold), tenTurns);
980+
governor.noteInferenceDone(
981+
inferenceDone(overThreshold + resumeDelta),
982+
tenTurns,
983+
);
984+
expect(
985+
governor.interceptActions(toolDone(), inferAction, capabilities),
986+
).toBeNull();
987+
988+
governor.noteInferenceDone(
989+
inferenceDoneWithTools(overThreshold + 2 * resumeDelta),
990+
tenTurns,
991+
);
992+
expect(
993+
governor.interceptActions(toolDone(), inferAction, capabilities),
994+
).not.toBeNull();
995+
});
996+
953997
test("does not fold under an outstanding tool batch, fires once it settles", () => {
954998
let continuations = 0;
955999
let nowMs = 40_000_000;

‎src/agent/compaction.ts‎

Lines changed: 34 additions & 27 deletions
Original file line numberDiff line numberDiff line change
@@ -317,6 +317,29 @@ export function createCompactionGovernor(
317317
return true;
318318
}
319319

320+
function inboundText(event: ReactorInboundEvent): string {
321+
if (event.type !== "message.received") return "";
322+
return typeof event.message.content === "string"
323+
? event.message.content
324+
: "";
325+
}
326+
327+
// Idle empty compact needs a meter-only re-entry; a raced operator message
328+
// needs a follow-up infer. Shared by the threshold idle path and TTL fold.
329+
function issueIdleFold(
330+
content: string,
331+
capabilities: ReactorCapabilities,
332+
reason: string,
333+
): ReactorAction[] {
334+
if (content.length > 0) postCompactInfer = true;
335+
else postCompactMeter = true;
336+
issueThresholdCompact();
337+
return [
338+
capabilities.compact(COMPACTOR_NAME, reason),
339+
...continuationActions(capabilities),
340+
];
341+
}
342+
320343
function interceptIdleContinuation(
321344
event: ReactorInboundEvent,
322345
capabilities: ReactorCapabilities,
@@ -329,40 +352,26 @@ export function createCompactionGovernor(
329352
}
330353
idlePending = false;
331354
pending = false;
332-
const content =
333-
typeof event.message.content === "string" ? event.message.content : "";
334355
// The reactor delivers no event after compact, so always request a
335356
// continuation to re-enter decide against the shrunk turns:
336357
// - raced operator content → re-infer to answer it
337358
// - empty synthetic continuation → meter-only sync (no infer)
338-
if (content.length > 0) {
339-
postCompactInfer = true;
340-
} else {
341-
postCompactMeter = true;
342-
}
343-
issueThresholdCompact();
344-
return [
345-
capabilities.compact(COMPACTOR_NAME, "context-threshold"),
346-
...continuationActions(capabilities),
347-
];
359+
return issueIdleFold(
360+
inboundText(event),
361+
capabilities,
362+
"context-threshold",
363+
);
348364
}
349365
// Unarmed idle re-entry past the provider TTL: same fold, same
350366
// keep-recent tail, same cap — but a "cache-ttl-recompress" reason so the
351367
// fold is attributable. No arming: every live re-entry re-checks the
352368
// window, so a sub-agent stall ping or operator message is the trigger.
353369
if (!isTtlRecompressDue(now())) return null;
354-
const content =
355-
typeof event.message.content === "string" ? event.message.content : "";
356-
if (content.length > 0) {
357-
postCompactInfer = true;
358-
} else {
359-
postCompactMeter = true;
360-
}
361-
issueThresholdCompact();
362-
return [
363-
capabilities.compact(COMPACTOR_NAME, "cache-ttl-recompress"),
364-
...continuationActions(capabilities),
365-
];
370+
return issueIdleFold(
371+
inboundText(event),
372+
capabilities,
373+
"cache-ttl-recompress",
374+
);
366375
}
367376

368377
// A context-overflow inference error would otherwise terminate the loop
@@ -396,9 +405,7 @@ export function createCompactionGovernor(
396405
event: ReactorInboundEvent,
397406
): "infer" | "meter" | null {
398407
if (event.type !== "message.received") return null;
399-
const content =
400-
typeof event.message.content === "string" ? event.message.content : "";
401-
if (content.length > 0) return null;
408+
if (inboundText(event).length > 0) return null;
402409
if (postCompactInfer) {
403410
postCompactInfer = false;
404411
return "infer";

‎src/provider/cache-ttl.test.ts‎

Lines changed: 4 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -21,7 +21,7 @@ describe("cacheTtlMsFor", () => {
2121
);
2222
});
2323

24-
test("disables idle recompress for local inference and unknown models", () => {
24+
test("disables idle recompress for local inference and missing model ids", () => {
2525
expect(cacheTtlMsFor("ollama/llama3.1")).toBeUndefined();
2626
expect(cacheTtlMsFor(undefined)).toBeUndefined();
2727
expect(cacheTtlMsFor("")).toBeUndefined();
@@ -41,5 +41,8 @@ describe("cacheTtlMsFor", () => {
4141
test("assumes OpenAI-style economics for unrecognized providers", () => {
4242
expect(cacheTtlMsFor("bifrost/some-model")).toBe(10 * MINUTE_MS);
4343
expect(cacheTtlMsFor("totally-new-provider/model-x")).toBe(10 * MINUTE_MS);
44+
// A truly unknown id (no provider segment, no family substring) still
45+
// gets the 10-minute default — not the local-inference disable.
46+
expect(cacheTtlMsFor("unknown-id")).toBe(10 * MINUTE_MS);
4447
});
4548
});

0 commit comments

Comments
 (0)