Skip to content

Commit b692de0

Browse files
committed
fix(compaction): limit idle recompress to the Anthropic cache
Guessed windows for other providers fold a cache that is still warm. Only Anthropic publishes a 5-minute expiry this client actually uses.
1 parent 1ae2496 commit b692de0

4 files changed

Lines changed: 99 additions & 158 deletions

File tree

‎docs/ARCHITECTURE.md‎

Lines changed: 4 additions & 8 deletions
Original file line numberDiff line numberDiff line change
@@ -193,14 +193,10 @@ When a cycle's input tokens cross a threshold, the director compacts the inferen
193193
- **Idle (end-of-turn)** — An interactive turn can end with a reply and then sit idle with no tool batch to intercept; the governor requests a continuation at that pause and compacts when it arrives. An operator message that races the continuation still compacts first, then re-enters inference to answer it.
194194
- **Idle recompress past the provider TTL** — The fold is a re-compress, not a cache play: provider KV caches expire on their own schedule, so a session compresses _after_ that expiry (the next turn is a cheaper write and later reads compound on the shrunk context) instead of only at 60% tokens. It is not limited to under-threshold sessions: once growth hysteresis has cleared `pending`, an over-threshold session in that gap can fire the same fold. Any live re-entry past the window fires it — an empty stall ping re-enters meter-only, a raced operator message re-infers — under reason `cache-ttl-recompress` through the same compactor, so the fresh tail stays raw exactly as in the threshold path. The interval follows provider cache economics (`src/provider/cache-ttl.ts`), not a global N minutes:
195195

196-
| Provider segment | Idle recompress allowed after | Why |
197-
| ----------------------------------------------------------------- | ----------------------------- | ------------------------------------------------------------------------------------ |
198-
| `anthropic`, `zen-messages`, `opencode-go-messages` | 5 min | Default ephemeral cache TTL; hour-long breakpoints are never set |
199-
| `openai-responses`, `codex-responses`, `openai-compatible`, `xai` | 10 min | In-memory prefixes typically evicted after 5–10 min idle; xAI assumed same economics |
200-
| `gemini` | 15 min | No published implicit-cache eviction window; conservative |
201-
| `deepseek` | 60 min | On-disk context cache persists for hours-to-days |
202-
| `ollama` | never | Local inference has no remote cache to expire |
203-
| anything else | 10 min | Assumed OpenAI-style in-memory economics; tune per upstream |
196+
| Provider segment | Idle recompress allowed after | Why |
197+
| -------------------------------------------------------------- | ----------------------------- | ---------------------------------------------------------------------------------------------------------------- |
198+
| `anthropic`, `zen-messages`, `opencode-go-messages` | 5 min | Published default ephemeral cache TTL; the 1-hour TTL is never set |
199+
| anything else, including OpenAI, xAI, Gemini, DeepSeek, ollama | never | No published 5-minute expiry, or no remote cache. Guessing a shorter window folds a cache that may still be warm |
204200

205201
The `ollama` table key is the bare provider segment; the harness stamps slash-form on `sourceId` (`ollama/default` / `ollama/<instance>`), with `provider: openai-compatible` (`buildOpenAISource`) and a bare model (`llama3` / `qwen3`). Idle-recompress disable matches via `isOllamaProviderId` on that `sourceId` (`src/provider/ollama.ts`), not by looking up slash-form as a table key.
206202

‎src/agent/compaction.test.ts‎

Lines changed: 38 additions & 30 deletions
Original file line numberDiff line numberDiff line change
@@ -67,6 +67,7 @@ function turnsOfLength(count: number, textLength: number): ConversationTurn[] {
6767
function inferenceDone(
6868
input: number,
6969
text = "",
70+
provider = "p",
7071
): Extract<ReactorInboundEvent, { type: "inference.done" }> {
7172
return {
7273
type: "inference.done",
@@ -75,7 +76,7 @@ function inferenceDone(
7576
content: text.length > 0 ? [{ type: "text", text }] : [],
7677
},
7778
usage: usage(input),
78-
source: { sourceId: "s", provider: "p", model: "m" },
79+
source: { sourceId: "s", provider, model: "m" },
7980
} as unknown as Extract<ReactorInboundEvent, { type: "inference.done" }>;
8081
}
8182

@@ -946,7 +947,11 @@ describe("provider-aware idle recompress (CL-8745)", () => {
946947
): Extract<ReactorInboundEvent, { type: "inference.done" }> {
947948
const source =
948949
typeof modelOrSource === "string"
949-
? { sourceId: "s", provider: "p", model: modelOrSource }
950+
? {
951+
sourceId: "s",
952+
provider: modelOrSource.split("/")[0] ?? "p",
953+
model: modelOrSource,
954+
}
950955
: {
951956
sourceId: modelOrSource.sourceId ?? "s",
952957
provider: modelOrSource.provider ?? "p",
@@ -1021,11 +1026,12 @@ describe("provider-aware idle recompress (CL-8745)", () => {
10211026
expect(governor.resumeAfterCompact(emptyMessage())).toBe("meter");
10221027
});
10231028

1024-
test("production LastCycleSource: anthropic fires at 5m, codex stays quiet until 10m", () => {
1029+
test("production LastCycleSource: anthropic fires at 5m, codex never does", () => {
10251030
// Harness stamps { sourceId, provider, model } with a bare model, not
10261031
// slash-form "anthropic/claude-opus-4-6". Anthropic's 5-minute window
1027-
// must come from provider, not a dummy model string; Codex must not
1028-
// inherit that 5-minute fire from sourceId "codex/work".
1032+
// must come from provider, not a dummy model string. Codex has no
1033+
// published 5-minute expiry, so it must stay quiet past the old 10-minute
1034+
// guess as well.
10291035
let nowMs = 15_000_000;
10301036
const clock = () => nowMs;
10311037
const anthropic = createCompactionGovernor(() => undefined, "", [], clock);
@@ -1057,13 +1063,13 @@ describe("provider-aware idle recompress (CL-8745)", () => {
10571063
codex.interceptIdleContinuation(emptyMessage(), capabilities),
10581064
).toBeNull();
10591065

1060-
nowMs += 5 * MINUTE_MS;
1066+
nowMs += 30 * MINUTE_MS;
10611067
expect(
10621068
codex.interceptIdleContinuation(emptyMessage(), capabilities),
1063-
).toEqual(ttlCompact);
1069+
).toBeNull();
10641070
});
10651071

1066-
test("follows provider economics: deepseek waits out its long window, ollama never fires", () => {
1072+
test("does not idle-recompress DeepSeek or ollama", () => {
10671073
let nowMs = 20_000_000;
10681074
const clock = () => nowMs;
10691075
const deepseek = createCompactionGovernor(() => undefined, "", [], clock);
@@ -1077,30 +1083,20 @@ describe("provider-aware idle recompress (CL-8745)", () => {
10771083
tenTurns,
10781084
);
10791085

1080-
// Past Anthropic/OpenAI windows but inside DeepSeek's hour: neither fires.
1081-
nowMs += 30 * MINUTE_MS;
1086+
nowMs += 90 * MINUTE_MS;
10821087
expect(
10831088
deepseek.interceptIdleContinuation(emptyMessage(), capabilities),
10841089
).toBeNull();
10851090
expect(
10861091
local.interceptIdleContinuation(emptyMessage(), capabilities),
10871092
).toBeNull();
1088-
1089-
// Past DeepSeek's hour: recompress fires; local inference still never
1090-
// does — no remote cache means no cache benefit.
1091-
nowMs += 31 * MINUTE_MS;
1092-
expect(
1093-
deepseek.interceptIdleContinuation(emptyMessage(), capabilities),
1094-
).toEqual(ttlCompact);
1095-
expect(
1096-
local.interceptIdleContinuation(emptyMessage(), capabilities),
1097-
).toBeNull();
10981093
});
10991094

11001095
test("production ollama LastCycleSource never fires cache-ttl-recompress", () => {
11011096
// Harness stamps { sourceId, provider, model } with a bare model. Ollama is
11021097
// buildOpenAISource: sourceId "ollama/default", provider openai-compatible,
1103-
// model llama3. Keying TTL only off model would take the 10-minute default.
1098+
// Keying TTL off the bare model would miss the ollama sourceId. Local
1099+
// inference stays disabled.
11041100
let nowMs = 25_000_000;
11051101
const governor = createCompactionGovernor(
11061102
() => undefined,
@@ -1120,7 +1116,7 @@ describe("provider-aware idle recompress (CL-8745)", () => {
11201116
tenTurns,
11211117
);
11221118

1123-
nowMs += 11 * MINUTE_MS;
1119+
nowMs += 5 * MINUTE_MS + 1;
11241120
expect(
11251121
governor.interceptIdleContinuation(emptyMessage(), capabilities),
11261122
).toBeNull();
@@ -1142,8 +1138,8 @@ describe("provider-aware idle recompress (CL-8745)", () => {
11421138
() => nowMs,
11431139
);
11441140
governor.noteInferenceDone(inferenceDone(overThreshold), tenTurns);
1145-
// The "m" fixture model carries the 10-minute default TTL; advance past it.
1146-
nowMs += 11 * MINUTE_MS;
1141+
// Fixture provider "p" has no TTL. Threshold arming still owns this session.
1142+
nowMs += 5 * MINUTE_MS + 1;
11471143
// Threshold arming owns the over-threshold session: no TTL double-fold.
11481144
expect(
11491145
governor.interceptIdleContinuation(emptyMessage(), capabilities),
@@ -1165,18 +1161,24 @@ describe("provider-aware idle recompress (CL-8745)", () => {
11651161
[],
11661162
() => nowMs,
11671163
);
1168-
governor.noteInferenceDone(inferenceDone(overThreshold), tenTurns);
1164+
governor.noteInferenceDone(
1165+
inferenceDone(overThreshold, "", "anthropic"),
1166+
tenTurns,
1167+
);
11691168
expect(
11701169
governor.interceptActions(toolDone(), inferAction, capabilities),
11711170
).not.toBeNull();
11721171
expect(governor.resumeAfterCompact(emptyMessage())).toBe("infer");
11731172

1174-
governor.noteInferenceDone(inferenceDone(overThreshold), tenTurns);
1173+
governor.noteInferenceDone(
1174+
inferenceDone(overThreshold, "", "anthropic"),
1175+
tenTurns,
1176+
);
11751177
expect(
11761178
governor.interceptActions(toolDone(), inferAction, capabilities),
11771179
).toBeNull();
11781180

1179-
nowMs += 11 * MINUTE_MS;
1181+
nowMs += 5 * MINUTE_MS + 1;
11801182
expect(
11811183
governor.interceptIdleContinuation(emptyMessage(), capabilities),
11821184
).toEqual(ttlCompact);
@@ -1193,14 +1195,20 @@ describe("provider-aware idle recompress (CL-8745)", () => {
11931195
[],
11941196
() => nowMs,
11951197
);
1196-
governor.noteInferenceDone(inferenceDone(overThreshold), tenTurns);
1198+
governor.noteInferenceDone(
1199+
inferenceDone(overThreshold, "", "anthropic"),
1200+
tenTurns,
1201+
);
11971202
expect(
11981203
governor.interceptActions(toolDone(), inferAction, capabilities),
11991204
).not.toBeNull();
12001205
expect(governor.resumeAfterCompact(emptyMessage())).toBe("infer");
12011206

1202-
governor.noteInferenceDone(inferenceDone(overThreshold), tenTurns);
1203-
nowMs += 11 * MINUTE_MS;
1207+
governor.noteInferenceDone(
1208+
inferenceDone(overThreshold, "", "anthropic"),
1209+
tenTurns,
1210+
);
1211+
nowMs += 5 * MINUTE_MS + 1;
12041212
expect(
12051213
governor.interceptIdleContinuation(emptyMessage(), capabilities),
12061214
).toEqual(ttlCompact);

‎src/provider/cache-ttl.test.ts‎

Lines changed: 19 additions & 30 deletions
Original file line numberDiff line numberDiff line change
@@ -4,24 +4,21 @@ import { cacheTtlMsFor } from "./cache-ttl.js";
44
const MINUTE_MS = 60_000;
55

66
describe("cacheTtlMsFor", () => {
7-
test("maps providers to their documented cache TTL windows", () => {
7+
test("allows idle recompress only for the published Anthropic 5-minute window", () => {
88
expect(cacheTtlMsFor("anthropic/claude-opus-4-6")).toBe(5 * MINUTE_MS);
9-
expect(cacheTtlMsFor("openai-responses/gpt-5.6")).toBe(10 * MINUTE_MS);
10-
expect(cacheTtlMsFor("codex-responses/gpt-5.6")).toBe(10 * MINUTE_MS);
11-
expect(cacheTtlMsFor("openai-compatible/custom")).toBe(10 * MINUTE_MS);
12-
expect(cacheTtlMsFor("xai/thegreataxios")).toBe(10 * MINUTE_MS);
13-
expect(cacheTtlMsFor("gemini/gemini-3-pro")).toBe(15 * MINUTE_MS);
14-
expect(cacheTtlMsFor("deepseek/deepseek-chat")).toBe(60 * MINUTE_MS);
15-
});
16-
17-
test("covers Anthropic-protocol adapters with the 5-minute window", () => {
189
expect(cacheTtlMsFor("zen-messages/claude-opus-4-6")).toBe(5 * MINUTE_MS);
1910
expect(cacheTtlMsFor("opencode-go-messages/claude-opus-4-6")).toBe(
2011
5 * MINUTE_MS,
2112
);
2213
});
2314

24-
test("disables idle recompress for local inference and missing model ids", () => {
15+
test("disables providers whose expiry is unpublished or longer than 5 minutes", () => {
16+
expect(cacheTtlMsFor("openai-responses/gpt-5.6")).toBeUndefined();
17+
expect(cacheTtlMsFor("codex-responses/gpt-5.6")).toBeUndefined();
18+
expect(cacheTtlMsFor("openai-compatible/custom")).toBeUndefined();
19+
expect(cacheTtlMsFor("xai/thegreataxios")).toBeUndefined();
20+
expect(cacheTtlMsFor("gemini/gemini-3-pro")).toBeUndefined();
21+
expect(cacheTtlMsFor("deepseek/deepseek-chat")).toBeUndefined();
2522
expect(cacheTtlMsFor("ollama/llama3.1")).toBeUndefined();
2623
expect(cacheTtlMsFor(undefined)).toBeUndefined();
2724
expect(cacheTtlMsFor("")).toBeUndefined();
@@ -37,7 +34,7 @@ describe("cacheTtlMsFor", () => {
3734
).toBeUndefined();
3835
});
3936

40-
test("maps bare LastCycleSource ids through provider and family", () => {
37+
test("maps a bare Anthropic LastCycleSource and leaves Codex quiet", () => {
4138
expect(
4239
cacheTtlMsFor({
4340
provider: "anthropic",
@@ -49,25 +46,17 @@ describe("cacheTtlMsFor", () => {
4946
provider: "codex-responses",
5047
model: "gpt-5.6-luna",
5148
}),
52-
).toBe(10 * MINUTE_MS);
53-
});
54-
55-
test("falls back to model family for unrecognized provider prefixes", () => {
56-
expect(cacheTtlMsFor("proxy-acme/grok-4")).toBe(10 * MINUTE_MS);
57-
expect(cacheTtlMsFor("proxy-acme/gemini-3-pro")).toBe(15 * MINUTE_MS);
58-
expect(cacheTtlMsFor("proxy-acme/ollama-qwen")).toBeUndefined();
59-
// An exact provider-segment match wins over the model family: an
60-
// openai-compatible account fronting Claude keeps the generic window.
61-
expect(cacheTtlMsFor("openai-compatible/claude-opus-4-6")).toBe(
62-
10 * MINUTE_MS,
63-
);
49+
).toBeUndefined();
6450
});
6551

66-
test("assumes OpenAI-style economics for unrecognized providers", () => {
67-
expect(cacheTtlMsFor("bifrost/some-model")).toBe(10 * MINUTE_MS);
68-
expect(cacheTtlMsFor("totally-new-provider/model-x")).toBe(10 * MINUTE_MS);
69-
// A truly unknown id (no provider segment, no family substring) still
70-
// gets the 10-minute default — not the local-inference disable.
71-
expect(cacheTtlMsFor("unknown-id")).toBe(10 * MINUTE_MS);
52+
test("does not inherit a window from the model family or an unknown provider", () => {
53+
expect(cacheTtlMsFor("proxy-acme/grok-4")).toBeUndefined();
54+
expect(cacheTtlMsFor("proxy-acme/gemini-3-pro")).toBeUndefined();
55+
expect(cacheTtlMsFor("proxy-acme/claude-opus-4-6")).toBeUndefined();
56+
// The provider segment wins: an openai-compatible account fronting
57+
// Claude is not the Anthropic messages protocol.
58+
expect(cacheTtlMsFor("openai-compatible/claude-opus-4-6")).toBeUndefined();
59+
expect(cacheTtlMsFor("bifrost/some-model")).toBeUndefined();
60+
expect(cacheTtlMsFor("unknown-id")).toBeUndefined();
7261
});
7362
});

0 commit comments

Comments
 (0)