Skip to content

Commit 18f315f

Browse files
committed
fix(compaction): key cache ttl off last-cycle source identity
Production LastCycleSource stamps a bare model and openai-compatible provider. Keying TTL only off model treated local Ollama as a 10-minute cache window.
1 parent 76d8b81 commit 18f315f

5 files changed

Lines changed: 119 additions & 14 deletions

File tree

‎docs/ARCHITECTURE.md‎

Lines changed: 3 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -191,8 +191,7 @@ When a cycle's input tokens cross a threshold, the director compacts the inferen
191191

192192
- **Threshold at a tool pause** — Once over threshold, the follow-up `infer` after a tool batch is swapped for a `compact` cycle, and inference resumes via a host continuation message. After a compact that remains over the high watermark, the governor uses **growth hysteresis** (wait for usage to grow by ~10% of the window) instead of re-arming on every cycle; dropping under 60% is not required.
193193
- **Idle (end-of-turn)** — An interactive turn can end with a reply and then sit idle with no tool batch to intercept; the governor requests a continuation at that pause and compacts when it arrives. An operator message that races the continuation still compacts first, then re-enters inference to answer it.
194-
- **Idle recompress past the provider TTL** — The fold is a re-compress, not a cache play: provider KV caches expire on their own schedule, so a session compresses _after_ that expiry (the next turn is a cheaper write and later reads compound on the shrunk context) instead of only at 60% tokens. It is not limited to under-threshold sessions: once growth hysteresis has cleared `pending`, an over-threshold session in that gap can fire the same fold. Any live re-entry past the window fires it — an empty stall ping re-enters meter-only, a raced operator message re-infers — under reason `cache-ttl-recompress` through the same compactor, so the fresh tail stays raw exactly as in the threshold path. The interval follows provider cache economics (`src/provider/cache-ttl.ts`), not a global N minutes.
195-
- **Overflow recovery** — A `context_overflow` inference error would otherwise become a terminal error reply; the governor compacts and retries instead, bounded so a history the compactor cannot shrink does not loop forever. Overflow ignores hysteresis for the compact itself.
194+
- **Idle recompress past the provider TTL** — The fold is a re-compress, not a cache play: provider KV caches expire on their own schedule, so a session compresses _after_ that expiry (the next turn is a cheaper write and later reads compound on the shrunk context) instead of only at 60% tokens. It is not limited to under-threshold sessions: once growth hysteresis has cleared `pending`, an over-threshold session in that gap can fire the same fold. Any live re-entry past the window fires it — an empty stall ping re-enters meter-only, a raced operator message re-infers — under reason `cache-ttl-recompress` through the same compactor, so the fresh tail stays raw exactly as in the threshold path. The interval follows provider cache economics (`src/provider/cache-ttl.ts`), not a global N minutes:
196195

197196
| Provider segment | Idle recompress allowed after | Why |
198197
| ----------------------------------------------------------------- | ----------------------------- | ------------------------------------------------------------------------------------ |
@@ -203,6 +202,8 @@ When a cycle's input tokens cross a threshold, the director compacts the inferen
203202
| `ollama` | never | Local inference has no remote cache to expire |
204203
| anything else | 10 min | Assumed OpenAI-style in-memory economics; tune per upstream |
205204

205+
- **Overflow recovery** — A `context_overflow` inference error would otherwise become a terminal error reply; the governor compacts and retries instead, bounded so a history the compactor cannot shrink does not loop forever. Overflow ignores hysteresis for the compact itself.
206+
206207
The TTL window is measured from the last `inference.done` (every inference rewrites the provider's prefix cache) with one fire per window per compact of any kind, and it never fires with a tool batch outstanding — stall pings can arrive mid-work. The consecutive-compact cap shared with the threshold path still bounds compact→infer→compact. Prompt-caching the summary call itself is an explicit non-goal: the fold carries no cache options.
207208

208209
The compaction control flow is shaped by a reactor invariant: a `compact` action runs in its own cycle (it cannot be paired with `infer`), and **the reactor delivers no event after a compact cycle**. A director that simply emitted `compact` in place of the follow-up `infer` would leave the loop idle forever — the cause of an earlier stall. Instead the governor pairs `compact` with an emit action (`custom.compaction.continue`), and the host answers that emission by delivering a content-less inbound message (`buildCompactionContinuationMessage()`) through the serial operation queue, generation-guarded like every other deliver. A hop superseded because interrupt already bumped generation and enqueued a rebuild is re-queued onto the replacement agent rather than consumed on the outgoing liveAgent. That message adds no turn (`createInboundTurn` returns `null` for empty content) but re-enters the loop, where the director issues the follow-up `infer` against the freshly truncated history. Each emission is answered at most once (a replayed duplicate of an already-answered emission is ignored), and an unsolicited continuation with neither resume flag set answers `wait` rather than burning a billable inference. The legacy `requestContinuation` closure survives only on the sub-agent path, which delivers directly with no host emit hop; idle arming keeps the two channels exclusive (closure fires and the arming flag stays false, or no closure and the flag reports the arming) so a caller honoring both cannot double-deliver.

‎src/agent/compaction.test.ts‎

Lines changed: 37 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -791,9 +791,15 @@ describe("provider-aware idle recompress (CL-8745)", () => {
791791
const MINUTE_MS = 60_000;
792792

793793
function ttlInferenceDone(
794-
model: string,
794+
modelOrSource:
795+
| string
796+
| { sourceId: string; provider: string; model: string },
795797
withTools: boolean,
796798
): Extract<ReactorInboundEvent, { type: "inference.done" }> {
799+
const source =
800+
typeof modelOrSource === "string"
801+
? { sourceId: "s", provider: "p", model: modelOrSource }
802+
: modelOrSource;
797803
return {
798804
type: "inference.done",
799805
turn: {
@@ -810,7 +816,7 @@ describe("provider-aware idle recompress (CL-8745)", () => {
810816
: [{ type: "text", text: "ok" }],
811817
},
812818
usage: usage(1000),
813-
source: { sourceId: "s", provider: "p", model },
819+
source,
814820
} as unknown as Extract<ReactorInboundEvent, { type: "inference.done" }>;
815821
}
816822

@@ -894,6 +900,35 @@ describe("provider-aware idle recompress (CL-8745)", () => {
894900
).toBeNull();
895901
});
896902

903+
test("production ollama LastCycleSource never fires cache-ttl-recompress", () => {
904+
// Harness stamps { sourceId, provider, model } with a bare model. Ollama is
905+
// buildOpenAISource: sourceId "ollama/default", provider openai-compatible,
906+
// model llama3. Keying TTL only off model would take the 10-minute default.
907+
let nowMs = 25_000_000;
908+
const governor = createCompactionGovernor(
909+
() => undefined,
910+
"",
911+
[],
912+
() => nowMs,
913+
);
914+
governor.noteInferenceDone(
915+
ttlInferenceDone(
916+
{
917+
sourceId: "ollama/default",
918+
provider: "openai-compatible",
919+
model: "llama3",
920+
},
921+
false,
922+
),
923+
tenTurns,
924+
);
925+
926+
nowMs += 11 * MINUTE_MS;
927+
expect(
928+
governor.interceptIdleContinuation(emptyMessage(), capabilities),
929+
).toBeNull();
930+
});
931+
897932
test("stays inert with no observed cache write", () => {
898933
const governor = createCompactionGovernor(() => undefined);
899934
expect(

‎src/agent/compaction.ts‎

Lines changed: 9 additions & 3 deletions
Original file line numberDiff line numberDiff line change
@@ -1,5 +1,6 @@
11
import type {
22
ConversationTurn,
3+
LastCycleSource,
34
ReactorAction,
45
ReactorCapabilities,
56
ReactorInboundEvent,
@@ -81,9 +82,13 @@ export function createCompactionGovernor(
8182
// can flag the number as approximate instead of implying provider-grade
8283
// precision.
8384
let usingEstimate = false;
84-
// Model of the last inference.done turn, kept for live re-checks between
85-
// inference cycles (see interceptActions) where the event carries no model.
85+
// Last-cycle source of the last inference.done, kept for live re-checks
86+
// between inference cycles (see interceptActions) where the event carries
87+
// no source. Threshold sizing still keys off `model`; TTL identity needs
88+
// `sourceId` / `provider` as well — production LastCycleSource stamps a
89+
// bare model, and Ollama is `openai-compatible` with id `ollama/…`.
8690
let lastModel: string | undefined;
91+
let lastCycleSource: LastCycleSource | undefined;
8792
let turnCount = 0;
8893
// Wall-clock of the last inference.done: the provider (re)wrote its prefix
8994
// cache for this session on that turn, so the provider TTL window in
@@ -189,6 +194,7 @@ export function createCompactionGovernor(
189194
consecutiveThresholdCompacts = 0;
190195
}
191196
syncFromTurns(turns);
197+
lastCycleSource = event.source;
192198
lastModel = event.source?.model;
193199
lastCacheWriteAt = now();
194200
// The terminal reply ends the previous tool batch (its results are
@@ -308,7 +314,7 @@ export function createCompactionGovernor(
308314
if (turnCount <= MIN_TURNS_TO_COMPACT) return false;
309315
if (atThresholdCompactCap()) return false;
310316
if (outstandingToolCalls > 0) return false;
311-
const ttl = cacheTtlMsFor(lastModel);
317+
const ttl = cacheTtlMsFor(lastCycleSource);
312318
if (ttl === undefined) return false;
313319
if (lastCacheWriteAt === undefined) return false;
314320
if (nowMs - lastCacheWriteAt < ttl) return false;

‎src/provider/cache-ttl.test.ts‎

Lines changed: 25 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -27,6 +27,31 @@ describe("cacheTtlMsFor", () => {
2727
expect(cacheTtlMsFor("")).toBeUndefined();
2828
});
2929

30+
test("keys ollama off production LastCycleSource, not a slash-form model", () => {
31+
expect(
32+
cacheTtlMsFor({
33+
sourceId: "ollama/default",
34+
provider: "openai-compatible",
35+
model: "llama3",
36+
}),
37+
).toBeUndefined();
38+
});
39+
40+
test("maps bare LastCycleSource ids through provider and family", () => {
41+
expect(
42+
cacheTtlMsFor({
43+
provider: "anthropic",
44+
model: "claude-opus-4-6",
45+
}),
46+
).toBe(5 * MINUTE_MS);
47+
expect(
48+
cacheTtlMsFor({
49+
provider: "codex-responses",
50+
model: "gpt-5.6-luna",
51+
}),
52+
).toBe(10 * MINUTE_MS);
53+
});
54+
3055
test("falls back to model family for unrecognized provider prefixes", () => {
3156
expect(cacheTtlMsFor("proxy-acme/grok-4")).toBe(10 * MINUTE_MS);
3257
expect(cacheTtlMsFor("proxy-acme/gemini-3-pro")).toBe(15 * MINUTE_MS);

‎src/provider/cache-ttl.ts‎

Lines changed: 45 additions & 7 deletions
Original file line numberDiff line numberDiff line change
@@ -23,13 +23,20 @@
2323
// 60 min.
2424
// - ollama: local inference has no remote cache to expire, so an idle
2525
// recompress would burn local compute for no cache benefit. Disabled.
26+
// Production LastCycleSource is `{ sourceId, provider, model }` with a
27+
// bare model; Ollama is `buildOpenAISource` (`provider: openai-compatible`,
28+
// `sourceId` like `ollama/default`, model `llama3` / `qwen3`). Identity is
29+
// `sourceId` via `isOllamaProviderId`, not a slash-form string the harness
30+
// never stamps.
2631
// - Unknown/custom providers (bifrost proxy, `openai-compatible` fronting an
2732
// unlisted upstream, unrecognized strings): assumed OpenAI-style in-memory
2833
// economics. 10 min.
2934
//
3035
// Deliberate non-goal: the summary call itself carries no `cache_control` —
3136
// prompt-caching the fold is not attempted.
3237

38+
import { isOllamaProviderId } from "./ollama.js";
39+
3340
const MINUTE_MS = 60_000;
3441

3542
const PROVIDER_TTLS_MS: Record<string, number | undefined> = {
@@ -65,6 +72,12 @@ const FAMILY_TTLS_MS: readonly (readonly [string, number | undefined])[] = [
6572
["ollama", undefined],
6673
];
6774

75+
type CacheTtlIdentity = {
76+
sourceId?: string;
77+
provider?: string;
78+
model?: string;
79+
};
80+
6881
// Canonical provider segment of a `provider:model` string: the account or
6982
// adapter name before the first "/" (custom names like `xai/thegreataxios`
7083
// carry the provider there), else the head before ":". Mirrors the
@@ -78,13 +91,7 @@ function canonicalSegment(model: string): string {
7891
return colon >= 0 ? head.slice(0, colon) : head;
7992
}
8093

81-
/**
82-
* Milliseconds of provider-cache idle after which a recompress is allowed,
83-
* or `undefined` when the model is undefined/empty or the provider has no
84-
* remote cache (local inference). Never throws; unknown strings get the
85-
* default 10-minute window.
86-
*/
87-
export function cacheTtlMsFor(model: string | undefined): number | undefined {
94+
function ttlForModelString(model: string | undefined): number | undefined {
8895
if (model === undefined) return undefined;
8996
const segment = canonicalSegment(model);
9097
if (segment.length === 0) return undefined;
@@ -96,3 +103,34 @@ export function cacheTtlMsFor(model: string | undefined): number | undefined {
96103
}
97104
return DEFAULT_TTL_MS;
98105
}
106+
107+
/**
108+
* Milliseconds of provider-cache idle after which a recompress is allowed,
109+
* or `undefined` when the identity is missing/empty or the provider has no
110+
* remote cache (local inference). Never throws; unknown strings get the
111+
* default 10-minute window.
112+
*
113+
* Accepts a slash-form `provider/model` string or a LastCycleSource-shaped
114+
* identity. Production events stamp a bare `model`; Ollama is recognized
115+
* from `sourceId` (`isOllamaProviderId`), not from a combined string the
116+
* harness never stamps.
117+
*/
118+
export function cacheTtlMsFor(
119+
identity: string | CacheTtlIdentity | undefined,
120+
): number | undefined {
121+
if (identity === undefined) return undefined;
122+
if (typeof identity === "string") return ttlForModelString(identity);
123+
124+
if (identity.sourceId !== undefined && isOllamaProviderId(identity.sourceId))
125+
return undefined;
126+
if (identity.provider !== undefined && isOllamaProviderId(identity.provider))
127+
return undefined;
128+
129+
if (identity.provider !== undefined) {
130+
const provider = identity.provider.toLowerCase();
131+
if (Object.hasOwn(PROVIDER_TTLS_MS, provider))
132+
return PROVIDER_TTLS_MS[provider];
133+
}
134+
135+
return ttlForModelString(identity.model);
136+
}

0 commit comments

Comments
 (0)