From b6d0faf01283d9b04cde66f4344346a13775654f Mon Sep 17 00:00:00 2001 From: Sawyer Date: Mon, 21 Sep 2026 11:32:48 -0700 Subject: [PATCH 1/6] feat(compaction): idle recompress past the provider cache ttl --- docs/ARCHITECTURE.md | 12 ++ src/agent/compaction.test.ts | 227 +++++++++++++++++++++++++++++++++ src/agent/compaction.ts | 102 +++++++++++++-- src/provider/cache-ttl.test.ts | 45 +++++++ src/provider/cache-ttl.ts | 97 ++++++++++++++ 5 files changed, 473 insertions(+), 10 deletions(-) create mode 100644 src/provider/cache-ttl.test.ts create mode 100644 src/provider/cache-ttl.ts diff --git a/docs/ARCHITECTURE.md b/docs/ARCHITECTURE.md index 97bacb163..d6b66deca 100644 --- a/docs/ARCHITECTURE.md +++ b/docs/ARCHITECTURE.md @@ -191,8 +191,20 @@ When a cycle's input tokens cross a threshold, the director compacts the inferen - **Threshold at a tool pause** — Once over threshold, the follow-up `infer` after a tool batch is swapped for a `compact` cycle, and inference resumes via a host continuation message. After a compact that remains over the high watermark, the governor uses **growth hysteresis** (wait for usage to grow by ~10% of the window) instead of re-arming on every cycle; dropping under 60% is not required. - **Idle (end-of-turn)** — An interactive turn can end with a reply and then sit idle with no tool batch to intercept; the governor requests a continuation at that pause and compacts when it arrives. An operator message that races the continuation still compacts first, then re-enters inference to answer it. +- **Idle recompress past the provider TTL** — The fold is a re-compress, not a cache play: provider KV caches expire on their own schedule, so an under-threshold session compresses _after_ that expiry (the next turn is a cheaper write and later reads compound on the shrunk context) instead of only at 60% tokens. Any live re-entry past the window fires it — an empty stall ping re-enters meter-only, a raced operator message re-infers — under reason `cache-ttl-recompress` through the same compactor, so the fresh tail stays raw exactly as in the threshold path. The interval follows provider cache economics (`src/provider/cache-ttl.ts`), not a global N minutes: - **Overflow recovery** — A `context_overflow` inference error would otherwise become a terminal error reply; the governor compacts and retries instead, bounded so a history the compactor cannot shrink does not loop forever. Overflow ignores hysteresis for the compact itself. +| Provider segment | Idle recompress allowed after | Why | +| ----------------------------------------------------------------- | ----------------------------- | ------------------------------------------------------------------------------------ | +| `anthropic`, `zen-messages`, `opencode-go-messages` | 5 min | Default ephemeral cache TTL; hour-long breakpoints are never set | +| `openai-responses`, `codex-responses`, `openai-compatible`, `xai` | 10 min | In-memory prefixes typically evicted after 5–10 min idle; xAI assumed same economics | +| `gemini` | 15 min | No published implicit-cache eviction window; conservative | +| `deepseek` | 60 min | On-disk context cache persists for hours-to-days | +| `ollama` | never | Local inference has no remote cache to expire | +| anything else | 10 min | Assumed OpenAI-style in-memory economics; tune per upstream | + +The TTL window is measured from the last `inference.done` (every inference rewrites the provider's prefix cache) with one fire per window per compact of any kind, and it never fires with a tool batch outstanding — stall pings can arrive mid-work. The consecutive-compact cap shared with the threshold path still bounds compact→infer→compact. Prompt-caching the summary call itself is an explicit non-goal: the fold carries no cache options. + The compaction control flow is shaped by a reactor invariant: a `compact` action runs in its own cycle (it cannot be paired with `infer`), and **the reactor delivers no event after a compact cycle**. A director that simply emitted `compact` in place of the follow-up `infer` would leave the loop idle forever — the cause of an earlier stall. Instead the governor pairs `compact` with an emit action (`custom.compaction.continue`), and the host answers that emission by delivering a content-less inbound message (`buildCompactionContinuationMessage()`) through the serial operation queue, generation-guarded like every other deliver. A hop superseded because interrupt already bumped generation and enqueued a rebuild is re-queued onto the replacement agent rather than consumed on the outgoing liveAgent. That message adds no turn (`createInboundTurn` returns `null` for empty content) but re-enters the loop, where the director issues the follow-up `infer` against the freshly truncated history. Each emission is answered at most once (a replayed duplicate of an already-answered emission is ignored), and an unsolicited continuation with neither resume flag set answers `wait` rather than burning a billable inference. The legacy `requestContinuation` closure survives only on the sub-agent path, which delivers directly with no host emit hop; idle arming keeps the two channels exclusive (closure fires and the arming flag stays false, or no closure and the flag reports the arming) so a caller honoring both cannot double-deliver. Compaction replaces older turns with a structured, workflow-aware summary rather than a stats blob: sections for **What Happened / What We're Doing / Relevant Links / Action Items / Next Steps**, with the active workflow and step woven in so compacting mid-`/build` or mid-`/plan` preserves the contract. The summary is produced by a one-shot model call; on any failure it falls back to a deterministic summary so a compaction cycle never breaks the session. diff --git a/src/agent/compaction.test.ts b/src/agent/compaction.test.ts index 44e9a8808..d19929870 100644 --- a/src/agent/compaction.test.ts +++ b/src/agent/compaction.test.ts @@ -786,3 +786,230 @@ describe("compaction governor", () => { expect(continuations).toBe(1); }); }); + +describe("provider-aware idle recompress (CL-8745)", () => { + const MINUTE_MS = 60_000; + + function ttlInferenceDone( + model: string, + withTools: boolean, + ): Extract { + return { + type: "inference.done", + turn: { + role: "assistant", + content: withTools + ? [ + { + type: "tool_call", + id: "c1", + name: "read_file", + arguments: { path: "a.ts" }, + }, + ] + : [{ type: "text", text: "ok" }], + }, + usage: usage(1000), + source: { sourceId: "s", provider: "p", model }, + } as unknown as Extract; + } + + function racedMessage(): ReactorInboundEvent { + return { + type: "message.received", + message: { content: "next question" }, + } as ReactorInboundEvent; + } + + const ttlCompact = [ + { + type: "compact", + compactor: "pruning-compactor", + reason: "cache-ttl-recompress", + }, + ] as ReactorAction[]; + + test("fires past the provider TTL while under threshold, meter-only on empty", () => { + let continuations = 0; + let nowMs = 10_000_000; + const governor = createCompactionGovernor( + () => continuations++, + "", + [], + () => nowMs, + ); + governor.noteInferenceDone( + ttlInferenceDone("anthropic/claude-opus-4-6", false), + tenTurns, + ); + + // Inside the 5-minute Anthropic window: no fire. + nowMs += 4 * MINUTE_MS; + expect( + governor.interceptIdleContinuation(emptyMessage(), capabilities), + ).toBeNull(); + + // Past the window: the same fold as the threshold path (same compactor, + // so the fresh tail stays raw) with an attributable reason. + nowMs += MINUTE_MS + 1; + expect( + governor.interceptIdleContinuation(emptyMessage(), capabilities), + ).toEqual(ttlCompact); + expect(continuations).toBe(1); + // Empty continuation adopts the shrunk turns without a new inference. + expect(governor.resumeAfterCompact(emptyMessage())).toBe("meter"); + }); + + test("follows provider economics: deepseek waits out its long window, ollama never fires", () => { + let nowMs = 20_000_000; + const clock = () => nowMs; + const deepseek = createCompactionGovernor(() => undefined, "", [], clock); + const local = createCompactionGovernor(() => undefined, "", [], clock); + deepseek.noteInferenceDone( + ttlInferenceDone("deepseek/deepseek-chat", false), + tenTurns, + ); + local.noteInferenceDone( + ttlInferenceDone("ollama/llama3.1", false), + tenTurns, + ); + + // Past Anthropic/OpenAI windows but inside DeepSeek's hour: neither fires. + nowMs += 30 * MINUTE_MS; + expect( + deepseek.interceptIdleContinuation(emptyMessage(), capabilities), + ).toBeNull(); + expect( + local.interceptIdleContinuation(emptyMessage(), capabilities), + ).toBeNull(); + + // Past DeepSeek's hour: recompress fires; local inference still never + // does — no remote cache means no cache benefit. + nowMs += 31 * MINUTE_MS; + expect( + deepseek.interceptIdleContinuation(emptyMessage(), capabilities), + ).toEqual(ttlCompact); + expect( + local.interceptIdleContinuation(emptyMessage(), capabilities), + ).toBeNull(); + }); + + test("stays inert with no observed cache write", () => { + const governor = createCompactionGovernor(() => undefined); + expect( + governor.interceptIdleContinuation(emptyMessage(), capabilities), + ).toBeNull(); + }); + + test("defers to the armed threshold path while over threshold", () => { + let nowMs = 30_000_000; + const governor = createCompactionGovernor( + () => undefined, + "", + [], + () => nowMs, + ); + governor.noteInferenceDone(inferenceDone(overThreshold), tenTurns); + // The "m" fixture model carries the 10-minute default TTL; advance past it. + nowMs += 11 * MINUTE_MS; + // Threshold arming owns the over-threshold session: no TTL double-fold. + expect( + governor.interceptIdleContinuation(emptyMessage(), capabilities), + ).toBeNull(); + expect( + governor.interceptIdleContinuation(racedMessage(), capabilities), + ).toBeNull(); + }); + + test("does not fold under an outstanding tool batch, fires once it settles", () => { + let continuations = 0; + let nowMs = 40_000_000; + const governor = createCompactionGovernor( + () => continuations++, + "", + [], + () => nowMs, + ); + governor.noteInferenceDone( + ttlInferenceDone("anthropic/claude-opus-4-6", true), + tenTurns, + ); + nowMs += 6 * MINUTE_MS; + expect( + governor.interceptIdleContinuation(emptyMessage(), capabilities), + ).toBeNull(); + // The batch settles (threshold path uninvolved: under threshold). + expect( + governor.interceptActions(toolDone(), inferAction, capabilities), + ).toBeNull(); + expect( + governor.interceptIdleContinuation(emptyMessage(), capabilities), + ).toEqual(ttlCompact); + expect(continuations).toBe(1); + }); + + test("one fire per window, then the shared consecutive-compact cap stops the spiral", () => { + let continuations = 0; + let nowMs = 50_000_000; + const governor = createCompactionGovernor( + () => continuations++, + "", + [], + () => nowMs, + ); + governor.noteInferenceDone( + ttlInferenceDone("anthropic/claude-opus-4-6", false), + tenTurns, + ); + + nowMs += 5 * MINUTE_MS + 1; + expect( + governor.interceptIdleContinuation(emptyMessage(), capabilities), + ).toEqual(ttlCompact); + expect(governor.resumeAfterCompact(emptyMessage())).toBe("meter"); + governor.notePostCompact(tenTurns); + + // Inside the next window: no second fire. + nowMs += 4 * MINUTE_MS; + expect( + governor.interceptIdleContinuation(emptyMessage(), capabilities), + ).toBeNull(); + + nowMs += MINUTE_MS + 1; + expect( + governor.interceptIdleContinuation(emptyMessage(), capabilities), + ).toEqual(ttlCompact); + expect(governor.resumeAfterCompact(emptyMessage())).toBe("meter"); + governor.notePostCompact(tenTurns); + + // Cap reached: the third window stays quiet. + nowMs += 5 * MINUTE_MS + 1; + expect( + governor.interceptIdleContinuation(emptyMessage(), capabilities), + ).toBeNull(); + expect(continuations).toBe(2); + }); + + test("a raced operator message past the TTL still folds, then re-infers", () => { + let continuations = 0; + let nowMs = 60_000_000; + const governor = createCompactionGovernor( + () => continuations++, + "", + [], + () => nowMs, + ); + governor.noteInferenceDone( + ttlInferenceDone("anthropic/claude-opus-4-6", false), + tenTurns, + ); + nowMs += 6 * MINUTE_MS; + expect( + governor.interceptIdleContinuation(racedMessage(), capabilities), + ).toEqual(ttlCompact); + expect(continuations).toBe(1); + // The follow-up empty continuation carries the infer that answers the + // raced question (mirrors the threshold raced path). + expect(governor.resumeAfterCompact(emptyMessage())).toBe("infer"); + }); +}); diff --git a/src/agent/compaction.ts b/src/agent/compaction.ts index 9f3fab454..c8aa08d3e 100644 --- a/src/agent/compaction.ts +++ b/src/agent/compaction.ts @@ -10,6 +10,7 @@ import { compactionThresholdFor, contextTokensFromUsage, } from "../provider/context-window.js"; +import { cacheTtlMsFor } from "../provider/cache-ttl.js"; import { COMPACTOR_KEEP_RECENT_TURNS, assistantTextIsCompactSpacerEcho, @@ -64,6 +65,7 @@ export function createCompactionGovernor( requestContinuation?: () => void, systemPrompt = "", toolDefinitions: readonly ToolDefinition[] = [], + now: () => number = Date.now, ) { let pending = false; let idlePending = false; @@ -83,6 +85,22 @@ export function createCompactionGovernor( // inference cycles (see interceptActions) where the event carries no model. let lastModel: string | undefined; let turnCount = 0; + // Wall-clock of the last inference.done: the provider (re)wrote its prefix + // cache for this session on that turn, so the provider TTL window in + // provider/cache-ttl.ts is measured from here. Stamped on every + // inference.done — estimated-usage providers wrote a cache entry too. + let lastCacheWriteAt: number | undefined; + // Wall-clock of the last issued compact of any kind. A fresh fold rewrites + // the session prefix, so the TTL recompress must not fire again inside the + // same window even when its own cache write has not been observed yet (the + // summary call bypasses this governor). + let lastCompactAt: number | undefined; + // Tool calls issued by the last inference.done and not yet settled. A TTL + // recompress must never fold while a batch is outstanding — the stall ping + // that triggers it can arrive mid-work, and folding under it would rewrite + // turns the pending results still belong to. Assigned (not incremented) on + // every inference.done so the serial loop self-heals a miscount. + let outstandingToolCalls = 0; // Growth hysteresis after a compact that remained over the high watermark: // snapshot the post-compact infer's usage, then do not re-arm until usage // grows by resumeDelta. Cleared once usage drops back to or under high. @@ -122,6 +140,7 @@ export function createCompactionGovernor( function noteCompactIssued(): void { awaitingPostCompactMeasurement = true; + lastCompactAt = now(); } function atThresholdCompactCap(): boolean { @@ -171,6 +190,16 @@ export function createCompactionGovernor( } syncFromTurns(turns); lastModel = event.source?.model; + lastCacheWriteAt = now(); + // The terminal reply ends the previous tool batch (its results are + // already in the turns) and opens the batch the reply just issued. A TTL + // recompress must never fold while a batch is outstanding — the stall + // ping that triggers it can arrive mid-work, and folding under it would + // rewrite turns the pending results still belong to. Assigned, not + // incremented, so the serial loop self-heals a miscount. + outstandingToolCalls = event.turn.content.filter( + (block) => block.type === "tool_call", + ).length; const reportedTokens = contextTokensFromUsage(event.usage); usingEstimate = reportedTokens <= 0; const contextTokens = usingEstimate ? estimate.tokens : reportedTokens; @@ -210,6 +239,7 @@ export function createCompactionGovernor( capabilities: ReactorCapabilities, ): ReactorAction[] | null { if (event.type !== "tool.done") return null; + if (outstandingToolCalls > 0) outstandingToolCalls -= 1; if (!pending && !(usingEstimate && isOverThreshold(estimate.tokens))) return null; if (!actions.some((a) => a.type === "infer")) return null; @@ -256,23 +286,71 @@ export function createCompactionGovernor( return true; } + // Provider-aware idle recompress (CL-8745): the fold is a re-compress, not + // a cache play. Provider KV caches expire on their own schedule + // (provider/cache-ttl.ts); compressing after that expiry makes the next + // turn a cheaper write and later reads compound on the shrunk context. This + // fires under threshold — the threshold path above owns over-threshold — + // on any live re-entry once `now - lastCacheWrite >= ttl`. Guards, in + // order: threshold arming defers (pending), the fresh-tail floor (turns at + // or under it are all kept, so a fold would shrink nothing), the + // consecutive-compact cap (existing death-spiral bound, shared with the + // threshold path), providers with no TTL (unknown model, local inference), + // no observed cache write yet, the TTL window itself, and one fire per + // window (a fresh fold rewrites the prefix; the summary call bypasses this + // governor so lastCacheWriteAt cannot observe it — lastCompactAt covers + // that). Never fires with a tool batch outstanding: the stall ping that + // triggers this can arrive mid-work. + function isTtlRecompressDue(nowMs: number): boolean { + if (pending) return false; + if (turnCount <= MIN_TURNS_TO_COMPACT) return false; + if (atThresholdCompactCap()) return false; + if (outstandingToolCalls > 0) return false; + const ttl = cacheTtlMsFor(lastModel); + if (ttl === undefined) return false; + if (lastCacheWriteAt === undefined) return false; + if (nowMs - lastCacheWriteAt < ttl) return false; + if (lastCompactAt !== undefined && nowMs - lastCompactAt < ttl) + return false; + return true; + } + function interceptIdleContinuation( event: ReactorInboundEvent, capabilities: ReactorCapabilities, ): ReactorAction[] | null { - if (!idlePending || event.type !== "message.received") return null; - if (atThresholdCompactCap()) { + if (event.type !== "message.received") return null; + if (idlePending) { + if (atThresholdCompactCap()) { + idlePending = false; + return null; + } idlePending = false; - return null; + pending = false; + const content = + typeof event.message.content === "string" ? event.message.content : ""; + // The reactor delivers no event after compact, so always request a + // continuation to re-enter decide against the shrunk turns: + // - raced operator content → re-infer to answer it + // - empty synthetic continuation → meter-only sync (no infer) + if (content.length > 0) { + postCompactInfer = true; + } else { + postCompactMeter = true; + } + issueThresholdCompact(); + return [ + capabilities.compact(COMPACTOR_NAME, "context-threshold"), + ...continuationActions(capabilities), + ]; } - idlePending = false; - pending = false; + // Unarmed idle re-entry past the provider TTL: same fold, same + // keep-recent tail, same cap — but a "cache-ttl-recompress" reason so the + // fold is attributable. No arming: every live re-entry re-checks the + // window, so a sub-agent stall ping or operator message is the trigger. + if (!isTtlRecompressDue(now())) return null; const content = typeof event.message.content === "string" ? event.message.content : ""; - // The reactor delivers no event after compact, so always request a - // continuation to re-enter decide against the shrunk turns: - // - raced operator content → re-infer to answer it - // - empty synthetic continuation → meter-only sync (no infer) if (content.length > 0) { postCompactInfer = true; } else { @@ -280,7 +358,7 @@ export function createCompactionGovernor( } issueThresholdCompact(); return [ - capabilities.compact(COMPACTOR_NAME, "context-threshold"), + capabilities.compact(COMPACTOR_NAME, "cache-ttl-recompress"), ...continuationActions(capabilities), ]; } @@ -336,6 +414,10 @@ export function createCompactionGovernor( function notePostCompact(turns: readonly ConversationTurn[]): void { syncFromTurns(turns); usingEstimate = true; + // A fold rewrites the turns: results already applied vanish from the + // live set, and post-compact stall pings (empty continuations) carry no + // tool traffic. Reset so a stale count cannot pin the TTL window shut. + outstandingToolCalls = 0; } // True while the governor expects the host to answer a continuation emit. diff --git a/src/provider/cache-ttl.test.ts b/src/provider/cache-ttl.test.ts new file mode 100644 index 000000000..f3d1a4bf9 --- /dev/null +++ b/src/provider/cache-ttl.test.ts @@ -0,0 +1,45 @@ +import { describe, expect, test } from "bun:test"; +import { cacheTtlMsFor } from "./cache-ttl.js"; + +const MINUTE_MS = 60_000; + +describe("cacheTtlMsFor", () => { + test("maps providers to their documented cache TTL windows", () => { + expect(cacheTtlMsFor("anthropic/claude-opus-4-6")).toBe(5 * MINUTE_MS); + expect(cacheTtlMsFor("openai-responses/gpt-5.6")).toBe(10 * MINUTE_MS); + expect(cacheTtlMsFor("codex-responses/gpt-5.6")).toBe(10 * MINUTE_MS); + expect(cacheTtlMsFor("openai-compatible/custom")).toBe(10 * MINUTE_MS); + expect(cacheTtlMsFor("xai/thegreataxios")).toBe(10 * MINUTE_MS); + expect(cacheTtlMsFor("gemini/gemini-3-pro")).toBe(15 * MINUTE_MS); + expect(cacheTtlMsFor("deepseek/deepseek-chat")).toBe(60 * MINUTE_MS); + }); + + test("covers Anthropic-protocol adapters with the 5-minute window", () => { + expect(cacheTtlMsFor("zen-messages/claude-opus-4-6")).toBe(5 * MINUTE_MS); + expect(cacheTtlMsFor("opencode-go-messages/claude-opus-4-6")).toBe( + 5 * MINUTE_MS, + ); + }); + + test("disables idle recompress for local inference and unknown models", () => { + expect(cacheTtlMsFor("ollama/llama3.1")).toBeUndefined(); + expect(cacheTtlMsFor(undefined)).toBeUndefined(); + expect(cacheTtlMsFor("")).toBeUndefined(); + }); + + test("falls back to model family for unrecognized provider prefixes", () => { + expect(cacheTtlMsFor("proxy-acme/grok-4")).toBe(10 * MINUTE_MS); + expect(cacheTtlMsFor("proxy-acme/gemini-3-pro")).toBe(15 * MINUTE_MS); + expect(cacheTtlMsFor("proxy-acme/ollama-qwen")).toBeUndefined(); + // An exact provider-segment match wins over the model family: an + // openai-compatible account fronting Claude keeps the generic window. + expect(cacheTtlMsFor("openai-compatible/claude-opus-4-6")).toBe( + 10 * MINUTE_MS, + ); + }); + + test("assumes OpenAI-style economics for unrecognized providers", () => { + expect(cacheTtlMsFor("bifrost/some-model")).toBe(10 * MINUTE_MS); + expect(cacheTtlMsFor("totally-new-provider/model-x")).toBe(10 * MINUTE_MS); + }); +}); diff --git a/src/provider/cache-ttl.ts b/src/provider/cache-ttl.ts new file mode 100644 index 000000000..5c9a33438 --- /dev/null +++ b/src/provider/cache-ttl.ts @@ -0,0 +1,97 @@ +// Provider prompt-cache TTLs for idle recompression (CL-8745). +// +// The fold is a re-compress, not a cache play: provider KV caches expire on +// their own schedule, and idle compression fires *after* that expiry so the +// next turn is a cheaper write (a smaller prefix to re-process) and later +// reads compound on the shrunk context. The interval therefore follows +// provider cache economics, not a single global N minutes. +// +// Mapping — idle recompress is allowed once `now - lastCacheWrite >= ttl`: +// - Anthropic (`anthropic`, plus `zen-messages` / `opencode-go-messages`, +// which speak the Anthropic messages protocol): default 5-minute ephemeral +// cache TTL. Hour-long breakpoints exist but we never set them. 5 min. +// - OpenAI family (`openai-responses`, `codex-responses`, +// `openai-compatible`): in-memory prefixes are typically evicted after +// 5-10 minutes idle (retained up to an hour off-peak); GPT-5.6+ families +// default to a 30-minute minimum TTL instead. 10 min splits the range. +// - xAI (`xai`, `grok-*`): per-server cache, no published TTL; assumed +// OpenAI-style in-memory economics. 10 min. +// - Gemini (`gemini`, `gemini-*`): implicit caching on 2.5+ with no published +// eviction window; explicit caches default to 1 hour. Conservative. 15 min. +// - DeepSeek (`deepseek`): on-disk context cache cleared only after +// hours-to-days idle, so an early recompress would fold a still-warm cache. +// 60 min. +// - ollama: local inference has no remote cache to expire, so an idle +// recompress would burn local compute for no cache benefit. Disabled. +// - Unknown/custom providers (bifrost proxy, `openai-compatible` fronting an +// unlisted upstream, unrecognized strings): assumed OpenAI-style in-memory +// economics. 10 min. +// +// Deliberate non-goal: the summary call itself carries no `cache_control` — +// prompt-caching the fold is not attempted. + +const MINUTE_MS = 60_000; + +const PROVIDER_TTLS_MS: Record = { + anthropic: 5 * MINUTE_MS, + "zen-messages": 5 * MINUTE_MS, + "opencode-go-messages": 5 * MINUTE_MS, + "openai-responses": 10 * MINUTE_MS, + "codex-responses": 10 * MINUTE_MS, + "openai-compatible": 10 * MINUTE_MS, + gemini: 15 * MINUTE_MS, + deepseek: 60 * MINUTE_MS, + xai: 10 * MINUTE_MS, + // No remote cache to expire; idle recompress would only burn local compute. + ollama: undefined, +}; + +const DEFAULT_TTL_MS = 10 * MINUTE_MS; + +// Model-family fallback for unrecognized provider prefixes (custom proxies, +// account names). An exact provider-segment match always wins — e.g. an +// `openai-compatible` account fronting Claude keeps the generic 10-minute +// window rather than Claude's 5. Checked in order; first substring wins. +const FAMILY_TTLS_MS: readonly (readonly [string, number | undefined])[] = [ + ["claude", 5 * MINUTE_MS], + ["anthropic", 5 * MINUTE_MS], + ["gpt", 10 * MINUTE_MS], + ["codex", 10 * MINUTE_MS], + ["openai", 10 * MINUTE_MS], + ["grok", 10 * MINUTE_MS], + ["xai", 10 * MINUTE_MS], + ["gemini", 15 * MINUTE_MS], + ["deepseek", 60 * MINUTE_MS], + ["ollama", undefined], +]; + +// Canonical provider segment of a `provider:model` string: the account or +// adapter name before the first "/" (custom names like `xai/thegreataxios` +// carry the provider there), else the head before ":". Mirrors the +// segmentation in provider/context-window.ts without its window-table +// fallback semantics. +function canonicalSegment(model: string): string { + const lower = model.toLowerCase(); + const slash = lower.indexOf("/"); + const head = slash >= 0 ? lower.slice(0, slash) : lower; + const colon = head.indexOf(":"); + return colon >= 0 ? head.slice(0, colon) : head; +} + +/** + * Milliseconds of provider-cache idle after which a recompress is allowed, + * or `undefined` when the model is unknown or the provider has no remote + * cache (local inference). Never throws; unknown strings get the default. + */ +export function cacheTtlMsFor(model: string | undefined): number | undefined { + if (model === undefined) return undefined; + const segment = canonicalSegment(model); + if (segment.length === 0) return undefined; + if (Object.hasOwn(PROVIDER_TTLS_MS, segment)) + return PROVIDER_TTLS_MS[segment]; + const lower = model.toLowerCase(); + for (const [family, ttl] of FAMILY_TTLS_MS) { + if (lower.includes(family)) return ttl; + } + return DEFAULT_TTL_MS; +} From 4487f8f904aa12271deb459e26d05cead595f517 Mon Sep 17 00:00:00 2001 From: Sawyer Date: Mon, 21 Sep 2026 16:30:04 -0700 Subject: [PATCH 2/6] test(compaction): characterize ttl fire in the hysteresis gap The TTL arm only defers while threshold arming is pending. After a compact that stays over the high watermark, growth hysteresis clears pending, so idle past the provider TTL still folds. Comments claimed otherwise; keep the behavior and document it. cacheTtlMsFor returns undefined only for missing, empty, or local models; unknown strings get the default window. --- src/agent/compaction.test.ts | 29 +++++++++++++++++++++++++++++ src/agent/compaction.ts | 18 ++++++++++-------- src/provider/cache-ttl.ts | 5 +++-- 3 files changed, 42 insertions(+), 10 deletions(-) diff --git a/src/agent/compaction.test.ts b/src/agent/compaction.test.ts index d19929870..5f8ca921b 100644 --- a/src/agent/compaction.test.ts +++ b/src/agent/compaction.test.ts @@ -921,6 +921,35 @@ describe("provider-aware idle recompress (CL-8745)", () => { ).toBeNull(); }); + test("fires cache-ttl-recompress in the hysteresis gap (over-threshold, no growth)", () => { + // Characterization, not a bug: after a threshold compact, a post-compact + // infer at the same usage clears `pending` via growth hysteresis, so the + // threshold path no longer owns the session. Idle past the provider TTL + // still folds — same window- and cap-bounded path as under-threshold. + let nowMs = 70_000_000; + const governor = createCompactionGovernor( + () => undefined, + "", + [], + () => nowMs, + ); + governor.noteInferenceDone(inferenceDone(overThreshold), tenTurns); + expect( + governor.interceptActions(toolDone(), inferAction, capabilities), + ).not.toBeNull(); + expect(governor.resumeAfterCompact(emptyMessage())).toBe("infer"); + + governor.noteInferenceDone(inferenceDone(overThreshold), tenTurns); + expect( + governor.interceptActions(toolDone(), inferAction, capabilities), + ).toBeNull(); + + nowMs += 11 * MINUTE_MS; + expect( + governor.interceptIdleContinuation(emptyMessage(), capabilities), + ).toEqual(ttlCompact); + }); + test("does not fold under an outstanding tool batch, fires once it settles", () => { let continuations = 0; let nowMs = 40_000_000; diff --git a/src/agent/compaction.ts b/src/agent/compaction.ts index c8aa08d3e..6e312640b 100644 --- a/src/agent/compaction.ts +++ b/src/agent/compaction.ts @@ -290,17 +290,19 @@ export function createCompactionGovernor( // a cache play. Provider KV caches expire on their own schedule // (provider/cache-ttl.ts); compressing after that expiry makes the next // turn a cheaper write and later reads compound on the shrunk context. This - // fires under threshold — the threshold path above owns over-threshold — - // on any live re-entry once `now - lastCacheWrite >= ttl`. Guards, in + // fires on any live re-entry once `now - lastCacheWrite >= ttl`, including + // over-threshold sessions whose threshold path has disarmed (`pending` is + // false via growth hysteresis — the threshold path owns only armed + // over-threshold; the fold is still window- and cap-bounded). Guards, in // order: threshold arming defers (pending), the fresh-tail floor (turns at // or under it are all kept, so a fold would shrink nothing), the // consecutive-compact cap (existing death-spiral bound, shared with the - // threshold path), providers with no TTL (unknown model, local inference), - // no observed cache write yet, the TTL window itself, and one fire per - // window (a fresh fold rewrites the prefix; the summary call bypasses this - // governor so lastCacheWriteAt cannot observe it — lastCompactAt covers - // that). Never fires with a tool batch outstanding: the stall ping that - // triggers this can arrive mid-work. + // threshold path), providers with no TTL (undefined/empty model, local + // inference), no observed cache write yet, the TTL window itself, and one + // fire per window (a fresh fold rewrites the prefix; the summary call + // bypasses this governor so lastCacheWriteAt cannot observe it — + // lastCompactAt covers that). Never fires with a tool batch outstanding: + // the stall ping that triggers this can arrive mid-work. function isTtlRecompressDue(nowMs: number): boolean { if (pending) return false; if (turnCount <= MIN_TURNS_TO_COMPACT) return false; diff --git a/src/provider/cache-ttl.ts b/src/provider/cache-ttl.ts index 5c9a33438..7b22da700 100644 --- a/src/provider/cache-ttl.ts +++ b/src/provider/cache-ttl.ts @@ -80,8 +80,9 @@ function canonicalSegment(model: string): string { /** * Milliseconds of provider-cache idle after which a recompress is allowed, - * or `undefined` when the model is unknown or the provider has no remote - * cache (local inference). Never throws; unknown strings get the default. + * or `undefined` when the model is undefined/empty or the provider has no + * remote cache (local inference). Never throws; unknown strings get the + * default 10-minute window. */ export function cacheTtlMsFor(model: string | undefined): number | undefined { if (model === undefined) return undefined; From 76d8b816b7e461a3528903fcb84e65757efaac0a Mon Sep 17 00:00:00 2001 From: Sawyer Date: Mon, 21 Sep 2026 16:51:56 -0700 Subject: [PATCH 3/6] fix(compaction): align ttl nits with hysteresis-gap behavior --- docs/ARCHITECTURE.md | 4 +-- src/agent/compaction.test.ts | 44 ++++++++++++++++++++++++ src/agent/compaction.ts | 61 +++++++++++++++++++--------------- src/provider/cache-ttl.test.ts | 5 ++- 4 files changed, 84 insertions(+), 30 deletions(-) diff --git a/docs/ARCHITECTURE.md b/docs/ARCHITECTURE.md index d6b66deca..e661dde30 100644 --- a/docs/ARCHITECTURE.md +++ b/docs/ARCHITECTURE.md @@ -187,11 +187,11 @@ The agent maintains an optional **`manage_tasks`** list (create/update via the h #### Context compaction (the compaction governor) -When a cycle's input tokens cross a threshold, the director compacts the inference-facing history (the full run is always retained in the context store). The threshold is **model-aware** — roughly 60% of the active model's real context window — so small-window models compact early enough to avoid provider context-overflow while large-window models do not compact prematurely. The compacted prefix is **append-only across passes**: the existing compacted user turn stays byte-identical; new folds become later summary turns with a harness-inserted assistant spacer (identified by reserved `model: "harness"`, plus a visible sentinel; persisted `[compaction]` tokens without a producer still freeze) between them so the prompt head can remain in the provider KV cache. Model-emitted copies of the spacer are not frozen. Spacer-only model replies are incomplete: ChatDirector nudges, then falls through loop-protection, workflow-idle, and open-task rails rather than empty-settling with work still open. The governor covers three cases: +When a cycle's input tokens cross a threshold, the director compacts the inference-facing history (the full run is always retained in the context store). The threshold is **model-aware** — roughly 60% of the active model's real context window — so small-window models compact early enough to avoid provider context-overflow while large-window models do not compact prematurely. The compacted prefix is **append-only across passes**: the existing compacted user turn stays byte-identical; new folds become later summary turns with a harness-inserted assistant spacer (identified by reserved `model: "harness"`, plus a visible sentinel; persisted `[compaction]` tokens without a producer still freeze) between them so the prompt head can remain in the provider KV cache. Model-emitted copies of the spacer are not frozen. Spacer-only model replies are incomplete: ChatDirector nudges, then falls through loop-protection, workflow-idle, and open-task rails rather than empty-settling with work still open. The governor covers four cases: - **Threshold at a tool pause** — Once over threshold, the follow-up `infer` after a tool batch is swapped for a `compact` cycle, and inference resumes via a host continuation message. After a compact that remains over the high watermark, the governor uses **growth hysteresis** (wait for usage to grow by ~10% of the window) instead of re-arming on every cycle; dropping under 60% is not required. - **Idle (end-of-turn)** — An interactive turn can end with a reply and then sit idle with no tool batch to intercept; the governor requests a continuation at that pause and compacts when it arrives. An operator message that races the continuation still compacts first, then re-enters inference to answer it. -- **Idle recompress past the provider TTL** — The fold is a re-compress, not a cache play: provider KV caches expire on their own schedule, so an under-threshold session compresses _after_ that expiry (the next turn is a cheaper write and later reads compound on the shrunk context) instead of only at 60% tokens. Any live re-entry past the window fires it — an empty stall ping re-enters meter-only, a raced operator message re-infers — under reason `cache-ttl-recompress` through the same compactor, so the fresh tail stays raw exactly as in the threshold path. The interval follows provider cache economics (`src/provider/cache-ttl.ts`), not a global N minutes: +- **Idle recompress past the provider TTL** — The fold is a re-compress, not a cache play: provider KV caches expire on their own schedule, so a session compresses _after_ that expiry (the next turn is a cheaper write and later reads compound on the shrunk context) instead of only at 60% tokens. It is not limited to under-threshold sessions: once growth hysteresis has cleared `pending`, an over-threshold session in that gap can fire the same fold. Any live re-entry past the window fires it — an empty stall ping re-enters meter-only, a raced operator message re-infers — under reason `cache-ttl-recompress` through the same compactor, so the fresh tail stays raw exactly as in the threshold path. The interval follows provider cache economics (`src/provider/cache-ttl.ts`), not a global N minutes. - **Overflow recovery** — A `context_overflow` inference error would otherwise become a terminal error reply; the governor compacts and retries instead, bounded so a history the compactor cannot shrink does not loop forever. Overflow ignores hysteresis for the compact itself. | Provider segment | Idle recompress allowed after | Why | diff --git a/src/agent/compaction.test.ts b/src/agent/compaction.test.ts index 5f8ca921b..86f1b184c 100644 --- a/src/agent/compaction.test.ts +++ b/src/agent/compaction.test.ts @@ -950,6 +950,50 @@ describe("provider-aware idle recompress (CL-8745)", () => { ).toEqual(ttlCompact); }); + test("after threshold compact and gap TTL, growth-armed compact stays blocked until a tool_call", () => { + // Threshold compact (consecutive=1) plus TTL fire in the hysteresis gap + // (consecutive=2) fills the shared cap. Later growth that would re-arm + // the threshold path stays blocked until a tool_call occupancy resets it. + let nowMs = 80_000_000; + const governor = createCompactionGovernor( + () => undefined, + "", + [], + () => nowMs, + ); + governor.noteInferenceDone(inferenceDone(overThreshold), tenTurns); + expect( + governor.interceptActions(toolDone(), inferAction, capabilities), + ).not.toBeNull(); + expect(governor.resumeAfterCompact(emptyMessage())).toBe("infer"); + + governor.noteInferenceDone(inferenceDone(overThreshold), tenTurns); + nowMs += 11 * MINUTE_MS; + expect( + governor.interceptIdleContinuation(emptyMessage(), capabilities), + ).toEqual(ttlCompact); + expect(governor.resumeAfterCompact(emptyMessage())).toBe("meter"); + + // Post-TTL snapshot, then growth past resumeDelta — pending re-arms, + // but the cap is full so interceptActions stays null. + governor.noteInferenceDone(inferenceDone(overThreshold), tenTurns); + governor.noteInferenceDone( + inferenceDone(overThreshold + resumeDelta), + tenTurns, + ); + expect( + governor.interceptActions(toolDone(), inferAction, capabilities), + ).toBeNull(); + + governor.noteInferenceDone( + inferenceDoneWithTools(overThreshold + 2 * resumeDelta), + tenTurns, + ); + expect( + governor.interceptActions(toolDone(), inferAction, capabilities), + ).not.toBeNull(); + }); + test("does not fold under an outstanding tool batch, fires once it settles", () => { let continuations = 0; let nowMs = 40_000_000; diff --git a/src/agent/compaction.ts b/src/agent/compaction.ts index 6e312640b..b8c3ef23d 100644 --- a/src/agent/compaction.ts +++ b/src/agent/compaction.ts @@ -317,6 +317,29 @@ export function createCompactionGovernor( return true; } + function inboundText(event: ReactorInboundEvent): string { + if (event.type !== "message.received") return ""; + return typeof event.message.content === "string" + ? event.message.content + : ""; + } + + // Idle empty compact needs a meter-only re-entry; a raced operator message + // needs a follow-up infer. Shared by the threshold idle path and TTL fold. + function issueIdleFold( + content: string, + capabilities: ReactorCapabilities, + reason: string, + ): ReactorAction[] { + if (content.length > 0) postCompactInfer = true; + else postCompactMeter = true; + issueThresholdCompact(); + return [ + capabilities.compact(COMPACTOR_NAME, reason), + ...continuationActions(capabilities), + ]; + } + function interceptIdleContinuation( event: ReactorInboundEvent, capabilities: ReactorCapabilities, @@ -329,40 +352,26 @@ export function createCompactionGovernor( } idlePending = false; pending = false; - const content = - typeof event.message.content === "string" ? event.message.content : ""; // The reactor delivers no event after compact, so always request a // continuation to re-enter decide against the shrunk turns: // - raced operator content → re-infer to answer it // - empty synthetic continuation → meter-only sync (no infer) - if (content.length > 0) { - postCompactInfer = true; - } else { - postCompactMeter = true; - } - issueThresholdCompact(); - return [ - capabilities.compact(COMPACTOR_NAME, "context-threshold"), - ...continuationActions(capabilities), - ]; + return issueIdleFold( + inboundText(event), + capabilities, + "context-threshold", + ); } // Unarmed idle re-entry past the provider TTL: same fold, same // keep-recent tail, same cap — but a "cache-ttl-recompress" reason so the // fold is attributable. No arming: every live re-entry re-checks the // window, so a sub-agent stall ping or operator message is the trigger. if (!isTtlRecompressDue(now())) return null; - const content = - typeof event.message.content === "string" ? event.message.content : ""; - if (content.length > 0) { - postCompactInfer = true; - } else { - postCompactMeter = true; - } - issueThresholdCompact(); - return [ - capabilities.compact(COMPACTOR_NAME, "cache-ttl-recompress"), - ...continuationActions(capabilities), - ]; + return issueIdleFold( + inboundText(event), + capabilities, + "cache-ttl-recompress", + ); } // A context-overflow inference error would otherwise terminate the loop @@ -396,9 +405,7 @@ export function createCompactionGovernor( event: ReactorInboundEvent, ): "infer" | "meter" | null { if (event.type !== "message.received") return null; - const content = - typeof event.message.content === "string" ? event.message.content : ""; - if (content.length > 0) return null; + if (inboundText(event).length > 0) return null; if (postCompactInfer) { postCompactInfer = false; return "infer"; diff --git a/src/provider/cache-ttl.test.ts b/src/provider/cache-ttl.test.ts index f3d1a4bf9..c11dc0780 100644 --- a/src/provider/cache-ttl.test.ts +++ b/src/provider/cache-ttl.test.ts @@ -21,7 +21,7 @@ describe("cacheTtlMsFor", () => { ); }); - test("disables idle recompress for local inference and unknown models", () => { + test("disables idle recompress for local inference and missing model ids", () => { expect(cacheTtlMsFor("ollama/llama3.1")).toBeUndefined(); expect(cacheTtlMsFor(undefined)).toBeUndefined(); expect(cacheTtlMsFor("")).toBeUndefined(); @@ -41,5 +41,8 @@ describe("cacheTtlMsFor", () => { test("assumes OpenAI-style economics for unrecognized providers", () => { expect(cacheTtlMsFor("bifrost/some-model")).toBe(10 * MINUTE_MS); expect(cacheTtlMsFor("totally-new-provider/model-x")).toBe(10 * MINUTE_MS); + // A truly unknown id (no provider segment, no family substring) still + // gets the 10-minute default — not the local-inference disable. + expect(cacheTtlMsFor("unknown-id")).toBe(10 * MINUTE_MS); }); }); From 18f315fb004e979128a48144e7da46ae3e1b5aa8 Mon Sep 17 00:00:00 2001 From: Sawyer Date: Mon, 21 Sep 2026 17:14:12 -0700 Subject: [PATCH 4/6] fix(compaction): key cache ttl off last-cycle source identity Production LastCycleSource stamps a bare model and openai-compatible provider. Keying TTL only off model treated local Ollama as a 10-minute cache window. --- docs/ARCHITECTURE.md | 5 ++-- src/agent/compaction.test.ts | 39 +++++++++++++++++++++++-- src/agent/compaction.ts | 12 ++++++-- src/provider/cache-ttl.test.ts | 25 ++++++++++++++++ src/provider/cache-ttl.ts | 52 +++++++++++++++++++++++++++++----- 5 files changed, 119 insertions(+), 14 deletions(-) diff --git a/docs/ARCHITECTURE.md b/docs/ARCHITECTURE.md index e661dde30..6c78e4f7d 100644 --- a/docs/ARCHITECTURE.md +++ b/docs/ARCHITECTURE.md @@ -191,8 +191,7 @@ When a cycle's input tokens cross a threshold, the director compacts the inferen - **Threshold at a tool pause** — Once over threshold, the follow-up `infer` after a tool batch is swapped for a `compact` cycle, and inference resumes via a host continuation message. After a compact that remains over the high watermark, the governor uses **growth hysteresis** (wait for usage to grow by ~10% of the window) instead of re-arming on every cycle; dropping under 60% is not required. - **Idle (end-of-turn)** — An interactive turn can end with a reply and then sit idle with no tool batch to intercept; the governor requests a continuation at that pause and compacts when it arrives. An operator message that races the continuation still compacts first, then re-enters inference to answer it. -- **Idle recompress past the provider TTL** — The fold is a re-compress, not a cache play: provider KV caches expire on their own schedule, so a session compresses _after_ that expiry (the next turn is a cheaper write and later reads compound on the shrunk context) instead of only at 60% tokens. It is not limited to under-threshold sessions: once growth hysteresis has cleared `pending`, an over-threshold session in that gap can fire the same fold. Any live re-entry past the window fires it — an empty stall ping re-enters meter-only, a raced operator message re-infers — under reason `cache-ttl-recompress` through the same compactor, so the fresh tail stays raw exactly as in the threshold path. The interval follows provider cache economics (`src/provider/cache-ttl.ts`), not a global N minutes. -- **Overflow recovery** — A `context_overflow` inference error would otherwise become a terminal error reply; the governor compacts and retries instead, bounded so a history the compactor cannot shrink does not loop forever. Overflow ignores hysteresis for the compact itself. +- **Idle recompress past the provider TTL** — The fold is a re-compress, not a cache play: provider KV caches expire on their own schedule, so a session compresses _after_ that expiry (the next turn is a cheaper write and later reads compound on the shrunk context) instead of only at 60% tokens. It is not limited to under-threshold sessions: once growth hysteresis has cleared `pending`, an over-threshold session in that gap can fire the same fold. Any live re-entry past the window fires it — an empty stall ping re-enters meter-only, a raced operator message re-infers — under reason `cache-ttl-recompress` through the same compactor, so the fresh tail stays raw exactly as in the threshold path. The interval follows provider cache economics (`src/provider/cache-ttl.ts`), not a global N minutes: | Provider segment | Idle recompress allowed after | Why | | ----------------------------------------------------------------- | ----------------------------- | ------------------------------------------------------------------------------------ | @@ -203,6 +202,8 @@ When a cycle's input tokens cross a threshold, the director compacts the inferen | `ollama` | never | Local inference has no remote cache to expire | | anything else | 10 min | Assumed OpenAI-style in-memory economics; tune per upstream | +- **Overflow recovery** — A `context_overflow` inference error would otherwise become a terminal error reply; the governor compacts and retries instead, bounded so a history the compactor cannot shrink does not loop forever. Overflow ignores hysteresis for the compact itself. + The TTL window is measured from the last `inference.done` (every inference rewrites the provider's prefix cache) with one fire per window per compact of any kind, and it never fires with a tool batch outstanding — stall pings can arrive mid-work. The consecutive-compact cap shared with the threshold path still bounds compact→infer→compact. Prompt-caching the summary call itself is an explicit non-goal: the fold carries no cache options. The compaction control flow is shaped by a reactor invariant: a `compact` action runs in its own cycle (it cannot be paired with `infer`), and **the reactor delivers no event after a compact cycle**. A director that simply emitted `compact` in place of the follow-up `infer` would leave the loop idle forever — the cause of an earlier stall. Instead the governor pairs `compact` with an emit action (`custom.compaction.continue`), and the host answers that emission by delivering a content-less inbound message (`buildCompactionContinuationMessage()`) through the serial operation queue, generation-guarded like every other deliver. A hop superseded because interrupt already bumped generation and enqueued a rebuild is re-queued onto the replacement agent rather than consumed on the outgoing liveAgent. That message adds no turn (`createInboundTurn` returns `null` for empty content) but re-enters the loop, where the director issues the follow-up `infer` against the freshly truncated history. Each emission is answered at most once (a replayed duplicate of an already-answered emission is ignored), and an unsolicited continuation with neither resume flag set answers `wait` rather than burning a billable inference. The legacy `requestContinuation` closure survives only on the sub-agent path, which delivers directly with no host emit hop; idle arming keeps the two channels exclusive (closure fires and the arming flag stays false, or no closure and the flag reports the arming) so a caller honoring both cannot double-deliver. diff --git a/src/agent/compaction.test.ts b/src/agent/compaction.test.ts index 86f1b184c..46aed9bc5 100644 --- a/src/agent/compaction.test.ts +++ b/src/agent/compaction.test.ts @@ -791,9 +791,15 @@ describe("provider-aware idle recompress (CL-8745)", () => { const MINUTE_MS = 60_000; function ttlInferenceDone( - model: string, + modelOrSource: + | string + | { sourceId: string; provider: string; model: string }, withTools: boolean, ): Extract { + const source = + typeof modelOrSource === "string" + ? { sourceId: "s", provider: "p", model: modelOrSource } + : modelOrSource; return { type: "inference.done", turn: { @@ -810,7 +816,7 @@ describe("provider-aware idle recompress (CL-8745)", () => { : [{ type: "text", text: "ok" }], }, usage: usage(1000), - source: { sourceId: "s", provider: "p", model }, + source, } as unknown as Extract; } @@ -894,6 +900,35 @@ describe("provider-aware idle recompress (CL-8745)", () => { ).toBeNull(); }); + test("production ollama LastCycleSource never fires cache-ttl-recompress", () => { + // Harness stamps { sourceId, provider, model } with a bare model. Ollama is + // buildOpenAISource: sourceId "ollama/default", provider openai-compatible, + // model llama3. Keying TTL only off model would take the 10-minute default. + let nowMs = 25_000_000; + const governor = createCompactionGovernor( + () => undefined, + "", + [], + () => nowMs, + ); + governor.noteInferenceDone( + ttlInferenceDone( + { + sourceId: "ollama/default", + provider: "openai-compatible", + model: "llama3", + }, + false, + ), + tenTurns, + ); + + nowMs += 11 * MINUTE_MS; + expect( + governor.interceptIdleContinuation(emptyMessage(), capabilities), + ).toBeNull(); + }); + test("stays inert with no observed cache write", () => { const governor = createCompactionGovernor(() => undefined); expect( diff --git a/src/agent/compaction.ts b/src/agent/compaction.ts index b8c3ef23d..13e378fb9 100644 --- a/src/agent/compaction.ts +++ b/src/agent/compaction.ts @@ -1,5 +1,6 @@ import type { ConversationTurn, + LastCycleSource, ReactorAction, ReactorCapabilities, ReactorInboundEvent, @@ -81,9 +82,13 @@ export function createCompactionGovernor( // can flag the number as approximate instead of implying provider-grade // precision. let usingEstimate = false; - // Model of the last inference.done turn, kept for live re-checks between - // inference cycles (see interceptActions) where the event carries no model. + // Last-cycle source of the last inference.done, kept for live re-checks + // between inference cycles (see interceptActions) where the event carries + // no source. Threshold sizing still keys off `model`; TTL identity needs + // `sourceId` / `provider` as well — production LastCycleSource stamps a + // bare model, and Ollama is `openai-compatible` with id `ollama/…`. let lastModel: string | undefined; + let lastCycleSource: LastCycleSource | undefined; let turnCount = 0; // Wall-clock of the last inference.done: the provider (re)wrote its prefix // cache for this session on that turn, so the provider TTL window in @@ -189,6 +194,7 @@ export function createCompactionGovernor( consecutiveThresholdCompacts = 0; } syncFromTurns(turns); + lastCycleSource = event.source; lastModel = event.source?.model; lastCacheWriteAt = now(); // The terminal reply ends the previous tool batch (its results are @@ -308,7 +314,7 @@ export function createCompactionGovernor( if (turnCount <= MIN_TURNS_TO_COMPACT) return false; if (atThresholdCompactCap()) return false; if (outstandingToolCalls > 0) return false; - const ttl = cacheTtlMsFor(lastModel); + const ttl = cacheTtlMsFor(lastCycleSource); if (ttl === undefined) return false; if (lastCacheWriteAt === undefined) return false; if (nowMs - lastCacheWriteAt < ttl) return false; diff --git a/src/provider/cache-ttl.test.ts b/src/provider/cache-ttl.test.ts index c11dc0780..0bf58fbf4 100644 --- a/src/provider/cache-ttl.test.ts +++ b/src/provider/cache-ttl.test.ts @@ -27,6 +27,31 @@ describe("cacheTtlMsFor", () => { expect(cacheTtlMsFor("")).toBeUndefined(); }); + test("keys ollama off production LastCycleSource, not a slash-form model", () => { + expect( + cacheTtlMsFor({ + sourceId: "ollama/default", + provider: "openai-compatible", + model: "llama3", + }), + ).toBeUndefined(); + }); + + test("maps bare LastCycleSource ids through provider and family", () => { + expect( + cacheTtlMsFor({ + provider: "anthropic", + model: "claude-opus-4-6", + }), + ).toBe(5 * MINUTE_MS); + expect( + cacheTtlMsFor({ + provider: "codex-responses", + model: "gpt-5.6-luna", + }), + ).toBe(10 * MINUTE_MS); + }); + test("falls back to model family for unrecognized provider prefixes", () => { expect(cacheTtlMsFor("proxy-acme/grok-4")).toBe(10 * MINUTE_MS); expect(cacheTtlMsFor("proxy-acme/gemini-3-pro")).toBe(15 * MINUTE_MS); diff --git a/src/provider/cache-ttl.ts b/src/provider/cache-ttl.ts index 7b22da700..a3803b52f 100644 --- a/src/provider/cache-ttl.ts +++ b/src/provider/cache-ttl.ts @@ -23,6 +23,11 @@ // 60 min. // - ollama: local inference has no remote cache to expire, so an idle // recompress would burn local compute for no cache benefit. Disabled. +// Production LastCycleSource is `{ sourceId, provider, model }` with a +// bare model; Ollama is `buildOpenAISource` (`provider: openai-compatible`, +// `sourceId` like `ollama/default`, model `llama3` / `qwen3`). Identity is +// `sourceId` via `isOllamaProviderId`, not a slash-form string the harness +// never stamps. // - Unknown/custom providers (bifrost proxy, `openai-compatible` fronting an // unlisted upstream, unrecognized strings): assumed OpenAI-style in-memory // economics. 10 min. @@ -30,6 +35,8 @@ // Deliberate non-goal: the summary call itself carries no `cache_control` — // prompt-caching the fold is not attempted. +import { isOllamaProviderId } from "./ollama.js"; + const MINUTE_MS = 60_000; const PROVIDER_TTLS_MS: Record = { @@ -65,6 +72,12 @@ const FAMILY_TTLS_MS: readonly (readonly [string, number | undefined])[] = [ ["ollama", undefined], ]; +type CacheTtlIdentity = { + sourceId?: string; + provider?: string; + model?: string; +}; + // Canonical provider segment of a `provider:model` string: the account or // adapter name before the first "/" (custom names like `xai/thegreataxios` // carry the provider there), else the head before ":". Mirrors the @@ -78,13 +91,7 @@ function canonicalSegment(model: string): string { return colon >= 0 ? head.slice(0, colon) : head; } -/** - * Milliseconds of provider-cache idle after which a recompress is allowed, - * or `undefined` when the model is undefined/empty or the provider has no - * remote cache (local inference). Never throws; unknown strings get the - * default 10-minute window. - */ -export function cacheTtlMsFor(model: string | undefined): number | undefined { +function ttlForModelString(model: string | undefined): number | undefined { if (model === undefined) return undefined; const segment = canonicalSegment(model); if (segment.length === 0) return undefined; @@ -96,3 +103,34 @@ export function cacheTtlMsFor(model: string | undefined): number | undefined { } return DEFAULT_TTL_MS; } + +/** + * Milliseconds of provider-cache idle after which a recompress is allowed, + * or `undefined` when the identity is missing/empty or the provider has no + * remote cache (local inference). Never throws; unknown strings get the + * default 10-minute window. + * + * Accepts a slash-form `provider/model` string or a LastCycleSource-shaped + * identity. Production events stamp a bare `model`; Ollama is recognized + * from `sourceId` (`isOllamaProviderId`), not from a combined string the + * harness never stamps. + */ +export function cacheTtlMsFor( + identity: string | CacheTtlIdentity | undefined, +): number | undefined { + if (identity === undefined) return undefined; + if (typeof identity === "string") return ttlForModelString(identity); + + if (identity.sourceId !== undefined && isOllamaProviderId(identity.sourceId)) + return undefined; + if (identity.provider !== undefined && isOllamaProviderId(identity.provider)) + return undefined; + + if (identity.provider !== undefined) { + const provider = identity.provider.toLowerCase(); + if (Object.hasOwn(PROVIDER_TTLS_MS, provider)) + return PROVIDER_TTLS_MS[provider]; + } + + return ttlForModelString(identity.model); +} From f0057e48c9465cfc0d9113e1e29a1a298c28f047 Mon Sep 17 00:00:00 2001 From: Sawyer Date: Mon, 21 Sep 2026 17:26:28 -0700 Subject: [PATCH 5/6] test(compaction): cover production last-cycle source ttl windows --- docs/ARCHITECTURE.md | 2 ++ src/agent/compaction.test.ts | 51 +++++++++++++++++++++++++++++++++--- 2 files changed, 50 insertions(+), 3 deletions(-) diff --git a/docs/ARCHITECTURE.md b/docs/ARCHITECTURE.md index 6c78e4f7d..9cc10db40 100644 --- a/docs/ARCHITECTURE.md +++ b/docs/ARCHITECTURE.md @@ -202,6 +202,8 @@ When a cycle's input tokens cross a threshold, the director compacts the inferen | `ollama` | never | Local inference has no remote cache to expire | | anything else | 10 min | Assumed OpenAI-style in-memory economics; tune per upstream | +Production identity for the `ollama` row is `sourceId` `ollama` / `ollama/` via `isOllamaProviderId` (`src/provider/ollama.ts`), not a bare model id: Ollama is `buildOpenAISource` (`provider: openai-compatible`, model `llama3` / `qwen3`). Slash-form `ollama/…` is a table key, not what the harness stamps. + - **Overflow recovery** — A `context_overflow` inference error would otherwise become a terminal error reply; the governor compacts and retries instead, bounded so a history the compactor cannot shrink does not loop forever. Overflow ignores hysteresis for the compact itself. The TTL window is measured from the last `inference.done` (every inference rewrites the provider's prefix cache) with one fire per window per compact of any kind, and it never fires with a tool batch outstanding — stall pings can arrive mid-work. The consecutive-compact cap shared with the threshold path still bounds compact→infer→compact. Prompt-caching the summary call itself is an explicit non-goal: the fold carries no cache options. diff --git a/src/agent/compaction.test.ts b/src/agent/compaction.test.ts index 46aed9bc5..f94738334 100644 --- a/src/agent/compaction.test.ts +++ b/src/agent/compaction.test.ts @@ -793,13 +793,17 @@ describe("provider-aware idle recompress (CL-8745)", () => { function ttlInferenceDone( modelOrSource: | string - | { sourceId: string; provider: string; model: string }, + | { sourceId?: string; provider?: string; model?: string }, withTools: boolean, ): Extract { const source = typeof modelOrSource === "string" ? { sourceId: "s", provider: "p", model: modelOrSource } - : modelOrSource; + : { + sourceId: modelOrSource.sourceId ?? "s", + provider: modelOrSource.provider ?? "p", + model: modelOrSource.model ?? "m", + }; return { type: "inference.done", turn: { @@ -845,7 +849,10 @@ describe("provider-aware idle recompress (CL-8745)", () => { () => nowMs, ); governor.noteInferenceDone( - ttlInferenceDone("anthropic/claude-opus-4-6", false), + ttlInferenceDone( + { provider: "anthropic", model: "claude-opus-4-6" }, + false, + ), tenTurns, ); @@ -866,6 +873,44 @@ describe("provider-aware idle recompress (CL-8745)", () => { expect(governor.resumeAfterCompact(emptyMessage())).toBe("meter"); }); + test("production LastCycleSource: anthropic fires at 5m, codex stays quiet until 10m", () => { + // Harness stamps { sourceId, provider, model } with a bare model, not + // slash-form "anthropic/claude-opus-4-6". Anthropic's 5-minute window + // must come from provider, not a dummy model string; Codex must not + // inherit that 5-minute fire from sourceId "codex/work". + let nowMs = 15_000_000; + const clock = () => nowMs; + const anthropic = createCompactionGovernor(() => undefined, "", [], clock); + const codex = createCompactionGovernor(() => undefined, "", [], clock); + anthropic.noteInferenceDone( + ttlInferenceDone( + { provider: "anthropic", model: "claude-opus-4-6" }, + false, + ), + tenTurns, + ); + codex.noteInferenceDone( + ttlInferenceDone( + { sourceId: "codex/work", provider: "codex-responses" }, + false, + ), + tenTurns, + ); + + nowMs += 5 * MINUTE_MS + 1; + expect( + anthropic.interceptIdleContinuation(emptyMessage(), capabilities), + ).toEqual(ttlCompact); + expect( + codex.interceptIdleContinuation(emptyMessage(), capabilities), + ).toBeNull(); + + nowMs += 5 * MINUTE_MS; + expect( + codex.interceptIdleContinuation(emptyMessage(), capabilities), + ).toEqual(ttlCompact); + }); + test("follows provider economics: deepseek waits out its long window, ollama never fires", () => { let nowMs = 20_000_000; const clock = () => nowMs; From 3b718f9a951cd4997b3f0262f051ec4ae2044149 Mon Sep 17 00:00:00 2001 From: Sawyer Date: Mon, 21 Sep 2026 17:37:19 -0700 Subject: [PATCH 6/6] docs(compaction): correct ollama ttl identity vs table key The table key is bare ollama; the harness stamps slash-form on sourceId. Stamp a production-looking Codex model on the governor fixture instead of the helper default. --- docs/ARCHITECTURE.md | 2 +- src/agent/compaction.test.ts | 6 +++++- 2 files changed, 6 insertions(+), 2 deletions(-) diff --git a/docs/ARCHITECTURE.md b/docs/ARCHITECTURE.md index 9cc10db40..139c12e9a 100644 --- a/docs/ARCHITECTURE.md +++ b/docs/ARCHITECTURE.md @@ -202,7 +202,7 @@ When a cycle's input tokens cross a threshold, the director compacts the inferen | `ollama` | never | Local inference has no remote cache to expire | | anything else | 10 min | Assumed OpenAI-style in-memory economics; tune per upstream | -Production identity for the `ollama` row is `sourceId` `ollama` / `ollama/` via `isOllamaProviderId` (`src/provider/ollama.ts`), not a bare model id: Ollama is `buildOpenAISource` (`provider: openai-compatible`, model `llama3` / `qwen3`). Slash-form `ollama/…` is a table key, not what the harness stamps. +The `ollama` table key is the bare provider segment; the harness stamps slash-form on `sourceId` (`ollama/default` / `ollama/`), with `provider: openai-compatible` (`buildOpenAISource`) and a bare model (`llama3` / `qwen3`). Idle-recompress disable matches via `isOllamaProviderId` on that `sourceId` (`src/provider/ollama.ts`), not by looking up slash-form as a table key. - **Overflow recovery** — A `context_overflow` inference error would otherwise become a terminal error reply; the governor compacts and retries instead, bounded so a history the compactor cannot shrink does not loop forever. Overflow ignores hysteresis for the compact itself. diff --git a/src/agent/compaction.test.ts b/src/agent/compaction.test.ts index f94738334..1d86af73d 100644 --- a/src/agent/compaction.test.ts +++ b/src/agent/compaction.test.ts @@ -891,7 +891,11 @@ describe("provider-aware idle recompress (CL-8745)", () => { ); codex.noteInferenceDone( ttlInferenceDone( - { sourceId: "codex/work", provider: "codex-responses" }, + { + sourceId: "codex/work", + provider: "codex-responses", + model: "gpt-5.6-luna", + }, false, ), tenTurns,