Skip to content

Commit d3780b4

Browse files
committed
fix(session): keep cache expiry from compacting stored turns
The settings fixture covers anthropicCachePrompt. Unused governor clocks are gone.
1 parent 7ba4087 commit d3780b4

4 files changed

Lines changed: 7 additions & 34 deletions

File tree

‎docs/ARCHITECTURE.md‎

Lines changed: 1 addition & 8 deletions
Original file line numberDiff line numberDiff line change
@@ -191,20 +191,13 @@ When a cycle's input tokens cross a threshold, the director compacts the inferen
191191

192192
- **Threshold at a tool pause** — Once over threshold, the follow-up `infer` after a tool batch is swapped for a `compact` cycle, and inference resumes via a host continuation message. After a compact that remains over the high watermark, the governor uses **growth hysteresis** (wait for usage to grow by ~10% of the window) instead of re-arming on every cycle; dropping under 60% is not required.
193193
- **Idle (end-of-turn)** — An interactive turn can end with a reply and then sit idle with no tool batch to intercept; the governor requests a continuation at that pause and compacts when it arrives. An operator message that races the continuation still compacts first, then re-enters inference to answer it.
194-
- **Idle recompress past the provider TTL** — The fold is a re-compress, not a cache play: provider KV caches expire on their own schedule, so a session compresses _after_ that expiry (the next turn is a cheaper write and later reads compound on the shrunk context) instead of only at 60% tokens. It is not limited to under-threshold sessions: once growth hysteresis has cleared `pending`, an over-threshold session in that gap can fire the same fold. Any live re-entry past the window fires it — an empty stall ping re-enters meter-only, a raced operator message re-infers — under reason `cache-ttl-recompress` through the same compactor, so the fresh tail stays raw exactly as in the threshold path. The interval follows provider cache economics (`src/provider/cache-ttl.ts`), not a global N minutes:
195-
196-
| Provider segment | Idle recompress allowed after | Why |
197-
| -------------------------------------------------------------- | ----------------------------- | ---------------------------------------------------------------------------------------------------------------- |
198-
| `anthropic`, `zen-messages`, `opencode-go-messages` | 5 min | Published default ephemeral cache TTL; the 1-hour TTL is never set |
199-
| anything else, including OpenAI, xAI, Gemini, DeepSeek, ollama | never | No published 5-minute expiry, or no remote cache. Guessing a shorter window folds a cache that may still be warm |
194+
- **Idle recompress past the provider TTL** — Cache expiry does not compact. Stored turns stay. When settings set `anthropicCachePrompt`, and the saved Anthropic-protocol cache write is at least 5 minutes old, the pre-inference context transform stubs old tool-result bodies on the outgoing prompt only. `turns.jsonl` is not rewritten. The setting defaults off.
200195

201196
The `ollama` table key is the bare provider segment; the harness stamps slash-form on `sourceId` (`ollama/default` / `ollama/<instance>`), with `provider: openai-compatible` (`buildOpenAISource`) and a bare model (`llama3` / `qwen3`). Idle-recompress disable matches via `isOllamaProviderId` on that `sourceId` (`src/provider/ollama.ts`), not by looking up slash-form as a table key.
202197

203198
- **Operator `/compact`** — A first-class slash command folds now without waiting for the occupancy governor. Optional trailing instructions go to the summarizer (empty uses the default structured fold) and are stored on the compact record so later auto-folds still see them. Director rebuild (`/model`, interrupt, resume) restores those instructions from the latest compact record's parameters. An in-flight turn uses the same pair-safe compact-then-continue hop as threshold compact; an idle compact shows the fold and does not start a new turn.
204199
- **Overflow recovery** — A `context_overflow` inference error would otherwise become a terminal error reply; the governor compacts and retries instead, bounded so a history the compactor cannot shrink does not loop forever. Overflow ignores hysteresis for the compact itself.
205200

206-
The TTL window is measured from the last `inference.done` (every inference rewrites the provider's prefix cache) with one fire per window per compact of any kind, and it never fires with a tool batch outstanding — stall pings can arrive mid-work. The consecutive-compact cap shared with the threshold path still bounds compact→infer→compact. Prompt-caching the summary call itself is an explicit non-goal: the fold carries no cache options. That write time is stored on the session run record for Anthropic-protocol identities only. A new process restores it and, once the stamp is at least 5 minutes old and the stored turn count is above the compact floor, issues the same `cache-ttl-recompress` fold before the first infer so the request reads the trimmed turns.
207-
208201
The compaction control flow is shaped by a reactor invariant: a `compact` action runs in its own cycle (it cannot be paired with `infer`), and **the reactor delivers no event after a compact cycle**. A director that simply emitted `compact` in place of the follow-up `infer` would leave the loop idle forever — the cause of an earlier stall. Instead the governor pairs `compact` with an emit action (`custom.compaction.continue`), and the host answers that emission by delivering a content-less inbound message (`buildCompactionContinuationMessage()`) through the serial operation queue, generation-guarded like every other deliver. A hop superseded because interrupt already bumped generation and enqueued a rebuild is re-queued onto the replacement agent rather than consumed on the outgoing liveAgent. That message adds no turn (`createInboundTurn` returns `null` for empty content) but re-enters the loop, where the director issues the follow-up `infer` against the freshly truncated history. Each emission is answered at most once (a replayed duplicate of an already-answered emission is ignored), and an unsolicited continuation with neither resume flag set answers `wait` rather than burning a billable inference. The legacy `requestContinuation` closure survives only on the sub-agent path, which delivers directly with no host emit hop; idle arming keeps the two channels exclusive (closure fires and the arming flag stays false, or no closure and the flag reports the arming) so a caller honoring both cannot double-deliver.
209202

210203
Compaction replaces older turns with a structured, workflow-aware summary rather than a stats blob: sections for **What Happened / What We're Doing / Relevant Links / Action Items / Next Steps**, with the active workflow and step woven in so compacting mid-`/build` or mid-`/plan` preserves the contract. The summary is produced by a one-shot model call; on any failure it falls back to a deterministic summary so a compaction cycle never breaks the session.

‎src/agent/compaction.test.ts‎

Lines changed: 0 additions & 8 deletions
Original file line numberDiff line numberDiff line change
@@ -984,14 +984,6 @@ describe("provider-aware idle recompress (CL-8745)", () => {
984984
} as ReactorInboundEvent;
985985
}
986986

987-
const ttlCompact = [
988-
{
989-
type: "compact",
990-
compactor: "pruning-compactor",
991-
reason: "cache-ttl-recompress",
992-
},
993-
] as ReactorAction[];
994-
995987
test("fires past the provider TTL while under threshold, meter-only on empty", () => {
996988
let continuations = 0;
997989
let nowMs = 10_000_000;

‎src/agent/compaction.ts‎

Lines changed: 5 additions & 18 deletions
Original file line numberDiff line numberDiff line change
@@ -135,7 +135,7 @@ export function createCompactionGovernor(
135135
requestContinuation?: () => void,
136136
systemPrompt = "",
137137
toolDefinitions: readonly ToolDefinition[] = [],
138-
now: () => number = Date.now,
138+
_now: () => number = Date.now,
139139
) {
140140
let pending = false;
141141
let idlePending = false;
@@ -163,16 +163,6 @@ export function createCompactionGovernor(
163163
let usingEstimate = false;
164164
let lastModel: string | undefined;
165165
let turnCount = 0;
166-
// Wall-clock of the last inference.done: the provider (re)wrote its prefix
167-
// cache for this session on that turn, so the provider TTL window in
168-
// provider/cache-ttl.ts is measured from here. Stamped on every
169-
// inference.done — estimated-usage providers wrote a cache entry too.
170-
let lastCacheWriteAt: number | undefined;
171-
// Wall-clock of the last issued compact of any kind. A fresh fold rewrites
172-
// the session prefix, so the TTL recompress must not fire again inside the
173-
// same window even when its own cache write has not been observed yet (the
174-
// summary call bypasses this governor).
175-
let lastCompactAt: number | undefined;
176166
// Tool calls issued by the last inference.done and not yet settled. A TTL
177167
// recompress must never fold while a batch is outstanding — the stall ping
178168
// that triggers it can arrive mid-work, and folding under it would rewrite
@@ -218,7 +208,6 @@ export function createCompactionGovernor(
218208

219209
function noteCompactIssued(): void {
220210
awaitingPostCompactMeasurement = true;
221-
lastCompactAt = now();
222211
}
223212

224213
function atThresholdCompactCap(): boolean {
@@ -278,7 +267,6 @@ export function createCompactionGovernor(
278267
}
279268
syncFromTurns(turns);
280269
lastModel = event.source?.model;
281-
lastCacheWriteAt = now();
282270
// The terminal reply ends the previous tool batch (its results are
283271
// already in the turns) and opens the batch the reply just issued. A TTL
284272
// recompress must never fold while a batch is outstanding — the stall
@@ -543,17 +531,16 @@ export function createCompactionGovernor(
543531
extraInstructions = trimmed;
544532
}
545533

546-
// A new process has no in-memory cache write. Seed the stored stamp, the
547-
// identity that wrote it, and the stored turns so the next message.received
548-
// can fold before its infer. Turn count is the stored set: the inbound
549-
// message is appended after this, and the floor is measured without it.
534+
// A new process has no in-memory turn count. Seed the stored turns so
535+
// threshold and manual compact see the resumed history. The cache-write
536+
// time lives on the run record for the prompt transform, not here.
550537
function restoreCacheWrite(args: {
551538
at: number;
552539
source: LastCycleSource;
553540
turns: readonly ConversationTurn[];
554541
}): void {
542+
void args.at;
555543
syncFromTurns(args.turns);
556-
lastCacheWriteAt = args.at;
557544
lastModel = args.source.model;
558545
}
559546

‎tests/unit/config.test.ts‎

Lines changed: 1 addition & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -324,6 +324,7 @@ test("loadSettings cannot silently drop a known optional key", async () => {
324324
recentModels: [{ provider: "p", model: "m" }],
325325
favoriteModels: [{ provider: "p", model: "m" }],
326326
dangerouslySkipPermissions: true,
327+
anthropicCachePrompt: true,
327328
};
328329
await writeFile(globalPath, JSON.stringify(fixture));
329330
const loaded = await loadSettings(globalPath);

0 commit comments

Comments
 (0)