Skip to content

Commit b4fab39

Browse files
feat(compaction): idle recompress past the provider cache ttl (#1151)
* feat(compaction): idle recompress past the provider cache ttl * test(compaction): characterize ttl fire in the hysteresis gap The TTL arm only defers while threshold arming is pending. After a compact that stays over the high watermark, growth hysteresis clears pending, so idle past the provider TTL still folds. Comments claimed otherwise; keep the behavior and document it. cacheTtlMsFor returns undefined only for missing, empty, or local models; unknown strings get the default window. * fix(compaction): align ttl nits with hysteresis-gap behavior * fix(compaction): key cache ttl off last-cycle source identity Production LastCycleSource stamps a bare model and openai-compatible provider. Keying TTL only off model treated local Ollama as a 10-minute cache window. * test(compaction): cover production last-cycle source ttl windows * docs(compaction): correct ollama ttl identity vs table key The table key is bare ollama; the harness stamps slash-form on sourceId. Stamp a production-looking Codex model on the governor fixture instead of the helper default.
1 parent 39849ac commit b4fab39

5 files changed

Lines changed: 732 additions & 27 deletions

File tree

‎docs/ARCHITECTURE.md‎

Lines changed: 16 additions & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -187,12 +187,27 @@ The agent maintains an optional **`manage_tasks`** list (create/update via the h
187187

188188
#### Context compaction (the compaction governor)
189189

190-
When a cycle's input tokens cross a threshold, the director compacts the inference-facing history (the full run is always retained in the context store). The threshold is **model-aware** — roughly 60% of the active model's real context window — so small-window models compact early enough to avoid provider context-overflow while large-window models do not compact prematurely. The compacted prefix is **append-only across passes**: the existing compacted user turn stays byte-identical; new folds become later summary turns with a harness-inserted assistant spacer (identified by reserved `model: "harness"`, plus a visible sentinel; persisted `[compaction]` tokens without a producer still freeze) between them so the prompt head can remain in the provider KV cache. Model-emitted copies of the spacer are not frozen. Spacer-only model replies are incomplete: ChatDirector nudges, then falls through loop-protection, workflow-idle, and open-task rails rather than empty-settling with work still open. The governor covers three cases:
190+
When a cycle's input tokens cross a threshold, the director compacts the inference-facing history (the full run is always retained in the context store). The threshold is **model-aware** — roughly 60% of the active model's real context window — so small-window models compact early enough to avoid provider context-overflow while large-window models do not compact prematurely. The compacted prefix is **append-only across passes**: the existing compacted user turn stays byte-identical; new folds become later summary turns with a harness-inserted assistant spacer (identified by reserved `model: "harness"`, plus a visible sentinel; persisted `[compaction]` tokens without a producer still freeze) between them so the prompt head can remain in the provider KV cache. Model-emitted copies of the spacer are not frozen. Spacer-only model replies are incomplete: ChatDirector nudges, then falls through loop-protection, workflow-idle, and open-task rails rather than empty-settling with work still open. The governor covers four cases:
191191

192192
- **Threshold at a tool pause** — Once over threshold, the follow-up `infer` after a tool batch is swapped for a `compact` cycle, and inference resumes via a host continuation message. After a compact that remains over the high watermark, the governor uses **growth hysteresis** (wait for usage to grow by ~10% of the window) instead of re-arming on every cycle; dropping under 60% is not required.
193193
- **Idle (end-of-turn)** — An interactive turn can end with a reply and then sit idle with no tool batch to intercept; the governor requests a continuation at that pause and compacts when it arrives. An operator message that races the continuation still compacts first, then re-enters inference to answer it.
194+
- **Idle recompress past the provider TTL** — The fold is a re-compress, not a cache play: provider KV caches expire on their own schedule, so a session compresses _after_ that expiry (the next turn is a cheaper write and later reads compound on the shrunk context) instead of only at 60% tokens. It is not limited to under-threshold sessions: once growth hysteresis has cleared `pending`, an over-threshold session in that gap can fire the same fold. Any live re-entry past the window fires it — an empty stall ping re-enters meter-only, a raced operator message re-infers — under reason `cache-ttl-recompress` through the same compactor, so the fresh tail stays raw exactly as in the threshold path. The interval follows provider cache economics (`src/provider/cache-ttl.ts`), not a global N minutes:
195+
196+
| Provider segment | Idle recompress allowed after | Why |
197+
| ----------------------------------------------------------------- | ----------------------------- | ------------------------------------------------------------------------------------ |
198+
| `anthropic`, `zen-messages`, `opencode-go-messages` | 5 min | Default ephemeral cache TTL; hour-long breakpoints are never set |
199+
| `openai-responses`, `codex-responses`, `openai-compatible`, `xai` | 10 min | In-memory prefixes typically evicted after 5–10 min idle; xAI assumed same economics |
200+
| `gemini` | 15 min | No published implicit-cache eviction window; conservative |
201+
| `deepseek` | 60 min | On-disk context cache persists for hours-to-days |
202+
| `ollama` | never | Local inference has no remote cache to expire |
203+
| anything else | 10 min | Assumed OpenAI-style in-memory economics; tune per upstream |
204+
205+
The `ollama` table key is the bare provider segment; the harness stamps slash-form on `sourceId` (`ollama/default` / `ollama/<instance>`), with `provider: openai-compatible` (`buildOpenAISource`) and a bare model (`llama3` / `qwen3`). Idle-recompress disable matches via `isOllamaProviderId` on that `sourceId` (`src/provider/ollama.ts`), not by looking up slash-form as a table key.
206+
194207
- **Overflow recovery** — A `context_overflow` inference error would otherwise become a terminal error reply; the governor compacts and retries instead, bounded so a history the compactor cannot shrink does not loop forever. Overflow ignores hysteresis for the compact itself.
195208

209+
The TTL window is measured from the last `inference.done` (every inference rewrites the provider's prefix cache) with one fire per window per compact of any kind, and it never fires with a tool batch outstanding — stall pings can arrive mid-work. The consecutive-compact cap shared with the threshold path still bounds compact→infer→compact. Prompt-caching the summary call itself is an explicit non-goal: the fold carries no cache options.
210+
196211
The compaction control flow is shaped by a reactor invariant: a `compact` action runs in its own cycle (it cannot be paired with `infer`), and **the reactor delivers no event after a compact cycle**. A director that simply emitted `compact` in place of the follow-up `infer` would leave the loop idle forever — the cause of an earlier stall. Instead the governor pairs `compact` with an emit action (`custom.compaction.continue`), and the host answers that emission by delivering a content-less inbound message (`buildCompactionContinuationMessage()`) through the serial operation queue, generation-guarded like every other deliver. A hop superseded because interrupt already bumped generation and enqueued a rebuild is re-queued onto the replacement agent rather than consumed on the outgoing liveAgent. That message adds no turn (`createInboundTurn` returns `null` for empty content) but re-enters the loop, where the director issues the follow-up `infer` against the freshly truncated history. Each emission is answered at most once (a replayed duplicate of an already-answered emission is ignored), and an unsolicited continuation with neither resume flag set answers `wait` rather than burning a billable inference. The legacy `requestContinuation` closure survives only on the sub-agent path, which delivers directly with no host emit hop; idle arming keeps the two channels exclusive (closure fires and the arming flag stays false, or no closure and the flag reports the arming) so a caller honoring both cannot double-deliver.
197212

198213
Compaction replaces older turns with a structured, workflow-aware summary rather than a stats blob: sections for **What Happened / What We're Doing / Relevant Links / Action Items / Next Steps**, with the active workflow and step woven in so compacting mid-`/build` or mid-`/plan` preserves the contract. The summary is produced by a one-shot model call; on any failure it falls back to a deterministic summary so a compaction cycle never breaks the session.

0 commit comments

Comments
 (0)