You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
feat(compaction): idle recompress past the provider cache ttl (#1151)
* feat(compaction): idle recompress past the provider cache ttl
* test(compaction): characterize ttl fire in the hysteresis gap
The TTL arm only defers while threshold arming is pending. After a
compact that stays over the high watermark, growth hysteresis clears
pending, so idle past the provider TTL still folds. Comments claimed
otherwise; keep the behavior and document it.
cacheTtlMsFor returns undefined only for missing, empty, or local
models; unknown strings get the default window.
* fix(compaction): align ttl nits with hysteresis-gap behavior
* fix(compaction): key cache ttl off last-cycle source identity
Production LastCycleSource stamps a bare model and openai-compatible
provider. Keying TTL only off model treated local Ollama as a 10-minute
cache window.
* test(compaction): cover production last-cycle source ttl windows
* docs(compaction): correct ollama ttl identity vs table key
The table key is bare ollama; the harness stamps slash-form on
sourceId. Stamp a production-looking Codex model on the governor
fixture instead of the helper default.
Copy file name to clipboardExpand all lines: docs/ARCHITECTURE.md
+16-1Lines changed: 16 additions & 1 deletion
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -187,12 +187,27 @@ The agent maintains an optional **`manage_tasks`** list (create/update via the h
187
187
188
188
#### Context compaction (the compaction governor)
189
189
190
-
When a cycle's input tokens cross a threshold, the director compacts the inference-facing history (the full run is always retained in the context store). The threshold is **model-aware** — roughly 60% of the active model's real context window — so small-window models compact early enough to avoid provider context-overflow while large-window models do not compact prematurely. The compacted prefix is **append-only across passes**: the existing compacted user turn stays byte-identical; new folds become later summary turns with a harness-inserted assistant spacer (identified by reserved `model: "harness"`, plus a visible sentinel; persisted `[compaction]` tokens without a producer still freeze) between them so the prompt head can remain in the provider KV cache. Model-emitted copies of the spacer are not frozen. Spacer-only model replies are incomplete: ChatDirector nudges, then falls through loop-protection, workflow-idle, and open-task rails rather than empty-settling with work still open. The governor covers three cases:
190
+
When a cycle's input tokens cross a threshold, the director compacts the inference-facing history (the full run is always retained in the context store). The threshold is **model-aware** — roughly 60% of the active model's real context window — so small-window models compact early enough to avoid provider context-overflow while large-window models do not compact prematurely. The compacted prefix is **append-only across passes**: the existing compacted user turn stays byte-identical; new folds become later summary turns with a harness-inserted assistant spacer (identified by reserved `model: "harness"`, plus a visible sentinel; persisted `[compaction]` tokens without a producer still freeze) between them so the prompt head can remain in the provider KV cache. Model-emitted copies of the spacer are not frozen. Spacer-only model replies are incomplete: ChatDirector nudges, then falls through loop-protection, workflow-idle, and open-task rails rather than empty-settling with work still open. The governor covers four cases:
191
191
192
192
-**Threshold at a tool pause** — Once over threshold, the follow-up `infer` after a tool batch is swapped for a `compact` cycle, and inference resumes via a host continuation message. After a compact that remains over the high watermark, the governor uses **growth hysteresis** (wait for usage to grow by ~10% of the window) instead of re-arming on every cycle; dropping under 60% is not required.
193
193
-**Idle (end-of-turn)** — An interactive turn can end with a reply and then sit idle with no tool batch to intercept; the governor requests a continuation at that pause and compacts when it arrives. An operator message that races the continuation still compacts first, then re-enters inference to answer it.
194
+
-**Idle recompress past the provider TTL** — The fold is a re-compress, not a cache play: provider KV caches expire on their own schedule, so a session compresses _after_ that expiry (the next turn is a cheaper write and later reads compound on the shrunk context) instead of only at 60% tokens. It is not limited to under-threshold sessions: once growth hysteresis has cleared `pending`, an over-threshold session in that gap can fire the same fold. Any live re-entry past the window fires it — an empty stall ping re-enters meter-only, a raced operator message re-infers — under reason `cache-ttl-recompress` through the same compactor, so the fresh tail stays raw exactly as in the threshold path. The interval follows provider cache economics (`src/provider/cache-ttl.ts`), not a global N minutes:
|`anthropic`, `zen-messages`, `opencode-go-messages`| 5 min | Default ephemeral cache TTL; hour-long breakpoints are never set |
199
+
|`openai-responses`, `codex-responses`, `openai-compatible`, `xai`| 10 min | In-memory prefixes typically evicted after 5–10 min idle; xAI assumed same economics |
200
+
|`gemini`| 15 min | No published implicit-cache eviction window; conservative |
201
+
|`deepseek`| 60 min | On-disk context cache persists for hours-to-days |
202
+
|`ollama`| never | Local inference has no remote cache to expire |
203
+
| anything else | 10 min | Assumed OpenAI-style in-memory economics; tune per upstream |
204
+
205
+
The `ollama` table key is the bare provider segment; the harness stamps slash-form on `sourceId` (`ollama/default` / `ollama/<instance>`), with `provider: openai-compatible` (`buildOpenAISource`) and a bare model (`llama3` / `qwen3`). Idle-recompress disable matches via `isOllamaProviderId` on that `sourceId` (`src/provider/ollama.ts`), not by looking up slash-form as a table key.
206
+
194
207
-**Overflow recovery** — A `context_overflow` inference error would otherwise become a terminal error reply; the governor compacts and retries instead, bounded so a history the compactor cannot shrink does not loop forever. Overflow ignores hysteresis for the compact itself.
195
208
209
+
The TTL window is measured from the last `inference.done` (every inference rewrites the provider's prefix cache) with one fire per window per compact of any kind, and it never fires with a tool batch outstanding — stall pings can arrive mid-work. The consecutive-compact cap shared with the threshold path still bounds compact→infer→compact. Prompt-caching the summary call itself is an explicit non-goal: the fold carries no cache options.
210
+
196
211
The compaction control flow is shaped by a reactor invariant: a `compact` action runs in its own cycle (it cannot be paired with `infer`), and **the reactor delivers no event after a compact cycle**. A director that simply emitted `compact` in place of the follow-up `infer` would leave the loop idle forever — the cause of an earlier stall. Instead the governor pairs `compact` with an emit action (`custom.compaction.continue`), and the host answers that emission by delivering a content-less inbound message (`buildCompactionContinuationMessage()`) through the serial operation queue, generation-guarded like every other deliver. A hop superseded because interrupt already bumped generation and enqueued a rebuild is re-queued onto the replacement agent rather than consumed on the outgoing liveAgent. That message adds no turn (`createInboundTurn` returns `null` for empty content) but re-enters the loop, where the director issues the follow-up `infer` against the freshly truncated history. Each emission is answered at most once (a replayed duplicate of an already-answered emission is ignored), and an unsolicited continuation with neither resume flag set answers `wait` rather than burning a billable inference. The legacy `requestContinuation` closure survives only on the sub-agent path, which delivers directly with no host emit hop; idle arming keeps the two channels exclusive (closure fires and the arming flag stays false, or no closure and the flag reports the arming) so a caller honoring both cannot double-deliver.
197
212
198
213
Compaction replaces older turns with a structured, workflow-aware summary rather than a stats blob: sections for **What Happened / What We're Doing / Relevant Links / Action Items / Next Steps**, with the active workflow and step woven in so compacting mid-`/build` or mid-`/plan` preserves the contract. The summary is produced by a one-shot model call; on any failure it falls back to a deterministic summary so a compaction cycle never breaks the session.
0 commit comments