You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Copy file name to clipboardExpand all lines: docs/ARCHITECTURE.md
+4-8Lines changed: 4 additions & 8 deletions
Display the source diff
Display the rich diff
Original file line number
Diff line number
Diff line change
@@ -193,14 +193,10 @@ When a cycle's input tokens cross a threshold, the director compacts the inferen
193
193
-**Idle (end-of-turn)** — An interactive turn can end with a reply and then sit idle with no tool batch to intercept; the governor requests a continuation at that pause and compacts when it arrives. An operator message that races the continuation still compacts first, then re-enters inference to answer it.
194
194
-**Idle recompress past the provider TTL** — The fold is a re-compress, not a cache play: provider KV caches expire on their own schedule, so a session compresses _after_ that expiry (the next turn is a cheaper write and later reads compound on the shrunk context) instead of only at 60% tokens. It is not limited to under-threshold sessions: once growth hysteresis has cleared `pending`, an over-threshold session in that gap can fire the same fold. Any live re-entry past the window fires it — an empty stall ping re-enters meter-only, a raced operator message re-infers — under reason `cache-ttl-recompress` through the same compactor, so the fresh tail stays raw exactly as in the threshold path. The interval follows provider cache economics (`src/provider/cache-ttl.ts`), not a global N minutes:
|`anthropic`, `zen-messages`, `opencode-go-messages`| 5 min | Default ephemeral cache TTL; hour-long breakpoints are never set |
199
-
|`openai-responses`, `codex-responses`, `openai-compatible`, `xai`| 10 min | In-memory prefixes typically evicted after 5–10 min idle; xAI assumed same economics |
200
-
|`gemini`| 15 min | No published implicit-cache eviction window; conservative |
201
-
|`deepseek`| 60 min | On-disk context cache persists for hours-to-days |
202
-
|`ollama`| never | Local inference has no remote cache to expire |
203
-
| anything else | 10 min | Assumed OpenAI-style in-memory economics; tune per upstream |
|`anthropic`, `zen-messages`, `opencode-go-messages`| 5 min | Published default ephemeral cache TTL; the 1-hour TTL is never set |
199
+
| anything else, including OpenAI, xAI, Gemini, DeepSeek, ollama | never | No published 5-minute expiry, or no remote cache. Guessing a shorter window folds a cache that may still be warm |
204
200
205
201
The `ollama` table key is the bare provider segment; the harness stamps slash-form on `sourceId` (`ollama/default` / `ollama/<instance>`), with `provider: openai-compatible` (`buildOpenAISource`) and a bare model (`llama3` / `qwen3`). Idle-recompress disable matches via `isOllamaProviderId` on that `sourceId` (`src/provider/ollama.ts`), not by looking up slash-form as a table key.
0 commit comments