fix(provider): stable history-trim cut keeps prefix-cache hits at the context ceiling - #238
Conversation
|
Added a Reproducing the issue section to the description: the production observation path (watch [history-trim] vs [usage] cacheHit lines at the ceiling) plus a ~20-line deterministic script that needs no API key — on stock the cut moves 5/5 follow-up turns, on this PR 0/5. Verification section now spells out the two regression tests, the head-to-head run numbers, and the compiled-artifact smoke test with the zero-headroom control. |
|
Restructured the verification part around real-environment evidence instead of tests: the issue was diagnosed from the reporter's live request log, and the fix is verified on the same surface. Added the runtime protocol with three acceptance checkpoints -- (1) trim landing <= budget - headroom (<= 595,534 for the 613,952 budget; pre-fix landings were 607,325-613,634), (2) no re-trim on the following turns (pre-fix: 207 trims across ~12 h), (3) cacheHit stays ~99% between trims (pre-fix: 99.8%/11.4% alternation with cachedTokens pinned at 69,888). The compiled artifact re-run with these exact production parameters lands at 594,832 tokens and survives 10 follow-up turns without re-trimming. |
|
Updated the description with live deployment evidence: the patched build is now running in the reporter's environment, and its runtime request log shows 153 requests, 98.2% average cacheHit, 147/153 at >= 99%, only two distinct prefix hashes (one per conversation). No |
|
Updated the description — the first live ceiling crossing is now verified on the fixed build. A conversation grew to 625,118 tokens (budget 613,952) and produced 15 trims, every landing at 584,252–595,440 tokens (low-water 595,534: 15/15 PASS; pre-fix landings were 607,325–613,634). The cut held constant across 13 consecutive trims at 99.9–100% cache hits; after the initial trim it shifted only twice in ~31 min of ceiling traffic, each shift costing exactly one 12.3% request, recovered on the next turn. High-context requests (> 580K tokens): post-fix 55 requests at 95.2% avg with 3 sub-50 vs pre-fix 243 requests at 63.4% avg with 100 sub-50 (41%). Also corrected the checkpoint-2 wording: the invariant is a constant drop count (the trimmer still runs while the history exceeds the budget — VS Code re-supplies the full history each turn — but a constant drop count means the cut, and the sent prefix, did not move). |
…plied history) The existing tests cover the inline case (same array grows, no re-trim). Production differs: VS Code re-supplies the FULL history every turn, so the trimmer keeps running - what must stay constant is the drop count, i.e. the cut, so consecutive sent payloads are nested prefixes and the provider cache stays warm through the re-trims. The new test mirrors that shape with uneven unit sizes (every fifth turn ~6x larger, like production tool results): after the initial trim, two further turns assert the drop count is unchanged and each payload deep-equals the prefix of the previous one. Verified against the compiled pre-fix artifact: the same scenario fails there (drop count 12 -> 13, cut moved). Also records the first live ceiling crossing in the issue doc and changelog: 15 trims, every landing <= 595,534 tokens (budget 613,952), 13 consecutive trims at the identical cut with 99.9-100% cache hits, and high-context misses down from 100/243 pre-fix to 3/55.
|
Added the production-shape test as commit
Also refreshed the issue doc and changelog with the first live ceiling-crossing numbers (15 trims / all landings ≤ 595,534 / high-context misses 100-of-243 → 3-of-55). |
…low-up turns Live follow-up showed the low-water mark alone was not enough: the crossing pass stops at the first unit that reaches the mark, so the slack for the next turns is bounded by one unit's size. Sessions with ~2.7K-token units against ~1K of per-request growth moved the cut every 1-3 requests (12.4% cache hit each time, hourly hit rate down to 37-68%), while a small-unit session only looked stable by luck. Shifting the target does not help (the landing hugs whichever target within one unit - verified in simulation). The cut now continues to the next HISTORY_TRIM_CUT_STEP_TOKENS (32,768, capped at 10% of budget) boundary of dropped payload. That fixed grid decouples the landing from the crossing: each advance buys up to a full step of growth. Simulation with production parameters: cut moves 84/300 -> 12/300 in the evening regime and 244/300 -> 14/300 in the smaller-unit regime. Adds a discriminating unit test (the same follow-up turn moves the cut without the step pass).
|
You were right that the evening still had misses — I dug into the live logs and the first fix turned out to be incomplete. New commit What the evening logs showed. From ~14h local the hourly hit rate fell to 37–68%, with a 12.4% miss on every drop-count change (the drop count walked 214 → 242 in ~50 minutes — the cut moved on nearly every turn). Why the low-water mark alone wasn't enough. The crossing pass stops at the first unit that reaches the mark, so the slack left for the next turns is bounded by that one unit's size. Sessions whose units (~2.7K tokens) are close to their per-turn growth (~1K) run out of slack almost every turn. The morning had only looked stable by luck — its dropped units happened to include large tool results (the first landing left 11K of slack; the afternoon didn't). Shifting the target doesn't help: the landing hugs whichever target within one unit — verified in simulation (unchanged move counts with the target shifted up to 98K lower). The fix — cut-step alignment. After the low-water crossing, the cut continues to the next The stability is now independent of unit size, and a discriminating unit test pins it (the same follow-up turn moves the cut without the step pass: 68 → 70 vs 79 → 79 dropped units). |
|
Been through this end to end and it looks good @lecommander. The simulation numbers, the discriminating test, and the two live runs (morning + evening) give me enough confidence. Small-budget safety is already bounded by the 10% cap, so I don't see a blocker. Merging this as-is. On the Nice one, thanks for digging through the evening logs on your own. Merging now. |
Summary
Keeps the provider prefix cache warm when a session sits at the context ceiling. The history trimmer previously cut to a minimal fit; at the ceiling that moves the cut point on nearly every following turn — and since the provider prefix cache only reuses the bytes before the first changed message, each moved cut re-billed the entire conversation at full input price.
Measured before this change (OpenCode Go,
deepseek-v4.1-flash, chat-completions, 614K-token session): cache hits alternated between 99.8% and 11.4% in the same conversation. Every 11.4% miss reportedcachedTokens=69888— exactly system prompt + tools. 207 trims fired across a ~12-hour span and the session ended at ~45% token efficiency.The fix landed in two rounds, both driven by live logs: a low-water mark (first commit) kept the cut stable while the slack lasted — but the crossing stops at the first unit below the mark, so the slack is always at most one unit wide, and sessions whose units are small relative to per-turn growth kept moving the cut. A cut-step alignment (second commit) makes the cut advance on a fixed grid of dropped payload, so each advance buys up to a full step of growth.
Problem
trimOldMessagesToFitContextdrops the oldest droppable units until the payload fits and then stops. Landing just under the budget leaves ~zero slack, so the next turn's growth (one user turn is easily 2–3K tokens) exceeds it again and the next trim moves the cut to a different position. Log correlation (trim drop count -> next hit rate):The cache-key layer is not at fault — it was verified working in the same session; the misses track the cut movement exactly.
The residual failure mode (found the same day, evening). The first fix looked solved in the morning — 13 consecutive trims at the identical cut, 99.9% hits — but the evening dropped to 37–68% per hour with a 12.4% miss on every drop-count change (drop counts walked 214 → 242 in ~50 minutes):
Those sessions had ~2.7K-token units against ~1K of growth per request, so the one-unit slack ran out almost every turn. Shifting the target does not help (the landing hugs whichever target within one unit — verified in simulation: unchanged move counts with the target shifted up to 98K below the mark); advancing the cut on a fixed step grid does.
Reproducing the issue
A. In production (what happened here). Run a long conversation on a large-context model until the history trimmer activates, then watch the OpenCode output channel (or
%APPDATA%\Code\logs\<session>\...\*-OpenCode.log):deepseek-v4.1-flashthat iscontextWindow − maxOutputTokens − 2,048 = 613,952tokens (in the reported session: ~750 messages, ~610K tokens).[history-trim] Dropped N old message(s) to fit context window (budget=613952, …)starts appearing, andNincrements by 2–4 nearly every turn.Nincrements, the next[usage]line shows the collapse — on the 11% missescachedTokensis a constant equal to system prompt + tools (69,888 here) whilepromptTokensstays ~610–615K:B. Deterministically, without an API key (no network). The same scenario can be reproduced against any build in a few lines — trim an over-budget history once, then append five normal follow-up turns (~1.7K tokens each) and check whether the cut (the first surviving message) stays put:
Output on stock vs this PR (same scenario,
npm run compilefirst):The stock landing leaves only 1,051 tokens of slack — smaller than one normal follow-up turn (~1.7K) — so every following turn re-trims and moves the cut. The patched landing leaves 17,701 tokens of slack, so the next five turns leave the cut untouched.
Changes
src/config.ts:HISTORY_TRIM_HEADROOM_RATIO(3%),HISTORY_TRIM_HEADROOM_MIN_TOKENS(8,192 — capped at 10% of the budget so a small budget is never dominated),HISTORY_TRIM_HEADROOM_MAX_TOKENS(32,768),HISTORY_TRIM_CUT_STEP_TOKENS(32,768 — capped at 10% of the budget).src/provider/historyTrim.ts: when a trim is unavoidable, the trimmer keeps dropping (same unit granularity andisUnsafeToolGroupsafety rules) until the payload is at or below the low-water markbudget - headroom, then on to the next cut-step boundary of dropped payload. The step is what makes the stability robust: the crossing alone lands within one unit of the mark, while the fixed grid gives each advance up to a full step of room for the following turns — one cold prefix per epoch instead of one on nearly every turn.src/test/messages.test.ts: four new tests (see below).docs/issues/101-20260920-history-trim-cache-hysteresis.md+ CHANGELOG entry.Verifying this in a real environment
This fix is verified on the same surface the issue was diagnosed from: the extension's own request log in a live VS Code session. No test harness needed — the log lines are the acceptance criteria.
Protocol (any machine, any account):
%APPDATA%\Code\logs\<session>\...\*-OpenCode.log) and keep one chat going until the trimmer activates — for a 1M-window model that iscontextWindow − maxOutputTokens − 2,048(budget=613952fordeepseek-v4.1-flash).[history-trim] … estimated payload now ~X tokensafter the trim must satisfyX ≤ budget − headroom. With this session's budget that is ≤ 595,534 tokens (headroom 18,418). Pre-fix, trims landed at 607,325–613,634 — minimal fit, 300–6,600 tokens of slack, repeatedly overflowing on the next turn.[history-trim]lines must stay constant across following turns. The trimmer still runs while the history exceeds the budget (VS Code supplies the full history each turn), but a constant drop count means the cut — and therefore the sent prefix — did not move: consecutive payloads are nested prefixes. Pre-fix, the drop count incremented almost every turn (207 trims in a ~12-hour window), moving the cut and cold-caching the provider prefix each time.[usage] … cacheHit=must stay in the ~99% band between trims instead of alternating. Pre-fix signature: 99.8% → 11.4% → 11.4% → 100.0% → 11.4%, withcachedTokenspinned at 69,888 (system + tools) on every 11% miss.Pre-fix baseline (the reporter's live log, last four trims before upgrading):
Post-fix, running the compiled artifact with these exact production parameters (
budget=613952,maxBytes=2762784, ~750 messages, ~1.4K tokens of growth per follow-up turn):Live deployment status — real traffic on the patched build; first ceiling crossing verified (updated 2026-09-20 12:55 local). A conversation on the fixed build grew past the ceiling —
deepseek-v4.1-flash, 122K → 625,118 tokens against the 613,952 budget — producing the first live trims. All three checkpoints pass:Pre-merge checks
npm run lint— all 7 checks pass (editorconfig, ESLint, markdown, prettier, shell, TypeScript, tests); 468/468 unit tests. CI on this PR:buildand GitGuardian both green.src/test/messages.test.ts, describetrimOldMessagesToFitContext — cache-stable headroom):drops below the low-water mark so the next turn does not re-trim— at a 100K budget the trim must land at or below 91,808 tokens while anchor and current prompt are preserved. Fails on stock (lands at ~99K).keeps the cut stable across following turns (no re-trim → warm prefix cache)— appends five follow-up turns after the initial trim and assertsremoved == 0and an unchanged first surviving message for each. Fails on stock.holds the cut when the full history is re-supplied each turn (production shape)— mirrors the live trace (uneven unit sizes; the full history returns every turn, so the trimmer keeps running): the drop count stays constant and each sent payload deep-equals the prefix of the previous one. Fails on the pre-fix artifact (drop count 12 → 13, cut moved).holds the cut across a normal follow-up turn via the cut-step grid— a ~1.1K-token follow-up turn does not move the cut; without the step pass the same scenario moves it (68 → 70 vs 79 → 79 dropped units — verified against the pre-step build).Trade-off
A trim now sheds a bounded amount of extra old context (at most the headroom plus one cut step; overshoot at most one unit) in exchange for a cut point that stays still. At the ceiling this is strictly cheaper: one ~11% request per epoch instead of nearly every request missing.
Notes
HISTORY_BYTES_PER_TOKEN.fitspath without theisUnsafeToolGroupguard (the guard only runs on the not-yet-fitting path). The new low-water pass always applies the guard. Flagging in case you want a follow-up — I kept this PR surgical.