Skip to content

fix(provider): stable history-trim cut keeps prefix-cache hits at the context ceiling - #238

Merged
ltmoerdani merged 3 commits into
ltmoerdani:mainfrom
lecommander:feat/history-trim-cache-hysteresis
Sep 23, 2026
Merged

ltmoerdani merged 3 commits into
ltmoerdani:mainfrom
lecommander:feat/history-trim-cache-hysteresis

Conversation

@lecommander

@lecommander lecommander commented Sep 20, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Keeps the provider prefix cache warm when a session sits at the context ceiling. The history trimmer previously cut to a minimal fit; at the ceiling that moves the cut point on nearly every following turn — and since the provider prefix cache only reuses the bytes before the first changed message, each moved cut re-billed the entire conversation at full input price.

Measured before this change (OpenCode Go, deepseek-v4.1-flash, chat-completions, 614K-token session): cache hits alternated between 99.8% and 11.4% in the same conversation. Every 11.4% miss reported cachedTokens=69888 — exactly system prompt + tools. 207 trims fired across a ~12-hour span and the session ended at ~45% token efficiency.

The fix landed in two rounds, both driven by live logs: a low-water mark (first commit) kept the cut stable while the slack lasted — but the crossing stops at the first unit below the mark, so the slack is always at most one unit wide, and sessions whose units are small relative to per-turn growth kept moving the cut. A cut-step alignment (second commit) makes the cut advance on a fixed grid of dropped payload, so each advance buys up to a full step of growth.

Problem

trimOldMessagesToFitContext drops the oldest droppable units until the payload fits and then stops. Landing just under the budget leaves ~zero slack, so the next turn's growth (one user turn is easily 2–3K tokens) exceeds it again and the next trim moves the cut to a different position. Log correlation (trim drop count -> next hit rate):

05:35  trim=245  -> 99.8%   (cut unchanged)
05:36  trim=249  -> 11.4%   (cut moved)
05:37  trim=251  -> 11.4%   (cut moved)
05:37  trim=253  -> 100.0%  (retry re-primes the new cut)
06:00  trim=317  -> 99.8% x4 (cut stable again)
06:02  trim=319  -> 11.3%   (cut moved)

The cache-key layer is not at fault — it was verified working in the same session; the misses track the cut movement exactly.

The residual failure mode (found the same day, evening). The first fix looked solved in the morning — 13 consecutive trims at the identical cut, 99.9% hits — but the evening dropped to 37–68% per hour with a 12.4% miss on every drop-count change (drop counts walked 214 → 242 in ~50 minutes):

17:47 TRIM dropped=214 -> hit 12.4%    (cut moved)
17:47 TRIM dropped=214 -> hit 100.0%   (same cut)
17:48 TRIM dropped=216 -> hit 12.4%    (cut moved again)
18:24 TRIM dropped=222 -> hit 12.4%
18:25 TRIM dropped=224 -> hit 12.4%

Those sessions had ~2.7K-token units against ~1K of growth per request, so the one-unit slack ran out almost every turn. Shifting the target does not help (the landing hugs whichever target within one unit — verified in simulation: unchanged move counts with the target shifted up to 98K below the mark); advancing the cut on a fixed step grid does.

Reproducing the issue

A. In production (what happened here). Run a long conversation on a large-context model until the history trimmer activates, then watch the OpenCode output channel (or %APPDATA%\Code\logs\<session>\...\*-OpenCode.log):

  1. Keep one chat going until it exceeds the trim budget — for deepseek-v4.1-flash that is contextWindow − maxOutputTokens − 2,048 = 613,952 tokens (in the reported session: ~750 messages, ~610K tokens).
  2. [history-trim] Dropped N old message(s) to fit context window (budget=613952, …) starts appearing, and N increments by 2–4 nearly every turn.
  3. Each time N increments, the next [usage] line shows the collapse — on the 11% misses cachedTokens is a constant equal to system prompt + tools (69,888 here) while promptTokens stays ~610–615K:
05:36:19 [history-trim] Dropped 249 ...   -> [usage] prompt=614077 cached=69888  cacheHit=11.4%
05:36:47 [history-trim] Dropped 251 ...   -> [usage] prompt=612777 cached=69888  cacheHit=11.4%
05:37:35 [history-trim] Dropped 253 ...   -> [usage] prompt=614971 cached=614784 cacheHit=100.0%  (retry re-primes the new cut)

B. Deterministically, without an API key (no network). The same scenario can be reproduced against any build in a few lines — trim an over-budget history once, then append five normal follow-up turns (~1.7K tokens each) and check whether the cut (the first surviving message) stays put:

const { trimOldMessagesToFitContext } = require("./out/provider/historyTrim.js");
const BUDGET = 100_000;
const messages = [{ role: "user", content: "ANCHOR " + "a".repeat(200) }];
for (let i = 0; i < 120; i++) messages.push({ role: "user", content: `turn ${i} ` + "p".repeat(4000) });
messages.push({ role: "user", content: "CURRENT " + "c".repeat(200) });

const t1 = trimOldMessagesToFitContext(messages, BUDGET, Number.MAX_SAFE_INTEGER);
let cut = messages[1].content, moves = 0;
for (let turn = 0; turn < 5; turn++) {
  messages.push({ role: "assistant", content: `reply ${turn} ` + "r".repeat(3000) });
  messages.push({ role: "user", content: `follow-up ${turn} ` + "q".repeat(3000) });
  const t = trimOldMessagesToFitContext(messages, BUDGET, Number.MAX_SAFE_INTEGER);
  if (t.removed > 0 && messages[1].content !== cut) { moves++; cut = messages[1].content; }
}
console.log("turn 1: finalTokens =", t1.finalTokens, "(slack:", BUDGET - t1.finalTokens, "); cut moves in next 5 turns:", moves);

Output on stock vs this PR (same scenario, npm run compile first):

[stock, pre-fix]    turn 1: finalTokens=98949 (slack: 1051);   cut moves in next 5 turns: 5
[patched, this PR]  turn 1: finalTokens=82299 (slack: 17701);  cut moves in next 5 turns: 0

The stock landing leaves only 1,051 tokens of slack — smaller than one normal follow-up turn (~1.7K) — so every following turn re-trims and moves the cut. The patched landing leaves 17,701 tokens of slack, so the next five turns leave the cut untouched.

Changes

  • src/config.ts: HISTORY_TRIM_HEADROOM_RATIO (3%), HISTORY_TRIM_HEADROOM_MIN_TOKENS (8,192 — capped at 10% of the budget so a small budget is never dominated), HISTORY_TRIM_HEADROOM_MAX_TOKENS (32,768), HISTORY_TRIM_CUT_STEP_TOKENS (32,768 — capped at 10% of the budget).
  • src/provider/historyTrim.ts: when a trim is unavoidable, the trimmer keeps dropping (same unit granularity and isUnsafeToolGroup safety rules) until the payload is at or below the low-water mark budget - headroom, then on to the next cut-step boundary of dropped payload. The step is what makes the stability robust: the crossing alone lands within one unit of the mark, while the fixed grid gives each advance up to a full step of room for the following turns — one cold prefix per epoch instead of one on nearly every turn.
  • src/test/messages.test.ts: four new tests (see below).
  • docs/issues/101-20260920-history-trim-cache-hysteresis.md + CHANGELOG entry.

Verifying this in a real environment

This fix is verified on the same surface the issue was diagnosed from: the extension's own request log in a live VS Code session. No test harness needed — the log lines are the acceptance criteria.

Protocol (any machine, any account):

  1. Open the OpenCode output channel (or %APPDATA%\Code\logs\<session>\...\*-OpenCode.log) and keep one chat going until the trimmer activates — for a 1M-window model that is contextWindow − maxOutputTokens − 2,048 (budget=613952 for deepseek-v4.1-flash).
  2. Checkpoint 1 — landing value. The first [history-trim] … estimated payload now ~X tokens after the trim must satisfy X ≤ budget − headroom. With this session's budget that is ≤ 595,534 tokens (headroom 18,418). Pre-fix, trims landed at 607,325–613,634 — minimal fit, 300–6,600 tokens of slack, repeatedly overflowing on the next turn.
  3. Checkpoint 2 — cut stability. The drop count in the [history-trim] lines must stay constant across following turns. The trimmer still runs while the history exceeds the budget (VS Code supplies the full history each turn), but a constant drop count means the cut — and therefore the sent prefix — did not move: consecutive payloads are nested prefixes. Pre-fix, the drop count incremented almost every turn (207 trims in a ~12-hour window), moving the cut and cold-caching the provider prefix each time.
  4. Checkpoint 3 — cache recovery. [usage] … cacheHit= must stay in the ~99% band between trims instead of alternating. Pre-fix signature: 99.8% → 11.4% → 11.4% → 100.0% → 11.4%, with cachedTokens pinned at 69,888 (system + tools) on every 11% miss.

Pre-fix baseline (the reporter's live log, last four trims before upgrading):

06:13:39 [history-trim] Dropped 333 ... ~613634 tokens   (slack 318)
06:15:02 [history-trim] Dropped 340 ... ~613288 tokens   (slack 664)
06:15:03 [history-trim] Dropped 343 ... ~607428 tokens   (slack 6,524)
06:16:10 [history-trim] Dropped 343 ... ~607325 tokens   (re-trimmed again, same cut)

Post-fix, running the compiled artifact with these exact production parameters (budget=613952, maxBytes=2762784, ~750 messages, ~1.4K tokens of growth per follow-up turn):

first trim: dropped 38 messages -> landed at 594,832 tokens / 2,163,043 bytes   (<= 595,534 low-water: PASS)
next 10 follow-up turns: re-trims = 0/10                                          (cut stable: PASS)

Live deployment status — real traffic on the patched build; first ceiling crossing verified (updated 2026-09-20 12:55 local). A conversation on the fixed build grew past the ceiling — deepseek-v4.1-flash, 122K → 625,118 tokens against the 613,952 budget — producing the first live trims. All three checkpoints pass:

all traffic:                    517 requests   avg cacheHit 98.9%   506/517 at >= 99%    6 sub-50
  pre-fix, same environment:    918 requests   avg 88.8%                                111 sub-50
high-context requests (> 580K tokens):
  post-fix:                      55 requests   avg 95.2%                                  3 sub-50
  pre-fix:                      243 requests   avg 63.4%                                100 sub-50 (41%)
15 history-trims: every landing 584,252 - 595,440 tokens                  (low-water 595,534: 15/15 PASS)
  pre-fix landings: 607,325 - 613,634
cut stability: 13 consecutive trims held the identical cut (dropped=38) at 99.9-100% hits; after
               the initial trim the cut shifted only twice in ~31 min of ceiling traffic
               (38 -> 40 -> 44), each shift costing exactly one 12.3% request, recovered next turn
  • Checkpoint 1 — landing: 15/15 live trims landed at or below the low-water mark.
  • Checkpoint 2 — cut stability: the sent payload stayed a nested prefix of the previous one for 13 consecutive turns (constant drop count, growing tail), so the provider prefix cache stayed warm through every one of those trims; the cut moved only twice after the initial trim.
  • Checkpoint 3 — recovery: hits return to 99.9-100% on the request immediately after each cut shift. The three sub-50 requests in the high-context sample are exactly the requests that followed a cut shift — against 100 fully-missed requests in the comparable pre-fix window.
  • Evening regression (same day) — the residual case: from ~14h local the hourly hit rate fell to 37–68%, a 12.4% miss on every drop-count change (214 → 242 drop count walk in ~50 minutes), because those turns' units (~2.7K tokens) approxed their per-request growth (~1K) — the one-unit slack ran out almost every turn. The morning's stability was luck (large tool-result units). The cut-step commit addresses this; the simulation below and the new unit test pin it.

Pre-merge checks

  1. Repo gate. npm run lint — all 7 checks pass (editorconfig, ESLint, markdown, prettier, shell, TypeScript, tests); 468/468 unit tests. CI on this PR: build and GitGuardian both green.
  2. Regression tests (src/test/messages.test.ts, describe trimOldMessagesToFitContext — cache-stable headroom):
    • drops below the low-water mark so the next turn does not re-trim — at a 100K budget the trim must land at or below 91,808 tokens while anchor and current prompt are preserved. Fails on stock (lands at ~99K).
    • keeps the cut stable across following turns (no re-trim → warm prefix cache) — appends five follow-up turns after the initial trim and asserts removed == 0 and an unchanged first surviving message for each. Fails on stock.
    • holds the cut when the full history is re-supplied each turn (production shape) — mirrors the live trace (uneven unit sizes; the full history returns every turn, so the trimmer keeps running): the drop count stays constant and each sent payload deep-equals the prefix of the previous one. Fails on the pre-fix artifact (drop count 12 → 13, cut moved).
    • holds the cut across a normal follow-up turn via the cut-step grid — a ~1.1K-token follow-up turn does not move the cut; without the step pass the same scenario moves it (68 → 70 vs 79 → 79 dropped units — verified against the pre-step build).
  3. Head-to-head reproduction (section B above): stock moves the cut 5/5 turns; this PR 0/5. Simulation with production parameters (300 requests, budget 613,952, byte cap 2,762,784): cut moves fall from 84/300 (28%) to 12/300 (4%) in the large-unit regime and from 244/300 (81%) to 14/300 (5%) in the small-unit regime — the step makes stability independent of unit size.
  4. Artifact smoke test against a patched extension build: trim lands at 82,299 vs the 91,808 low-water mark (100K budget); cut survives five turns; and a control run with the headroom forced to zero still re-trims and moves the cut — proving the behavioral delta is the headroom plus step, nothing else.

Trade-off

A trim now sheds a bounded amount of extra old context (at most the headroom plus one cut step; overshoot at most one unit) in exchange for a cut point that stays still. At the ceiling this is strictly cheaper: one ~11% request per epoch instead of nearly every request missing.

Notes

  • Behavior only changes when a trim is already triggered and the targets are reachable through safe units; histories under the budget are untouched (early return), and tiny budgets keep their previous behavior (both margins are capped at 10% of the budget).
  • The byte ceiling gets the same treatment via HISTORY_BYTES_PER_TOKEN.
  • Observation, not changed here: the minimal-fit loop commits a drop on its fits path without the isUnsafeToolGroup guard (the guard only runs on the not-yet-fitting path). The new low-water pass always applies the guard. Flagging in case you want a follow-up — I kept this PR surgical.

@lecommander

Copy link
Copy Markdown
Contributor Author

Added a Reproducing the issue section to the description: the production observation path (watch [history-trim] vs [usage] cacheHit lines at the ceiling) plus a ~20-line deterministic script that needs no API key — on stock the cut moves 5/5 follow-up turns, on this PR 0/5. Verification section now spells out the two regression tests, the head-to-head run numbers, and the compiled-artifact smoke test with the zero-headroom control.

@lecommander

lecommander commented Sep 20, 2026 •

Copy link
Copy Markdown
Contributor Author

Restructured the verification part around real-environment evidence instead of tests: the issue was diagnosed from the reporter's live request log, and the fix is verified on the same surface. Added the runtime protocol with three acceptance checkpoints -- (1) trim landing <= budget - headroom (<= 595,534 for the 613,952 budget; pre-fix landings were 607,325-613,634), (2) no re-trim on the following turns (pre-fix: 207 trims across ~12 h), (3) cacheHit stays ~99% between trims (pre-fix: 99.8%/11.4% alternation with cachedTokens pinned at 69,888). The compiled artifact re-run with these exact production parameters lands at 594,832 tokens and survives 10 follow-up turns without re-trimming.

@lecommander

Copy link
Copy Markdown
Contributor Author

Updated the description with live deployment evidence: the patched build is now running in the reporter's environment, and its runtime request log shows 153 requests, 98.2% average cacheHit, 147/153 at >= 99%, only two distinct prefix hashes (one per conversation).

No [history-trim] events yet -- both live conversations are at ~122-374K tokens, under the 613,952 budget, so the trim-dependent checkpoints (landing <= 595,534, no re-trim on following turns) remain pending on the first conversation that crosses it.

@lecommander

Copy link
Copy Markdown
Contributor Author

Updated the description — the first live ceiling crossing is now verified on the fixed build.

A conversation grew to 625,118 tokens (budget 613,952) and produced 15 trims, every landing at 584,252–595,440 tokens (low-water 595,534: 15/15 PASS; pre-fix landings were 607,325–613,634). The cut held constant across 13 consecutive trims at 99.9–100% cache hits; after the initial trim it shifted only twice in ~31 min of ceiling traffic, each shift costing exactly one 12.3% request, recovered on the next turn.

High-context requests (> 580K tokens): post-fix 55 requests at 95.2% avg with 3 sub-50 vs pre-fix 243 requests at 63.4% avg with 100 sub-50 (41%).

Also corrected the checkpoint-2 wording: the invariant is a constant drop count (the trimmer still runs while the history exceeds the budget — VS Code re-supplies the full history each turn — but a constant drop count means the cut, and the sent prefix, did not move).

…plied history)

The existing tests cover the inline case (same array grows, no re-trim). Production differs: VS Code re-supplies the FULL history every turn, so the trimmer keeps running - what must stay constant is the drop count, i.e. the cut, so consecutive sent payloads are nested prefixes and the provider cache stays warm through the re-trims.

The new test mirrors that shape with uneven unit sizes (every fifth turn ~6x larger, like production tool results): after the initial trim, two further turns assert the drop count is unchanged and each payload deep-equals the prefix of the previous one. Verified against the compiled pre-fix artifact: the same scenario fails there (drop count 12 -> 13, cut moved).

Also records the first live ceiling crossing in the issue doc and changelog: 15 trims, every landing <= 595,534 tokens (budget 613,952), 13 consecutive trims at the identical cut with 99.9-100% cache hits, and high-context misses down from 100/243 pre-fix to 3/55.
@lecommander

Copy link
Copy Markdown
Contributor Author

Added the production-shape test as commit bf2f167:

  • The existing tests cover the inline case (same array grows → no re-trim). Production differs: VS Code re-supplies the full history every turn, so the trimmer keeps running — what must stay constant is the drop count (the cut), so consecutive sent payloads are nested prefixes and the provider cache stays warm through the re-trims.
  • The new test mirrors that shape with uneven unit sizes (every fifth turn ~6x larger, like production tool results): after the initial trim, two further turns assert the drop count is unchanged and each payload deep-equals the prefix of the previous one.
  • Verified against the compiled pre-fix artifact: the same scenario fails there (drop count 12 → 13, cut moved) — so the test discriminates, same as the other two.
  • npm run lint clean, 467/467 tests, CI green on the new commit.

Also refreshed the issue doc and changelog with the first live ceiling-crossing numbers (15 trims / all landings ≤ 595,534 / high-context misses 100-of-243 → 3-of-55).

…low-up turns

Live follow-up showed the low-water mark alone was not enough: the crossing pass stops at the first unit that reaches the mark, so the slack for the next turns is bounded by one unit's size. Sessions with ~2.7K-token units against ~1K of per-request growth moved the cut every 1-3 requests (12.4% cache hit each time, hourly hit rate down to 37-68%), while a small-unit session only looked stable by luck. Shifting the target does not help (the landing hugs whichever target within one unit - verified in simulation).

The cut now continues to the next HISTORY_TRIM_CUT_STEP_TOKENS (32,768, capped at 10% of budget) boundary of dropped payload. That fixed grid decouples the landing from the crossing: each advance buys up to a full step of growth. Simulation with production parameters: cut moves 84/300 -> 12/300 in the evening regime and 244/300 -> 14/300 in the smaller-unit regime. Adds a discriminating unit test (the same follow-up turn moves the cut without the step pass).
@lecommander

Copy link
Copy Markdown
Contributor Author

You were right that the evening still had misses — I dug into the live logs and the first fix turned out to be incomplete. New commit ef7f31d closes the gap.

What the evening logs showed. From ~14h local the hourly hit rate fell to 37–68%, with a 12.4% miss on every drop-count change (the drop count walked 214 → 242 in ~50 minutes — the cut moved on nearly every turn).

Why the low-water mark alone wasn't enough. The crossing pass stops at the first unit that reaches the mark, so the slack left for the next turns is bounded by that one unit's size. Sessions whose units (~2.7K tokens) are close to their per-turn growth (~1K) run out of slack almost every turn. The morning had only looked stable by luck — its dropped units happened to include large tool results (the first landing left 11K of slack; the afternoon didn't). Shifting the target doesn't help: the landing hugs whichever target within one unit — verified in simulation (unchanged move counts with the target shifted up to 98K lower).

The fix — cut-step alignment. After the low-water crossing, the cut continues to the next HISTORY_TRIM_CUT_STEP_TOKENS (32,768; capped at 10% of budget) boundary of dropped payload. That fixed grid decouples the landing from the crossing, so each advance buys up to a full step of growth. Simulation with production parameters (300 requests, budget 613,952):

evening regime (~2.7K units, ~1K growth):  cut moves 84/300 -> 12/300   (every 23 requests)
smaller-unit regime (~800 units, ~1K):     cut moves 244/300 -> 14/300  (every 21 requests)

The stability is now independent of unit size, and a discriminating unit test pins it (the same follow-up turn moves the cut without the step pass: 68 → 70 vs 79 → 79 dropped units). npm run lint clean, 468/468 tests, CI green on the new commit. The description now documents both rounds, the evening trace, and the simulation.

@ltmoerdani

Copy link
Copy Markdown
Owner

Been through this end to end and it looks good @lecommander. The simulation numbers, the discriminating test, and the two live runs (morning + evening) give me enough confidence. Small-budget safety is already bounded by the 10% cap, so I don't see a blocker.

Merging this as-is. On the isUnsafeToolGroup gap you flagged in the minimal-fit loop: that one is pre-existing and outside this PR's scope, so I'd keep it separate. If you're up for a small follow-up PR for it, that'd be welcome, just outline the failure mode first so we know what we're fixing.

Nice one, thanks for digging through the evening logs on your own. Merging now.

@ltmoerdani
ltmoerdani merged commit 0a6d9ed into ltmoerdani:main Sep 23, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants