Skip to content

fix(context-budget): measure against the real window with an explicit token budget - #49

Open
TastyTom13 wants to merge 4 commits into
mainfrom
fm/firstmate-context-nudge-real-window
Open

TastyTom13 wants to merge 4 commits into
mainfrom
fm/firstmate-context-nudge-real-window

Conversation

@TastyTom13

@TastyTom13 TastyTom13 commented Oct 2, 2026 •

Copy link
Copy Markdown
Owner

Intent

Fix the fork-only context nudge so it measures against the real Claude Code compaction window (FM_CONTEXT_WINDOW, else CLAUDE_CODE_AUTO_COMPACT_WINDOW, else 200000) and nudges early against an explicit FM_CONTEXT_BUDGET token budget (default 500000, clamped to the window) with bands at 40/60 percent of the budget overridable in tokens, printing tokens, percent of window, percent of budget, and band, keeping the throttle and quiet-moment rule unchanged; docs explain the budget sits below the window per the captain's 2026-10-02 decision.

What Changed

  • bin/fm-context-budget.sh now resolves the real window from --window, then FM_CONTEXT_WINDOW, then a numeric CLAUDE_CODE_AUTO_COMPACT_WINDOW, then 200000. It adds an FM_CONTEXT_BUDGET token budget (default 500000, clamped to the window). The 40/60 percent bands are now token thresholds taken from that budget, and FM_CONTEXT_NUDGE_SUGGEST and FM_CONTEXT_NUDGE_NOW override them in tokens. The printed line shows tokens, percent of the window (with its source), percent of the budget, and the band. The throttle step is now budget percent divided by 20. The quiet-moment rule and the throttle record format are unchanged.
  • docs/configuration.md, docs/turnend-guard.md, and the comments in bin/fm-turnend-guard.sh describe the window and budget precedence, the new env vars, and the budget-based steps. The docs say the budget sits below the window per the captain's 2026-10-02 decision, so the nudge fires before Claude Code's hard auto-compaction.
  • tests/fm-context-budget.test.sh adds 78 lines of cases for the new behaviour.

Risk Assessment

✅ Low: The change is confined to one advisory script, its docs and its tests, it matches every point of the stated intent, and the earlier environment-leak finding is fixed by a single top-of-file unset.

Testing

I ran the budget test file (all pass) and drove the real script by hand on synthetic transcripts. The printed line shows tokens, percent of window, percent of budget, source and band. Window precedence, clamp, 40/60 token bands and overrides all behave. The nudge throttle fires once per 20 percent budget step and rejects repeats. Docs explain the budget sits below the window. The change has no UI, so there are no screenshots. The Stop hook was covered by the test file, not by a live Claude session, so that scenario is marked untested.

  • Live validation: ✅ go - 7 of 8 scenarios driven live against the product
Scenario Result Live Evidence
Default window 200000: a 500000 budget is clamped to the window and the line says so; bands quiet/next/now at 35/45/65 percent ✅ pass live live-script-run.txt section A
Window from CLAUDE_CODE_AUTO_COMPACT_WINDOW=1M with default 500k budget prints window percent and budget percent and band ✅ pass live live-script-run.txt section B
FM_CONTEXT_WINDOW beats CLAUDE_CODE_AUTO_COMPACT_WINDOW and the line names the source ✅ pass live live-script-run.txt section C
Explicit FM_CONTEXT_BUDGET=300000 moves bands to 40/60 percent of that budget ✅ pass live live-script-run.txt section D
Token overrides FM_CONTEXT_NUDGE_SUGGEST/NOW change the bands; invalid values fall back to defaults; --percent prints percent of window ✅ pass live live-script-run.txt sections E, F, G
Adversarial: nudge throttle stays silent below the band, repeats within a step, and announces once on each new step or band upgrade ✅ pass live live-script-run.txt section H (rc=1,1,0,1,0 with record s1 2 next then s1 3 now)
Tests are isolated from the caller's FM_CONTEXT_* environment (review round 1 fix) ✅ pass live test-run.txt; tests/fm-context-budget.test.sh unsets the vars at the top and all 21 pass
Turn-end guard Stop hook prints the nudge once per step and only at quiet moments ⏸️ untested no The prior payload did not establish a live result. It relied on the repo's hermetic hook tests, which run the real guard script, and no live Claude session was driven.
Evidence: Live script transcript

Source: Live script transcript

== A default window 200000, budget clamped
context 70k tokens, 35% of the 200k window (default), 35% of the 200k budget (budget 500k clamped to the window) - quiet
context 90k tokens, 45% of the 200k window (default), 45% of the 200k budget (budget 500k clamped to the window) - suggest /stow at the next quiet moment
context 130k tokens, 65% of the 200k window (default), 65% of the 200k budget (budget 500k clamped to the window) - suggest /stow now
context 180k tokens, 90% of the 200k window (default), 90% of the 200k budget (budget 500k clamped to the window) - suggest /stow now
== B window from CLAUDE_CODE_AUTO_COMPACT_WINDOW=1000000 budget default 500k
context 150k tokens, 15% of the 1000k window (CLAUDE_CODE_AUTO_COMPACT_WINDOW), 30% of the 500k budget - quiet
context 250k tokens, 25% of the 1000k window (CLAUDE_CODE_AUTO_COMPACT_WINDOW), 50% of the 500k budget - suggest /stow at the next quiet moment
context 350k tokens, 35% of the 1000k window (CLAUDE_CODE_AUTO_COMPACT_WINDOW), 70% of the 500k budget - suggest /stow now
== C FM_CONTEXT_WINDOW beats compact window
context 250k tokens, 41% of the 600k window (FM_CONTEXT_WINDOW), 50% of the 500k budget - suggest /stow at the next quiet moment
== D explicit budget 300000 window 1M
context 100k tokens, 10% of the 1000k window (FM_CONTEXT_WINDOW), 33% of the 300k budget - quiet
context 150k tokens, 15% of the 1000k window (FM_CONTEXT_WINDOW), 50% of the 300k budget - suggest /stow at the next quiet moment
context 200k tokens, 20% of the 1000k window (FM_CONTEXT_WINDOW), 66% of the 300k budget - suggest /stow now
== E token threshold overrides
context 150k tokens, 15% of the 1000k window (FM_CONTEXT_WINDOW), 30% of the 500k budget - suggest /stow now
== F invalid values fall back
context 150k tokens, 30% of the 500k window (CLAUDE_CODE_AUTO_COMPACT_WINDOW), 30% of the 500k budget - quiet
== G percent mode = percent of window
15
== H nudge throttle (window 1M, budget 500k, steps of 100k)
cat: /var/folders/08/gk0k1cz96snd794rnghc_98r0000gn/T/tmp.9EDc3RPzDe/state/.context-budget-nudged: No such file or directory
rc=1 tokens=150000 record=
cat: /var/folders/08/gk0k1cz96snd794rnghc_98r0000gn/T/tmp.9EDc3RPzDe/state/.context-budget-nudged: No such file or directory
rc=1 tokens=150000 record=
context 250k tokens, 25% of the 1000k window (FM_CONTEXT_WINDOW), 50% of the 500k budget - suggest /stow at the next quiet moment
rc=0 tokens=250000 record=s1 2 next
rc=1 tokens=250000 record=s1 2 next
context 350k tokens, 35% of the 1000k window (FM_CONTEXT_WINDOW), 70% of the 500k budget - suggest /stow now
rc=0 tokens=350000 record=s1 3 now
Evidence: Test file run

Source: Test file run

ok - estimator: percentage comes from the newest usage record, not the first
ok - estimator: widens past a usage-free tail to find the newest usage record
ok - estimator: empty auto-detection projects stay silent
ok - estimator: verdict bands are quiet under 40, next quiet moment to 60, now above 60
ok - estimator: falls back to transcript bytes / 4 when no usage record is readable
ok - estimator: an unreadable transcript prints nothing and exits 1
ok - estimator: window comes from --window, FM_CONTEXT_WINDOW, the compaction window, then 200000
ok - estimator: budget bands sit at 40 and 60 percent of the budget, overridable, clamped to the window
ok - nudge: steps and the once-per-step throttle follow the budget
ok - nudge: silent below 40 percent and records nothing
ok - nudge: the 60-to-61 band upgrade is announced once
ok - nudge: a downward band change permits the next upward crossing
ok - nudge: announces at most once per 20 percent step and again on the next step
ok - nudge: a different session id starts its own step count
ok - nudge: an unwritable state directory stays silent
ok - nudge: compaction steps the record down and the next real crossing announces again
ok - hook: an idle Claude turn end prints the suggestion once per step
ok - hook: no suggestion below 40 percent of the context window
ok - hook: the nudge waits for a low-disruption moment and then lands
ok - hook: persistent parent replies do not count as active task status
ok - hook: only the Claude Stop path prints the suggestion

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

🔧 **Review** - 1 issue found → auto-fixed ✅
  • ⚠️ tests/fm-context-budget.test.sh:158 - The new budget tests rely on the caller's environment. test_budget_bands_and_clamp (tests/fm-context-budget.test.sh:158-185) and test_nudge_steps_follow_the_budget (:188-203) pass only --window 600000. They assume the default 500000 budget and the default 40/60 bands, but nothing unsets FM_CONTEXT_BUDGET, FM_CONTEXT_NUDGE_SUGGEST or FM_CONTEXT_NUDGE_NOW. The script now reads all three, so a shell that exports one of them gives different bands and the tests fail. Same class, same file: the hook tests (:368 and :376) pin FM_CONTEXT_WINDOW=200000 but inherit FM_CONTEXT_BUDGET and the NUDGE_* thresholds, so a lower exported budget changes when the guard nudges. The 'nudge()' helper (:207-209) also passes only --window. The fix is to run these calls under env -u FM_CONTEXT_BUDGET -u FM_CONTEXT_NUDGE_SUGGEST -u FM_CONTEXT_NUDGE_NOW -u CLAUDE_CODE_AUTO_COMPACT_WINDOW, or to unset them once at the top of the test file.

🔧 Fix applied.
✅ Re-checked - no issues remain.

✅ **Test** - passed

✅ No issues found.

  • Live validation: ✅ go - 7 of 8 scenarios driven live against the product
Scenario Result Live Evidence
Default window 200000: a 500000 budget is clamped to the window and the line says so; bands quiet/next/now at 35/45/65 percent ✅ pass live live-script-run.txt section A
Window from CLAUDE_CODE_AUTO_COMPACT_WINDOW=1M with default 500k budget prints window percent and budget percent and band ✅ pass live live-script-run.txt section B
FM_CONTEXT_WINDOW beats CLAUDE_CODE_AUTO_COMPACT_WINDOW and the line names the source ✅ pass live live-script-run.txt section C
Explicit FM_CONTEXT_BUDGET=300000 moves bands to 40/60 percent of that budget ✅ pass live live-script-run.txt section D
Token overrides FM_CONTEXT_NUDGE_SUGGEST/NOW change the bands; invalid values fall back to defaults; --percent prints percent of window ✅ pass live live-script-run.txt sections E, F, G
Adversarial: nudge throttle stays silent below the band, repeats within a step, and announces once on each new step or band upgrade ✅ pass live live-script-run.txt section H (rc=1,1,0,1,0 with record s1 2 next then s1 3 now)
Tests are isolated from the caller's FM_CONTEXT_* environment (review round 1 fix) ✅ pass live test-run.txt; tests/fm-context-budget.test.sh unsets the vars at the top and all 21 pass
Turn-end guard Stop hook prints the nudge once per step and only at quiet moments ⏸️ untested no The prior payload did not establish a live result. It relied on the repo's hermetic hook tests, which run the real guard script, and no live Claude session was driven.
  • bash tests/fm-context-budget.test.sh (21 ok, includes the review-round env isolation and the real turn-end guard hook tests)
  • Hand-driven bin/fm-context-budget.sh against synthetic Claude transcripts: default window clamp, CLAUDE_CODE_AUTO_COMPACT_WINDOW window, FM_CONTEXT_WINDOW precedence, explicit FM_CONTEXT_BUDGET, token-threshold overrides, invalid values, --percent, --nudge throttle
  • Read the docs diff (docs/configuration.md, docs/turnend-guard.md) for the budget-below-window explanation and 2026-10-02 decision
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

@TastyTom13
TastyTom13 force-pushed the fm/firstmate-context-nudge-real-window branch from fcc0ff7 to 8a0037f Compare October 2, 2026 14:43
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant