Distributional backtest: sampling-only vs persistence-aware forecasts - #87
Merged
Merged
Conversation
Expanding-window backtest over FY2016-24 targets (424 state-years): three Normal predictive constructions for next-year state rates, scored with CRPS, pinball loss, log score, and interval coverage. The sampling-only construction (the published-CI construction read as a forecast) covers 48% at nominal 90% and loses every target year; adding the AR(1) innovation to the variance alone (no shrinkage) reaches 81% at about double the predictive SD; the full persistence model reaches 85% and wins the large-movement years. National year level carried identically for all three. A penalty translation prices the gap: centered on FY2025 officials, delay-aware FY2028 bills on official issuance put the median state's most-likely-bucket probability at 0.84 under sampling-only versus 0.59 with persistence in the variance; the national bill SD rises from $0.84B to $1.01B under independent state draws (a lower bound absent cross-state process correlation). Artifact locks cover score domains, per-year winner consistency, bucket-probability coherence, live input hashes, and exact raw regeneration. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…n-correct bills Three blockers from the review, all fixed: 1. FY2028 semantics: the penalty translation now mirrors the simulator's verified election machinery — FY2028 keys to the elected minimum of the locked FY2025 rate and the simulated FY2026 measurement, zero when either crosses the delay test; FY2029 keys to FY2026 alone. Illinois (locked 14.67, delayed) now prices exactly zero with zero variance; a statutory lock enforces it. 2. Steel-manned static: a static_fair construction carries the anchor year's own sampling error (var = 2v), and sensitivity rows rerun the static family at the published design-SE scale (1.1pp). At that scale with anchor error the static family is roughly calibrated (90% coverage at 1.56pp width); the persistence model reaches the same calibration sharper (CRPS 0.878 vs 0.901), and the naive single-SE forward read covers 48-64%. 3. Disclosures: the reconstructed-to-official variance transfer is stated as an additional assumption beyond the additive-wedge level evidence, and the national SD claims no cross-state correlation rather than a lower bound. 18 tests green including exact regeneration and the new statutory and sensitivity locks. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This was referenced Sep 1, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Scores three predictive constructions for next-year state error rates on held-out years (FY2016–24, 424 state-years): the sampling-only construction (published-CI style, read as a forecast), a no-shrinkage variant with the AR(1) innovation added to the variance, and the full persistence model. Proper scoring rules throughout (CRPS, pinball over 19 quantiles, log score) plus interval coverage.
Result: sampling-only covers 48% at nominal 90% and wins zero of eight years; persistence-in-the-variance alone reaches 81% at ~double the predictive SD; the full model reaches 85% and takes the large-movement years. The penalty translation prices it: median state most-likely-bucket probability 0.84 → 0.59, national FY2028 bill SD $0.84B → $1.01B (independent draws; correlation would widen further).
Deterministic, seeded; artifact + generated memo + locks (score domains, winner consistency, bucket coherence, live input hashes incl. movement/issuance/persistence artifacts, exact regeneration — 16 tests green with test_persistence).
🤖 Generated with Claude Code