Conversation
The lab's 2026-10-03 revision (PR #92) withdrew the reading that the 37 non-reproduced replay cases bound computation-side error. The colorado_replay_reconciliation section of analysis/cause_shares.json still carried it, so: - rename the replay outcomes explained_input_facts -> reproduced and computation_side_upper_bound -> not_reproduced, single-sourced in REPLAY_OUTCOMES; - state only the test in reference.classification, and raise if a replay row's within5 flag disagrees with abs(engine_on_original - RAWBEN) <= 5; - qualify committed_prose_discrepancy.claim_location with commit 755a7f3 and point claim_status at the lab's correction. Regenerate every artifact that pins the changed hashes, transitively, with its own generator: coding_consistency, engine_scenario_data (and the app.js pin, with an ASSET_V bump), persistence, persistence_backtest, interventions, uhip_decomposition, fixed_donor_decomposition and adoption_national. Data leaves are unchanged; only input hashes and the recorded macOS version move. Tests lock the labels and classification, check the committed replay against the documented test, assert the partition identities, compare the reconciliation with the lab's independently computed claims_audit.json, and property-test _replay_metric's additivity. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Contributor
Author
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Aligns
analysis/cause_shares.pyandanalysis/cause_shares.jsonwith the corrected amterr lab reading from #92. That PR withdrew the claim that the 37 replay cases that do not reproduce the issued benefit are an "upper bound on computation-side error". Of the 26 Colorado cases with a layer-2 computational finding, 13 reproduce, 10 do not and 3 were not replayed, and the solver moved nothing in 17 of the 37 (paper/snapshot/labs/amterr/ANALYSIS.md, "Layer 3: engine replay / Results";claims_audit.jsoncase_level).Stacked on #92. This PR's base is
amterr-lab-claims-audituntil #92 merges. After that I rebase ontomainand retarget.What changes in
colorado_replay_reconciliationexplained_input_facts,computation_side_upper_boundreproduced,not_reproducedreference.classificationcommitted_prose_discrepancy.claim_locationpaper/snapshot/labs/amterr/ANALYSIS.md:L62-L67… :L62-L67 at commit 755a7f3179e964dccdf6dc2f7266e7c4758a0141committed_prose_discrepancy.claim_statusclaims_audit.jsoncase_level.broad_coded_misses)The labels are single-sourced in
REPLAY_OUTCOMES. The generator now raises if any replay row'swithin5flag disagrees withabs(engine_on_original - RAWBEN) <= 5, so the classification string is checked against the data rather than asserted. All 88 values under the renamed keys are unchanged.No other code reads the old keys. A grep of
app/,analysis/,paper/,paper-causal/,tests/,scripts/anddocs/finds them only in these two files.Hash-chain regeneration
cause_shares.jsonrecords the sha256 of its generator, and other artifacts pin the JSON's sha256, some of them transitively. Each artifact below was regenerated with its own generator, in dependency order. A leaf-by-leaf diff confirms that no data value changed. The only changes are input hashes and, where an artifact records it,environment.platform.analysis/cause_shares.jsonprovenance.generator.sha256analysis/coding_consistency.jsonprovenance.reference_artifact_sha256app/public/engine_scenario_data.jsonprovenance.cause_shares_sha256app/public/app.jsADOPT_DATA_SHA256pin;ASSET_V20260820b → 20261003a (withindex.html), per the cache-buster rule, so browsers do not pair a cached payload with the new pinanalysis/adoption_national.jsonprovenance.engine_scenario_data_sha256analysis/persistence_results.jsoninput_hashes.coding_consistency;environment.platformanalysis/persistence_backtest_results.jsoninput_hashes.{coding_consistency,persistence_results};environment.platformanalysis/interventions_results.jsoninput_hashes.{coding_consistency,persistence_results}analysis/uhip_decomposition_results.jsoninput_hashes.{cause_shares,coding_consistency};environment.platformanalysis/fixed_donor_decomposition_results.jsoninput_hashes.{cause_shares,coding_consistency,joint_fit_results};environment.platformEnvironment strings. I regenerated on Python 3.14.4 with numpy 2.5.1, pandas 2.3.3 and pyreadstat 1.3.5 (the recorded versions), so
pythonand package strings are unchanged.environment.platformmoves frommacOS-26.5.1-…tomacOS-26.6.2-…because the Mac's OS was updated; this is environment noise, not data. A side effect: the two local raw-regeneration tests that failed on cleanmainonly because of that string (test_fixed_donor_decomposition,test_uhip_decomposition) now pass on this machine. They still skip in CI.The memos the generators also write (PERSISTENCE.md, PERSISTENCE_BACKTEST.md, INTERVENTIONS.md, UHIP_DECOMPOSITION.md, FIXED_DONOR.md) came out byte-identical.
build_interventions_app_data.pyreproducesapp/public/interventions_data.jsonbyte for byte.Left as historical records (deliberately not changed). Both still contain the old
cause_shares.jsonhashc5b16617…:analysis/NATURE_REPORT.mdis the dated 2026-08-11 lane report for PR Commit the layer-2 finding-nature tabulation; quote it in the paper #39. Its gate table records that the artifact hashed toc5b16617…on two runs that day, which remains true. It has no generator and no test reads it.paper-causal/_fixed_donor_results.readonly.jsonis the frozen input snapshot from PR The migrations paper: three SNAP eligibility-system replacements in the QC record (revision 1) #74 (2026-08-16). It has never been refreshed: onmainit already differs fromanalysis/fixed_donor_decomposition_results.jsonin three input hashes (coding_consistency,joint_fit_results,system_migrations), and no test compares the two.tests/test_migrations_figures.pypins the snapshot's own hash and skips without matplotlib.Coordination with the manuscript task
The manuscript prose (
paper/index.qmd,paper/FACTS.mdD4) still calls the 37 a "computation-side upper bound". That text belongs to the separate manuscript task ("Correct FSBEN and reconstruction claims in the snap-qc-sim paper"), so this PR does not touchpaper/. Neither file names the renamed keys, so nothing there is forced. FACTS K2 citesanalysis/cause_shares.json colorado_replay_reconciliationby section name, which still resolves.Invariants and tests (
tests/test_cause_shares.py)Invariants, each executed:
within5 == (abs(engine_on_original - rawben) <= 5). The generator enforces this too.{reproduced, not_reproduced}, and the section contains no "upper bound" or "computation_side" text.cause_shares.pyreads the FY2024 SAV; the lab'saudit_claims.pyreads the May 2026 CSV posting. They agree on 283/246/37 cases, on dollars to the cent ($99,128,122.62 / $71,388,303.72 / $27,739,818.90), on the 76/21 above-threshold split, on 283/283 solver-engine concordance, and on the broad residual (10 cases, $4,469,987.48).claim_locationnames a full 40-hex commit. The cited ANALYSIS.md section and figures exist in the revised file, and lines 62–67 at that commit hold the "$3.28M/yr = 3.3%" claim. The git-history check skips in shallow CI checkouts._replay_metricis additive over any two-way partition of a slice, for arbitrary weights and amounts.Against the pre-change JSON, four of the new tests fail; against the new JSON all 14 tests in the file pass.
Verification
Full local suite and independent review: in progress; results will be added here.
axiom: n/a: analysis artifact labels, regenerated provenance hashes and tests; no policy encoding.
🤖 Generated with Claude Code