Correct the amterr lab's layer-3 claims and make it rerunnable - #92
Conversation
ANALYSIS.md, recomputed from the committed replay and the FY2024 QC file: - The 10 broad-coded non-reproduced cases carry $4.47M (4.5% of replayed error dollars); the old $3.28M (3.3%) does not reproduce. - The engine equals FSBEN in 7 of the 10, all solver no_change rows where the replayed input is the file's input. FSBEN is Mathematica's calculated benefit. The old example case, 202312-40441, carries cause 15 and is outside the 10. - 7 of the 10 broad-coded findings concern the computation; they carry 1.8% of Colorado error dollars (0.18 points of 9.97; 0.16 above the $56 threshold), with codes 10 and 20 and no 17 or 19. - The solver moves only the ELEMENT1 input; it moved nothing in 17 of the 37 non-reproduced cases and in 16 of the 246 reproduced ones. - Of 14 reproduced software-coded cases, 10 moved the software-coded element, 2 the same unearned-income total through another finding, 2 rent. The combined "$7.6M / 6.7%" figure is withdrawn. - National issuance is $91.4B (the old $88.8B was weighted FSBEN); national codes 20/21 and 10/22 are $106.9M and $1,188M. - States the weight vintage (May 2026 posting) and adds August 2026 figures. Reproducibility: - audit_claims.py regenerates every figure into claims_audit.json under both postings, and regenerates native_decomposition.json and ../phase_a_classification.json (both match the committed copies). - amterr_replay.py and reconstruct_co_fy2024.R take AMTERR_LAB_DIR and SNAP_QC_REPO; no user-specific paths remain. - README.md pins giannella/snap_qc, axiom-oracles, rulespec-us, axiom-rules-engine and both QC postings. Reruns at those pins reproduced both solver CSVs and amterr_replay_results.json byte for byte. - tests/test_amterr_lab.py: committed-file checks, Hypothesis property tests, and a data-gated regeneration test. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…lver text - ANALYSIS.md no longer calls the 37 non-reproduced cases an upper bound on computation-side error. Of the 26 cases with a layer-2 computational finding, 13 reproduce ($3.6M), 10 do not ($3.5M) and 3 were not replayed; two reproduced cases carry 366/56/17 (computer programming error). audit_claims.py now records both sets. - Solver description follows reconstruct_co_fy2024.R: income steps stop on passing RAWBEN; rent, utility and deduction steps also stop within $3; the utility snap applies to every util_up/util_down row. - FSBEN citation: listed as constructed on PDF p. 73, legend p. 75, formula p. 91. - claims_audit.json states the moved-input convention the code uses. - amterr_replay.py docstring no longer asserts what a match or miss means. - Tests compare the sha256 values the audit records with the live files, and match whole sentences and table rows of ANALYSIS.md. - README: download and hash-check steps for both postings; --frozen. - FACTS K2: the 4.5% share attaches to $4.47M. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ssertion The reviewer approved 8786e7d and suggested two optional precision fixes, applied here verbatim: which steps stop at zero or at the shelter cap, and a test asserting 26 - 13 - 10 = 3 computational-finding cases not replayed. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Independent review (Opus 5.5 via Subfleet, read-only):
The reviewer had no shell in either round and checked by reading the code, artifacts, R script and tech doc, with hand arithmetic from the replay weights. Tests were run by the author: 19 pass with the QC data; 18 pass and 1 skips without it; |
|
Merged (squash, c4333b3) at reviewed head a112f25. Gates at that head: |
The lab's 2026-10-03 revision (PR #92) withdrew the reading that the 37 non-reproduced replay cases bound computation-side error. The colorado_replay_reconciliation section of analysis/cause_shares.json still carried it, so: - rename the replay outcomes explained_input_facts -> reproduced and computation_side_upper_bound -> not_reproduced, single-sourced in REPLAY_OUTCOMES; - state only the test in reference.classification, and raise if a replay row's within5 flag disagrees with abs(engine_on_original - RAWBEN) <= 5; - qualify committed_prose_discrepancy.claim_location with commit 755a7f3 and point claim_status at the lab's correction. Regenerate every artifact that pins the changed hashes, transitively, with its own generator: coding_consistency, engine_scenario_data (and the app.js pin, with an ASSET_V bump), persistence, persistence_backtest, interventions, uhip_decomposition, fixed_donor_decomposition and adoption_national. Data leaves are unchanged; only input hashes and the recorded macOS version move. Tests lock the labels and classification, check the committed replay against the documented test, assert the partition identities, compare the reconciliation with the lab's independently computed claims_audit.json, and property-test _replay_metric's additivity. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
) * Retire the "upper bound" label in the cause_shares replay crosswalk The lab's 2026-10-03 revision (PR #92) withdrew the reading that the 37 non-reproduced replay cases bound computation-side error. The colorado_replay_reconciliation section of analysis/cause_shares.json still carried it, so: - rename the replay outcomes explained_input_facts -> reproduced and computation_side_upper_bound -> not_reproduced, single-sourced in REPLAY_OUTCOMES; - state only the test in reference.classification, and raise if a replay row's within5 flag disagrees with abs(engine_on_original - RAWBEN) <= 5; - qualify committed_prose_discrepancy.claim_location with commit 755a7f3 and point claim_status at the lab's correction. Regenerate every artifact that pins the changed hashes, transitively, with its own generator: coding_consistency, engine_scenario_data (and the app.js pin, with an ASSET_V bump), persistence, persistence_backtest, interventions, uhip_decomposition, fixed_donor_decomposition and adoption_national. Data leaves are unchanged; only input hashes and the recorded macOS version move. Tests lock the labels and classification, check the committed replay against the documented test, assert the partition identities, compare the reconciliation with the lab's independently computed claims_audit.json, and property-test _replay_metric's additivity. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Test that the generator rejects a within5 flag contradicting the test Addresses the independent review: no test exercised the new guard in colorado_replay_reconciliation, so removing it left the suite green. Three synthetic one-row replays (a $100 gap flagged reproduced, a $3 gap flagged not reproduced, a missing engine benefit flagged reproduced) now must raise. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Test that the within5 guard accepts matching flags at the $5 boundary The rejection cases alone would not notice an le -> lt slip, and the committed replay has no gap of exactly $5. Three consistent one-row replays (gap of exactly $5 flagged reproduced, $6 gap and missing engine benefit flagged not reproduced) must now get past the guard to the next call. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Corrects the error-case lab at
paper/snapshot/labs/amterrand makes it rerunnable. Context: TheAxiomFoundation/axiom.org#297 rewrote the Colorado SNAP QC report after recomputing this lab; this PR fixes the lab itself. Each finding below was recomputed from the committed replay and the FY2024 QC file before the text changed.What was wrong, and what the text says now
no_changerows, where the replayed input is the file's own input. FSBEN is Mathematica's calculated benefit (tech doc PDF pp. 75, 91).no_change, 3 eligible but unmoved) and in 16 of the 246 reproduced cases. The upper-bound reading is withdrawn: of the 26 cases with a layer-2 computational finding, 13 reproduce ($3.6M), 10 do not ($3.5M) and 3 were not replayed._OLDpair); case-level results are identical and dollars move slightly (data entry $18.47M → $18.43M). A table gives both.Reproducibility
audit_claims.py→claims_audit.jsonrecomputes every figure under both postings. It also regeneratesnative_decomposition.jsonand../phase_a_classification.jsonfrom the May posting and checks them against the committed copies (both match).phase_a_classification.jsonwas already committed one directory up; ANALYSIS.md now cites that path.amterr_replay.pyandreconstruct_co_fy2024.RtakeAMTERR_LAB_DIR(default: the script's directory) andSNAP_QC_REPO. No user-specific paths remain. No computation changed.README.mdgives pins and commands: giannella/snap_qc741e10bf, axiom-oraclesd34aa6fa, rulespec-usb53ce208, axiom-rules-enginede0efdc7, both QC postings by sha256.co_fy2024_reconstruction.csv,fy2024_reconstruction_national.csvandamterr_replay_results.jsonbyte for byte. The July engine binary was not recorded; the rerun used the 2026-08-06 build of the same commit attested inpaper/snapshot/cert/CERT_REPORT.md.Invariants and tests (
tests/test_amterr_lab.py)Invariants the audit must satisfy for any input:
Tests: committed-file checks of those identities; Hypothesis property tests on the classifier and partitions (synthetic cases); a differential check that ANALYSIS.md quotes the audit's figures; a regeneration test against both raw postings, which skips where the QC files are absent (CI). The tests also compare the sha256 values the audit records with the live files, so editing an audited artifact fails in CI. Locally: 19 pass with the data, 18 pass and 1 skips without it.
hypothesisjoins thedevextra.Review
Round 1 of the independent review (Opus 5.5 via Subfleet) requested changes: the "upper bound" sentence still overclaimed, the no-data tests did not lock the audited files' hashes, and four smaller points (FSBEN page citation, the audit's stated moved-input convention, the solver description, the replay script's docstring). Commit 8786e7d addresses each; round 2 is running.
Notes
analysis/cause_shares.jsonstill cites "ANALYSIS.md:L62-L67" for the $3.28M claim. That line range is the pre-revision file (commit 755a7f3). I left the artifact unchanged because six downstream artifacts pin its hash; FACTS K2 now carries the commit-qualified pointer.test_fixed_donor_decomposition,test_uhip_decompositionraw regeneration) fail on this Mac on cleanmainbecause the committed artifacts record an older macOS version inenvironment.platform. Unrelated to this PR; they skip in CI.paper/index.qmd) and FACTS D4/C4 are out of scope here and flagged separately.axiom: n/a: analysis documentation, audit script and tests; no policy encoding.
🤖 Generated with Claude Code