Skip to content

Correct the amterr lab's layer-3 claims and make it rerunnable - #92

Merged
MaxGhenis merged 3 commits into
mainfrom
amterr-lab-claims-audit
Oct 3, 2026
Merged

MaxGhenis merged 3 commits into
mainfrom
amterr-lab-claims-audit

Conversation

@MaxGhenis

@MaxGhenis MaxGhenis commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

Corrects the error-case lab at paper/snapshot/labs/amterr and makes it rerunnable. Context: TheAxiomFoundation/axiom.org#297 rewrote the Colorado SNAP QC report after recomputing this lab; this PR fixes the lab itself. Each finding below was recomputed from the committed replay and the FY2024 QC file before the text changed.

What was wrong, and what the text says now

July text Recomputed
"10 of those 37 … $3.28M/yr = 3.3% of replayed error dollars" 10 cases, $4.47M (4.5%). $3.28M does not reproduce (HWGT × AMTERR gives $4.47M; HWGT × |engine − RAWBEN| gives $2.91M).
"on the facts the agency recorded, the verified engine returns the reviewer-certified correct benefit" The engine equals FSBEN in 7 of the 10. All 7 are solver no_change rows, where the replayed input is the file's own input. FSBEN is Mathematica's calculated benefit (tech doc PDF pp. 75, 91).
Example case 202312-40441 Carries cause 15 only, so it is outside the 10.
All 10 framed as computation 7 of the 10 broad-coded findings concern the computation (4 × 520/80, 520/75, 520/98, 362/52). The others are 331/44/17, 350/44/17, 161/6/10. The 7 carry codes 10 (5 cases) and 20 (2); none carries 17 or 19.
"no single-variable original value + correct math reproduces the issuance"; the 37 as an "upper bound on computation-side" The solver moves only the ELEMENT1 input. It moved nothing in 17 of the 37 (14 no_change, 3 eligible but unmoved) and in 16 of the 246 reproduced cases. The upper-bound reading is withdrawn: of the 26 cases with a layer-2 computational finding, 13 reproduce ($3.6M), 10 do not ($3.5M) and 3 were not replayed.
Broad class ≈ "~1.0 point of the rate" 10.5% of Colorado error dollars (1.05 points of 9.97), of which 5.7% (0.56) reproduces from a single input, 4.0% (0.40) does not, 0.9% (0.09) was not replayed. The 7 candidates are 1.8% (0.18 points; 0.16 on dollars above the $56 threshold).
14 software cases = "automation fed itself the wrong input"; 2 = "computation logic itself wrong" Of the 14, 10 moved the software-coded element ($1.8M), 2 moved the same unearned-income total through another finding ($1.7M), 2 moved rent ($0.8M). In the 2 that do not reproduce, the software finding is on an element the solver never moved. No software-coded finding is on rent; one mass-change finding is on the medical deduction. The combined "$7.6M ≈ 6.7%" figure is withdrawn.
"$88.8B issuance" nationally $91.4B (sum HWGT × RAWBEN); $88.8B was sum HWGT × FSBEN.
National codes 20/21 $107.1M; 10/22 $1,210M (14.9%) $106.9M; $1,188M (14.6%).
FY2025 "being measured now" FY2025 was published June 24, 2026: Colorado 10.09%.
Weight vintage unstated May 2026 posting. The August 2026 posting changed only HWGT/FYWGT (and the _OLD pair); case-level results are identical and dollars move slightly (data entry $18.47M → $18.43M). A table gives both.

Reproducibility

  • audit_claims.py → claims_audit.json recomputes every figure under both postings. It also regenerates native_decomposition.json and ../phase_a_classification.json from the May posting and checks them against the committed copies (both match). phase_a_classification.json was already committed one directory up; ANALYSIS.md now cites that path.
  • Paths: amterr_replay.py and reconstruct_co_fy2024.R take AMTERR_LAB_DIR (default: the script's directory) and SNAP_QC_REPO. No user-specific paths remain. No computation changed.
  • README.md gives pins and commands: giannella/snap_qc 741e10bf, axiom-oracles d34aa6fa, rulespec-us b53ce208, axiom-rules-engine de0efdc7, both QC postings by sha256.
  • Reruns at those pins (2026-10-03) reproduced co_fy2024_reconstruction.csv, fy2024_reconstruction_national.csv and amterr_replay_results.json byte for byte. The July engine binary was not recorded; the rerun used the 2026-08-06 build of the same commit attested in paper/snapshot/cert/CERT_REPORT.md.

Invariants and tests (tests/test_amterr_lab.py)

Invariants the audit must satisfy for any input:

  • reproduced + not reproduced = replayed, in cases and dollars; each cause class splits exactly into reproduced / not reproduced / not replayed;
  • computation candidates ⊆ broad-coded misses ⊆ non-reproduced cases, with dollars ordered the same way;
  • an unmoved non-reproduced row replays the file's inputs, so the engine equals FSBEN;
  • layer-2 classes partition the error cases and their dollars;
  • shares lie in [0, 1], are unchanged by rescaling weights, and points = share × 9.97;
  • any-presence membership is monotone in the code set;
  • case-level results and case counts are identical under both postings.

Tests: committed-file checks of those identities; Hypothesis property tests on the classifier and partitions (synthetic cases); a differential check that ANALYSIS.md quotes the audit's figures; a regeneration test against both raw postings, which skips where the QC files are absent (CI). The tests also compare the sha256 values the audit records with the live files, so editing an audited artifact fails in CI. Locally: 19 pass with the data, 18 pass and 1 skips without it. hypothesis joins the dev extra.

Review

Round 1 of the independent review (Opus 5.5 via Subfleet) requested changes: the "upper bound" sentence still overclaimed, the no-data tests did not lock the audited files' hashes, and four smaller points (FSBEN page citation, the audit's stated moved-input convention, the solver description, the replay script's docstring). Commit 8786e7d addresses each; round 2 is running.

Notes

  • analysis/cause_shares.json still cites "ANALYSIS.md:L62-L67" for the $3.28M claim. That line range is the pre-revision file (commit 755a7f3). I left the artifact unchanged because six downstream artifacts pin its hash; FACTS K2 now carries the commit-qualified pointer.
  • Two data-gated tests (test_fixed_donor_decomposition, test_uhip_decomposition raw regeneration) fail on this Mac on clean main because the committed artifacts record an older macOS version in environment.platform. Unrelated to this PR; they skip in CI.
  • The manuscript's layer-3 paragraph (paper/index.qmd) and FACTS D4/C4 are out of scope here and flagged separately.

axiom: n/a: analysis documentation, audit script and tests; no policy encoding.

🤖 Generated with Claude Code

MaxGhenis and others added 3 commits October 3, 2026 14:27
ANALYSIS.md, recomputed from the committed replay and the FY2024 QC file:
- The 10 broad-coded non-reproduced cases carry $4.47M (4.5% of replayed
  error dollars); the old $3.28M (3.3%) does not reproduce.
- The engine equals FSBEN in 7 of the 10, all solver no_change rows where
  the replayed input is the file's input. FSBEN is Mathematica's calculated
  benefit. The old example case, 202312-40441, carries cause 15 and is
  outside the 10.
- 7 of the 10 broad-coded findings concern the computation; they carry 1.8%
  of Colorado error dollars (0.18 points of 9.97; 0.16 above the $56
  threshold), with codes 10 and 20 and no 17 or 19.
- The solver moves only the ELEMENT1 input; it moved nothing in 17 of the
  37 non-reproduced cases and in 16 of the 246 reproduced ones.
- Of 14 reproduced software-coded cases, 10 moved the software-coded
  element, 2 the same unearned-income total through another finding, 2
  rent. The combined "$7.6M / 6.7%" figure is withdrawn.
- National issuance is $91.4B (the old $88.8B was weighted FSBEN); national
  codes 20/21 and 10/22 are $106.9M and $1,188M.
- States the weight vintage (May 2026 posting) and adds August 2026 figures.

Reproducibility:
- audit_claims.py regenerates every figure into claims_audit.json under
  both postings, and regenerates native_decomposition.json and
  ../phase_a_classification.json (both match the committed copies).
- amterr_replay.py and reconstruct_co_fy2024.R take AMTERR_LAB_DIR and
  SNAP_QC_REPO; no user-specific paths remain.
- README.md pins giannella/snap_qc, axiom-oracles, rulespec-us,
  axiom-rules-engine and both QC postings. Reruns at those pins reproduced
  both solver CSVs and amterr_replay_results.json byte for byte.
- tests/test_amterr_lab.py: committed-file checks, Hypothesis property
  tests, and a data-gated regeneration test.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…lver text

- ANALYSIS.md no longer calls the 37 non-reproduced cases an upper bound on
  computation-side error. Of the 26 cases with a layer-2 computational
  finding, 13 reproduce ($3.6M), 10 do not ($3.5M) and 3 were not replayed;
  two reproduced cases carry 366/56/17 (computer programming error).
  audit_claims.py now records both sets.
- Solver description follows reconstruct_co_fy2024.R: income steps stop on
  passing RAWBEN; rent, utility and deduction steps also stop within $3;
  the utility snap applies to every util_up/util_down row.
- FSBEN citation: listed as constructed on PDF p. 73, legend p. 75,
  formula p. 91.
- claims_audit.json states the moved-input convention the code uses.
- amterr_replay.py docstring no longer asserts what a match or miss means.
- Tests compare the sha256 values the audit records with the live files,
  and match whole sentences and table rows of ANALYSIS.md.
- README: download and hash-check steps for both postings; --frozen.
- FACTS K2: the 4.5% share attaches to $4.47M.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ssertion

The reviewer approved 8786e7d and suggested two optional precision fixes,
applied here verbatim: which steps stop at zero or at the shelter cap, and
a test asserting 26 - 13 - 10 = 3 computational-finding cases not replayed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@MaxGhenis

Copy link
Copy Markdown
Contributor Author

Independent review (Opus 5.5 via Subfleet, read-only):

  • Round 1 on 93bc46a: REQUEST CHANGES. Two majors (the "upper bound on computation-side" sentence still overclaimed; the no-data tests did not lock the audited files' hashes) and four minors (FSBEN page citation, the audit's stated moved-input convention, solver description, replay docstring).
  • Round 2 on 8786e7d: APPROVE. All six findings and the nits resolved; no blocking findings. Two optional precision suggestions were applied verbatim in a112f25.

The reviewer had no shell in either round and checked by reading the code, artifacts, R script and tech doc, with hand arithmetic from the replay weights. Tests were run by the author: 19 pass with the QC data; 18 pass and 1 skips without it; audit_claims.py --check is current. Out-of-scope follow-up the reviewer flagged (paper/FACTS.md D4, paper/index.qmd, and the classification string in analysis/cause_shares.*) is tracked separately.

@MaxGhenis
MaxGhenis merged commit c4333b3 into main Oct 3, 2026
1 check passed
@MaxGhenis
MaxGhenis deleted the amterr-lab-claims-audit branch October 3, 2026 20:55
@MaxGhenis

Copy link
Copy Markdown
Contributor Author

Merged (squash, c4333b3) at reviewed head a112f25. Gates at that head: test check passed (gh pr checks exit 0), MERGEABLE, not a draft, no CHANGES_REQUESTED. Independent review (Opus 5.5 via Subfleet; reports archived at ~/reviews/snap-qc-sim-pr92-2026-10-03/ on Max's Mac): round 1 REQUEST CHANGES on 93bc46a; round 2 APPROVE on 8786e7d; delta confirmation APPROVE on a112f25 (that commit applies round 2's two optional suggestions verbatim; the author confirmed the 8786e7d..a112f25 diff contains only those two edits). Follow-ups: manuscript, README and simulator copies (task_99975d8c), the cause_shares upper-bound labels (task_4d6d9ce4), local macOS-version test failures (task_095ddfeb).

MaxGhenis added a commit that referenced this pull request Oct 3, 2026
The lab's 2026-10-03 revision (PR #92) withdrew the reading that the 37
non-reproduced replay cases bound computation-side error. The
colorado_replay_reconciliation section of analysis/cause_shares.json still
carried it, so:

- rename the replay outcomes explained_input_facts -> reproduced and
  computation_side_upper_bound -> not_reproduced, single-sourced in
  REPLAY_OUTCOMES;
- state only the test in reference.classification, and raise if a replay
  row's within5 flag disagrees with abs(engine_on_original - RAWBEN) <= 5;
- qualify committed_prose_discrepancy.claim_location with commit 755a7f3
  and point claim_status at the lab's correction.

Regenerate every artifact that pins the changed hashes, transitively, with
its own generator: coding_consistency, engine_scenario_data (and the
app.js pin, with an ASSET_V bump), persistence, persistence_backtest,
interventions, uhip_decomposition, fixed_donor_decomposition and
adoption_national. Data leaves are unchanged; only input hashes and the
recorded macOS version move.

Tests lock the labels and classification, check the committed replay
against the documented test, assert the partition identities, compare the
reconciliation with the lab's independently computed claims_audit.json,
and property-test _replay_metric's additivity.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
MaxGhenis added a commit that referenced this pull request Oct 4, 2026
)

* Retire the "upper bound" label in the cause_shares replay crosswalk

The lab's 2026-10-03 revision (PR #92) withdrew the reading that the 37
non-reproduced replay cases bound computation-side error. The
colorado_replay_reconciliation section of analysis/cause_shares.json still
carried it, so:

- rename the replay outcomes explained_input_facts -> reproduced and
  computation_side_upper_bound -> not_reproduced, single-sourced in
  REPLAY_OUTCOMES;
- state only the test in reference.classification, and raise if a replay
  row's within5 flag disagrees with abs(engine_on_original - RAWBEN) <= 5;
- qualify committed_prose_discrepancy.claim_location with commit 755a7f3
  and point claim_status at the lab's correction.

Regenerate every artifact that pins the changed hashes, transitively, with
its own generator: coding_consistency, engine_scenario_data (and the
app.js pin, with an ASSET_V bump), persistence, persistence_backtest,
interventions, uhip_decomposition, fixed_donor_decomposition and
adoption_national. Data leaves are unchanged; only input hashes and the
recorded macOS version move.

Tests lock the labels and classification, check the committed replay
against the documented test, assert the partition identities, compare the
reconciliation with the lab's independently computed claims_audit.json,
and property-test _replay_metric's additivity.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Test that the generator rejects a within5 flag contradicting the test

Addresses the independent review: no test exercised the new guard in
colorado_replay_reconciliation, so removing it left the suite green. Three
synthetic one-row replays (a $100 gap flagged reproduced, a $3 gap flagged
not reproduced, a missing engine benefit flagged reproduced) now must raise.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Test that the within5 guard accepts matching flags at the $5 boundary

The rejection cases alone would not notice an le -> lt slip, and the
committed replay has no gap of exactly $5. Three consistent one-row replays
(gap of exactly $5 flagged reproduced, $6 gap and missing engine benefit
flagged not reproduced) must now get past the guard to the next call.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant