Skip to content

Correct retired SNAP QC claims in the paper, README, FACTS and simulator - #95

Open
MaxGhenis wants to merge 9 commits into
mainfrom
retire-qc-claims
Open

MaxGhenis wants to merge 9 commits into
mainfrom
retire-qc-claims

Conversation

@MaxGhenis

@MaxGhenis MaxGhenis commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

TheAxiomFoundation/axiom.org#297 retires several claims about USDA's SNAP Quality Control file and the Colorado error-case replay. A sweep (findings off-18 and oogs-5 to oogs-14) found copies here: in the working paper, the README, the fact catalog and the live simulator. This PR corrects all 15 locations, plus the copies of the same claims that sit next to them. It builds on #92 (the lab's claims_audit.json, the source for every replay number) and #94 (which retired the "upper bound" label in analysis/cause_shares.json).

axiom: n/a: manuscript, README and simulator copy only; no policy encoding changed.

Before and after

Line numbers are at 755a7f3, before #92. Numbers are sourced in the next section.

# Where Before After
1 app/public/index.html:133 (live simulator, embedded on policyengine.org); the same panel's FSBEN line in app.js "FSBEN is recomputed by the QC Minimodel used by the Food and Nutrition Administration (formerly FNS) — a closed model" "Mathematica computes the QC file's formula benefit (FSBEN) for USDA from each edited case record with the benefit formula of the QC Minimodel, one of the SNAP microsimulation models of the Food and Nutrition Administration (formerly FNS)". "A closed model" is cut: the tech doc doesn't say so, and it documents the model in chapter IV.
2 paper/index.qmd:108 "parity certifies agreement with the agency's own computational canon, not independent adjudication" "parity certifies agreement with the benefit computation recorded in the file; it does not adjudicate case outcomes". "Internally consistent case records" becomes "edited case records" (59 of 856 Colorado cases end more than $5 from BENFIX).
3 :299 "inputs and the agency's own computed outputs" "edited inputs and the benefit chain Mathematica computes for USDA from the edited record" (the codebook marks FSGRINC, FSSTDDED, FSSLTDED, FSNETINC, BENMAX and FSBEN as constructed)
4 :349 FSBEN "is computed by the FNS QC Minimodel — the agency's own benefit-calculation software"; editing adjusts deductions "until the calculated benefit matches the raw benefit within $5" Mathematica computes FSBEN for USDA with the QC Minimodel's benefit formula, "one of FNS's SNAP microsimulation models". Steps 12–13 are described as the tech doc gives them, and the sequence "does not always close the gap: in Colorado, FSBEN ends within $5 of BENFIX in 797 of 856 cases"
5 :362 "agrees exactly with the agency's own computational canon" "agrees exactly with the benefit computation recorded in the edited file"
6 :368 "the counterparty (the agency's own canon)" "the counterparty (a computation Mathematica runs for USDA on the edited case records)"
7 :1102 "against the agency's own computational canon — where the bar can be" "against the benefit computation recorded in the file, where the bar can be"
8 :25 (abstract) "explains 78.4% … — and 91.4% … — as correct arithmetic on wrong facts" "For 283 of the 305 Colorado cases with a recorded payment deviation, a public solver changes the input named by each case's first finding, where it can; run on those inputs, the verified engine reproduces the issued benefit within $5 for 78.4% … and 91.4% …, which is consistent with correct arithmetic on a wrong input." "Pricing that decomposition" becomes "Pricing a cause-code decomposition", since the pricing uses cause codes.
9 :447 "the agency's arithmetic was correct, applied to wrong facts" "which is consistent with correct arithmetic applied to a wrong input" (the solver moves one input)
10 :443 the solver "exploits the file's disclosure-protection perturbation structure to recover original values" What reconstruct_co_fy2024.R does. It starts from the file's edited inputs and moves only the ELEMENT1 input. Household size moves once, by one person: one fewer for natures 12, 14 or 16, one more for nature 7; the nature code sets the direction. Income, rent, utility and deduction inputs move in $3 steps toward RAWBEN until a stop rule binds, and a stepped utility amount is then reset to a commonly reported value. Nothing moves when the first finding falls outside those codes. A footnote gives the stop rules and the snap rule exactly. For example, income steps have no within-$3 stop, so where RAWBEN is the maximum allotment income-lowering steps run to zero income or the step limit; the snap picks the common amount nearest the stepped one. The layer-3 table cell names no stepping or reset rule: "A public solver changes the input named by each case's first finding, where it can".
11 :456 "33 of the 246 explained cases have deviations of $5 or less, mechanically within the comparison tolerance" The count stands; only "explained" is retired. The body now says: "In 230 of the 246 the solver moved an input; in the other 16 it took no step, and the issued benefit was already within $5 of FSBEN, so the match restates the parity result. In 19 of the 230 the issued benefit was also within $5 of FSBEN before the solver moved, so those matches did not need the move." The 33 are AMTERR ≤ $5. In all 33, and in 2 more with AMTERR of $6, RAWBEN started within $5 of FSBEN, so none of those 35 matches needed a move (FACTS H8).
12 :711 (fig-co caption) "official rate sits 0.03 points below the 15%-share boundary" "official fiscal 2024 rate sits 0.03 points below the 15%-share boundary, and its fiscal 2025 rate, 10.09%, sits above it"
13 README.md:31 "official rate 9.97%, 0.03 points from the 15% boundary" "official FY 2024 rate 9.97%, 0.03 points below the 10% rate where the 15% cost share begins, …; its FY 2025 rate, 10.09%, is above it"
14 paper/FACTS.md D4 "explains 246/283 … as correct arithmetic on wrong facts; … 37 cases form the computation-side upper bound" It states what the solver does and that a match is consistent with a wrong input, gives the 230/16 and 20/17 (14 + 3) splits and the 13/10/3 split of the 26 computational-finding cases, and records the old wording as SUPERSEDED
15 paper/FACTS.md C4 tech-doc errata "documented in the lab analysis and report page" WITHDRAWN. No source records them: neither version of ANALYSIS.md (27e1b09, c4333b3), the axiom.org report (original and corrected text) nor the axiom-oracles SNAP QC playbook; GitHub issue searches in both orgs find none; round-1 and round-2 referees couldn't find the "reported upstream" thread. The manuscript's own copies are cut: the abstract, the introduction ("defects on three sides" becomes "two sides") and @sec-oracle-scope ("errata … (reported upstream)").

Copies of the same retired claims next to these, corrected for consistency:

  • Replay "upper bound", which ANALYSIS.md (Correct the amterr lab's layer-3 claims and make it rerunnable #92) and cause_shares.json (Retire the "upper bound" label in the cause_shares replay crosswalk #94) both withdrew. It appeared in the layer-3 row of tbl-decompose, the 21.6% "computation-side upper bound for official errors" in the replay paragraph, "bounded above by the replay's 21.6% case share", and the limitations' "residual class is an upper bound". The table cell now says the replay does not identify the computation share and gives the reproduction rates. A new paragraph says neither outcome identifies a computation error: 17 of 37 misses moved nothing (14 no_change + 3 stopped), and of the 26 computational-finding cases, 13 reproduce, 10 do not and 3 were not replayed. Because the replay no longer estimates a computation share, the intro's "three nested estimates" becomes two estimates plus a replay that tests one input per case; the section opener's "three nested answers" becomes "three layers of evidence"; and the caption drops "nested".
  • The codebook row for FSBEN ("recomputed by the FNS QC Minimodel"), the roadmap's "Minimodel-canon scope point", the limitations' "Minimodel-computed chain", and the Disclosure's "every parity claim replays against the government's own recorded values" (now "the benefit computation recorded in the QC file").
  • The simulator's engine panel. It said "the Minimodel's full-formula recomputation (FSBEN)" (app.js); it now says Mathematica computes it for USDA with the QC Minimodel's benefit formula. The heading "Axiom rules engine vs the FNA QC Minimodel" and the matching hint now name the QC file's benefit chain. The same wording in analysis/engine_comparison.py and its generated ENGINE_COMPARISON.md is updated too: regenerating from the May posting reproduces engine_data.json byte for byte (sha256 2517d26e…, the pinned value) and the report with only that line changed. ASSET_V is bumped to 20261004a, since app.js changed.
  • @sec-oracle-scope no longer calls the gap between issued amounts and the computed chain "a finding about somebody's rules". The intro asks about "computation failure versus information failure". references.bib drops "pre-edit" from the solver note.
  • The intro's "defects on three sides" now says the process "surfaced two defects in the encodings under test and left a residue in issued amounts". The replay no longer supports the information-failure conclusion, which rests on layers 1 and 2. The layer-3 table cell gives the miss counts: of 37, 17 moved nothing and 10 carry a computational finding, 8 of them both.
  • FACTS F2 (the 0.03pp margin) now carries the FY2025 rate. FACTS C7 is new and states what the parity target is, with tech-doc pages. FACTS G gains four prohibitions, one per retired claim.
  • app/public/paper/: re-rendered (Quarto 1.9.36, the same version as revision 9). The wrapper now reads revision 10 · 2026-10-03, with a one-line note on what changed (it also still said "revision-7"), and its iframe cache key is r10-20261003. The manuscript date: is now 2026-10-03. The clean render also drops a "Notebooks" nav that revision 9 picked up from untracked files and that linked a local /Users/maxghenis/.cache/... path.

Where every number comes from

Each figure below was read this session from the source named.

Figure Source
283, 305, 22, 246, 230, 16, 37, 17 (14 + 3), 26 / 13 / 10 / 3, 283 solver-engine agreement paper/snapshot/labs/amterr/claims_audit.json (case_level.replay, layer2_computational_findings, by_posting.may2026.replay, totals.CO.error_cases)
76/97 = 78.4%, 170/186 = 91.4%, 246/283 = 86.9% analysis/cause_shares.json colorado_replay_reconciliation.slices (#94 key names), cross-checked against the per-case replay rows with AMTERR > $56
33 cases with AMTERR ≤ $5, 16 of them unmoved (FACTS H8 only) amterr_replay_results.json and claims_audit.json replay_inputs_unchanged_keys
797 of 856 within $5, 743 exact Recomputed from USDA's FY2024 QC file (August 2026 posting, sha256 e871a8e9…; the May posting gives the same, since only the weights differ). Matches the axiom.org ledger.
9.97% (FY2024), 10.09% (FY2025, table dated 2026-06-24), 0.03 points claims_audit.json inputs.official_rates_percent (FNS PER tables)
$3 steps, 1,000 steps, stop rules, ELEMENT1, household size ± 1 (element 150, natures 7/12/14/16), utility snap, 2 snap-only moves reconstruct_co_fy2024.R (read this session; it has no perturbation logic); claims_audit.json moved_with_zero_correctedamount_keys
17 unmoved misses, 10 with a computational finding, 8 both claims_audit.json not_reproduced_cases and not_reproduced_with_computational_finding_keys
Mathematica for USDA; "one of FNA's SNAP microsimulation models"; constructed variables; Steps 12–13 Tech doc, August 2026 posting: title page and acknowledgments (Mathematica for FNA, contract 12-3198-24-Q-0029), PDF p.15, pp.36–37, p.65 (the Minimodel's FSBEN calculation points to the codebook entry), p.71, p.73, p.91

The paper keeps its stated convention of "FNS" (line 88). The simulator already says FNA.

Invariants the corrected text relies on

tests/test_retired_claims.py checks each of these. It needs only committed files, except for the 797-of-856 check, which needs the QC file and skips without it.

  1. Replay partition. reproduced + not reproduced = replayed (246 + 37 = 283 = solver-engine agreement). Moved + unmoved = reproduced (230 + 16). Above-threshold + sub-threshold = replayed (97 + 186), and their reproduced counts sum to 246.
  2. Mechanical matches. Every unmoved reproduced case has engine = FSBEN and |RAWBEN − FSBEN| ≤ $5 (16 of 16). This is what licenses "the match restates the parity result".
  3. Differential check. The threshold split in cause_shares.json equals the split recomputed from the per-case replay rows (AMTERR > $56: 97 cases, 76 reproduced; ≤ $56: 186 cases, 170 reproduced).
  4. Margin direction. FY2024 CO 9.97 < 10 ≤ FY2025 CO 10.09, so "below" and "above" are both true. Every "0.03 points", "0.03-point" or "0.03pp" in the manuscript, README and FACTS has "10.09%" within 200 characters.
  5. Posting invariance. 797/856 and 743 are the same under both postings.
  6. Prose lock. Whole sentences in the manuscript are rebuilt from the artifacts, so a changed artifact fails the test. Hypothesis properties check that changing any single count changes the required prose (the lock can't pass on stale text), that rendering is deterministic, and that complementary shares sum to 100 ± 0.1.
  7. Retired wording. It is absent from paper/index.qmd, the rendered HTML and PDF (the PDF check uses pdftotext), README.md, app/public/index.html and app/public/app.js. FACTS keeps it only inside SUPERSEDED, WITHDRAWN and prohibition rows.

Headline findings: queued for Max

No headline number changes. Parity (6,081 of 6,194), the 46–50% tier odds, the $7.7B → $7.6B / $6.9B pricing, 18 of 53 tier changes, and 78.4% / 91.4% are all unchanged.

What the abstract and introduction claim does change:

  • The abstract's replay sentence goes from "explains … as correct arithmetic on wrong facts" to "reproduces … consistent with correct arithmetic on a wrong input" (item 8, as specified).
  • The abstract's errata clause is cut (item 15: no source exists).
  • The introduction lists two estimates of the computation share plus a replay test, where it listed three nested estimates. Layer 3 no longer bounds the computation share; Correct the amterr lab's layer-3 claims and make it rerunnable #92 had already withdrawn that bound in the lab analysis.
  • The conclusion's "most measured error among eligible cases is information failure" now rests on layers 1 and 2 (86.0% "other input" in layer 2; 7.4–10.6% computing-apparatus codes in layer 1).

Both review rounds judged that these alter a headline finding (round 1: "narrowly"). The task's rule sends a change that alters a headline finding to Max, so the merge is queued as decision d933 instead of self-merged. The production deploy, which outlives the merge, is decision d934.

After merge

  • The paper and simulator only change live after a manual Vercel deploy. The snap-qc-sim Vercel project has no git integration. Production deploys are CLI uploads, and the last one, at 2026-08-22 04:22 EDT, matches Cause-coded computation-error lever #86. The project's root is . with output public, so the upload is the app/ folder (from a checkout linked to the project, vercel deploy app --prod). That deploy also publishes Retire the "upper bound" label in the cause_shares replay crosswalk #94's app change (engine_scenario_data.json).
  • Copy site_libs/ before deploying from a fresh checkout. app/public/paper/web/site_libs/ is gitignored, and the live site serves it from the deployer's local copy. A deploy from a fresh clone would serve the manuscript without its Quarto CSS and JS. Copy paper/out/site_libs there after quarto render first.
  • The re-render is already in this PR. No further render is needed. ASSET_V is 20261004a. The paper's iframe key r10-20261003 is new in this PR and nothing has been deployed under it, so it needs no bump.
  • policyengine.org updates on its own. It embeds the simulator in an iframe (policyengine-app-v2 apps.json), so it follows the Vercel deploy with no PR there.

Not in this PR

Tests

Reviews (Opus, Subfleet), archived at ~/reviews/snap-qc-sim-pr95/:

  • Round 1 (review-r1.md): REQUEST_CHANGES with 11 findings. 8de4b13 addressed them; round 2 graded three of them partial.

  • Round 2 (review-r2.md): REQUEST_CHANGES on three precision points. Those were the simulator's Minimodel wording, the utility snap and the household-size natures. It also raised two recommended fixes and five nits, and confirmed the 797/743/856 and 8-case recounts and that the committed HTML is byte-identical to a fresh render. All of it is addressed in b8a35e7.

  • Round 3 (review-r3.md): REQUEST_CHANGES on two solver sentences, checked against a replica that reproduces the committed solver output exactly. The utility snap anchors on the stepped amount, not the file's, and income steps stop only once the benefit passes RAWBEN. It also found the "so nothing moved" causal reading wrong, a JS test lock that couldn't fire, and nits. All addressed in 9e34b2e. Its out-of-scope findings (the same snap wording in ANALYSIS.md; 6 misses the snap created; capped-benefit cases weakly identified) are chipped as task_1431362e.

  • Round 4 (review-r4.md): REQUEST_CHANGES on three short wording points. "Toward the issued amount" was false for household-size moves, whose direction comes from the nature code. D4 applied the shelter rule to all steps. The maximum-allotment sentence holds only for income-lowering steps. It confirmed the footnote clause by clause against a replica (262/262 stepped inputs, 15/15 snaps), the byte-identical render, and the JS mutation check. Addressed in the round-4 commit.

  • Round 5 (review-r5.md): it resolved all round-4 items. It asked to change one sentence: the layer-3 cell's parenthetical omitted the utility reset, which sets the final input in all 15 utility cases. The cell now claims no mechanism. Addressed in the round-5 commit.

  • Round 6 (review-r6.md) confirmed the table-cell fix and ran a final read of the whole diff, with two parallel readers checking 61 numbers. It asked to change one sentence (U1): the 16-versus-33 reconciliation undercounted the matches that needed no move. It also recommended five small changes: "within $5 of" for "yields", the proration note on the 59, "at most one input", the intro's test description, and the tech-doc citation. All of these are addressed in the round-6 commit.

  • Round 7 (review-r7.md) resolved U1 and recommendations 1–6, mutation-checked the new number locks, and asked for two fixes. V1: the paper's "keeps FNS, the name on every document the fiscal 2024 data cite" became false once the citation moved to the FNA-lettered August posting; it now reads "the agency's name during fiscal 2024". V2: the amterr_replay.py docstring still called the solver's inputs "pre-edit original values" and no chip covered it; it is corrected here. Both are addressed in the round-7 commit.

  • Round 8 (review-r8.md) resolved all four round-7 items: it confirmed the FNS clause against the FY2024 rate table's letterhead and the text of both tech-doc postings, checked the docstring against the code and data, and mutation-checked the C7 lock. It found one stale line: the lab README quoted the print label the round-7 commit renamed. That line is fixed and locked by a test in e72957b.

  • Round 9 (review-r9.md): APPROVE at e72957b. It confirmed the README fix and mutation-tested the new lock (7 mutations, all caught). It rebuilt the engine at the pinned commits and re-ran amterr_replay.py: the output is byte-identical to the committed amterr_replay_results.json, and the README's two quoted lines appear word for word. It also recomputed 305/856, 283, 97/186, 78.4%, 91.4% and 86.9% from the QC file. CI passes (401 passed, 22 data-gated skips), and the PR is MERGEABLE and CLEAN.

  • tests/test_retired_claims.py: 22 tests collected, all passing locally. The Hypothesis tests run with deadline=None, since they hit deadlines under load. The 797-of-856 lock is data-gated and needs the QC CSV, so it skips in CI; the PDF check needs pdftotext and skips without it.

  • Targeted runs at b8a35e7: test_retired_claims, test_paper_embed, test_asset_versioning, test_engine_comparison, test_engine_lever, test_events_section and test_migrations_wrapper (51 passed). test_amterr_lab and test_preregistration also passed at 8de4b13.

  • ruff check on CI's paths is clean.

  • Full local suite at b8a35e7 (9e34b2e changes only prose, the render and the test file): 417 passed, 3 skipped, 2 failed. Both failures fail only at $.environment.python (recorded 3.14.4, local 3.14.7); they fail on clean main too and skip in CI. Details are in the PR comment.

🤖 Generated with Claude Code

MaxGhenis and others added 3 commits October 3, 2026 22:27
FSBEN: Mathematica computes it for USDA from each edited case record with
the QC Minimodel's benefit formula (tech doc, August 2026 posting). The
paper and simulator no longer call it "the agency's own computational
canon", "the agency's own benefit-calculation software" or "a closed
model". The Step 12-13 reconciliation is described as the tech doc gives
it; in Colorado FSBEN ends within $5 of BENFIX in 797 of 856 cases.

Replay: the solver starts from the file's edited inputs and moves only the
ELEMENT1 input in $3 steps. A match (246 of 283) is consistent with
correct arithmetic on a wrong input; 230 moved an input and 16 moved
nothing. A miss does not establish a computation error: 17 of the 37
moved nothing, and of 26 cases with a computational finding 13 reproduce.
The "computation-side upper bound" reading and the disclosure-perturbation
description are withdrawn, matching ANALYSIS.md (#92) and cause_shares.json
(#94). The superseded "33 of 246" is reconciled in FACTS H8.

Colorado's 0.03-point FY 2024 margin now carries its FY 2025 rate, 10.09%.
FACTS C4 (tech-doc errata) is withdrawn: no source records the errata, so
the abstract, introduction and oracle section no longer claim them.

Re-rendered app/public/paper/web (revision 10) and bumped the wrapper's
iframe version. tests/test_retired_claims.py ties every corrected figure
to claims_audit.json, cause_shares.json and the replay rows and keeps the
retired wording out of the living files.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Disclosure: parity replays "against the benefit computation recorded in
  the QC file", replacing "the government's own recorded values" (a missed
  paraphrase of the retired claim).
- Intro: the issuance residue is "a residue in issued amounts", no longer
  a defect of state systems; "two nested estimates" drops "nested".
- The replay no longer supports the information-failure conclusion;
  layers 1 and 2 carry it.
- Replay paragraph: "a wrong input" (the solver moves one input), and the
  utility-amount snap after the $3 steps is described.
- Layer-3 table cell: of the 37 misses, 17 moved nothing and 10 carry a
  computational finding (8 are both), from claims_audit.json.
- "305 cases with a recorded payment deviation", matching section 5;
  limitations say "deviation cases"; "Minimodel-computed chain" becomes
  "computed benefit chain".
- README and FACTS F2 name the 10% rate where the 15% share begins.
- Wrapper says six rounds of adversarial review, as the paper does.
- Tests: retire "agency's own" and "government's own", check the PDF
  (pdftotext), cover the "0.03pp" form in FACTS, and lock the 8-case
  overlap. Re-rendered app/public/paper/web.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Simulator: the engine panel's FSBEN line now says Mathematica computes
  it for USDA with the QC Minimodel's benefit formula ("the Minimodel's
  full-formula recomputation" retired); the panel heading and hint name
  the QC file's benefit chain. ASSET_V 20261004a, since app.js changed.
  The engine-comparison generator and its report carry the same wording
  (engine_data.json regenerates byte-identically, sha256 2517d26e).
- Solver description: household size moves only for element 150 with
  natures 7, 12, 14 or 16; the stop rules include a $0 benefit; every
  case labeled a utility error is snapped (to zero if nothing lies
  below), which moves 2 amounts the steps left unchanged. The 14 misses
  are cases whose first finding falls outside the solver's codes.
- Oracle scope: the issuance gap is no longer called a finding about
  somebody's rules.
- Intro: "computation failure versus information failure".
- FACTS: D4 mirrors the solver description; H8 drops its "X, never Y";
  C7 cites PDF p.71 for FSGRINC and FSNETINC.
- references.bib: drops "pre-edit" from the solver note.
- Tests: Hypothesis deadline=None (deadline flakes under load); scan
  app.js and retire the Minimodel recomputation phrasings.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@MaxGhenis

Copy link
Copy Markdown
Contributor Author

Full local suite at b8a35e7 (macOS, both FY2024 QC postings present, so the data-gated tests ran): 417 passed, 3 skipped and 2 failed in 48 minutes.

The 2 failures are tests/test_fixed_donor_decomposition.py::test_raw_regeneration_matches_committed_artifact and tests/test_uhip_decomposition.py::test_raw_regeneration_matches_committed_artifact. Both fail on clean main too, and both fail only at $.environment.python: the committed artifacts record 3.14.4 and this machine runs 3.14.7. They skip in CI and touch nothing this PR changes.

MaxGhenis and others added 6 commits October 4, 2026 08:52
- Solver description: the body now says what moves and where (household
  size once, by one person, for element 150 with natures 7/12/14/16;
  income, rent, utility and deductions in $3 steps toward RAWBEN), and a
  footnote gives the stop rules as reconstruct_co_fy2024.R applies them:
  income steps stop only once the benefit passes RAWBEN, so they run to
  zero income or the step limit where RAWBEN is the maximum allotment;
  the cap is the shelter deduction's; the shelter-at-zero and
  uncapped-benefit-below-zero rules are included. The utility snap picks
  the common amount nearest the stepped amount, from those above or
  below the file's UTIL by the sign of RAWBEN - FSBEN. FACTS D4 mirrors
  it.
- The 16 unmoved matches: "it took no step, and the issued benefit was
  already within $5 of FSBEN", replacing a "so nothing moved" causal
  reading the code does not support.
- Abstract and table: the solver moves inputs "toward" the issued amount.
- Oracle scope: "reaching parity surfaced two defects in the encodings",
  matching FACTS C3. FACTS C1: "computed benefit chain".
- Tests: the JS scan rejoins concatenated string literals, so the retired
  app.js sentence (split across a template-literal join) is caught; a
  unit test covers the join.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Abstract: the solver "changes" the named input; "toward the issued
  amount" was false for household-size moves, whose direction comes from
  the nature code (2 of 7 Colorado moves go away from RAWBEN). The table
  cell keeps "toward" for the $3-stepped inputs only, and the body says
  which natures remove or add a person.
- FACTS D4: the shelter-deduction stop rule applies to rent and utility
  steps only, matching the footnote.
- Footnote and D4: the maximum-allotment sentence is about
  income-lowering steps; "all steps stop" replaces "every step stops".

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The cell's parenthetical ("other inputs in $3 steps toward the issued
amount") left out the utility reset, which sets the final input in all 15
Colorado utility cases. The cell now says only that the solver changes
the named input where it can, as the abstract does; the body and its
footnote carry the mechanism.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- The old "33 within comparison tolerance mechanically" was right in
  substance: all 33, and 2 more with AMTERR of $6, started with RAWBEN
  within $5 of FSBEN, so none of those 35 matches needed a move. The
  solver took no step in 16 and moved an input anyway in 19. The paper
  now says so for the 19, and FACTS H8 retires only the word "explained".
- "yields a benefit within $5 of the issued benefit" (only 51 of 246
  matches are exact).
- 797 of 856: 21 of the other 59 are prorated allotments (ALLADJ 2),
  which the full-month FSBEN does not reflect; FACTS C7 mirrors it.
- Intro: the replay tests "changing the input named by each case's first
  finding"; limitations: the solver moves "at most one" input; the table
  caption matches the section's "three layers of evidence".
- references.bib: the tech doc is the August 2026 Mathematica report by
  Leftin et al., not an FNS 2025 publication.
- Tests lock the 35/16/19 split, the new sentence and the 21 prorated
  cases.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- The paper keeps FNS as "the agency's name during fiscal 2024"; "the name
  on every document the fiscal 2024 data cite" became false once the
  citation moved to the FNA-lettered August 2026 tech doc.
- amterr_replay.py's docstring no longer calls the solver's inputs
  "pre-edit original values"; the engine_on_original field name stays,
  since the pinned results JSON uses it.
- FACTS H8: rent's "closeness stop is $3". The C7 proration count is now
  locked in the data-gated test.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
dfc1ddc renamed amterr_replay.py's summary label from engine(original)
to engine(solver inputs); the lab README still quoted the old one. It now
quotes the new label, and a test builds that line from claims_audit.json
and checks the README and the script agree.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@MaxGhenis

Copy link
Copy Markdown
Contributor Author

Status at head e72957b. The PR is ready to merge. The merge is waiting on Max's decision.

  • Independent review: 9 rounds of Opus on Subfleet. Round 9 approves at e72957b. All reviews are archived in ~/reviews/snap-qc-sim-pr95/ (review-r1.md to review-r9.md, plus replica scripts in r3-checks/, r4-checks/, r8-checks/ and r9-checks/).
  • Checks: gh pr checks exits 0, and the PR is MERGEABLE and CLEAN. The full local suite passed at b8a35e7, except the two known environment-version regeneration tests.
  • Why Max decides: every review round judged that the abstract and introduction changes alter a headline finding. The replay moves from "explains" to "consistent with", the errata claim is withdrawn, and the intro loses its "three nested estimates". No headline number changes. The merge is queued as d933.
  • After merge: the live site changes only with a manual Vercel deploy of app/, queued as d934. Copy paper/out/site_libs into app/public/paper/web/ first. Details are under "After merge" in the PR body.
  • Follow-ups chipped:
  • No dependencies: Correct the amterr lab's layer-3 claims and make it rerunnable #92 and Retire the "upper bound" label in the cause_shares replay crosswalk #94 are merged, and this branch is based on Retire the "upper bound" label in the cause_shares replay crosswalk #94's merge commit.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant