Conversation
FSBEN: Mathematica computes it for USDA from each edited case record with the QC Minimodel's benefit formula (tech doc, August 2026 posting). The paper and simulator no longer call it "the agency's own computational canon", "the agency's own benefit-calculation software" or "a closed model". The Step 12-13 reconciliation is described as the tech doc gives it; in Colorado FSBEN ends within $5 of BENFIX in 797 of 856 cases. Replay: the solver starts from the file's edited inputs and moves only the ELEMENT1 input in $3 steps. A match (246 of 283) is consistent with correct arithmetic on a wrong input; 230 moved an input and 16 moved nothing. A miss does not establish a computation error: 17 of the 37 moved nothing, and of 26 cases with a computational finding 13 reproduce. The "computation-side upper bound" reading and the disclosure-perturbation description are withdrawn, matching ANALYSIS.md (#92) and cause_shares.json (#94). The superseded "33 of 246" is reconciled in FACTS H8. Colorado's 0.03-point FY 2024 margin now carries its FY 2025 rate, 10.09%. FACTS C4 (tech-doc errata) is withdrawn: no source records the errata, so the abstract, introduction and oracle section no longer claim them. Re-rendered app/public/paper/web (revision 10) and bumped the wrapper's iframe version. tests/test_retired_claims.py ties every corrected figure to claims_audit.json, cause_shares.json and the replay rows and keeps the retired wording out of the living files. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Disclosure: parity replays "against the benefit computation recorded in the QC file", replacing "the government's own recorded values" (a missed paraphrase of the retired claim). - Intro: the issuance residue is "a residue in issued amounts", no longer a defect of state systems; "two nested estimates" drops "nested". - The replay no longer supports the information-failure conclusion; layers 1 and 2 carry it. - Replay paragraph: "a wrong input" (the solver moves one input), and the utility-amount snap after the $3 steps is described. - Layer-3 table cell: of the 37 misses, 17 moved nothing and 10 carry a computational finding (8 are both), from claims_audit.json. - "305 cases with a recorded payment deviation", matching section 5; limitations say "deviation cases"; "Minimodel-computed chain" becomes "computed benefit chain". - README and FACTS F2 name the 10% rate where the 15% share begins. - Wrapper says six rounds of adversarial review, as the paper does. - Tests: retire "agency's own" and "government's own", check the PDF (pdftotext), cover the "0.03pp" form in FACTS, and lock the 8-case overlap. Re-rendered app/public/paper/web. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Simulator: the engine panel's FSBEN line now says Mathematica computes
it for USDA with the QC Minimodel's benefit formula ("the Minimodel's
full-formula recomputation" retired); the panel heading and hint name
the QC file's benefit chain. ASSET_V 20261004a, since app.js changed.
The engine-comparison generator and its report carry the same wording
(engine_data.json regenerates byte-identically, sha256 2517d26e).
- Solver description: household size moves only for element 150 with
natures 7, 12, 14 or 16; the stop rules include a $0 benefit; every
case labeled a utility error is snapped (to zero if nothing lies
below), which moves 2 amounts the steps left unchanged. The 14 misses
are cases whose first finding falls outside the solver's codes.
- Oracle scope: the issuance gap is no longer called a finding about
somebody's rules.
- Intro: "computation failure versus information failure".
- FACTS: D4 mirrors the solver description; H8 drops its "X, never Y";
C7 cites PDF p.71 for FSGRINC and FSNETINC.
- references.bib: drops "pre-edit" from the solver note.
- Tests: Hypothesis deadline=None (deadline flakes under load); scan
app.js and retire the Minimodel recomputation phrasings.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Contributor
Author
|
Full local suite at b8a35e7 (macOS, both FY2024 QC postings present, so the data-gated tests ran): 417 passed, 3 skipped and 2 failed in 48 minutes. The 2 failures are |
- Solver description: the body now says what moves and where (household size once, by one person, for element 150 with natures 7/12/14/16; income, rent, utility and deductions in $3 steps toward RAWBEN), and a footnote gives the stop rules as reconstruct_co_fy2024.R applies them: income steps stop only once the benefit passes RAWBEN, so they run to zero income or the step limit where RAWBEN is the maximum allotment; the cap is the shelter deduction's; the shelter-at-zero and uncapped-benefit-below-zero rules are included. The utility snap picks the common amount nearest the stepped amount, from those above or below the file's UTIL by the sign of RAWBEN - FSBEN. FACTS D4 mirrors it. - The 16 unmoved matches: "it took no step, and the issued benefit was already within $5 of FSBEN", replacing a "so nothing moved" causal reading the code does not support. - Abstract and table: the solver moves inputs "toward" the issued amount. - Oracle scope: "reaching parity surfaced two defects in the encodings", matching FACTS C3. FACTS C1: "computed benefit chain". - Tests: the JS scan rejoins concatenated string literals, so the retired app.js sentence (split across a template-literal join) is caught; a unit test covers the join. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Abstract: the solver "changes" the named input; "toward the issued amount" was false for household-size moves, whose direction comes from the nature code (2 of 7 Colorado moves go away from RAWBEN). The table cell keeps "toward" for the $3-stepped inputs only, and the body says which natures remove or add a person. - FACTS D4: the shelter-deduction stop rule applies to rent and utility steps only, matching the footnote. - Footnote and D4: the maximum-allotment sentence is about income-lowering steps; "all steps stop" replaces "every step stops". Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The cell's parenthetical ("other inputs in $3 steps toward the issued
amount") left out the utility reset, which sets the final input in all 15
Colorado utility cases. The cell now says only that the solver changes
the named input where it can, as the abstract does; the body and its
footnote carry the mechanism.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- The old "33 within comparison tolerance mechanically" was right in substance: all 33, and 2 more with AMTERR of $6, started with RAWBEN within $5 of FSBEN, so none of those 35 matches needed a move. The solver took no step in 16 and moved an input anyway in 19. The paper now says so for the 19, and FACTS H8 retires only the word "explained". - "yields a benefit within $5 of the issued benefit" (only 51 of 246 matches are exact). - 797 of 856: 21 of the other 59 are prorated allotments (ALLADJ 2), which the full-month FSBEN does not reflect; FACTS C7 mirrors it. - Intro: the replay tests "changing the input named by each case's first finding"; limitations: the solver moves "at most one" input; the table caption matches the section's "three layers of evidence". - references.bib: the tech doc is the August 2026 Mathematica report by Leftin et al., not an FNS 2025 publication. - Tests lock the 35/16/19 split, the new sentence and the 21 prorated cases. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- The paper keeps FNS as "the agency's name during fiscal 2024"; "the name on every document the fiscal 2024 data cite" became false once the citation moved to the FNA-lettered August 2026 tech doc. - amterr_replay.py's docstring no longer calls the solver's inputs "pre-edit original values"; the engine_on_original field name stays, since the pinned results JSON uses it. - FACTS H8: rent's "closeness stop is $3". The C7 proration count is now locked in the data-gated test. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
dfc1ddc renamed amterr_replay.py's summary label from engine(original) to engine(solver inputs); the lab README still quoted the old one. It now quotes the new label, and a test builds that line from claims_audit.json and checks the README and the script agree. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Contributor
Author
|
Status at head e72957b. The PR is ready to merge. The merge is waiting on Max's decision.
|
This was referenced Oct 4, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
TheAxiomFoundation/axiom.org#297 retires several claims about USDA's SNAP Quality Control file and the Colorado error-case replay. A sweep (findings off-18 and oogs-5 to oogs-14) found copies here: in the working paper, the README, the fact catalog and the live simulator. This PR corrects all 15 locations, plus the copies of the same claims that sit next to them. It builds on #92 (the lab's
claims_audit.json, the source for every replay number) and #94 (which retired the "upper bound" label inanalysis/cause_shares.json).axiom: n/a: manuscript, README and simulator copy only; no policy encoding changed.
Before and after
Line numbers are at 755a7f3, before #92. Numbers are sourced in the next section.
app/public/index.html:133(live simulator, embedded on policyengine.org); the same panel's FSBEN line inapp.jspaper/index.qmd:108:299:349:362:368:1102:25(abstract):447:443reconstruct_co_fy2024.Rdoes. It starts from the file's edited inputs and moves only the ELEMENT1 input. Household size moves once, by one person: one fewer for natures 12, 14 or 16, one more for nature 7; the nature code sets the direction. Income, rent, utility and deduction inputs move in $3 steps toward RAWBEN until a stop rule binds, and a stepped utility amount is then reset to a commonly reported value. Nothing moves when the first finding falls outside those codes. A footnote gives the stop rules and the snap rule exactly. For example, income steps have no within-$3 stop, so where RAWBEN is the maximum allotment income-lowering steps run to zero income or the step limit; the snap picks the common amount nearest the stepped one. The layer-3 table cell names no stepping or reset rule: "A public solver changes the input named by each case's first finding, where it can".:456:711(fig-co caption)README.md:31paper/FACTS.mdD4paper/FACTS.mdC4Copies of the same retired claims next to these, corrected for consistency:
cause_shares.json(Retire the "upper bound" label in the cause_shares replay crosswalk #94) both withdrew. It appeared in the layer-3 row oftbl-decompose, the 21.6% "computation-side upper bound for official errors" in the replay paragraph, "bounded above by the replay's 21.6% case share", and the limitations' "residual class is an upper bound". The table cell now says the replay does not identify the computation share and gives the reproduction rates. A new paragraph says neither outcome identifies a computation error: 17 of 37 misses moved nothing (14no_change+ 3 stopped), and of the 26 computational-finding cases, 13 reproduce, 10 do not and 3 were not replayed. Because the replay no longer estimates a computation share, the intro's "three nested estimates" becomes two estimates plus a replay that tests one input per case; the section opener's "three nested answers" becomes "three layers of evidence"; and the caption drops "nested".app.js); it now says Mathematica computes it for USDA with the QC Minimodel's benefit formula. The heading "Axiom rules engine vs the FNA QC Minimodel" and the matching hint now name the QC file's benefit chain. The same wording inanalysis/engine_comparison.pyand its generatedENGINE_COMPARISON.mdis updated too: regenerating from the May posting reproducesengine_data.jsonbyte for byte (sha256 2517d26e…, the pinned value) and the report with only that line changed. ASSET_V is bumped to20261004a, sinceapp.jschanged.references.bibdrops "pre-edit" from the solver note.app/public/paper/: re-rendered (Quarto 1.9.36, the same version as revision 9). The wrapper now reads revision 10 · 2026-10-03, with a one-line note on what changed (it also still said "revision-7"), and its iframe cache key isr10-20261003. The manuscriptdate:is now 2026-10-03. The clean render also drops a "Notebooks" nav that revision 9 picked up from untracked files and that linked a local/Users/maxghenis/.cache/...path.Where every number comes from
Each figure below was read this session from the source named.
paper/snapshot/labs/amterr/claims_audit.json(case_level.replay,layer2_computational_findings,by_posting.may2026.replay,totals.CO.error_cases)analysis/cause_shares.jsoncolorado_replay_reconciliation.slices(#94 key names), cross-checked against the per-case replay rows with AMTERR > $56amterr_replay_results.jsonandclaims_audit.jsonreplay_inputs_unchanged_keysclaims_audit.jsoninputs.official_rates_percent(FNS PER tables)reconstruct_co_fy2024.R(read this session; it has no perturbation logic);claims_audit.jsonmoved_with_zero_correctedamount_keysclaims_audit.jsonnot_reproduced_casesandnot_reproduced_with_computational_finding_keysThe paper keeps its stated convention of "FNS" (line 88). The simulator already says FNA.
Invariants the corrected text relies on
tests/test_retired_claims.pychecks each of these. It needs only committed files, except for the 797-of-856 check, which needs the QC file and skips without it.cause_shares.jsonequals the split recomputed from the per-case replay rows (AMTERR > $56: 97 cases, 76 reproduced; ≤ $56: 186 cases, 170 reproduced).paper/index.qmd, the rendered HTML and PDF (the PDF check usespdftotext),README.md,app/public/index.htmlandapp/public/app.js. FACTS keeps it only inside SUPERSEDED, WITHDRAWN and prohibition rows.Headline findings: queued for Max
No headline number changes. Parity (6,081 of 6,194), the 46–50% tier odds, the $7.7B → $7.6B / $6.9B pricing, 18 of 53 tier changes, and 78.4% / 91.4% are all unchanged.
What the abstract and introduction claim does change:
Both review rounds judged that these alter a headline finding (round 1: "narrowly"). The task's rule sends a change that alters a headline finding to Max, so the merge is queued as decision d933 instead of self-merged. The production deploy, which outlives the merge, is decision d934.
After merge
snap-qc-simVercel project has no git integration. Production deploys are CLI uploads, and the last one, at 2026-08-22 04:22 EDT, matches Cause-coded computation-error lever #86. The project's root is.with outputpublic, so the upload is theapp/folder (from a checkout linked to the project,vercel deploy app --prod). That deploy also publishes Retire the "upper bound" label in the cause_shares replay crosswalk #94's app change (engine_scenario_data.json).site_libs/before deploying from a fresh checkout.app/public/paper/web/site_libs/is gitignored, and the live site serves it from the deployer's local copy. A deploy from a fresh clone would serve the manuscript without its Quarto CSS and JS. Copypaper/out/site_libsthere afterquarto renderfirst.20261004a. The paper's iframe keyr10-20261003is new in this PR and nothing has been deployed under it, so it needs no bump.apps.json), so it follows the Vercel deploy with no PR there.Not in this PR
docs/v2-error-model.md. It carries the same stale replay reading and is covered by chip task_7739a01c from the Retire the "upper bound" label in the cause_shares replay crosswalk #94 session.Tests
Reviews (Opus, Subfleet), archived at
~/reviews/snap-qc-sim-pr95/:Round 1 (
review-r1.md): REQUEST_CHANGES with 11 findings. 8de4b13 addressed them; round 2 graded three of them partial.Round 2 (
review-r2.md): REQUEST_CHANGES on three precision points. Those were the simulator's Minimodel wording, the utility snap and the household-size natures. It also raised two recommended fixes and five nits, and confirmed the 797/743/856 and 8-case recounts and that the committed HTML is byte-identical to a fresh render. All of it is addressed in b8a35e7.Round 3 (
review-r3.md): REQUEST_CHANGES on two solver sentences, checked against a replica that reproduces the committed solver output exactly. The utility snap anchors on the stepped amount, not the file's, and income steps stop only once the benefit passes RAWBEN. It also found the "so nothing moved" causal reading wrong, a JS test lock that couldn't fire, and nits. All addressed in 9e34b2e. Its out-of-scope findings (the same snap wording in ANALYSIS.md; 6 misses the snap created; capped-benefit cases weakly identified) are chipped as task_1431362e.Round 4 (
review-r4.md): REQUEST_CHANGES on three short wording points. "Toward the issued amount" was false for household-size moves, whose direction comes from the nature code. D4 applied the shelter rule to all steps. The maximum-allotment sentence holds only for income-lowering steps. It confirmed the footnote clause by clause against a replica (262/262 stepped inputs, 15/15 snaps), the byte-identical render, and the JS mutation check. Addressed in the round-4 commit.Round 5 (
review-r5.md): it resolved all round-4 items. It asked to change one sentence: the layer-3 cell's parenthetical omitted the utility reset, which sets the final input in all 15 utility cases. The cell now claims no mechanism. Addressed in the round-5 commit.Round 6 (
review-r6.md) confirmed the table-cell fix and ran a final read of the whole diff, with two parallel readers checking 61 numbers. It asked to change one sentence (U1): the 16-versus-33 reconciliation undercounted the matches that needed no move. It also recommended five small changes: "within $5 of" for "yields", the proration note on the 59, "at most one input", the intro's test description, and the tech-doc citation. All of these are addressed in the round-6 commit.Round 7 (
review-r7.md) resolved U1 and recommendations 1–6, mutation-checked the new number locks, and asked for two fixes. V1: the paper's "keeps FNS, the name on every document the fiscal 2024 data cite" became false once the citation moved to the FNA-lettered August posting; it now reads "the agency's name during fiscal 2024". V2: theamterr_replay.pydocstring still called the solver's inputs "pre-edit original values" and no chip covered it; it is corrected here. Both are addressed in the round-7 commit.Round 8 (
review-r8.md) resolved all four round-7 items: it confirmed the FNS clause against the FY2024 rate table's letterhead and the text of both tech-doc postings, checked the docstring against the code and data, and mutation-checked the C7 lock. It found one stale line: the lab README quoted the print label the round-7 commit renamed. That line is fixed and locked by a test in e72957b.Round 9 (
review-r9.md): APPROVE at e72957b. It confirmed the README fix and mutation-tested the new lock (7 mutations, all caught). It rebuilt the engine at the pinned commits and re-ranamterr_replay.py: the output is byte-identical to the committedamterr_replay_results.json, and the README's two quoted lines appear word for word. It also recomputed 305/856, 283, 97/186, 78.4%, 91.4% and 86.9% from the QC file. CI passes (401 passed, 22 data-gated skips), and the PR is MERGEABLE and CLEAN.tests/test_retired_claims.py: 22 tests collected, all passing locally. The Hypothesis tests run withdeadline=None, since they hit deadlines under load. The 797-of-856 lock is data-gated and needs the QC CSV, so it skips in CI; the PDF check needspdftotextand skips without it.Targeted runs at b8a35e7:
test_retired_claims,test_paper_embed,test_asset_versioning,test_engine_comparison,test_engine_lever,test_events_sectionandtest_migrations_wrapper(51 passed).test_amterr_labandtest_preregistrationalso passed at 8de4b13.ruff checkon CI's paths is clean.Full local suite at b8a35e7 (9e34b2e changes only prose, the render and the test file): 417 passed, 3 skipped, 2 failed. Both failures fail only at
$.environment.python(recorded 3.14.4, local 3.14.7); they fail on cleanmaintoo and skip in CI. Details are in the PR comment.🤖 Generated with Claude Code