Conversation
ANALYSIS.md's method section now states the solver's rules as reconstruct_co_fy2024.R applies them, matching the paper's corrected footnote in #95: - The utility reset takes the commonly reported amount nearest the solver's stepped amount. Candidates are amounts that more than 5 filtered cases of any review status report in the state and calendar year, above the file's UTIL for util_up and below it for util_down. The stepped amount is kept if none lies above; the amount is set to 0 if none lies below. - Rent and utility steps also stop at a zero shelter deduction when lowering, and income-raising steps stop when the uncapped benefit falls below zero. "every step stops" now reads "all steps stop". - Household size takes its direction from the nature code, so 2 of the 7 Colorado moves take the benefit away from RAWBEN. "In 20 the solver moved an input and stopped short" was true of 9 of the 20. In 6 the $3 steps reached within $3 of RAWBEN and the utility reset then moved the input off that match; in 5 the household-size move missed. The results now also flag 34 weakly identified matches: RAWBEN is the maximum allotment and the solver lowered income, so every income at or below break-even reproduces it. audit_claims.py ports the solver's benefit formula and fails unless it reproduces all 283 recorded benefits. It also re-applies the utility reset, which reproduces all 15 final amounts. Each new count lands in claims_audit.json, and test_amterr_lab.py locks the sentences that quote them. New property tests check the two facts the caveat rests on: the benefit never exceeds the maximum allotment, and it falls as income rises. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Contributor
Author
|
Compatibility with #95: |
An adversarial pass (8 independent agents, each with its own port of the
solver) confirmed every count in the previous commit. It also found that
the weak-identification caveat understated its own argument.
- The caveat now covers every match on a flat stretch of the benefit
formula, where the match bounds the moved input on one side only: 60 of
the 230 moved matches, 31 of them above the threshold.
- 53 are at or within $5 of the maximum allotment. That is the 34 income
cuts plus 19 where every further rent, utility or medical amount also
reproduces RAWBEN.
- 4 are at the $23 minimum benefit.
- 3 are rent increases at the shelter-deduction cap.
- "Break-even point" usually means the income at which the benefit phases
out. It now reads "the point where net income reaches zero": the solver
pays the maximum exactly when net income is zero or less, which a new
property test checks.
- "ran on to zero income" now reads "ran that income down to $0", since
other income can remain.
- The stop rules now say "at or above the maximum allotment" and add the
mirror rule at the minimum benefit.
- The correctedamount note now covers household-size moves (always 0) and
all 15 utility rows (always the pre-reset change).
- The audit check is stated as covering two parts of the account, which is
what it does.
- The revision note quotes the old reset rule correctly.
- New notes:
- 202405-40908's reset leaves the benefit farther from RAWBEN than FSBEN.
- 2 of the 10 computational misses are reset-off cases.
audit_claims.py classifies each match's flat stretch by pushing the moved
input without limit in the solver's direction (flat_stretch). The new
counts are in claims_audit.json, and test_amterr_lab.py locks them.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
MaxGhenis
added a commit
that referenced
this pull request
Oct 4, 2026
MaxGhenis
added a commit
that referenced
this pull request
Oct 4, 2026
… paper This follows #96's second round, after 8 independent agents checked the claims. - The replay paragraph now gives all 60 weakly identified moved matches, 31 above the threshold: - 53 at or within $5 of the maximum allotment; - 4 at the minimum benefit; - 3 at the shelter-deduction cap. Of the 53, 34 are income cuts at the maximum. The mechanism is now stated as "up to the point where net income reaches zero" rather than "break-even". - The 37 misses are now each accounted for once: 17 unmoved, 6 reset-off, 9 stopped short and 5 household-size misses. - It notes that 2 of the 10 computational misses are reset-off cases. - "A stepped utility amount is then reset" now reads "the utility amount", since two rows were reset without a step. - The footnote and D4 now say "at or above the maximum allotment" and add the mirror rule at the minimum benefit. - D5 and D6 are rewritten to match, and the locks follow. A new identity checks that the misses partition the 37. Re-rendered with Quarto 1.9.36. The PDF is still 27 pages. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The independent Opus review of 00de6bf asked for two required fixes and three recommended ones. All are verified against the solver's formula. - The 3 shelter-cap matches don't reach the cap: the solver stopped $11 to $14 short of $672 on the within-$3 rule. The text now says they "stop within $14 of the shelter-deduction cap", and that past the cap the benefit stays within $5 of RAWBEN. The audit records the gap. - 2 of the 60 (202403-40765 and 202404-40803) reproduce at every utility amount. "On one side only" now reads "on one side at most", and these two are named. The audit pushes each input the opposite way too (opposite_push_holds, bounded_on_neither_side_keys). - The caveat now says "by the solver's formula": the engine was not run at the pushed amounts. The "farther from RAWBEN" list now fails the build unless the engine agrees with the solver's benefit. - Pushes that stay within $5 without reaching a flat stretch are counted and named (unnamed_push_holds_keys: 4 already at $0, 3 whose band runs down to $0). ANALYSIS.md says they are not counted. - More of each lock is derived from the audit: the at_max composition (one household-size match, one unmoved), the $23 minimum, the step limit and the correctedamount "2 of them". Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
MaxGhenis
added a commit
that referenced
this pull request
Oct 4, 2026
MaxGhenis
added a commit
that referenced
this pull request
Oct 4, 2026
Mirrors #96's review-r1 fixes and #97's review-r1 suggestions: - The weak-identification sentence now says the issued benefit lies within $5 of a stretch where the solver's formula is flat. The 3 shelter cases are now "rent past the shelter-deduction cap", since the solver stopped $11 to $14 short of it. The sentence also says "by that formula", "on one side at most" and "34 of the 53". - FACTS D5 gives the push the audit uses ($100,000 or $0), notes the engine was not run there, and names the 2 matches bounded on neither side and the 7 matches it doesn't count. - "Of the 10, 2 are …", so no sentence opens with a numeral. Re-rendered; the PDF is still 27 pages. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Brings the amterr lab's
ANALYSIS.mdinto line withreconstruct_co_fy2024.R, the same way #95 corrected the paper's version (its[^solver-rules]footnote and FACTS row D4). Every new count comes fromclaims_audit.json, whichaudit_claims.pynow regenerates with a port of the solver's own benefit formula.What changed in ANALYSIS.md
Method section, against the R code
R:232-247). The solver takes the commonly reported amount nearest its stepped amount. Candidates are amounts that more than 5 filtered cases of any review status report in the same state and calendar year. They lie above the file's UTIL forutil_up(RAWBEN > FSBEN) and below it forutil_down. The stepped amount is kept if none lies above; the amount is set to 0 if none lies below. The old text said "nearest value above (or below) the file's UTIL". That rule gives the final amount in 8 of the 15 utility rows; the corrected rule gives all 15.R:179-180). Rent and utility steps also stop at a zero shelter deduction when lowering. Also added fromR:152: income-raising steps stop when the uncapped benefit falls below zero. The maximum-allotment sentence is added too.R:124-131). The nature code sets the direction: 12/14/16 removes a person, 7 adds one. So 2 of the 7 Colorado moves take the benefit away from RAWBEN (202310-40265, 202402-40618).Results
at_maxflag marks 54 of the 246 matches: 52 of the 53, one household-size match and one unmoved match.Round 2 (00de6bf)
Before review, 8 independent agents each ported the solver from the R code and tried to refute 62f0bb9. Every count held. Their wording and completeness findings are fixed in 00de6bf:
correctedamountnote, which now covers household-size moves (always 0) and all 15 utility rows (always the pre-reset change);Their results:
~/reviews/snap-qc-sim-pr96/skeptic-workflow-result.json.Review round 1 (561f34d)
The Opus review of 00de6bf (
~/reviews/snap-qc-sim-pr96/review-r1.md) requested changes. All five findings are fixed in 561f34d:audit_claims.py
solver_benefitportscalculate_raw_benefits.classify_solver_movesraises unless the port reproducesrawben_recreatedfor all 283 rows. It then computes the benefit after the steps but before the reset, and sorts each moved miss intosteps_stopped_short,reset_off_after_step_matchorhousehold_size.utility_reset_checkrebuilds the reset's candidate pool from the posting:It then re-applies the rule.
flat_stretchpushes each matched input without limit in the solver's direction. It names the stretch (maximum allotment, minimum benefit or shelter cap) where the pushed benefit still reproduces RAWBEN. The benefit is monotone in each input, so every amount in between reproduces RAWBEN too.New
claims_audit.jsonblocks:case_level.solver_outcomes(includingweakly_identified) andcase_level.utility_reset_rule. Each case row also gets amoved_miss_kindfield. No existing value changed.Inputs added from the posting:
BENMAX,FSNELDERandFSNDIS(the shelter cap is lifted for an elderly or disabled member).Constants, checked against snap_qc at the README pin 741e10bf:
Verification
~/reviews/snap-qc-sim-pr95/r4-checks/), repointed to this checkout:tests/test_amterr_lab.py: 24 passed with both postings present, including the regeneration test and 600 property examples.tests/test_cause_shares.pyandtests/test_paper_embed.py: 23 passed.audit_claims.py --checkreports the audit is current.tests/test_retired_claims.pyexists only on Correct retired SNAP QC claims in the paper, README, FACTS and simulator #95's branch, so it was run on this branch merged with Correct retired SNAP QC claims in the paper, README, FACTS and simulator #95 (result in a comment below).Invariants (tested)
test_solver_outcome_counts_are_consistent).at_maxflag, and all sit on the cap stretch. The stretches partition the 60.Relation to #95
There are no file overlaps with #95. #95 touches the lab's README and
amterr_replay.py; this PR touches neither. The paper's replay paragraph exists only on #95's branch, so the paper caveat goes in a separate PR stacked on #95.axiom: n/a: lab documentation and audit tooling; no policy encoding changes
🤖 Generated with Claude Code