Skip to content

Correct the amterr lab's solver description against the R code - #96

Open
MaxGhenis wants to merge 3 commits into
mainfrom
amterr-lab-solver-description
Open

MaxGhenis wants to merge 3 commits into
mainfrom
amterr-lab-solver-description

Conversation

@MaxGhenis

@MaxGhenis MaxGhenis commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Brings the amterr lab's ANALYSIS.md into line with reconstruct_co_fy2024.R, the same way #95 corrected the paper's version (its [^solver-rules] footnote and FACTS row D4). Every new count comes from claims_audit.json, which audit_claims.py now regenerates with a port of the solver's own benefit formula.

What changed in ANALYSIS.md

Method section, against the R code

  • Utility reset (R:232-247). The solver takes the commonly reported amount nearest its stepped amount. Candidates are amounts that more than 5 filtered cases of any review status report in the same state and calendar year. They lie above the file's UTIL for util_up (RAWBEN > FSBEN) and below it for util_down. The stepped amount is kept if none lies above; the amount is set to 0 if none lies below. The old text said "nearest value above (or below) the file's UTIL". That rule gives the final amount in 8 of the 15 utility rows; the corrected rule gives all 15.
  • Shelter stop (R:179-180). Rent and utility steps also stop at a zero shelter deduction when lowering. Also added from R:152: income-raising steps stop when the uncapped benefit falls below zero. The maximum-allotment sentence is added too.
  • Wording. "every step stops" now reads "all steps stop".
  • Household size (R:124-131). The nature code sets the direction: 12/14/16 removes a person, 7 adds one. So 2 of the 7 Colorado moves take the benefit away from RAWBEN (202310-40265, 202402-40618).

Results

  • "In 20 the solver moved an input and stopped short" was true of 9 of the 20.
    • In 6, the $3 steps reached within $3 of RAWBEN and the utility reset then moved the input off that match: 202310-40297, 202312-40456, 202312-40513, 202402-40603, 202405-40908 and 202409-41321.
    • In 5, the household-size move missed.
  • New caveat: 60 weakly identified matches. In each, the issued benefit lies within $5 of a stretch where the solver's benefit formula is flat in the moved input, so by that formula the match bounds the input on one side at most. 2 of the 60 (202403-40765, 202404-40803) are bounded on neither side. 31 of the 60 are above the $56 threshold.
    • 53 are at or within $5 of the maximum allotment. In 34 of them RAWBEN is the maximum and the solver lowered an income. The solver pays the maximum exactly when net income is ≤ 0, so every value of that income up to the point where net income reaches zero yields it. The steps ran that income down to $0 in 32 and to the step limit in 2. In the other 19, every further rent, utility or medical amount also reproduces RAWBEN.
    • 4 are at the $23 minimum benefit.
    • 3 are rent increases that stop within $14 of the shelter-deduction cap. Past the cap the benefit stays within $5 of RAWBEN.
    • Not counted: 7 pushes that stay within $5 but reach no flat stretch.
    • The R script's at_max flag marks 54 of the 246 matches: 52 of the 53, one household-size match and one unmoved match.
  • The broad-coded and software-coded sections now note that 202312-40456 and 202312-40513 are reset-off misses.

Round 2 (00de6bf)

Before review, 8 independent agents each ported the solver from the R code and tried to refute 62f0bb9. Every count held. Their wording and completeness findings are fixed in 00de6bf:

  • the 34 → 60 weak-identification scope above, and "break-even point" → "the point where net income reaches zero";
  • "that income down to $0" (other income can remain);
  • "at or above the maximum allotment", plus the mirror rule at the minimum benefit;
  • the correctedamount note, which now covers household-size moves (always 0) and all 15 utility rows (always the pre-reset change);
  • the audit is now described as checking "two parts of this account", which is what it checks;
  • a corrected quote of the old reset rule;
  • new notes: 202405-40908's reset leaves the benefit farther from RAWBEN than FSBEN, and 2 of the 10 computational misses are reset-off cases.

Their results: ~/reviews/snap-qc-sim-pr96/skeptic-workflow-result.json.

Review round 1 (561f34d)

The Opus review of 00de6bf (~/reviews/snap-qc-sim-pr96/review-r1.md) requested changes. All five findings are fixed in 561f34d:

  1. Shelter cap. The solver stopped $11–$14 short of the cap, so the text now reads "stop within $14".
  2. Neither side. 2 matches reproduce at any utility amount, so the text now reads "on one side at most" and names them.
  3. Solver's formula. The caveat is scoped to the solver's formula, and the build now fails unless the engine agrees with the solver on the "farther from RAWBEN" list.
  4. Unnamed pushes are now counted.
  5. Derived locks. More of each lock is derived from the audit.

audit_claims.py

  • solver_benefit ports calculate_raw_benefits. classify_solver_moves raises unless the port reproduces rawben_recreated for all 283 rows. It then computes the benefit after the steps but before the reset, and sorts each moved miss into steps_stopped_short, reset_off_after_step_match or household_size.

  • utility_reset_check rebuilds the reset's candidate pool from the posting:

    • 803 Colorado cases;
    • candidates {0, 91, 560} for 2023 and {0, 91, 356, 560} for 2024.

    It then re-applies the rule.

  • flat_stretch pushes each matched input without limit in the solver's direction. It names the stretch (maximum allotment, minimum benefit or shelter cap) where the pushed benefit still reproduces RAWBEN. The benefit is monotone in each input, so every amount in between reproduces RAWBEN too.

  • New claims_audit.json blocks: case_level.solver_outcomes (including weakly_identified) and case_level.utility_reset_rule. Each case row also gets a moved_miss_kind field. No existing value changed.

  • Inputs added from the posting: BENMAX, FSNELDER and FSNDIS (the shelter cap is lifted for an elderly or disabled member).

  • Constants, checked against snap_qc at the README pin 741e10bf:

    • FY2024 maximum allotments, sizes 1–20;
    • the $672 shelter cap.

Verification

  • Replica differential. The round-3/4 replicas from the Correct retired SNAP QC claims in the paper, README, FACTS and simulator #95 reviews (~/reviews/snap-qc-sim-pr95/r4-checks/), repointed to this checkout:
    • They reproduce 262/262 stepped inputs, 15/15 utility amounts and 283/283 final benefits.
    • Their own step loop agrees with the audit on every new list: the 6 reset-off keys, the 9 stopped-short keys, the 5 household-size misses, the 35/34 at-maximum cases, 32 at $0 income and 2 at the step limit.
    • The audit derives these a different way. It recomputes the benefit at the stepped amount rather than re-running the loop.
  • Mutation checks. Each of these fails a test or the build:
    • restoring the retired sentence;
    • changing "2 of the 7";
    • swapping in the old nearest-to-file-UTIL rule;
    • a shelter cap of 671.
  • Tests:

Invariants (tested)

  • The ported benefit equals the solver's recorded benefit for every replayed row. The build fails otherwise.
  • The described reset rule reproduces every final utility amount (test_solver_outcome_counts_are_consistent).
  • The three moved-miss kinds partition the 20 moved misses. They are disjoint, and each case row carries its kind.
  • 32 + 2 = 34. All 34 are matches, all were moved, all carry the at_max flag, and all sit on the cap stretch. The stretches partition the 60.
  • Property tests (Hypothesis):
    • the benefit never exceeds the maximum allotment;
    • it is non-increasing in earned and unearned income;
    • it is non-decreasing in rent, the utility amount and each deduction;
    • every income at or below one that yields the maximum also yields it;
    • the benefit is the maximum exactly when net income ≤ 0, and net income never falls as income rises. This is the weak-identification mechanism.

Relation to #95

There are no file overlaps with #95. #95 touches the lab's README and amterr_replay.py; this PR touches neither. The paper's replay paragraph exists only on #95's branch, so the paper caveat goes in a separate PR stacked on #95.

axiom: n/a: lab documentation and audit tooling; no policy encoding changes

🤖 Generated with Claude Code

ANALYSIS.md's method section now states the solver's rules as
reconstruct_co_fy2024.R applies them, matching the paper's corrected
footnote in #95:

- The utility reset takes the commonly reported amount nearest the
  solver's stepped amount. Candidates are amounts that more than 5
  filtered cases of any review status report in the state and calendar
  year, above the file's UTIL for util_up and below it for util_down. The
  stepped amount is kept if none lies above; the amount is set to 0 if
  none lies below.
- Rent and utility steps also stop at a zero shelter deduction when
  lowering, and income-raising steps stop when the uncapped benefit falls
  below zero. "every step stops" now reads "all steps stop".
- Household size takes its direction from the nature code, so 2 of the 7
  Colorado moves take the benefit away from RAWBEN.

"In 20 the solver moved an input and stopped short" was true of 9 of the
20. In 6 the $3 steps reached within $3 of RAWBEN and the utility reset
then moved the input off that match; in 5 the household-size move missed.
The results now also flag 34 weakly identified matches: RAWBEN is the
maximum allotment and the solver lowered income, so every income at or
below break-even reproduces it.

audit_claims.py ports the solver's benefit formula and fails unless it
reproduces all 283 recorded benefits. It also re-applies the utility reset,
which reproduces all 15 final amounts. Each new count lands in
claims_audit.json, and test_amterr_lab.py locks the sentences that quote
them. New property tests check the two facts the caveat rests on: the
benefit never exceeds the maximum allotment, and it falls as income rises.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@MaxGhenis

Copy link
Copy Markdown
Contributor Author

Compatibility with #95: tests/test_retired_claims.py exists only on #95's branch, so I ran it on #95's head e72957b merged with this PR's 62f0bb9 (merge commit 93c439e, before any paper edit). uv run --frozen --extra dev --extra analysis pytest -q -p no:cacheprovider tests/test_retired_claims.py tests/test_paper_embed.py tests/test_amterr_lab.py: 49 passed, none skipped. The paper caveat that quotes this PR's new counts is #97, stacked on #95.

An adversarial pass (8 independent agents, each with its own port of the
solver) confirmed every count in the previous commit. It also found that
the weak-identification caveat understated its own argument.

- The caveat now covers every match on a flat stretch of the benefit
  formula, where the match bounds the moved input on one side only: 60 of
  the 230 moved matches, 31 of them above the threshold.
  - 53 are at or within $5 of the maximum allotment. That is the 34 income
    cuts plus 19 where every further rent, utility or medical amount also
    reproduces RAWBEN.
  - 4 are at the $23 minimum benefit.
  - 3 are rent increases at the shelter-deduction cap.
- "Break-even point" usually means the income at which the benefit phases
  out. It now reads "the point where net income reaches zero": the solver
  pays the maximum exactly when net income is zero or less, which a new
  property test checks.
- "ran on to zero income" now reads "ran that income down to $0", since
  other income can remain.
- The stop rules now say "at or above the maximum allotment" and add the
  mirror rule at the minimum benefit.
- The correctedamount note now covers household-size moves (always 0) and
  all 15 utility rows (always the pre-reset change).
- The audit check is stated as covering two parts of the account, which is
  what it does.
- The revision note quotes the old reset rule correctly.
- New notes:
  - 202405-40908's reset leaves the benefit farther from RAWBEN than FSBEN.
  - 2 of the 10 computational misses are reset-off cases.

audit_claims.py classifies each match's flat stretch by pushing the moved
input without limit in the solver's direction (flat_stretch). The new
counts are in claims_audit.json, and test_amterr_lab.py locks them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
MaxGhenis added a commit that referenced this pull request Oct 4, 2026
MaxGhenis added a commit that referenced this pull request Oct 4, 2026
… paper

This follows #96's second round, after 8 independent agents checked the
claims.

- The replay paragraph now gives all 60 weakly identified moved matches,
  31 above the threshold:
  - 53 at or within $5 of the maximum allotment;
  - 4 at the minimum benefit;
  - 3 at the shelter-deduction cap.

  Of the 53, 34 are income cuts at the maximum. The mechanism is now
  stated as "up to the point where net income reaches zero" rather than
  "break-even".
- The 37 misses are now each accounted for once: 17 unmoved, 6 reset-off,
  9 stopped short and 5 household-size misses.
- It notes that 2 of the 10 computational misses are reset-off cases.
- "A stepped utility amount is then reset" now reads "the utility amount",
  since two rows were reset without a step.
- The footnote and D4 now say "at or above the maximum allotment" and add
  the mirror rule at the minimum benefit.
- D5 and D6 are rewritten to match, and the locks follow. A new identity
  checks that the misses partition the 37.

Re-rendered with Quarto 1.9.36. The PDF is still 27 pages.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The independent Opus review of 00de6bf asked for two required fixes and
three recommended ones. All are verified against the solver's formula.

- The 3 shelter-cap matches don't reach the cap: the solver stopped $11
  to $14 short of $672 on the within-$3 rule. The text now says they "stop
  within $14 of the shelter-deduction cap", and that past the cap the
  benefit stays within $5 of RAWBEN. The audit records the gap.
- 2 of the 60 (202403-40765 and 202404-40803) reproduce at every utility
  amount. "On one side only" now reads "on one side at most", and these
  two are named. The audit pushes each input the opposite way too
  (opposite_push_holds, bounded_on_neither_side_keys).
- The caveat now says "by the solver's formula": the engine was not run
  at the pushed amounts. The "farther from RAWBEN" list now fails the
  build unless the engine agrees with the solver's benefit.
- Pushes that stay within $5 without reaching a flat stretch are counted
  and named (unnamed_push_holds_keys: 4 already at $0, 3 whose band runs
  down to $0). ANALYSIS.md says they are not counted.
- More of each lock is derived from the audit: the at_max composition
  (one household-size match, one unmoved), the $23 minimum, the step limit
  and the correctedamount "2 of them".

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
MaxGhenis added a commit that referenced this pull request Oct 4, 2026
MaxGhenis added a commit that referenced this pull request Oct 4, 2026
Mirrors #96's review-r1 fixes and #97's review-r1 suggestions:
- The weak-identification sentence now says the issued benefit lies
  within $5 of a stretch where the solver's formula is flat. The 3
  shelter cases are now "rent past the shelter-deduction cap", since the
  solver stopped $11 to $14 short of it. The sentence also says "by that
  formula", "on one side at most" and "34 of the 53".
- FACTS D5 gives the push the audit uses ($100,000 or $0), notes the
  engine was not run there, and names the 2 matches bounded on neither
  side and the 7 matches it doesn't count.
- "Of the 10, 2 are …", so no sentence opens with a numeral.

Re-rendered; the PDF is still 27 pages.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant