Skip to content

Flag the replay's weakly identified matches and reset-off misses in the paper - #97

Open
MaxGhenis wants to merge 9 commits into
retire-qc-claimsfrom
amterr-replay-caveat-paper
Open

MaxGhenis wants to merge 9 commits into
retire-qc-claimsfrom
amterr-replay-caveat-paper

Conversation

@MaxGhenis

@MaxGhenis MaxGhenis commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Stacked on #95 (base retire-qc-claims, final at e72957b), with #96 merged in. The replay paragraph this edits exists only on #95's branch, and the counts it quotes come from #96's claims_audit.json. Merge order: #96 → #95 → this PR, which GitHub retargets to main once #95 merges. Until then, the diff against #95 also shows #96's lab files. The paper changes are commits 1e2c45d, 773e170 and a945cde. a945cde applies #96's review-r1 fixes and this PR's review-r1 suggestions (~/reviews/snap-qc-sim-pr97/review-r1.md, which approved the content).

What it changes in the paper

Replay paragraph (#sec-decompose)

  • Weak identification, after the threshold rates: "In 60 of the 230 moved matches, 31 of them above the threshold, the issued benefit lies within $5 of a stretch where the solver's benefit formula is flat in the moved input (the maximum allotment in 53, the minimum benefit in 4 and rent past the shelter-deduction cap in 3). By that formula every amount of the input on the stretch reproduces the issued benefit, so the match bounds the input on one side at most and is weakly identified. In 34 of the 53 the issued benefit is the maximum allotment and the solver lowered an income; the solver pays the maximum whenever net income is zero or less, so every value of that income up to the point where net income reaches zero yields it."

  • Accounting of the 37 misses. The solver's-limits sentence now covers each one exactly once:

    • 17 where nothing moved, which the paper already said;
    • "In 6 more, its $3 steps reached within $3 of the issued benefit before the utility reset moved the input off that match";
    • 9 where a stop rule ended the steps short;
    • 5 household-size misses.

    This is round 3's R7. It goes beyond the brief, which asked for the weak-identification caveat only.

  • "Of the 10 [computational misses], 2 are among the 6 the utility reset moved off a match."

  • "a stepped utility amount is then reset" now reads "the utility amount is then reset". Two rows were reset without a step.

Footnote and FACTS D4. The stop rules now say "at or above the maximum allotment" (202311-40369 has RAWBEN above it). They also add the mirror rule: at or below a one- or two-person unit's minimum benefit, income-raising steps run until the uncapped benefit falls below zero.

New FACTS rows

  • D5: weak identification (60 = 53 + 4 + 3; 34 income cuts at the maximum; the at_max flag's 54).
  • D6: what the solver did in the 20 moved misses, the 3 moves that end farther from RAWBEN than FSBEN, and the audit's two checks of D4.

Abstract

The abstract is unchanged. Its claim, that the engine "reproduces the issued benefit within $5 for 78.4% … and 91.4% … which is consistent with correct arithmetic on a wrong input", still holds for every weakly identified match. The caveat qualifies how much those matches show; it doesn't make the sentence false. So this PR needs no decision of its own. Its merge waits only on #95, which is Max's call (d933).

Rendering and revision

  • Re-rendered with cd paper && quarto render index.qmd (Quarto 1.9.36). paper/out/index.html and index.pdf are copied to app/public/paper/web/.
  • The HTML diff is the edited paragraphs, the footnote and the date. The PDF is still 27 pages.
  • Revision 10 is not yet published (production serves revision 9 until d934), so it keeps its number. Its date moves to 2026-10-04 in index.qmd and the wrapper, and the wrapper's cache key moves to r10-20261004.

Tests

  • tests/test_retired_claims.py:
    • replay_quotes locks the new sentences against claims_audit.json (case_level.solver_outcomes);
    • the Hypothesis count strategy draws the new counts;
    • test_replay_counts_partition checks that the stretches sum to 60, that 34 ≤ 53, and that 14 + 3 + 6 + 9 + 5 = 37 exactly;
    • test_facts_quotes_the_solver_outcomes locks D5 and D6.
  • Mutation checks. Each of these fails a lock:
    • 60 → 61 in the paragraph;
    • an altered miss-accounting sentence;
    • 4 → 5 in D5.
  • Result: tests/test_retired_claims.py, test_paper_embed.py, test_amterr_lab.py, test_asset_versioning.py and test_cause_shares.py give 73 passed, none skipped.

axiom: n/a: manuscript text and fact catalog; no policy encoding changes

🤖 Generated with Claude Code

MaxGhenis and others added 3 commits October 4, 2026 14:53
ANALYSIS.md's method section now states the solver's rules as
reconstruct_co_fy2024.R applies them, matching the paper's corrected
footnote in #95:

- The utility reset takes the commonly reported amount nearest the
  solver's stepped amount. Candidates are amounts that more than 5
  filtered cases of any review status report in the state and calendar
  year, above the file's UTIL for util_up and below it for util_down. The
  stepped amount is kept if none lies above; the amount is set to 0 if
  none lies below.
- Rent and utility steps also stop at a zero shelter deduction when
  lowering, and income-raising steps stop when the uncapped benefit falls
  below zero. "every step stops" now reads "all steps stop".
- Household size takes its direction from the nature code, so 2 of the 7
  Colorado moves take the benefit away from RAWBEN.

"In 20 the solver moved an input and stopped short" was true of 9 of the
20. In 6 the $3 steps reached within $3 of RAWBEN and the utility reset
then moved the input off that match; in 5 the household-size move missed.
The results now also flag 34 weakly identified matches: RAWBEN is the
maximum allotment and the solver lowered income, so every income at or
below break-even reproduces it.

audit_claims.py ports the solver's benefit formula and fails unless it
reproduces all 283 recorded benefits. It also re-applies the utility reset,
which reproduces all 15 final amounts. Each new count lands in
claims_audit.json, and test_amterr_lab.py locks the sentences that quote
them. New property tests check the two facts the caveat rests on: the
benefit never exceeds the maximum allotment, and it falls as income rises.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…he paper

The replay paragraph now says that 34 of the 230 moved matches, 19 of them
above the threshold, are weakly identified. In each, the issued benefit is
the maximum allotment and the solver lowered an income. Every income at or
below break-even yields the maximum, so these matches do not pin the
income down.

It also adds a sixth cause to the account of non-reproduced cases. In 6,
the $3 steps reached within $3 of the issued benefit, and the utility
reset then moved the input off that match.

New FACTS rows D5 and D6 record both, and test_retired_claims.py locks the
sentences against claims_audit.json (from #96). The paper is re-rendered.
Revision 10 is not yet published, so it keeps its number; its date and
the wrapper's cache key move to 2026-10-04. The abstract is unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
MaxGhenis and others added 6 commits October 4, 2026 15:28
An adversarial pass (8 independent agents, each with its own port of the
solver) confirmed every count in the previous commit. It also found that
the weak-identification caveat understated its own argument.

- The caveat now covers every match on a flat stretch of the benefit
  formula, where the match bounds the moved input on one side only: 60 of
  the 230 moved matches, 31 of them above the threshold.
  - 53 are at or within $5 of the maximum allotment. That is the 34 income
    cuts plus 19 where every further rent, utility or medical amount also
    reproduces RAWBEN.
  - 4 are at the $23 minimum benefit.
  - 3 are rent increases at the shelter-deduction cap.
- "Break-even point" usually means the income at which the benefit phases
  out. It now reads "the point where net income reaches zero": the solver
  pays the maximum exactly when net income is zero or less, which a new
  property test checks.
- "ran on to zero income" now reads "ran that income down to $0", since
  other income can remain.
- The stop rules now say "at or above the maximum allotment" and add the
  mirror rule at the minimum benefit.
- The correctedamount note now covers household-size moves (always 0) and
  all 15 utility rows (always the pre-reset change).
- The audit check is stated as covering two parts of the account, which is
  what it does.
- The revision note quotes the old reset rule correctly.
- New notes:
  - 202405-40908's reset leaves the benefit farther from RAWBEN than FSBEN.
  - 2 of the 10 computational misses are reset-off cases.

audit_claims.py classifies each match's flat stretch by pushing the moved
input without limit in the solver's direction (flat_stretch). The new
counts are in claims_audit.json, and test_amterr_lab.py locks them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… paper

This follows #96's second round, after 8 independent agents checked the
claims.

- The replay paragraph now gives all 60 weakly identified moved matches,
  31 above the threshold:
  - 53 at or within $5 of the maximum allotment;
  - 4 at the minimum benefit;
  - 3 at the shelter-deduction cap.

  Of the 53, 34 are income cuts at the maximum. The mechanism is now
  stated as "up to the point where net income reaches zero" rather than
  "break-even".
- The 37 misses are now each accounted for once: 17 unmoved, 6 reset-off,
  9 stopped short and 5 household-size misses.
- It notes that 2 of the 10 computational misses are reset-off cases.
- "A stepped utility amount is then reset" now reads "the utility amount",
  since two rows were reset without a step.
- The footnote and D4 now say "at or above the maximum allotment" and add
  the mirror rule at the minimum benefit.
- D5 and D6 are rewritten to match, and the locks follow. A new identity
  checks that the misses partition the 37.

Re-rendered with Quarto 1.9.36. The PDF is still 27 pages.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The independent Opus review of 00de6bf asked for two required fixes and
three recommended ones. All are verified against the solver's formula.

- The 3 shelter-cap matches don't reach the cap: the solver stopped $11
  to $14 short of $672 on the within-$3 rule. The text now says they "stop
  within $14 of the shelter-deduction cap", and that past the cap the
  benefit stays within $5 of RAWBEN. The audit records the gap.
- 2 of the 60 (202403-40765 and 202404-40803) reproduce at every utility
  amount. "On one side only" now reads "on one side at most", and these
  two are named. The audit pushes each input the opposite way too
  (opposite_push_holds, bounded_on_neither_side_keys).
- The caveat now says "by the solver's formula": the engine was not run
  at the pushed amounts. The "farther from RAWBEN" list now fails the
  build unless the engine agrees with the solver's benefit.
- Pushes that stay within $5 without reaching a flat stretch are counted
  and named (unnamed_push_holds_keys: 4 already at $0, 3 whose band runs
  down to $0). ANALYSIS.md says they are not counted.
- More of each lock is derived from the audit: the at_max composition
  (one household-size match, one unmoved), the $23 minimum, the step limit
  and the correctedamount "2 of them".

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Mirrors #96's review-r1 fixes and #97's review-r1 suggestions:
- The weak-identification sentence now says the issued benefit lies
  within $5 of a stretch where the solver's formula is flat. The 3
  shelter cases are now "rent past the shelter-deduction cap", since the
  solver stopped $11 to $14 short of it. The sentence also says "by that
  formula", "on one side at most" and "34 of the 53".
- FACTS D5 gives the push the audit uses ($100,000 or $0), notes the
  engine was not run there, and names the 2 matches bounded on neither
  side and the 7 matches it doesn't count.
- "Of the 10, 2 are …", so no sentence opens with a numeral.

Re-rendered; the PDF is still 27 pages.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant