From 62f0bb958e806924f9233c47a80174ca1f098d6e Mon Sep 17 00:00:00 2001
From: Max Ghenis
Date: Sun, 4 Oct 2026 14:53:14 -0400
Subject: [PATCH 1/6] Correct the amterr lab's solver description against the R
code
ANALYSIS.md's method section now states the solver's rules as
reconstruct_co_fy2024.R applies them, matching the paper's corrected
footnote in #95:
- The utility reset takes the commonly reported amount nearest the
solver's stepped amount. Candidates are amounts that more than 5
filtered cases of any review status report in the state and calendar
year, above the file's UTIL for util_up and below it for util_down. The
stepped amount is kept if none lies above; the amount is set to 0 if
none lies below.
- Rent and utility steps also stop at a zero shelter deduction when
lowering, and income-raising steps stop when the uncapped benefit falls
below zero. "every step stops" now reads "all steps stop".
- Household size takes its direction from the nature code, so 2 of the 7
Colorado moves take the benefit away from RAWBEN.
"In 20 the solver moved an input and stopped short" was true of 9 of the
20. In 6 the $3 steps reached within $3 of RAWBEN and the utility reset
then moved the input off that match; in 5 the household-size move missed.
The results now also flag 34 weakly identified matches: RAWBEN is the
maximum allotment and the solver lowered income, so every income at or
below break-even reproduces it.
audit_claims.py ports the solver's benefit formula and fails unless it
reproduces all 283 recorded benefits. It also re-applies the utility reset,
which reproduces all 15 final amounts. Each new count lands in
claims_audit.json, and test_amterr_lab.py locks the sentences that quote
them. New property tests check the two facts the caveat rests on: the
benefit never exceeds the maximum allotment, and it falls as income rises.
Co-Authored-By: Claude Opus 5.5
---
paper/snapshot/labs/amterr/ANALYSIS.md | 126 ++++++--
paper/snapshot/labs/amterr/audit_claims.py | 321 ++++++++++++++++++-
paper/snapshot/labs/amterr/claims_audit.json | 188 +++++++++++
tests/test_amterr_lab.py | 245 +++++++++++++-
4 files changed, 833 insertions(+), 47 deletions(-)
diff --git a/paper/snapshot/labs/amterr/ANALYSIS.md b/paper/snapshot/labs/amterr/ANALYSIS.md
index 4e23297..add8aa0 100644
--- a/paper/snapshot/labs/amterr/ANALYSIS.md
+++ b/paper/snapshot/labs/amterr/ANALYSIS.md
@@ -5,7 +5,10 @@ top of the 856/856 benefit-parity result (axiom-oracles#268) and the
Giannella/Molin raw-variable reconstruction (github.com/giannella/snap_qc,
run for FY2024 in this lab: `reconstruct_co_fy2024.R`). Revised 2026-10-03:
layer 3, the software-cause split and the cost-share comparison were
-recomputed and corrected, and the weight vintage is now stated (see
+recomputed and corrected, and the weight vintage is now stated. Revised
+2026-10-04: the solver description was corrected against the R code, and
+layer 3 now says what the solver's moves did in each non-reproduced case
+and flags the weakly identified matches (see
[revision history](#revision-history)). Every figure below that is computed
from the QC file is regenerated by `audit_claims.py` into
`claims_audit.json`; `README.md` gives the pins and commands to rerun the
@@ -92,26 +95,47 @@ caught in our own encoding.
The Giannella/Molin solver (FY2024 adaptation, smoothing off) starts from
the file's edited inputs (household size, earned and unearned income, rent,
utilities and deductions) and moves only the input named by the case's first
-finding (ELEMENT1): earned or unearned income, rent, the utility allowance,
-medical, dependent care or child-support deduction in $3 steps, or household
-size by one person for element 150 (natures 7, 12, 14, 16). Income steps
-stop once the recomputed benefit passes RAWBEN (the benefit the agency
-issued). Rent, utility and deduction steps also stop within $3 of it. Steps
-that lower an input stop at zero; rent and utility steps also stop when the
-shelter deduction reaches its cap; every step stops when the benefit reaches
-$0 or after 1,000 steps. Rows the solver labels
-`util_up` or `util_down` then have the utility allowance set to the nearest
-value above (or below) the file's UTIL that more than 5 filtered cases in
-that state and calendar year use; a `util_down` row with no such value is
-set to 0. When ELEMENT1 is outside those lists the inputs stay as they are
-(`correctednotes == "no_change"`). The Axiom
-engine, which reproduces all 856 Colorado FSBEN values at zero tolerance on
-the file's inputs (axiom-oracles#268), then computes the benefit on the
-solver's inputs, and the result is compared with RAWBEN at the file's own $5
-editing tolerance. Below, "moved" means at least one of the eight replayed
-inputs differs from the file value it started from. (The solver's
-`correctedamount` column misses two utility rows, 202403-40765 and
-202404-40803, because it is recorded before the utility snap.)
+finding (ELEMENT1).
+
+- **Household size** moves once, by one person, for element 150: one fewer
+ for natures 12, 14 or 16 and one more for nature 7. The nature code alone
+ sets the direction, so the move can take the benefit away from RAWBEN (the
+ benefit the agency issued). It does in 2 of the 7 Colorado household-size
+ moves, 202310-40265 and 202402-40618.
+- **Earned or unearned income, rent, the utility allowance, and the medical,
+ dependent care or child-support deduction** move in $3 steps toward
+ RAWBEN. Income steps stop once the recomputed benefit passes RAWBEN; rent,
+ utility and deduction steps also stop within $3 of it. Steps that lower an
+ input stop at zero; rent and utility steps also stop when the shelter
+ deduction reaches its cap (when raising) or zero (when lowering);
+ income-raising steps stop when the uncapped benefit falls below zero; and
+ all steps stop when the benefit reaches $0 or after 1,000 steps. Where
+ RAWBEN is the maximum allotment, which the benefit cannot pass,
+ income-lowering steps continue until the income reaches zero or the step
+ limit.
+- **The utility reset** then applies to every row the solver labels
+ `util_up` (RAWBEN above FSBEN) or `util_down` (RAWBEN below FSBEN). The
+ candidates are the utility amounts that more than 5 cases report in the
+ same state and calendar year, counting cases of any review status that
+ pass the solver's consistency filters (listed under Results). A `util_up`
+ row takes the candidate above the file's UTIL that lies nearest the
+ amount its steps reached, and keeps that stepped amount if no candidate
+ lies above. A `util_down` row takes the candidate below the file's UTIL
+ nearest its stepped amount, and is set to 0 if none lies below.
+- **Any other first finding** leaves the inputs as they are
+ (`correctednotes == "no_change"`).
+
+The Axiom engine, which reproduces all 856 Colorado FSBEN values at zero
+tolerance on the file's inputs (axiom-oracles#268), then computes the
+benefit on the solver's inputs, and the result is compared with RAWBEN at
+the file's own $5 editing tolerance. Below, "moved" means at least one of
+the eight replayed inputs differs from the file value it started from. (The
+solver's `correctedamount` column misses two utility rows, 202403-40765 and
+202404-40803, because it is recorded before the utility reset.)
+`audit_claims.py` checks this account against the solver's output. Its port
+of the solver's benefit formula reproduces all 283 recreated benefits, and
+the reset rule above reproduces the final utility amount in all 15 utility
+rows.
FSBEN is a constructed variable: the final benefit Mathematica's model
calculates from the edited inputs (technical documentation: listed as
@@ -127,10 +151,22 @@ certified to receive in the sample month (PDF p. 92).
- 246 of 283 (86.9% of cases, 72.0% of the $99.1M replayed error dollars)
reproduce the issued benefit within $5. In 230 the solver moved an input;
in 16 nothing moved and |RAWBEN − FSBEN| ≤ $5 already.
-- 37 do not reproduce. In 20 the solver moved an input and stopped short.
- In 17 nothing moved: 14 `no_change` rows and 3 rows where the solver's
- element applied but it stopped before moving. For those 17 the replayed
- input is the file's input, so the engine returns FSBEN.
+- 34 of the 230 are weakly identified. In each, RAWBEN is the maximum
+ allotment and the solver lowered an income. Every value of that income at
+ or below its break-even point yields the maximum, so the match does not
+ pin the income down. The steps, which cannot pass RAWBEN, ran on to zero
+ income in 32 and to the 1,000-step limit in 2. 19 of the 34 are above the
+ $56 threshold. The solver exports an `at_max` flag to mark such cases (the
+ uncapped benefit on the replayed inputs is at least the maximum allotment
+ less $5); it is set for 54 of the 246 matches, these 34 among them.
+- 37 do not reproduce. In 20 the solver moved an input: in 9 its $3 steps
+ stopped short of RAWBEN; in 6 they reached within $3 of it and the
+ utility reset then moved the input off that match (202310-40297,
+ 202312-40456, 202312-40513, 202402-40603, 202405-40908 and 202409-41321);
+ and in 5 the one-person household-size move missed. In 17 nothing moved:
+ 14 `no_change` rows and 3 rows where the solver's element applied but it
+ stopped before moving. For those 17 the replayed input is the file's
+ input, so the engine returns FSBEN.
- The solver's own recomputed benefit and the engine agree on the within-$5
classification for all 283 cases (246 reproduced, 37 not).
@@ -140,7 +176,8 @@ with an input error on that element; it does not identify which input the
agency had wrong, and it does not rule out a computation error that a moved
input absorbed. A non-reproduced case does not by itself show a computation
error: the solver tries one element in one direction, so multi-element and
-household-composition errors also land among the 37.
+household-composition errors also land among the 37, as do the 6 cases the
+utility reset moved off a match their steps had reached.
The replay outcome does not separate the layer-2 computational findings. Of
the 26 cases that carry one, 13 reproduce ($3.6M/yr, 3.2% of Colorado error
@@ -177,7 +214,9 @@ What the replay shows for these cases:
for those cases. It does not show what facts the agency recorded or what
the agency's system computed from them.
- In the other 3 the solver moved a different finding's input, and the
- engine's benefit matches neither RAWBEN nor FSBEN.
+ engine's benefit matches neither RAWBEN nor FSBEN. In 2 of them,
+ 202312-40456 and 202312-40513, the utility steps had reached within $3 of
+ RAWBEN before the reset moved the input off that match.
- 7 of the 10 broad-coded findings concern the benefit computation (520 in
six cases; 362/52 in one). These 7 computation candidates carry $2.0M/yr,
1.8% of Colorado error dollars. The other 3 are income or eligibility
@@ -225,8 +264,10 @@ for codes 20/21 and $1,210M (14.9%) for codes 10/22; neither reproduces.
2 ($0.8M) it moved rent.
- 2 do not reproduce ($2.4M/yr): 202312-40513 and 202401-40548. In both the
software-coded finding (331/44/17 and 350/44/17) is on an element the
- solver never moved; it moved ELEMENT1 (364 and 363). The replay does not
- show that the computation logic failed in either case.
+ solver never moved; it moved ELEMENT1 (364 and 363). In 202312-40513 its
+ utility steps had reached within $3 of RAWBEN before the reset moved the
+ input off that match. The replay does not show that the computation logic
+ failed in either case.
- Mass-change (19) findings in these 16 are on RSDI (4 cases), SSI (2) and
the medical deduction (1); programming (17) findings are on RSDI, SSI,
child support received and the child-support deduction. None is on rent
@@ -281,8 +322,9 @@ above the threshold. The 7 candidates are 7 sampled cases.
- The engine replay checks the benefit computation. Eligibility-side errors
(person included or excluded) enter only through the solver's household
size step.
-- The solver moves one element in $3 steps; $5 is the comparison tolerance.
- Exact-tolerance rates are not meaningful at this layer.
+- The solver moves one element: in $3 steps, then a reset for the utility
+ allowance, or by one person for household size. $5 is the comparison
+ tolerance. Exact-tolerance rates are not meaningful at this layer.
- FY2024 was evaluated through the compile-time COLA overlay at the nominal
period, as in the parity run; rulespec-us#759 (the dependency inversion
that would remove the overlay) was still open on 2026-10-03.
@@ -367,3 +409,25 @@ P.L. 119-21 sec. 10105. Rerun instructions and pins: `README.md`.
that FY2025 was still being measured.
- Paths in the scripts are configurable; `README.md` gives pins and
commands. `phase_a_classification.json` lives one directory up.
+- **2026-10-04.** The solver description now matches
+ `reconstruct_co_fy2024.R`. `audit_claims.py` ports the solver's benefit
+ formula and re-applies the utility reset, and the new counts below come
+ from it. Changes:
+ - Utility reset: the solver takes the candidate nearest the amount its
+ steps reached. The 2026-10-03 text said the candidate nearest the
+ file's UTIL; that rule gives the final amount in 8 of the 15 utility
+ rows. The candidates are counted over filtered cases of any review
+ status, and a `util_up` row with no candidate above keeps its stepped
+ amount.
+ - Stop rules: rent and utility steps also stop at a zero shelter
+ deduction when lowering, and income-raising steps stop when the
+ uncapped benefit falls below zero. Where RAWBEN is the maximum
+ allotment, income-lowering steps run on to zero income or the step
+ limit.
+ - Household size: the nature code sets the direction. The move takes the
+ benefit away from RAWBEN in 2 of the 7 Colorado cases.
+ - "In 20 the solver moved an input and stopped short" was true of 9 of
+ the 20. In 6 the steps reached within $3 of RAWBEN and the utility reset
+ moved the input off that match; in 5 the household-size move missed.
+ - Added: 34 of the 230 moved matches are weakly identified (RAWBEN at the
+ maximum allotment, income lowered).
diff --git a/paper/snapshot/labs/amterr/audit_claims.py b/paper/snapshot/labs/amterr/audit_claims.py
index ef76733..39fab01 100644
--- a/paper/snapshot/labs/amterr/audit_claims.py
+++ b/paper/snapshot/labs/amterr/audit_claims.py
@@ -8,7 +8,10 @@
Outputs ``claims_audit.json`` (this directory). The audit also regenerates
``native_decomposition.json`` and ``../phase_a_classification.json`` from the
-May 2026 posting and records whether they match the committed copies.
+May 2026 posting and records whether they match the committed copies. To check
+ANALYSIS.md's account of the solver, it ports the solver's benefit formula and
+utility reset from ``reconstruct_co_fy2024.R`` and fails unless the port
+reproduces every benefit the solver recorded.
The two postings differ only in the weight columns (HWGT, FYWGT, HWGT_OLD,
FYWGT_OLD); the audit re-verifies that, so every case-level result (which
@@ -140,6 +143,42 @@
366: "rawcsded",
150: "rawusize",
}
+# The solver's step size, step limit and within-$3 stop (income_shift,
+# max_iterations and diff_matches in reconstruct_co_fy2024.R).
+SOLVER_STEP = 3
+SOLVER_MAX_STEPS = 1000
+SOLVER_MATCH = 3
+# FY2024 excess shelter deduction cap (48 states and DC; snap_qc
+# additional_data/year_data.csv). The solver lifts it for units with an
+# elderly or disabled member (FSNELDER + FSNDIS > 0).
+SHELTER_CAP_FY2024 = 672
+# FY2024 maximum allotments by household size, 48 states and DC (snap_qc
+# additional_data/max_allotments.csv). The solver keeps only rows whose BENMAX
+# matches this table, which also bounds the utility reset's candidate pool.
+MAX_ALLOTMENT_FY2024 = {
+ **dict(enumerate((291, 535, 766, 973, 1155, 1386, 1532, 1751), start=1)),
+ **{size: 1751 + 219 * (size - 8) for size in range(9, 21)},
+}
+# The reset's candidates: amounts more than this many filtered cases report.
+UTILITY_CANDIDATE_MIN_CASES = 5
+UTILITY_RESET_NOTES = ("util_up", "util_down")
+# Labels adjust_income gives a moved income when RAWBEN exceeds the file-side
+# uncapped benefit, so its steps lower the income (_error: it reached zero).
+INCOME_LOWERING_NOTES = ("earn_down", "earn_error", "unearn_down", "unearn_error")
+# The solver's benefit inputs, as written to co_fy2024_reconstruction.csv.
+SOLVER_BENEFIT_INPUTS = (
+ "rawearn",
+ "rawunearn",
+ "rawrent",
+ "rawutil",
+ "rawmedded",
+ "rawdepded",
+ "rawcsded",
+ "rawstdded",
+ "rawhomeless_ded",
+ "rawbenmax",
+ "rawminimum_ben",
+)
FINDING_COLUMNS = tuple(
f"{root}{slot}"
@@ -155,6 +194,9 @@
"RAWBEN",
"FSBEN",
"HWGT",
+ "BENMAX",
+ "FSNELDER",
+ "FSNDIS",
*REPLAY_INPUTS.values(),
*FINDING_COLUMNS,
)
@@ -427,9 +469,17 @@ def load_replay() -> pd.DataFrame:
case_key(y, h)
for y, h in zip(reconstruction["YRMONTH"], reconstruction["HHLDNO"])
]
+ solver_columns = [c for c in SOLVER_BENEFIT_INPUTS if c not in REPLAY_INPUTS]
keep = reconstruction[
- ["key", "correctedamount", "ELEMENT1", *REPLAY_INPUTS]
- ].rename(columns={"ELEMENT1": "solver_element"})
+ [
+ "key",
+ "correctedamount",
+ "ELEMENT1",
+ *REPLAY_INPUTS,
+ *solver_columns,
+ "at_max",
+ ]
+ ].rename(columns={"ELEMENT1": "solver_element", "at_max": "solver_at_max"})
merged = replay.merge(keep, on="key", how="left", validate="1:1")
if merged["correctedamount"].isna().any():
raise AssertionError("replay rows missing from the reconstruction")
@@ -470,7 +520,229 @@ def join_replay(replay: pd.DataFrame, posting: pd.DataFrame) -> pd.DataFrame:
joined["solver_moved_input"] = changed.any(axis=1)
if (joined["solver_moved_input"] & joined["solver_no_change"]).any():
raise AssertionError("a no_change row has a changed input")
- return joined
+ return classify_solver_moves(joined)
+
+
+def solver_benefit(inputs: pd.DataFrame) -> pd.Series:
+ """The solver's benefit: calculate_raw_benefits in reconstruct_co_fy2024.R.
+
+ ``inputs`` carries ``SOLVER_BENEFIT_INPUTS`` and ``shelter_cap`` (infinite
+ for a unit with an elderly or disabled member). The operations and their
+ order follow the R function, so the floors land on the same doubles.
+ """
+ earn = inputs["rawearn"]
+ net_before_shelter = (earn + inputs["rawunearn"]) - (
+ earn * 0.2
+ + inputs["rawdepded"]
+ + inputs["rawmedded"]
+ + inputs["rawcsded"]
+ + inputs["rawstdded"]
+ )
+ half = np.maximum(net_before_shelter * 0.5, 0)
+ uncapped_shelter = np.floor((inputs["rawrent"] + inputs["rawutil"]) - half)
+ shelter = np.floor(
+ np.minimum(np.maximum(uncapped_shelter, 0), inputs["shelter_cap"])
+ )
+ net = np.floor(net_before_shelter - (shelter + inputs["rawhomeless_ded"]))
+ uncapped = np.floor(inputs["rawbenmax"] - 0.3 * net)
+ return np.minimum(
+ np.maximum(uncapped, inputs["rawminimum_ben"]), inputs["rawbenmax"]
+ )
+
+
+def classify_solver_moves(joined: pd.DataFrame) -> pd.DataFrame:
+ """Recompute what each move did, with the solver's own benefit formula.
+
+ * ``benefit_after_steps``: the benefit on the inputs the $3 steps reached,
+ before the utility reset (the stepped utility amount is the file's UTIL
+ plus ``correctedamount``, which the solver records before the reset).
+ * ``moved_miss_kind``, for a moved input that does not reproduce:
+ ``household_size`` (the one-person move), ``reset_off_after_step_match``
+ (the steps reached within $3 of RAWBEN and the reset moved the input off
+ that match) or ``steps_stopped_short`` (a stop rule ended the steps on
+ the starting side of RAWBEN, more than $3 from it).
+ """
+ out = joined.copy()
+ inputs = out[list(SOLVER_BENEFIT_INPUTS)].astype(float)
+ elderly_or_disabled = (out["FSNELDER"] + out["FSNDIS"]) > 0
+ inputs["shelter_cap"] = np.where(elderly_or_disabled, np.inf, SHELTER_CAP_FY2024)
+ final = solver_benefit(inputs)
+ if not final.eq(out["rawben_recreated"].astype(float)).all():
+ raise AssertionError("solver_benefit does not reproduce rawben_recreated")
+
+ notes = out["correctednotes"]
+ utility = notes.isin(UTILITY_RESET_NOTES)
+ stepped = inputs.assign(rawutil=out["UTIL"] + out["correctedamount"])
+ out["benefit_after_steps"] = final.where(~utility, solver_benefit(stepped))
+ rawben, fsben = out["RAWBEN"].astype(float), out["FSBEN"].astype(float)
+ steps_matched = (rawben - out["benefit_after_steps"]).abs() <= SOLVER_MATCH
+ household = notes.str.startswith("hhsize")
+ moved_miss = out["solver_moved_input"] & ~out["reproduced"]
+ kind = pd.Series(
+ np.select(
+ [
+ moved_miss & household,
+ moved_miss & ~household & steps_matched,
+ moved_miss & ~household & ~steps_matched,
+ ],
+ [
+ "household_size",
+ "reset_off_after_step_match",
+ "steps_stopped_short",
+ ],
+ default=None,
+ ),
+ index=out.index,
+ dtype=object,
+ )
+ if (kind.eq("reset_off_after_step_match") & ~utility).any():
+ raise AssertionError("a stepped non-utility miss ended within $3")
+ short = kind.eq("steps_stopped_short")
+ starting_side = np.sign(rawben - out["benefit_after_steps"]) == np.sign(
+ rawben - fsben
+ )
+ if not starting_side[short].all():
+ raise AssertionError("a miss counted as stopped short passed RAWBEN")
+ out["moved_miss_kind"] = kind
+ # The household-size move's direction comes from the nature code alone.
+ out["household_size_move_away_from_rawben"] = household & (
+ (final - fsben) * (rawben - fsben) < 0
+ )
+ out["income_lowered_rawben_at_maximum"] = notes.isin(
+ INCOME_LOWERING_NOTES
+ ) & rawben.eq(out["rawbenmax"].astype(float))
+ return out
+
+
+def solver_outcomes(joined: pd.DataFrame) -> dict[str, Any]:
+ """Counts behind ANALYSIS.md's account of what the solver's moves did."""
+ reproduced = joined["reproduced"]
+ keys = joined["key"]
+ notes = joined["correctednotes"]
+
+ household = notes.str.startswith("hhsize")
+ away = joined["household_size_move_away_from_rawben"]
+
+ utility = notes.isin(UTILITY_RESET_NOTES)
+ steps_matched = (
+ joined["RAWBEN"].astype(float) - joined["benefit_after_steps"]
+ ).abs() <= SOLVER_MATCH
+ reset_off = joined["moved_miss_kind"].eq("reset_off_after_step_match")
+
+ kinds = ("steps_stopped_short", "reset_off_after_step_match", "household_size")
+ moved_miss = joined["solver_moved_input"] & ~reproduced
+
+ at_maximum = joined["income_lowered_rawben_at_maximum"]
+ matched = at_maximum & reproduced
+ lowered = joined["rawearn"].where(notes.str.startswith("earn"), joined["rawunearn"])
+ at_zero = matched & lowered.eq(0)
+ at_limit = matched & ~at_zero
+ at_limit &= joined["correctedamount"].eq(-SOLVER_STEP * SOLVER_MAX_STEPS)
+ if (at_zero | at_limit).sum() != matched.sum():
+ raise AssertionError(
+ "an at-maximum income match ended before zero or the limit"
+ )
+ flagged = joined["solver_at_max"].astype(bool) & reproduced
+
+ return {
+ "solver_benefit_reproduces_rawben_recreated_n": len(joined),
+ "household_size": {
+ "n": int(household.sum()),
+ "reproduced_n": int((household & reproduced).sum()),
+ "away_from_rawben_keys": sorted(keys[away]),
+ },
+ "utility_reset": {
+ "n": int(utility.sum()),
+ "steps_within_3_of_rawben_n": int((utility & steps_matched).sum()),
+ "reproduced_n": int((utility & reproduced).sum()),
+ "reset_off_after_step_match_keys": sorted(keys[reset_off]),
+ },
+ "not_reproduced_solver_moved_input": {
+ "n": int(moved_miss.sum()),
+ **{
+ f"{kind}_keys": sorted(keys[joined["moved_miss_kind"].eq(kind)])
+ for kind in kinds
+ },
+ },
+ "income_lowered_rawben_at_maximum": {
+ "n": int(at_maximum.sum()),
+ "reproduced_n": int(matched.sum()),
+ "reproduced_above_threshold_n": int(
+ (matched & joined["AMTERR"].gt(THRESHOLD_FY2024)).sum()
+ ),
+ "reproduced_ended_at_zero_income_n": int(at_zero.sum()),
+ "reproduced_ended_at_step_limit_n": int(at_limit.sum()),
+ "reproduced_at_max_flag_n": int((matched & flagged).sum()),
+ "reproduced_keys": sorted(keys[matched]),
+ },
+ "reproduced_at_max_flag": {
+ "n": int(flagged.sum()),
+ "by_correctednotes": {
+ note: int(count)
+ for note, count in sorted(notes[flagged].value_counts().items())
+ },
+ },
+ }
+
+
+def utility_reset_check(frame: pd.DataFrame, joined: pd.DataFrame) -> dict[str, Any]:
+ """Re-apply the utility reset that ANALYSIS.md describes.
+
+ The candidate pool follows reconstruct_co_fy2024.R: Colorado cases of any
+ review status that pass its consistency filters (|RAWBEN - FSBEN| within $5
+ of AMTERR, RENT and UTIL present, BENMAX equal to the FY2024 table value),
+ grouped by calendar year. A candidate is a UTIL amount that more than five
+ of them report. Each util_up (util_down) row takes the candidate above
+ (below) the file's UTIL nearest the amount its steps reached, keeps that
+ amount when none lies above, and gets 0 when none lies below; ties go to
+ the smaller amount, as R's which.min over the sorted counts does. The
+ rule the 2026-10-03 text gave instead, the candidate nearest the file's
+ UTIL, is scored alongside it.
+ """
+ pool = frame[frame["STATE"] == COLORADO_FIPS]
+ consistent = ((pool["RAWBEN"] - pool["FSBEN"]).abs() - pool["AMTERR"]).abs() <= 5
+ pool = pool[consistent & pool["RENT"].notna() & pool["UTIL"].notna()]
+ table = pool["FSUSIZE"].fillna(0).map(MAX_ALLOTMENT_FY2024)
+ pool = pool[pool["BENMAX"].eq(table)]
+ counts = pool.groupby([pool["YRMONTH"] // 100, "UTIL"]).size()
+ candidates: dict[int, list[float]] = {}
+ for (year, amount), n in counts.items():
+ if n > UTILITY_CANDIDATE_MIN_CASES:
+ candidates.setdefault(int(year), []).append(float(amount))
+
+ rows = joined[joined["correctednotes"].isin(UTILITY_RESET_NOTES)]
+ by_rule, by_file_util = [], []
+ for _, row in rows.iterrows():
+ pool_year = np.array(sorted(candidates.get(int(row["YRMONTH"]) // 100, [])))
+ raising = row["correctednotes"] == "util_up"
+ file_util = float(row["UTIL"])
+ side = (
+ pool_year[pool_year > file_util]
+ if raising
+ else pool_year[pool_year < file_util]
+ )
+ stepped = file_util + float(row["correctedamount"])
+ if len(side) == 0:
+ rule = file_rule = stepped if raising else 0.0
+ else:
+ rule = float(side[np.argmin(np.abs(side - stepped))])
+ file_rule = float(side.min() if raising else side.max())
+ final = float(row["rawutil"])
+ by_rule.append(rule == final)
+ by_file_util.append(file_rule == final)
+ keys = list(rows["key"])
+ return {
+ "candidate_pool_n": len(pool),
+ "candidates_by_calendar_year": {
+ str(year): amounts for year, amounts in sorted(candidates.items())
+ },
+ "rows_n": len(rows),
+ "rule_reproduces_final_amount_n": int(sum(by_rule)),
+ "nearest_to_file_util_reproduces_final_amount_n": int(sum(by_file_util)),
+ "nearest_to_file_util_misses_keys": sorted(
+ k for k, hit in zip(keys, by_file_util) if not hit
+ ),
+ }
def case_findings(joined: pd.DataFrame, mask: pd.Series) -> list[dict[str, Any]]:
@@ -489,6 +761,7 @@ def case_findings(joined: pd.DataFrame, mask: pd.Series) -> list[dict[str, Any]]
"reproduced": bool(row["reproduced"]),
"correctednotes": row["correctednotes"],
"solver_moved_input": bool(row["solver_moved_input"]),
+ "moved_miss_kind": row["moved_miss_kind"],
"solver_element": int(row["solver_element"]),
"findings": [list(f) for f in found],
"broad_coded_findings": [list(f) for f in broad],
@@ -591,6 +864,7 @@ def case_level(joined: pd.DataFrame) -> dict[str, Any]:
snapped &= ~joined["correctednotes"].str.startswith("hhsize")
return {
"replay": replay_partition(joined, joined["solver_moved_input"]),
+ "solver_outcomes": solver_outcomes(joined),
"replay_inputs_unchanged_keys": unchanged,
"moved_with_zero_correctedamount_keys": sorted(joined.loc[snapped, "key"]),
"broad_coded_misses": {
@@ -878,6 +1152,17 @@ def weighted_results(frame: pd.DataFrame, replay: pd.DataFrame) -> dict[str, Any
# artifact
+def posting_case_level(frame: pd.DataFrame, replay: pd.DataFrame) -> dict[str, Any]:
+ """Every case-level block; build() checks it is the same under each posting."""
+ joined = join_replay(replay, frame)
+ return {
+ **case_level(joined),
+ "utility_reset_rule": utility_reset_check(frame, joined),
+ "layer2_computational_findings": computational_findings(frame),
+ "reconstruction": reconstruction_rows(),
+ }
+
+
def build(postings: Iterable[str] = tuple(POSTINGS)) -> dict[str, Any]:
postings = tuple(postings)
frames = {label: load_posting(label) for label in postings}
@@ -913,6 +1198,22 @@ def build(postings: Iterable[str] = tuple(POSTINGS)) -> dict[str, Any]:
+ ", ".join(REPLAY_INPUTS.values())
+ "; missing -> 0)"
),
+ "benefit_after_steps": (
+ "the solver's benefit formula (calculate_raw_benefits, ported as "
+ "solver_benefit and checked against rawben_recreated for every "
+ "row) on the inputs its $3 steps reached; for util_up and "
+ "util_down rows the utility amount is UTIL + correctedamount, "
+ "before the reset to a commonly reported amount"
+ ),
+ "moved_miss_kind": (
+ "for a moved input that does not reproduce: household_size; "
+ f"reset_off_after_step_match (benefit_after_steps within "
+ f"${SOLVER_MATCH} of RAWBEN); otherwise steps_stopped_short"
+ ),
+ "income_lowered_rawben_at_maximum": (
+ f"correctednotes in {list(INCOME_LOWERING_NOTES)} and RAWBEN "
+ "equal to the solver's maximum allotment (rawbenmax)"
+ ),
"computation_candidate": (
"not reproduced, and a finding carrying a code in "
f"{sorted(BROAD_CODES)} is computational under the layer-2 rule"
@@ -931,11 +1232,7 @@ def build(postings: Iterable[str] = tuple(POSTINGS)) -> dict[str, Any]:
phase_a_classification(legacy), PHASE_A_PATH
),
},
- "case_level": {
- **case_level(join_replay(replay, legacy)),
- "layer2_computational_findings": computational_findings(legacy),
- "reconstruction": reconstruction_rows(),
- },
+ "case_level": posting_case_level(legacy, replay),
"by_posting": {
label: weighted_results(frame, replay) for label, frame in frames.items()
},
@@ -943,11 +1240,7 @@ def build(postings: Iterable[str] = tuple(POSTINGS)) -> dict[str, Any]:
if {"may2026", "aug2026"} <= set(postings):
artifact["posting_diff"] = posting_diff()
for label in postings:
- level = {
- **case_level(join_replay(replay, frames[label])),
- "layer2_computational_findings": computational_findings(frames[label]),
- "reconstruction": artifact["case_level"]["reconstruction"],
- }
+ level = posting_case_level(frames[label], replay)
if level != artifact["case_level"]:
raise AssertionError(f"case-level results differ under {label}")
return artifact
diff --git a/paper/snapshot/labs/amterr/claims_audit.json b/paper/snapshot/labs/amterr/claims_audit.json
index 7be629d..5a95700 100644
--- a/paper/snapshot/labs/amterr/claims_audit.json
+++ b/paper/snapshot/labs/amterr/claims_audit.json
@@ -41,6 +41,9 @@
"code_class": "a case counts if any of AGENCY1-AGENCY9 is in the class (classes overlap)",
"reproduced": "abs(engine_on_original - RAWBEN) <= 5",
"solver_moved_input": "any replayed input (rawusize, rawearn, rawunearn, rawrent, rawutil, rawmedded, rawdepded, rawcsded) differs from its file field (FSUSIZE, FSEARN, FSUNEARN, RENT, UTIL, FSMEDDED, FSDEPDED, FSCSDED; missing -> 0)",
+ "benefit_after_steps": "the solver's benefit formula (calculate_raw_benefits, ported as solver_benefit and checked against rawben_recreated for every row) on the inputs its $3 steps reached; for util_up and util_down rows the utility amount is UTIL + correctedamount, before the reset to a commonly reported amount",
+ "moved_miss_kind": "for a moved input that does not reproduce: household_size; reset_off_after_step_match (benefit_after_steps within $3 of RAWBEN); otherwise steps_stopped_short",
+ "income_lowered_rawben_at_maximum": "correctednotes in ['earn_down', 'earn_error', 'unearn_down', 'unearn_error'] and RAWBEN equal to the solver's maximum allotment (rawbenmax)",
"computation_candidate": "not reproduced, and a finding carrying a code in [10, 17, 19, 20, 21, 22] is computational under the layer-2 rule",
"points_of_official_fy2024_rate": "share of file error dollars x 9.97; an approximation, since the official rate is regression-adjusted and counts only errors above the $56 threshold"
},
@@ -87,6 +90,116 @@
"util_up": 5
}
},
+ "solver_outcomes": {
+ "solver_benefit_reproduces_rawben_recreated_n": 283,
+ "household_size": {
+ "n": 7,
+ "reproduced_n": 2,
+ "away_from_rawben_keys": [
+ "202310-40265",
+ "202402-40618"
+ ]
+ },
+ "utility_reset": {
+ "n": 15,
+ "steps_within_3_of_rawben_n": 14,
+ "reproduced_n": 8,
+ "reset_off_after_step_match_keys": [
+ "202310-40297",
+ "202312-40456",
+ "202312-40513",
+ "202402-40603",
+ "202405-40908",
+ "202409-41321"
+ ]
+ },
+ "not_reproduced_solver_moved_input": {
+ "n": 20,
+ "steps_stopped_short_keys": [
+ "202310-40304",
+ "202311-40369",
+ "202401-40546",
+ "202401-40548",
+ "202405-40957",
+ "202406-41065",
+ "202408-41268",
+ "202409-41302",
+ "202409-41362"
+ ],
+ "reset_off_after_step_match_keys": [
+ "202310-40297",
+ "202312-40456",
+ "202312-40513",
+ "202402-40603",
+ "202405-40908",
+ "202409-41321"
+ ],
+ "household_size_keys": [
+ "202310-40265",
+ "202311-40407",
+ "202402-40618",
+ "202403-40738",
+ "202406-40998"
+ ]
+ },
+ "income_lowered_rawben_at_maximum": {
+ "n": 35,
+ "reproduced_n": 34,
+ "reproduced_above_threshold_n": 19,
+ "reproduced_ended_at_zero_income_n": 32,
+ "reproduced_ended_at_step_limit_n": 2,
+ "reproduced_at_max_flag_n": 34,
+ "reproduced_keys": [
+ "202310-40260",
+ "202310-40281",
+ "202310-40303",
+ "202311-40362",
+ "202311-40366",
+ "202311-40389",
+ "202312-40457",
+ "202312-40501",
+ "202312-40506",
+ "202312-40511",
+ "202401-40557",
+ "202401-40578",
+ "202401-40584",
+ "202401-40599",
+ "202402-40672",
+ "202403-40712",
+ "202403-40744",
+ "202404-40805",
+ "202404-40869",
+ "202405-40879",
+ "202405-40919",
+ "202405-40947",
+ "202405-40960",
+ "202406-41001",
+ "202406-41032",
+ "202406-41053",
+ "202406-41059",
+ "202406-41071",
+ "202407-41150",
+ "202407-41165",
+ "202407-41175",
+ "202408-41240",
+ "202408-41282",
+ "202409-41365"
+ ]
+ },
+ "reproduced_at_max_flag": {
+ "n": 54,
+ "by_correctednotes": {
+ "earn_down": 19,
+ "hhsize_up": 1,
+ "med_up": 1,
+ "rent_down": 1,
+ "rent_up": 14,
+ "unearn_down": 15,
+ "util_down": 1,
+ "util_up": 2
+ }
+ }
+ },
"replay_inputs_unchanged_keys": [
"202310-40314",
"202310-40336",
@@ -163,6 +276,7 @@
"reproduced": false,
"correctednotes": "no_change",
"solver_moved_input": false,
+ "moved_miss_kind": null,
"solver_element": 520,
"findings": [
[
@@ -190,6 +304,7 @@
"reproduced": false,
"correctednotes": "util_up",
"solver_moved_input": true,
+ "moved_miss_kind": "reset_off_after_step_match",
"solver_element": 364,
"findings": [
[
@@ -227,6 +342,7 @@
"reproduced": false,
"correctednotes": "no_change",
"solver_moved_input": false,
+ "moved_miss_kind": null,
"solver_element": 520,
"findings": [
[
@@ -254,6 +370,7 @@
"reproduced": false,
"correctednotes": "util_up",
"solver_moved_input": true,
+ "moved_miss_kind": "reset_off_after_step_match",
"solver_element": 364,
"findings": [
[
@@ -291,6 +408,7 @@
"reproduced": false,
"correctednotes": "no_change",
"solver_moved_input": false,
+ "moved_miss_kind": null,
"solver_element": 520,
"findings": [
[
@@ -318,6 +436,7 @@
"reproduced": false,
"correctednotes": "rent_down",
"solver_moved_input": true,
+ "moved_miss_kind": "steps_stopped_short",
"solver_element": 363,
"findings": [
[
@@ -355,6 +474,7 @@
"reproduced": false,
"correctednotes": "no_change",
"solver_moved_input": false,
+ "moved_miss_kind": null,
"solver_element": 362,
"findings": [
[
@@ -382,6 +502,7 @@
"reproduced": false,
"correctednotes": "no_change",
"solver_moved_input": false,
+ "moved_miss_kind": null,
"solver_element": 161,
"findings": [
[
@@ -414,6 +535,7 @@
"reproduced": false,
"correctednotes": "no_change",
"solver_moved_input": false,
+ "moved_miss_kind": null,
"solver_element": 520,
"findings": [
[
@@ -446,6 +568,7 @@
"reproduced": false,
"correctednotes": "no_change",
"solver_moved_input": false,
+ "moved_miss_kind": null,
"solver_element": 520,
"findings": [
[
@@ -925,6 +1048,7 @@
"reproduced": false,
"correctednotes": "hhsize_up",
"solver_moved_input": true,
+ "moved_miss_kind": "household_size",
"solver_element": 150,
"findings": [
[
@@ -951,6 +1075,7 @@
"reproduced": false,
"correctednotes": "util_up",
"solver_moved_input": true,
+ "moved_miss_kind": "reset_off_after_step_match",
"solver_element": 364,
"findings": [
[
@@ -972,6 +1097,7 @@
"reproduced": false,
"correctednotes": "earn_down",
"solver_moved_input": true,
+ "moved_miss_kind": "steps_stopped_short",
"solver_element": 311,
"findings": [
[
@@ -993,6 +1119,7 @@
"reproduced": false,
"correctednotes": "no_change",
"solver_moved_input": false,
+ "moved_miss_kind": null,
"solver_element": 520,
"findings": [
[
@@ -1020,6 +1147,7 @@
"reproduced": false,
"correctednotes": "earn_error",
"solver_moved_input": true,
+ "moved_miss_kind": "steps_stopped_short",
"solver_element": 312,
"findings": [
[
@@ -1056,6 +1184,7 @@
"reproduced": false,
"correctednotes": "hhsize_up",
"solver_moved_input": true,
+ "moved_miss_kind": "household_size",
"solver_element": 150,
"findings": [
[
@@ -1092,6 +1221,7 @@
"reproduced": false,
"correctednotes": "no_change",
"solver_moved_input": false,
+ "moved_miss_kind": null,
"solver_element": 520,
"findings": [
[
@@ -1113,6 +1243,7 @@
"reproduced": false,
"correctednotes": "util_up",
"solver_moved_input": true,
+ "moved_miss_kind": "reset_off_after_step_match",
"solver_element": 364,
"findings": [
[
@@ -1150,6 +1281,7 @@
"reproduced": false,
"correctednotes": "no_change",
"solver_moved_input": false,
+ "moved_miss_kind": null,
"solver_element": 520,
"findings": [
[
@@ -1177,6 +1309,7 @@
"reproduced": false,
"correctednotes": "no_change",
"solver_moved_input": false,
+ "moved_miss_kind": null,
"solver_element": 362,
"findings": [
[
@@ -1198,6 +1331,7 @@
"reproduced": false,
"correctednotes": "util_up",
"solver_moved_input": true,
+ "moved_miss_kind": "reset_off_after_step_match",
"solver_element": 364,
"findings": [
[
@@ -1235,6 +1369,7 @@
"reproduced": false,
"correctednotes": "no_change",
"solver_moved_input": false,
+ "moved_miss_kind": null,
"solver_element": 520,
"findings": [
[
@@ -1262,6 +1397,7 @@
"reproduced": false,
"correctednotes": "util_down",
"solver_moved_input": true,
+ "moved_miss_kind": "steps_stopped_short",
"solver_element": 364,
"findings": [
[
@@ -1283,6 +1419,7 @@
"reproduced": false,
"correctednotes": "rent_down",
"solver_moved_input": true,
+ "moved_miss_kind": "steps_stopped_short",
"solver_element": 363,
"findings": [
[
@@ -1320,6 +1457,7 @@
"reproduced": false,
"correctednotes": "rent_down",
"solver_moved_input": false,
+ "moved_miss_kind": null,
"solver_element": 363,
"findings": [
[
@@ -1346,6 +1484,7 @@
"reproduced": false,
"correctednotes": "no_change",
"solver_moved_input": false,
+ "moved_miss_kind": null,
"solver_element": 520,
"findings": [
[
@@ -1367,6 +1506,7 @@
"reproduced": false,
"correctednotes": "no_change",
"solver_moved_input": false,
+ "moved_miss_kind": null,
"solver_element": 150,
"findings": [
[
@@ -1393,6 +1533,7 @@
"reproduced": false,
"correctednotes": "util_up",
"solver_moved_input": true,
+ "moved_miss_kind": "reset_off_after_step_match",
"solver_element": 364,
"findings": [
[
@@ -1414,6 +1555,7 @@
"reproduced": false,
"correctednotes": "hhsize_down",
"solver_moved_input": true,
+ "moved_miss_kind": "household_size",
"solver_element": 150,
"findings": [
[
@@ -1435,6 +1577,7 @@
"reproduced": false,
"correctednotes": "no_change",
"solver_moved_input": false,
+ "moved_miss_kind": null,
"solver_element": 362,
"findings": [
[
@@ -1456,6 +1599,7 @@
"reproduced": false,
"correctednotes": "hhsize_up",
"solver_moved_input": true,
+ "moved_miss_kind": "household_size",
"solver_element": 150,
"findings": [
[
@@ -1487,6 +1631,7 @@
"reproduced": false,
"correctednotes": "no_change",
"solver_moved_input": false,
+ "moved_miss_kind": null,
"solver_element": 362,
"findings": [
[
@@ -1508,6 +1653,7 @@
"reproduced": false,
"correctednotes": "util_down",
"solver_moved_input": true,
+ "moved_miss_kind": "reset_off_after_step_match",
"solver_element": 364,
"findings": [
[
@@ -1529,6 +1675,7 @@
"reproduced": false,
"correctednotes": "no_change",
"solver_moved_input": false,
+ "moved_miss_kind": null,
"solver_element": 362,
"findings": [
[
@@ -1556,6 +1703,7 @@
"reproduced": false,
"correctednotes": "unearn_error",
"solver_moved_input": false,
+ "moved_miss_kind": null,
"solver_element": 350,
"findings": [
[
@@ -1582,6 +1730,7 @@
"reproduced": false,
"correctednotes": "earn_down",
"solver_moved_input": true,
+ "moved_miss_kind": "steps_stopped_short",
"solver_element": 311,
"findings": [
[
@@ -1603,6 +1752,7 @@
"reproduced": false,
"correctednotes": "no_change",
"solver_moved_input": false,
+ "moved_miss_kind": null,
"solver_element": 161,
"findings": [
[
@@ -1635,6 +1785,7 @@
"reproduced": false,
"correctednotes": "hhsize_down",
"solver_moved_input": true,
+ "moved_miss_kind": "household_size",
"solver_element": 150,
"findings": [
[
@@ -1656,6 +1807,7 @@
"reproduced": false,
"correctednotes": "earn_error",
"solver_moved_input": true,
+ "moved_miss_kind": "steps_stopped_short",
"solver_element": 312,
"findings": [
[
@@ -1682,6 +1834,7 @@
"reproduced": false,
"correctednotes": "rent_error",
"solver_moved_input": false,
+ "moved_miss_kind": null,
"solver_element": 363,
"findings": [
[
@@ -1708,6 +1861,7 @@
"reproduced": false,
"correctednotes": "no_change",
"solver_moved_input": false,
+ "moved_miss_kind": null,
"solver_element": 150,
"findings": [
[
@@ -1739,6 +1893,7 @@
"reproduced": false,
"correctednotes": "no_change",
"solver_moved_input": false,
+ "moved_miss_kind": null,
"solver_element": 520,
"findings": [
[
@@ -1771,6 +1926,7 @@
"reproduced": false,
"correctednotes": "rent_error",
"solver_moved_input": true,
+ "moved_miss_kind": "steps_stopped_short",
"solver_element": 363,
"findings": [
[
@@ -1797,6 +1953,7 @@
"reproduced": false,
"correctednotes": "no_change",
"solver_moved_input": false,
+ "moved_miss_kind": null,
"solver_element": 520,
"findings": [
[
@@ -1824,6 +1981,7 @@
"reproduced": false,
"correctednotes": "earn_error",
"solver_moved_input": true,
+ "moved_miss_kind": "steps_stopped_short",
"solver_element": 314,
"findings": [
[
@@ -1850,6 +2008,7 @@
"reproduced": false,
"correctednotes": "util_up",
"solver_moved_input": true,
+ "moved_miss_kind": "reset_off_after_step_match",
"solver_element": 364,
"findings": [
[
@@ -1876,6 +2035,7 @@
"reproduced": false,
"correctednotes": "earn_down",
"solver_moved_input": true,
+ "moved_miss_kind": "steps_stopped_short",
"solver_element": 311,
"findings": [
[
@@ -1898,6 +2058,34 @@
"broad_coded_finding_is_computational": false
}
],
+ "utility_reset_rule": {
+ "candidate_pool_n": 803,
+ "candidates_by_calendar_year": {
+ "2023": [
+ 0.0,
+ 91.0,
+ 560.0
+ ],
+ "2024": [
+ 0.0,
+ 91.0,
+ 356.0,
+ 560.0
+ ]
+ },
+ "rows_n": 15,
+ "rule_reproduces_final_amount_n": 15,
+ "nearest_to_file_util_reproduces_final_amount_n": 8,
+ "nearest_to_file_util_misses_keys": [
+ "202310-40297",
+ "202402-40603",
+ "202403-40709",
+ "202403-40716",
+ "202407-41181",
+ "202409-41321",
+ "202409-41353"
+ ]
+ },
"layer2_computational_findings": {
"cases": 26,
"findings": 29,
diff --git a/tests/test_amterr_lab.py b/tests/test_amterr_lab.py
index b128459..4879f4c 100644
--- a/tests/test_amterr_lab.py
+++ b/tests/test_amterr_lab.py
@@ -193,6 +193,77 @@ def test_case_counts_do_not_depend_on_posting():
assert other["layer2"][name]["n"] == metric["n"]
+def test_moved_misses_partition_by_what_the_solver_did():
+ outcomes = CASES["solver_outcomes"]
+ moved = outcomes["not_reproduced_solver_moved_input"]
+ kinds = ("steps_stopped_short", "reset_off_after_step_match", "household_size")
+ lists = [set(moved[f"{kind}_keys"]) for kind in kinds]
+ assert sum(len(keys) for keys in lists) == len(set().union(*lists))
+ assert moved["n"] == len(set().union(*lists))
+ assert moved["n"] == CASES["replay"]["not_reproduced_solver_moved_input_n"]
+ rows = {row["key"]: row for row in CASES["not_reproduced_cases"]}
+ for kind, keys in zip(kinds, lists):
+ for key in keys:
+ assert rows[key]["solver_moved_input"], key
+ assert rows[key]["moved_miss_kind"] == kind, key
+ unmoved = [row for row in rows.values() if not row["solver_moved_input"]]
+ assert all(row["moved_miss_kind"] is None for row in unmoved)
+ for key in moved["reset_off_after_step_match_keys"]:
+ assert rows[key]["correctednotes"] in audit.UTILITY_RESET_NOTES
+ for key in moved["household_size_keys"]:
+ assert rows[key]["correctednotes"].startswith("hhsize")
+ assert (
+ moved["reset_off_after_step_match_keys"]
+ == outcomes["utility_reset"]["reset_off_after_step_match_keys"]
+ )
+
+
+def test_solver_outcome_counts_are_consistent():
+ outcomes = CASES["solver_outcomes"]
+ assert (
+ outcomes["solver_benefit_reproduces_rawben_recreated_n"]
+ == (CASES["replay"]["n"])
+ )
+ household = outcomes["household_size"]
+ assert len(household["away_from_rawben_keys"]) <= household["n"]
+ assert household["reproduced_n"] <= household["n"]
+ missed = {row["key"] for row in CASES["not_reproduced_cases"]}
+ # A move away from RAWBEN cannot reproduce it.
+ assert set(household["away_from_rawben_keys"]) <= missed
+
+ utility = outcomes["utility_reset"]
+ reset_off = utility["reset_off_after_step_match_keys"]
+ assert len(reset_off) <= utility["steps_within_3_of_rawben_n"] <= utility["n"]
+ assert utility["reproduced_n"] <= utility["n"]
+ assert set(reset_off) <= missed
+
+ rule = CASES["utility_reset_rule"]
+ assert rule["rows_n"] == utility["n"]
+ assert rule["rule_reproduces_final_amount_n"] == rule["rows_n"]
+ assert rule["rows_n"] - rule[
+ "nearest_to_file_util_reproduces_final_amount_n"
+ ] == len(rule["nearest_to_file_util_misses_keys"])
+ for amounts in rule["candidates_by_calendar_year"].values():
+ assert amounts == sorted(set(amounts))
+
+ at_maximum = outcomes["income_lowered_rawben_at_maximum"]
+ matched = set(at_maximum["reproduced_keys"])
+ assert len(matched) == at_maximum["reproduced_n"] <= at_maximum["n"]
+ assert matched.isdisjoint(missed)
+ assert matched.isdisjoint(CASES["replay_inputs_unchanged_keys"])
+ assert (
+ at_maximum["reproduced_ended_at_zero_income_n"]
+ + at_maximum["reproduced_ended_at_step_limit_n"]
+ == at_maximum["reproduced_n"]
+ )
+ assert at_maximum["reproduced_above_threshold_n"] <= at_maximum["reproduced_n"]
+ # The solver's own at_max flag marks every one of them, and more.
+ flagged = outcomes["reproduced_at_max_flag"]
+ assert at_maximum["reproduced_at_max_flag_n"] == at_maximum["reproduced_n"]
+ assert at_maximum["reproduced_n"] <= flagged["n"] <= CASES["replay"]["reproduced_n"]
+ assert sum(flagged["by_correctednotes"].values()) == flagged["n"]
+
+
def _millions(value: float, digits: int) -> str:
return f"${value / 1e6:.{digits}f}M"
@@ -201,6 +272,11 @@ def _percent(value: float, digits: int = 1) -> str:
return f"{100 * value:.{digits}f}%"
+def _keys(keys: list[str]) -> str:
+ """Case keys as ANALYSIS.md lists them: "a, b and c"."""
+ return keys[0] if len(keys) == 1 else f"{', '.join(keys[:-1])} and {keys[-1]}"
+
+
def test_analysis_quotes_the_audit():
"""Whole sentences and table rows, so a figure cannot match elsewhere."""
text = " ".join((LAB / "ANALYSIS.md").read_text(encoding="utf-8").split())
@@ -209,6 +285,22 @@ def test_analysis_quotes_the_audit():
counts = CASES["replay"]
replay = may["replay"]
misses = replay["broad_coded_misses"]
+ outcomes = CASES["solver_outcomes"]
+ moved = outcomes["not_reproduced_solver_moved_input"]
+ reset_off = moved["reset_off_after_step_match_keys"]
+ household = outcomes["household_size"]
+ at_maximum = outcomes["income_lowered_rawben_at_maximum"]
+ rule = CASES["utility_reset_rule"]
+ broad_reset_off = [
+ row["key"]
+ for row in CASES["broad_coded_misses"]["cases"]
+ if row["moved_miss_kind"] == "reset_off_after_step_match"
+ ]
+ software_reset_off = [
+ row["key"]
+ for row in CASES["software_coded_replayed"]["cases"]
+ if not row["reproduced"] and row["key"] in reset_off
+ ]
quoted = [
(
f"{counts['reproduced_n']} of {counts['n']} "
@@ -221,8 +313,75 @@ def test_analysis_quotes_the_audit():
f"in {counts['reproduced_without_move_n']} nothing moved"
),
(
- f"In {counts['not_reproduced_solver_moved_input_n']} the solver moved an "
- "input and stopped short"
+ f"In {moved['n']} the solver moved an input: in "
+ f"{len(moved['steps_stopped_short_keys'])} its $3 steps stopped short of "
+ f"RAWBEN; in {len(reset_off)} they reached within $3 of it and the utility "
+ f"reset then moved the input off that match ({_keys(reset_off)}); and in "
+ f"{len(moved['household_size_keys'])} the one-person household-size move "
+ "missed."
+ ),
+ (
+ f"{at_maximum['reproduced_n']} of the "
+ f"{counts['reproduced_solver_moved_input_n']} are weakly identified."
+ ),
+ (
+ "The steps, which cannot pass RAWBEN, ran on to zero income in "
+ f"{at_maximum['reproduced_ended_at_zero_income_n']} and to the "
+ f"1,000-step limit in {at_maximum['reproduced_ended_at_step_limit_n']}. "
+ f"{at_maximum['reproduced_above_threshold_n']} of the "
+ f"{at_maximum['reproduced_n']} are above the ${audit.THRESHOLD_FY2024} "
+ "threshold."
+ ),
+ (
+ f"it is set for {outcomes['reproduced_at_max_flag']['n']} of the "
+ f"{counts['reproduced_n']} matches, these {at_maximum['reproduced_n']} "
+ "among them."
+ ),
+ (
+ f"It does in {len(household['away_from_rawben_keys'])} of the "
+ f"{household['n']} Colorado household-size moves, "
+ f"{_keys(household['away_from_rawben_keys'])}."
+ ),
+ (
+ "Its port of the solver's benefit formula reproduces all "
+ f"{outcomes['solver_benefit_reproduces_rawben_recreated_n']} recreated "
+ "benefits, and the reset rule above reproduces the final utility amount "
+ f"in all {rule['rule_reproduces_final_amount_n']} utility rows."
+ ),
+ (
+ f"as do the {len(reset_off)} cases the utility reset moved off a match "
+ "their steps had reached."
+ ),
+ (
+ f"In {len(broad_reset_off)} of them, {_keys(broad_reset_off)}, the utility "
+ "steps had reached within $3 of RAWBEN before the reset moved the input "
+ "off that match."
+ ),
+ (
+ f"In {_keys(software_reset_off)} its utility steps had reached within $3 "
+ "of RAWBEN before the reset moved the input off that match."
+ ),
+ (
+ "that rule gives the final amount in "
+ f"{rule['nearest_to_file_util_reproduces_final_amount_n']} of the "
+ f"{rule['rows_n']} utility rows."
+ ),
+ (
+ "The move takes the benefit away from RAWBEN in "
+ f"{len(household['away_from_rawben_keys'])} of the {household['n']} "
+ "Colorado cases."
+ ),
+ (
+ f'"In {moved["n"]} the solver moved an input and stopped short" was true '
+ f"of {len(moved['steps_stopped_short_keys'])} of the {moved['n']}. In "
+ f"{len(reset_off)} the steps reached within $3 of RAWBEN and the utility "
+ "reset moved the input off that match; in "
+ f"{len(moved['household_size_keys'])} the household-size move missed."
+ ),
+ (
+ f"Added: {at_maximum['reproduced_n']} of the "
+ f"{counts['reproduced_solver_moved_input_n']} moved matches are weakly "
+ "identified"
),
(
f"They carry {_millions(misses['dollars'], 2)}/yr: "
@@ -470,3 +629,85 @@ def test_any_presence_is_monotone_in_the_code_set(frame):
broad = audit.class_metric(errors, audit.any_code(errors, audit.BROAD_CODES), total)
assert strict["n"] <= broad["n"]
assert strict["dollars"] <= broad["dollars"] + 0.01
+
+
+def test_retired_solver_wording_is_gone():
+ """The 2026-10-03 solver wording survives only in the revision history."""
+ text = " ".join((LAB / "ANALYSIS.md").read_text(encoding="utf-8").split())
+ body = text.split("## Revision history")[0]
+ for phrase in (
+ "moved an input and stopped short",
+ "every step stops",
+ "value above (or below) the file's UTIL",
+ "utility snap",
+ ):
+ assert phrase not in body, phrase
+ assert "value above (or below) the file's UTIL" not in text
+
+
+# ---------------------------------------------------------------------------
+# properties of the ported solver benefit, which the weak-identification
+# caveat relies on: the benefit never exceeds the maximum allotment, falls as
+# income rises and rises with shelter costs and deductions.
+
+
+@st.composite
+def solver_units(draw):
+ """One unit's solver inputs; incomes and costs in whole dollars."""
+ dollars = st.integers(min_value=0, max_value=6000)
+ size = draw(st.integers(min_value=1, max_value=20))
+ return {
+ "rawearn": float(draw(dollars)),
+ "rawunearn": float(draw(dollars)),
+ "rawrent": float(draw(dollars)),
+ "rawutil": float(draw(st.sampled_from((0, 91, 356, 560)))),
+ "rawmedded": float(draw(st.integers(min_value=0, max_value=1500))),
+ "rawdepded": float(draw(st.integers(min_value=0, max_value=1500))),
+ "rawcsded": float(draw(st.integers(min_value=0, max_value=1500))),
+ "rawstdded": float(draw(st.sampled_from((198, 208, 244, 279)))),
+ "rawhomeless_ded": float(draw(st.sampled_from((0, 179)))),
+ "rawbenmax": float(audit.MAX_ALLOTMENT_FY2024[size]),
+ "rawminimum_ben": 23.0 if size < 3 else 0.0,
+ "shelter_cap": draw(
+ st.sampled_from((float(audit.SHELTER_CAP_FY2024), float("inf")))
+ ),
+ }
+
+
+def _benefits(unit: dict, column: str, values: list[float]) -> np.ndarray:
+ frame = pd.DataFrame([{**unit, column: value} for value in values])
+ return audit.solver_benefit(frame).to_numpy()
+
+
+@settings(max_examples=300, deadline=None)
+@given(
+ solver_units(),
+ st.sampled_from(("rawearn", "rawunearn")),
+ st.lists(st.integers(min_value=0, max_value=9000), min_size=2, max_size=30),
+)
+def test_benefit_falls_as_income_rises_and_stays_at_the_maximum_below_break_even(
+ unit, column, incomes
+):
+ values = sorted(float(v) for v in incomes)
+ benefits = _benefits(unit, column, values)
+ assert np.all(np.diff(benefits) <= 0)
+ assert np.all(benefits <= unit["rawbenmax"])
+ assert np.all(benefits >= min(unit["rawminimum_ben"], unit["rawbenmax"]))
+ # Every income at or below one that yields the maximum also yields it.
+ at_maximum = benefits == unit["rawbenmax"]
+ if at_maximum.any():
+ last = np.flatnonzero(at_maximum).max()
+ assert at_maximum[: last + 1].all()
+
+
+@settings(max_examples=300, deadline=None)
+@given(
+ solver_units(),
+ st.sampled_from(("rawrent", "rawutil", "rawmedded", "rawdepded", "rawcsded")),
+ st.lists(st.integers(min_value=0, max_value=6000), min_size=2, max_size=30),
+)
+def test_benefit_rises_with_shelter_costs_and_deductions(unit, column, amounts):
+ values = sorted(float(v) for v in amounts)
+ benefits = _benefits(unit, column, values)
+ assert np.all(np.diff(benefits) >= 0)
+ assert np.all(benefits <= unit["rawbenmax"])
From 1e2c45d27c77c6c80f2caf79f727a5c24c1c0bc6 Mon Sep 17 00:00:00 2001
From: Max Ghenis
Date: Sun, 4 Oct 2026 15:00:43 -0400
Subject: [PATCH 2/6] Flag the replay's weakly identified matches and reset-off
misses in the paper
The replay paragraph now says that 34 of the 230 moved matches, 19 of them
above the threshold, are weakly identified. In each, the issued benefit is
the maximum allotment and the solver lowered an income. Every income at or
below break-even yields the maximum, so these matches do not pin the
income down.
It also adds a sixth cause to the account of non-reproduced cases. In 6,
the $3 steps reached within $3 of the issued benefit, and the utility
reset then moved the input off that match.
New FACTS rows D5 and D6 record both, and test_retired_claims.py locks the
sentences against claims_audit.json (from #96). The paper is re-rendered.
Revision 10 is not yet published, so it keeps its number; its date and
the wrapper's cache key move to 2026-10-04. The abstract is unchanged.
Co-Authored-By: Claude Opus 5.5
---
app/public/paper/index.html | 10 ++--
app/public/paper/web/index.html | 14 +++---
app/public/paper/web/index.pdf | Bin 214069 -> 214845 bytes
paper/FACTS.md | 2 +
paper/index.qmd | 12 +++--
tests/test_retired_claims.py | 78 ++++++++++++++++++++++++++++++++
6 files changed, 101 insertions(+), 15 deletions(-)
diff --git a/app/public/paper/index.html b/app/public/paper/index.html
index 506aae1..55fd83f 100644
--- a/app/public/paper/index.html
+++ b/app/public/paper/index.html
@@ -76,7 +76,7 @@
Administrative quality control data as an oracle
failures, and the fiscal 2025 rates landing where the sampling-noise
mechanism says realized years should: 18 of 53 jurisdictions changed
cost-share tiers, only 7 of those changes beyond two-year sampling
- noise. This page embeds revision 10 (2026-10-03), which corrects how
+ noise. This page embeds revision 10 (2026-10-04), which corrects how
revision 9 described the QC file's computed benefit and the Colorado
error-case replay, adds Colorado's fiscal 2025 rate beside its fiscal
2024 margin, and withdraws a claim of errata in the federal technical
@@ -86,18 +86,18 @@
Administrative quality control data as an oracle
official-vs-file wedge decomposition, and the pre-registered
model-repair experiments.
- Revision 10 · 2026-10-03 · Max Ghenis, PolicyEngine
+ Revision 10 · 2026-10-04 · Max Ghenis, PolicyEngine