From 62f0bb958e806924f9233c47a80174ca1f098d6e Mon Sep 17 00:00:00 2001 From: Max Ghenis Date: Sun, 4 Oct 2026 14:53:14 -0400 Subject: [PATCH 1/3] Correct the amterr lab's solver description against the R code ANALYSIS.md's method section now states the solver's rules as reconstruct_co_fy2024.R applies them, matching the paper's corrected footnote in #95: - The utility reset takes the commonly reported amount nearest the solver's stepped amount. Candidates are amounts that more than 5 filtered cases of any review status report in the state and calendar year, above the file's UTIL for util_up and below it for util_down. The stepped amount is kept if none lies above; the amount is set to 0 if none lies below. - Rent and utility steps also stop at a zero shelter deduction when lowering, and income-raising steps stop when the uncapped benefit falls below zero. "every step stops" now reads "all steps stop". - Household size takes its direction from the nature code, so 2 of the 7 Colorado moves take the benefit away from RAWBEN. "In 20 the solver moved an input and stopped short" was true of 9 of the 20. In 6 the $3 steps reached within $3 of RAWBEN and the utility reset then moved the input off that match; in 5 the household-size move missed. The results now also flag 34 weakly identified matches: RAWBEN is the maximum allotment and the solver lowered income, so every income at or below break-even reproduces it. audit_claims.py ports the solver's benefit formula and fails unless it reproduces all 283 recorded benefits. It also re-applies the utility reset, which reproduces all 15 final amounts. Each new count lands in claims_audit.json, and test_amterr_lab.py locks the sentences that quote them. New property tests check the two facts the caveat rests on: the benefit never exceeds the maximum allotment, and it falls as income rises. Co-Authored-By: Claude Opus 5.5 --- paper/snapshot/labs/amterr/ANALYSIS.md | 126 ++++++-- paper/snapshot/labs/amterr/audit_claims.py | 321 ++++++++++++++++++- paper/snapshot/labs/amterr/claims_audit.json | 188 +++++++++++ tests/test_amterr_lab.py | 245 +++++++++++++- 4 files changed, 833 insertions(+), 47 deletions(-) diff --git a/paper/snapshot/labs/amterr/ANALYSIS.md b/paper/snapshot/labs/amterr/ANALYSIS.md index 4e23297..add8aa0 100644 --- a/paper/snapshot/labs/amterr/ANALYSIS.md +++ b/paper/snapshot/labs/amterr/ANALYSIS.md @@ -5,7 +5,10 @@ top of the 856/856 benefit-parity result (axiom-oracles#268) and the Giannella/Molin raw-variable reconstruction (github.com/giannella/snap_qc, run for FY2024 in this lab: `reconstruct_co_fy2024.R`). Revised 2026-10-03: layer 3, the software-cause split and the cost-share comparison were -recomputed and corrected, and the weight vintage is now stated (see +recomputed and corrected, and the weight vintage is now stated. Revised +2026-10-04: the solver description was corrected against the R code, and +layer 3 now says what the solver's moves did in each non-reproduced case +and flags the weakly identified matches (see [revision history](#revision-history)). Every figure below that is computed from the QC file is regenerated by `audit_claims.py` into `claims_audit.json`; `README.md` gives the pins and commands to rerun the @@ -92,26 +95,47 @@ caught in our own encoding. The Giannella/Molin solver (FY2024 adaptation, smoothing off) starts from the file's edited inputs (household size, earned and unearned income, rent, utilities and deductions) and moves only the input named by the case's first -finding (ELEMENT1): earned or unearned income, rent, the utility allowance, -medical, dependent care or child-support deduction in $3 steps, or household -size by one person for element 150 (natures 7, 12, 14, 16). Income steps -stop once the recomputed benefit passes RAWBEN (the benefit the agency -issued). Rent, utility and deduction steps also stop within $3 of it. Steps -that lower an input stop at zero; rent and utility steps also stop when the -shelter deduction reaches its cap; every step stops when the benefit reaches -$0 or after 1,000 steps. Rows the solver labels -`util_up` or `util_down` then have the utility allowance set to the nearest -value above (or below) the file's UTIL that more than 5 filtered cases in -that state and calendar year use; a `util_down` row with no such value is -set to 0. When ELEMENT1 is outside those lists the inputs stay as they are -(`correctednotes == "no_change"`). The Axiom -engine, which reproduces all 856 Colorado FSBEN values at zero tolerance on -the file's inputs (axiom-oracles#268), then computes the benefit on the -solver's inputs, and the result is compared with RAWBEN at the file's own $5 -editing tolerance. Below, "moved" means at least one of the eight replayed -inputs differs from the file value it started from. (The solver's -`correctedamount` column misses two utility rows, 202403-40765 and -202404-40803, because it is recorded before the utility snap.) +finding (ELEMENT1). + +- **Household size** moves once, by one person, for element 150: one fewer + for natures 12, 14 or 16 and one more for nature 7. The nature code alone + sets the direction, so the move can take the benefit away from RAWBEN (the + benefit the agency issued). It does in 2 of the 7 Colorado household-size + moves, 202310-40265 and 202402-40618. +- **Earned or unearned income, rent, the utility allowance, and the medical, + dependent care or child-support deduction** move in $3 steps toward + RAWBEN. Income steps stop once the recomputed benefit passes RAWBEN; rent, + utility and deduction steps also stop within $3 of it. Steps that lower an + input stop at zero; rent and utility steps also stop when the shelter + deduction reaches its cap (when raising) or zero (when lowering); + income-raising steps stop when the uncapped benefit falls below zero; and + all steps stop when the benefit reaches $0 or after 1,000 steps. Where + RAWBEN is the maximum allotment, which the benefit cannot pass, + income-lowering steps continue until the income reaches zero or the step + limit. +- **The utility reset** then applies to every row the solver labels + `util_up` (RAWBEN above FSBEN) or `util_down` (RAWBEN below FSBEN). The + candidates are the utility amounts that more than 5 cases report in the + same state and calendar year, counting cases of any review status that + pass the solver's consistency filters (listed under Results). A `util_up` + row takes the candidate above the file's UTIL that lies nearest the + amount its steps reached, and keeps that stepped amount if no candidate + lies above. A `util_down` row takes the candidate below the file's UTIL + nearest its stepped amount, and is set to 0 if none lies below. +- **Any other first finding** leaves the inputs as they are + (`correctednotes == "no_change"`). + +The Axiom engine, which reproduces all 856 Colorado FSBEN values at zero +tolerance on the file's inputs (axiom-oracles#268), then computes the +benefit on the solver's inputs, and the result is compared with RAWBEN at +the file's own $5 editing tolerance. Below, "moved" means at least one of +the eight replayed inputs differs from the file value it started from. (The +solver's `correctedamount` column misses two utility rows, 202403-40765 and +202404-40803, because it is recorded before the utility reset.) +`audit_claims.py` checks this account against the solver's output. Its port +of the solver's benefit formula reproduces all 283 recreated benefits, and +the reset rule above reproduces the final utility amount in all 15 utility +rows. FSBEN is a constructed variable: the final benefit Mathematica's model calculates from the edited inputs (technical documentation: listed as @@ -127,10 +151,22 @@ certified to receive in the sample month (PDF p. 92). - 246 of 283 (86.9% of cases, 72.0% of the $99.1M replayed error dollars) reproduce the issued benefit within $5. In 230 the solver moved an input; in 16 nothing moved and |RAWBEN − FSBEN| ≤ $5 already. -- 37 do not reproduce. In 20 the solver moved an input and stopped short. - In 17 nothing moved: 14 `no_change` rows and 3 rows where the solver's - element applied but it stopped before moving. For those 17 the replayed - input is the file's input, so the engine returns FSBEN. +- 34 of the 230 are weakly identified. In each, RAWBEN is the maximum + allotment and the solver lowered an income. Every value of that income at + or below its break-even point yields the maximum, so the match does not + pin the income down. The steps, which cannot pass RAWBEN, ran on to zero + income in 32 and to the 1,000-step limit in 2. 19 of the 34 are above the + $56 threshold. The solver exports an `at_max` flag to mark such cases (the + uncapped benefit on the replayed inputs is at least the maximum allotment + less $5); it is set for 54 of the 246 matches, these 34 among them. +- 37 do not reproduce. In 20 the solver moved an input: in 9 its $3 steps + stopped short of RAWBEN; in 6 they reached within $3 of it and the + utility reset then moved the input off that match (202310-40297, + 202312-40456, 202312-40513, 202402-40603, 202405-40908 and 202409-41321); + and in 5 the one-person household-size move missed. In 17 nothing moved: + 14 `no_change` rows and 3 rows where the solver's element applied but it + stopped before moving. For those 17 the replayed input is the file's + input, so the engine returns FSBEN. - The solver's own recomputed benefit and the engine agree on the within-$5 classification for all 283 cases (246 reproduced, 37 not). @@ -140,7 +176,8 @@ with an input error on that element; it does not identify which input the agency had wrong, and it does not rule out a computation error that a moved input absorbed. A non-reproduced case does not by itself show a computation error: the solver tries one element in one direction, so multi-element and -household-composition errors also land among the 37. +household-composition errors also land among the 37, as do the 6 cases the +utility reset moved off a match their steps had reached. The replay outcome does not separate the layer-2 computational findings. Of the 26 cases that carry one, 13 reproduce ($3.6M/yr, 3.2% of Colorado error @@ -177,7 +214,9 @@ What the replay shows for these cases: for those cases. It does not show what facts the agency recorded or what the agency's system computed from them. - In the other 3 the solver moved a different finding's input, and the - engine's benefit matches neither RAWBEN nor FSBEN. + engine's benefit matches neither RAWBEN nor FSBEN. In 2 of them, + 202312-40456 and 202312-40513, the utility steps had reached within $3 of + RAWBEN before the reset moved the input off that match. - 7 of the 10 broad-coded findings concern the benefit computation (520 in six cases; 362/52 in one). These 7 computation candidates carry $2.0M/yr, 1.8% of Colorado error dollars. The other 3 are income or eligibility @@ -225,8 +264,10 @@ for codes 20/21 and $1,210M (14.9%) for codes 10/22; neither reproduces. 2 ($0.8M) it moved rent. - 2 do not reproduce ($2.4M/yr): 202312-40513 and 202401-40548. In both the software-coded finding (331/44/17 and 350/44/17) is on an element the - solver never moved; it moved ELEMENT1 (364 and 363). The replay does not - show that the computation logic failed in either case. + solver never moved; it moved ELEMENT1 (364 and 363). In 202312-40513 its + utility steps had reached within $3 of RAWBEN before the reset moved the + input off that match. The replay does not show that the computation logic + failed in either case. - Mass-change (19) findings in these 16 are on RSDI (4 cases), SSI (2) and the medical deduction (1); programming (17) findings are on RSDI, SSI, child support received and the child-support deduction. None is on rent @@ -281,8 +322,9 @@ above the threshold. The 7 candidates are 7 sampled cases. - The engine replay checks the benefit computation. Eligibility-side errors (person included or excluded) enter only through the solver's household size step. -- The solver moves one element in $3 steps; $5 is the comparison tolerance. - Exact-tolerance rates are not meaningful at this layer. +- The solver moves one element: in $3 steps, then a reset for the utility + allowance, or by one person for household size. $5 is the comparison + tolerance. Exact-tolerance rates are not meaningful at this layer. - FY2024 was evaluated through the compile-time COLA overlay at the nominal period, as in the parity run; rulespec-us#759 (the dependency inversion that would remove the overlay) was still open on 2026-10-03. @@ -367,3 +409,25 @@ P.L. 119-21 sec. 10105. Rerun instructions and pins: `README.md`. that FY2025 was still being measured. - Paths in the scripts are configurable; `README.md` gives pins and commands. `phase_a_classification.json` lives one directory up. +- **2026-10-04.** The solver description now matches + `reconstruct_co_fy2024.R`. `audit_claims.py` ports the solver's benefit + formula and re-applies the utility reset, and the new counts below come + from it. Changes: + - Utility reset: the solver takes the candidate nearest the amount its + steps reached. The 2026-10-03 text said the candidate nearest the + file's UTIL; that rule gives the final amount in 8 of the 15 utility + rows. The candidates are counted over filtered cases of any review + status, and a `util_up` row with no candidate above keeps its stepped + amount. + - Stop rules: rent and utility steps also stop at a zero shelter + deduction when lowering, and income-raising steps stop when the + uncapped benefit falls below zero. Where RAWBEN is the maximum + allotment, income-lowering steps run on to zero income or the step + limit. + - Household size: the nature code sets the direction. The move takes the + benefit away from RAWBEN in 2 of the 7 Colorado cases. + - "In 20 the solver moved an input and stopped short" was true of 9 of + the 20. In 6 the steps reached within $3 of RAWBEN and the utility reset + moved the input off that match; in 5 the household-size move missed. + - Added: 34 of the 230 moved matches are weakly identified (RAWBEN at the + maximum allotment, income lowered). diff --git a/paper/snapshot/labs/amterr/audit_claims.py b/paper/snapshot/labs/amterr/audit_claims.py index ef76733..39fab01 100644 --- a/paper/snapshot/labs/amterr/audit_claims.py +++ b/paper/snapshot/labs/amterr/audit_claims.py @@ -8,7 +8,10 @@ Outputs ``claims_audit.json`` (this directory). The audit also regenerates ``native_decomposition.json`` and ``../phase_a_classification.json`` from the -May 2026 posting and records whether they match the committed copies. +May 2026 posting and records whether they match the committed copies. To check +ANALYSIS.md's account of the solver, it ports the solver's benefit formula and +utility reset from ``reconstruct_co_fy2024.R`` and fails unless the port +reproduces every benefit the solver recorded. The two postings differ only in the weight columns (HWGT, FYWGT, HWGT_OLD, FYWGT_OLD); the audit re-verifies that, so every case-level result (which @@ -140,6 +143,42 @@ 366: "rawcsded", 150: "rawusize", } +# The solver's step size, step limit and within-$3 stop (income_shift, +# max_iterations and diff_matches in reconstruct_co_fy2024.R). +SOLVER_STEP = 3 +SOLVER_MAX_STEPS = 1000 +SOLVER_MATCH = 3 +# FY2024 excess shelter deduction cap (48 states and DC; snap_qc +# additional_data/year_data.csv). The solver lifts it for units with an +# elderly or disabled member (FSNELDER + FSNDIS > 0). +SHELTER_CAP_FY2024 = 672 +# FY2024 maximum allotments by household size, 48 states and DC (snap_qc +# additional_data/max_allotments.csv). The solver keeps only rows whose BENMAX +# matches this table, which also bounds the utility reset's candidate pool. +MAX_ALLOTMENT_FY2024 = { + **dict(enumerate((291, 535, 766, 973, 1155, 1386, 1532, 1751), start=1)), + **{size: 1751 + 219 * (size - 8) for size in range(9, 21)}, +} +# The reset's candidates: amounts more than this many filtered cases report. +UTILITY_CANDIDATE_MIN_CASES = 5 +UTILITY_RESET_NOTES = ("util_up", "util_down") +# Labels adjust_income gives a moved income when RAWBEN exceeds the file-side +# uncapped benefit, so its steps lower the income (_error: it reached zero). +INCOME_LOWERING_NOTES = ("earn_down", "earn_error", "unearn_down", "unearn_error") +# The solver's benefit inputs, as written to co_fy2024_reconstruction.csv. +SOLVER_BENEFIT_INPUTS = ( + "rawearn", + "rawunearn", + "rawrent", + "rawutil", + "rawmedded", + "rawdepded", + "rawcsded", + "rawstdded", + "rawhomeless_ded", + "rawbenmax", + "rawminimum_ben", +) FINDING_COLUMNS = tuple( f"{root}{slot}" @@ -155,6 +194,9 @@ "RAWBEN", "FSBEN", "HWGT", + "BENMAX", + "FSNELDER", + "FSNDIS", *REPLAY_INPUTS.values(), *FINDING_COLUMNS, ) @@ -427,9 +469,17 @@ def load_replay() -> pd.DataFrame: case_key(y, h) for y, h in zip(reconstruction["YRMONTH"], reconstruction["HHLDNO"]) ] + solver_columns = [c for c in SOLVER_BENEFIT_INPUTS if c not in REPLAY_INPUTS] keep = reconstruction[ - ["key", "correctedamount", "ELEMENT1", *REPLAY_INPUTS] - ].rename(columns={"ELEMENT1": "solver_element"}) + [ + "key", + "correctedamount", + "ELEMENT1", + *REPLAY_INPUTS, + *solver_columns, + "at_max", + ] + ].rename(columns={"ELEMENT1": "solver_element", "at_max": "solver_at_max"}) merged = replay.merge(keep, on="key", how="left", validate="1:1") if merged["correctedamount"].isna().any(): raise AssertionError("replay rows missing from the reconstruction") @@ -470,7 +520,229 @@ def join_replay(replay: pd.DataFrame, posting: pd.DataFrame) -> pd.DataFrame: joined["solver_moved_input"] = changed.any(axis=1) if (joined["solver_moved_input"] & joined["solver_no_change"]).any(): raise AssertionError("a no_change row has a changed input") - return joined + return classify_solver_moves(joined) + + +def solver_benefit(inputs: pd.DataFrame) -> pd.Series: + """The solver's benefit: calculate_raw_benefits in reconstruct_co_fy2024.R. + + ``inputs`` carries ``SOLVER_BENEFIT_INPUTS`` and ``shelter_cap`` (infinite + for a unit with an elderly or disabled member). The operations and their + order follow the R function, so the floors land on the same doubles. + """ + earn = inputs["rawearn"] + net_before_shelter = (earn + inputs["rawunearn"]) - ( + earn * 0.2 + + inputs["rawdepded"] + + inputs["rawmedded"] + + inputs["rawcsded"] + + inputs["rawstdded"] + ) + half = np.maximum(net_before_shelter * 0.5, 0) + uncapped_shelter = np.floor((inputs["rawrent"] + inputs["rawutil"]) - half) + shelter = np.floor( + np.minimum(np.maximum(uncapped_shelter, 0), inputs["shelter_cap"]) + ) + net = np.floor(net_before_shelter - (shelter + inputs["rawhomeless_ded"])) + uncapped = np.floor(inputs["rawbenmax"] - 0.3 * net) + return np.minimum( + np.maximum(uncapped, inputs["rawminimum_ben"]), inputs["rawbenmax"] + ) + + +def classify_solver_moves(joined: pd.DataFrame) -> pd.DataFrame: + """Recompute what each move did, with the solver's own benefit formula. + + * ``benefit_after_steps``: the benefit on the inputs the $3 steps reached, + before the utility reset (the stepped utility amount is the file's UTIL + plus ``correctedamount``, which the solver records before the reset). + * ``moved_miss_kind``, for a moved input that does not reproduce: + ``household_size`` (the one-person move), ``reset_off_after_step_match`` + (the steps reached within $3 of RAWBEN and the reset moved the input off + that match) or ``steps_stopped_short`` (a stop rule ended the steps on + the starting side of RAWBEN, more than $3 from it). + """ + out = joined.copy() + inputs = out[list(SOLVER_BENEFIT_INPUTS)].astype(float) + elderly_or_disabled = (out["FSNELDER"] + out["FSNDIS"]) > 0 + inputs["shelter_cap"] = np.where(elderly_or_disabled, np.inf, SHELTER_CAP_FY2024) + final = solver_benefit(inputs) + if not final.eq(out["rawben_recreated"].astype(float)).all(): + raise AssertionError("solver_benefit does not reproduce rawben_recreated") + + notes = out["correctednotes"] + utility = notes.isin(UTILITY_RESET_NOTES) + stepped = inputs.assign(rawutil=out["UTIL"] + out["correctedamount"]) + out["benefit_after_steps"] = final.where(~utility, solver_benefit(stepped)) + rawben, fsben = out["RAWBEN"].astype(float), out["FSBEN"].astype(float) + steps_matched = (rawben - out["benefit_after_steps"]).abs() <= SOLVER_MATCH + household = notes.str.startswith("hhsize") + moved_miss = out["solver_moved_input"] & ~out["reproduced"] + kind = pd.Series( + np.select( + [ + moved_miss & household, + moved_miss & ~household & steps_matched, + moved_miss & ~household & ~steps_matched, + ], + [ + "household_size", + "reset_off_after_step_match", + "steps_stopped_short", + ], + default=None, + ), + index=out.index, + dtype=object, + ) + if (kind.eq("reset_off_after_step_match") & ~utility).any(): + raise AssertionError("a stepped non-utility miss ended within $3") + short = kind.eq("steps_stopped_short") + starting_side = np.sign(rawben - out["benefit_after_steps"]) == np.sign( + rawben - fsben + ) + if not starting_side[short].all(): + raise AssertionError("a miss counted as stopped short passed RAWBEN") + out["moved_miss_kind"] = kind + # The household-size move's direction comes from the nature code alone. + out["household_size_move_away_from_rawben"] = household & ( + (final - fsben) * (rawben - fsben) < 0 + ) + out["income_lowered_rawben_at_maximum"] = notes.isin( + INCOME_LOWERING_NOTES + ) & rawben.eq(out["rawbenmax"].astype(float)) + return out + + +def solver_outcomes(joined: pd.DataFrame) -> dict[str, Any]: + """Counts behind ANALYSIS.md's account of what the solver's moves did.""" + reproduced = joined["reproduced"] + keys = joined["key"] + notes = joined["correctednotes"] + + household = notes.str.startswith("hhsize") + away = joined["household_size_move_away_from_rawben"] + + utility = notes.isin(UTILITY_RESET_NOTES) + steps_matched = ( + joined["RAWBEN"].astype(float) - joined["benefit_after_steps"] + ).abs() <= SOLVER_MATCH + reset_off = joined["moved_miss_kind"].eq("reset_off_after_step_match") + + kinds = ("steps_stopped_short", "reset_off_after_step_match", "household_size") + moved_miss = joined["solver_moved_input"] & ~reproduced + + at_maximum = joined["income_lowered_rawben_at_maximum"] + matched = at_maximum & reproduced + lowered = joined["rawearn"].where(notes.str.startswith("earn"), joined["rawunearn"]) + at_zero = matched & lowered.eq(0) + at_limit = matched & ~at_zero + at_limit &= joined["correctedamount"].eq(-SOLVER_STEP * SOLVER_MAX_STEPS) + if (at_zero | at_limit).sum() != matched.sum(): + raise AssertionError( + "an at-maximum income match ended before zero or the limit" + ) + flagged = joined["solver_at_max"].astype(bool) & reproduced + + return { + "solver_benefit_reproduces_rawben_recreated_n": len(joined), + "household_size": { + "n": int(household.sum()), + "reproduced_n": int((household & reproduced).sum()), + "away_from_rawben_keys": sorted(keys[away]), + }, + "utility_reset": { + "n": int(utility.sum()), + "steps_within_3_of_rawben_n": int((utility & steps_matched).sum()), + "reproduced_n": int((utility & reproduced).sum()), + "reset_off_after_step_match_keys": sorted(keys[reset_off]), + }, + "not_reproduced_solver_moved_input": { + "n": int(moved_miss.sum()), + **{ + f"{kind}_keys": sorted(keys[joined["moved_miss_kind"].eq(kind)]) + for kind in kinds + }, + }, + "income_lowered_rawben_at_maximum": { + "n": int(at_maximum.sum()), + "reproduced_n": int(matched.sum()), + "reproduced_above_threshold_n": int( + (matched & joined["AMTERR"].gt(THRESHOLD_FY2024)).sum() + ), + "reproduced_ended_at_zero_income_n": int(at_zero.sum()), + "reproduced_ended_at_step_limit_n": int(at_limit.sum()), + "reproduced_at_max_flag_n": int((matched & flagged).sum()), + "reproduced_keys": sorted(keys[matched]), + }, + "reproduced_at_max_flag": { + "n": int(flagged.sum()), + "by_correctednotes": { + note: int(count) + for note, count in sorted(notes[flagged].value_counts().items()) + }, + }, + } + + +def utility_reset_check(frame: pd.DataFrame, joined: pd.DataFrame) -> dict[str, Any]: + """Re-apply the utility reset that ANALYSIS.md describes. + + The candidate pool follows reconstruct_co_fy2024.R: Colorado cases of any + review status that pass its consistency filters (|RAWBEN - FSBEN| within $5 + of AMTERR, RENT and UTIL present, BENMAX equal to the FY2024 table value), + grouped by calendar year. A candidate is a UTIL amount that more than five + of them report. Each util_up (util_down) row takes the candidate above + (below) the file's UTIL nearest the amount its steps reached, keeps that + amount when none lies above, and gets 0 when none lies below; ties go to + the smaller amount, as R's which.min over the sorted counts does. The + rule the 2026-10-03 text gave instead, the candidate nearest the file's + UTIL, is scored alongside it. + """ + pool = frame[frame["STATE"] == COLORADO_FIPS] + consistent = ((pool["RAWBEN"] - pool["FSBEN"]).abs() - pool["AMTERR"]).abs() <= 5 + pool = pool[consistent & pool["RENT"].notna() & pool["UTIL"].notna()] + table = pool["FSUSIZE"].fillna(0).map(MAX_ALLOTMENT_FY2024) + pool = pool[pool["BENMAX"].eq(table)] + counts = pool.groupby([pool["YRMONTH"] // 100, "UTIL"]).size() + candidates: dict[int, list[float]] = {} + for (year, amount), n in counts.items(): + if n > UTILITY_CANDIDATE_MIN_CASES: + candidates.setdefault(int(year), []).append(float(amount)) + + rows = joined[joined["correctednotes"].isin(UTILITY_RESET_NOTES)] + by_rule, by_file_util = [], [] + for _, row in rows.iterrows(): + pool_year = np.array(sorted(candidates.get(int(row["YRMONTH"]) // 100, []))) + raising = row["correctednotes"] == "util_up" + file_util = float(row["UTIL"]) + side = ( + pool_year[pool_year > file_util] + if raising + else pool_year[pool_year < file_util] + ) + stepped = file_util + float(row["correctedamount"]) + if len(side) == 0: + rule = file_rule = stepped if raising else 0.0 + else: + rule = float(side[np.argmin(np.abs(side - stepped))]) + file_rule = float(side.min() if raising else side.max()) + final = float(row["rawutil"]) + by_rule.append(rule == final) + by_file_util.append(file_rule == final) + keys = list(rows["key"]) + return { + "candidate_pool_n": len(pool), + "candidates_by_calendar_year": { + str(year): amounts for year, amounts in sorted(candidates.items()) + }, + "rows_n": len(rows), + "rule_reproduces_final_amount_n": int(sum(by_rule)), + "nearest_to_file_util_reproduces_final_amount_n": int(sum(by_file_util)), + "nearest_to_file_util_misses_keys": sorted( + k for k, hit in zip(keys, by_file_util) if not hit + ), + } def case_findings(joined: pd.DataFrame, mask: pd.Series) -> list[dict[str, Any]]: @@ -489,6 +761,7 @@ def case_findings(joined: pd.DataFrame, mask: pd.Series) -> list[dict[str, Any]] "reproduced": bool(row["reproduced"]), "correctednotes": row["correctednotes"], "solver_moved_input": bool(row["solver_moved_input"]), + "moved_miss_kind": row["moved_miss_kind"], "solver_element": int(row["solver_element"]), "findings": [list(f) for f in found], "broad_coded_findings": [list(f) for f in broad], @@ -591,6 +864,7 @@ def case_level(joined: pd.DataFrame) -> dict[str, Any]: snapped &= ~joined["correctednotes"].str.startswith("hhsize") return { "replay": replay_partition(joined, joined["solver_moved_input"]), + "solver_outcomes": solver_outcomes(joined), "replay_inputs_unchanged_keys": unchanged, "moved_with_zero_correctedamount_keys": sorted(joined.loc[snapped, "key"]), "broad_coded_misses": { @@ -878,6 +1152,17 @@ def weighted_results(frame: pd.DataFrame, replay: pd.DataFrame) -> dict[str, Any # artifact +def posting_case_level(frame: pd.DataFrame, replay: pd.DataFrame) -> dict[str, Any]: + """Every case-level block; build() checks it is the same under each posting.""" + joined = join_replay(replay, frame) + return { + **case_level(joined), + "utility_reset_rule": utility_reset_check(frame, joined), + "layer2_computational_findings": computational_findings(frame), + "reconstruction": reconstruction_rows(), + } + + def build(postings: Iterable[str] = tuple(POSTINGS)) -> dict[str, Any]: postings = tuple(postings) frames = {label: load_posting(label) for label in postings} @@ -913,6 +1198,22 @@ def build(postings: Iterable[str] = tuple(POSTINGS)) -> dict[str, Any]: + ", ".join(REPLAY_INPUTS.values()) + "; missing -> 0)" ), + "benefit_after_steps": ( + "the solver's benefit formula (calculate_raw_benefits, ported as " + "solver_benefit and checked against rawben_recreated for every " + "row) on the inputs its $3 steps reached; for util_up and " + "util_down rows the utility amount is UTIL + correctedamount, " + "before the reset to a commonly reported amount" + ), + "moved_miss_kind": ( + "for a moved input that does not reproduce: household_size; " + f"reset_off_after_step_match (benefit_after_steps within " + f"${SOLVER_MATCH} of RAWBEN); otherwise steps_stopped_short" + ), + "income_lowered_rawben_at_maximum": ( + f"correctednotes in {list(INCOME_LOWERING_NOTES)} and RAWBEN " + "equal to the solver's maximum allotment (rawbenmax)" + ), "computation_candidate": ( "not reproduced, and a finding carrying a code in " f"{sorted(BROAD_CODES)} is computational under the layer-2 rule" @@ -931,11 +1232,7 @@ def build(postings: Iterable[str] = tuple(POSTINGS)) -> dict[str, Any]: phase_a_classification(legacy), PHASE_A_PATH ), }, - "case_level": { - **case_level(join_replay(replay, legacy)), - "layer2_computational_findings": computational_findings(legacy), - "reconstruction": reconstruction_rows(), - }, + "case_level": posting_case_level(legacy, replay), "by_posting": { label: weighted_results(frame, replay) for label, frame in frames.items() }, @@ -943,11 +1240,7 @@ def build(postings: Iterable[str] = tuple(POSTINGS)) -> dict[str, Any]: if {"may2026", "aug2026"} <= set(postings): artifact["posting_diff"] = posting_diff() for label in postings: - level = { - **case_level(join_replay(replay, frames[label])), - "layer2_computational_findings": computational_findings(frames[label]), - "reconstruction": artifact["case_level"]["reconstruction"], - } + level = posting_case_level(frames[label], replay) if level != artifact["case_level"]: raise AssertionError(f"case-level results differ under {label}") return artifact diff --git a/paper/snapshot/labs/amterr/claims_audit.json b/paper/snapshot/labs/amterr/claims_audit.json index 7be629d..5a95700 100644 --- a/paper/snapshot/labs/amterr/claims_audit.json +++ b/paper/snapshot/labs/amterr/claims_audit.json @@ -41,6 +41,9 @@ "code_class": "a case counts if any of AGENCY1-AGENCY9 is in the class (classes overlap)", "reproduced": "abs(engine_on_original - RAWBEN) <= 5", "solver_moved_input": "any replayed input (rawusize, rawearn, rawunearn, rawrent, rawutil, rawmedded, rawdepded, rawcsded) differs from its file field (FSUSIZE, FSEARN, FSUNEARN, RENT, UTIL, FSMEDDED, FSDEPDED, FSCSDED; missing -> 0)", + "benefit_after_steps": "the solver's benefit formula (calculate_raw_benefits, ported as solver_benefit and checked against rawben_recreated for every row) on the inputs its $3 steps reached; for util_up and util_down rows the utility amount is UTIL + correctedamount, before the reset to a commonly reported amount", + "moved_miss_kind": "for a moved input that does not reproduce: household_size; reset_off_after_step_match (benefit_after_steps within $3 of RAWBEN); otherwise steps_stopped_short", + "income_lowered_rawben_at_maximum": "correctednotes in ['earn_down', 'earn_error', 'unearn_down', 'unearn_error'] and RAWBEN equal to the solver's maximum allotment (rawbenmax)", "computation_candidate": "not reproduced, and a finding carrying a code in [10, 17, 19, 20, 21, 22] is computational under the layer-2 rule", "points_of_official_fy2024_rate": "share of file error dollars x 9.97; an approximation, since the official rate is regression-adjusted and counts only errors above the $56 threshold" }, @@ -87,6 +90,116 @@ "util_up": 5 } }, + "solver_outcomes": { + "solver_benefit_reproduces_rawben_recreated_n": 283, + "household_size": { + "n": 7, + "reproduced_n": 2, + "away_from_rawben_keys": [ + "202310-40265", + "202402-40618" + ] + }, + "utility_reset": { + "n": 15, + "steps_within_3_of_rawben_n": 14, + "reproduced_n": 8, + "reset_off_after_step_match_keys": [ + "202310-40297", + "202312-40456", + "202312-40513", + "202402-40603", + "202405-40908", + "202409-41321" + ] + }, + "not_reproduced_solver_moved_input": { + "n": 20, + "steps_stopped_short_keys": [ + "202310-40304", + "202311-40369", + "202401-40546", + "202401-40548", + "202405-40957", + "202406-41065", + "202408-41268", + "202409-41302", + "202409-41362" + ], + "reset_off_after_step_match_keys": [ + "202310-40297", + "202312-40456", + "202312-40513", + "202402-40603", + "202405-40908", + "202409-41321" + ], + "household_size_keys": [ + "202310-40265", + "202311-40407", + "202402-40618", + "202403-40738", + "202406-40998" + ] + }, + "income_lowered_rawben_at_maximum": { + "n": 35, + "reproduced_n": 34, + "reproduced_above_threshold_n": 19, + "reproduced_ended_at_zero_income_n": 32, + "reproduced_ended_at_step_limit_n": 2, + "reproduced_at_max_flag_n": 34, + "reproduced_keys": [ + "202310-40260", + "202310-40281", + "202310-40303", + "202311-40362", + "202311-40366", + "202311-40389", + "202312-40457", + "202312-40501", + "202312-40506", + "202312-40511", + "202401-40557", + "202401-40578", + "202401-40584", + "202401-40599", + "202402-40672", + "202403-40712", + "202403-40744", + "202404-40805", + "202404-40869", + "202405-40879", + "202405-40919", + "202405-40947", + "202405-40960", + "202406-41001", + "202406-41032", + "202406-41053", + "202406-41059", + "202406-41071", + "202407-41150", + "202407-41165", + "202407-41175", + "202408-41240", + "202408-41282", + "202409-41365" + ] + }, + "reproduced_at_max_flag": { + "n": 54, + "by_correctednotes": { + "earn_down": 19, + "hhsize_up": 1, + "med_up": 1, + "rent_down": 1, + "rent_up": 14, + "unearn_down": 15, + "util_down": 1, + "util_up": 2 + } + } + }, "replay_inputs_unchanged_keys": [ "202310-40314", "202310-40336", @@ -163,6 +276,7 @@ "reproduced": false, "correctednotes": "no_change", "solver_moved_input": false, + "moved_miss_kind": null, "solver_element": 520, "findings": [ [ @@ -190,6 +304,7 @@ "reproduced": false, "correctednotes": "util_up", "solver_moved_input": true, + "moved_miss_kind": "reset_off_after_step_match", "solver_element": 364, "findings": [ [ @@ -227,6 +342,7 @@ "reproduced": false, "correctednotes": "no_change", "solver_moved_input": false, + "moved_miss_kind": null, "solver_element": 520, "findings": [ [ @@ -254,6 +370,7 @@ "reproduced": false, "correctednotes": "util_up", "solver_moved_input": true, + "moved_miss_kind": "reset_off_after_step_match", "solver_element": 364, "findings": [ [ @@ -291,6 +408,7 @@ "reproduced": false, "correctednotes": "no_change", "solver_moved_input": false, + "moved_miss_kind": null, "solver_element": 520, "findings": [ [ @@ -318,6 +436,7 @@ "reproduced": false, "correctednotes": "rent_down", "solver_moved_input": true, + "moved_miss_kind": "steps_stopped_short", "solver_element": 363, "findings": [ [ @@ -355,6 +474,7 @@ "reproduced": false, "correctednotes": "no_change", "solver_moved_input": false, + "moved_miss_kind": null, "solver_element": 362, "findings": [ [ @@ -382,6 +502,7 @@ "reproduced": false, "correctednotes": "no_change", "solver_moved_input": false, + "moved_miss_kind": null, "solver_element": 161, "findings": [ [ @@ -414,6 +535,7 @@ "reproduced": false, "correctednotes": "no_change", "solver_moved_input": false, + "moved_miss_kind": null, "solver_element": 520, "findings": [ [ @@ -446,6 +568,7 @@ "reproduced": false, "correctednotes": "no_change", "solver_moved_input": false, + "moved_miss_kind": null, "solver_element": 520, "findings": [ [ @@ -925,6 +1048,7 @@ "reproduced": false, "correctednotes": "hhsize_up", "solver_moved_input": true, + "moved_miss_kind": "household_size", "solver_element": 150, "findings": [ [ @@ -951,6 +1075,7 @@ "reproduced": false, "correctednotes": "util_up", "solver_moved_input": true, + "moved_miss_kind": "reset_off_after_step_match", "solver_element": 364, "findings": [ [ @@ -972,6 +1097,7 @@ "reproduced": false, "correctednotes": "earn_down", "solver_moved_input": true, + "moved_miss_kind": "steps_stopped_short", "solver_element": 311, "findings": [ [ @@ -993,6 +1119,7 @@ "reproduced": false, "correctednotes": "no_change", "solver_moved_input": false, + "moved_miss_kind": null, "solver_element": 520, "findings": [ [ @@ -1020,6 +1147,7 @@ "reproduced": false, "correctednotes": "earn_error", "solver_moved_input": true, + "moved_miss_kind": "steps_stopped_short", "solver_element": 312, "findings": [ [ @@ -1056,6 +1184,7 @@ "reproduced": false, "correctednotes": "hhsize_up", "solver_moved_input": true, + "moved_miss_kind": "household_size", "solver_element": 150, "findings": [ [ @@ -1092,6 +1221,7 @@ "reproduced": false, "correctednotes": "no_change", "solver_moved_input": false, + "moved_miss_kind": null, "solver_element": 520, "findings": [ [ @@ -1113,6 +1243,7 @@ "reproduced": false, "correctednotes": "util_up", "solver_moved_input": true, + "moved_miss_kind": "reset_off_after_step_match", "solver_element": 364, "findings": [ [ @@ -1150,6 +1281,7 @@ "reproduced": false, "correctednotes": "no_change", "solver_moved_input": false, + "moved_miss_kind": null, "solver_element": 520, "findings": [ [ @@ -1177,6 +1309,7 @@ "reproduced": false, "correctednotes": "no_change", "solver_moved_input": false, + "moved_miss_kind": null, "solver_element": 362, "findings": [ [ @@ -1198,6 +1331,7 @@ "reproduced": false, "correctednotes": "util_up", "solver_moved_input": true, + "moved_miss_kind": "reset_off_after_step_match", "solver_element": 364, "findings": [ [ @@ -1235,6 +1369,7 @@ "reproduced": false, "correctednotes": "no_change", "solver_moved_input": false, + "moved_miss_kind": null, "solver_element": 520, "findings": [ [ @@ -1262,6 +1397,7 @@ "reproduced": false, "correctednotes": "util_down", "solver_moved_input": true, + "moved_miss_kind": "steps_stopped_short", "solver_element": 364, "findings": [ [ @@ -1283,6 +1419,7 @@ "reproduced": false, "correctednotes": "rent_down", "solver_moved_input": true, + "moved_miss_kind": "steps_stopped_short", "solver_element": 363, "findings": [ [ @@ -1320,6 +1457,7 @@ "reproduced": false, "correctednotes": "rent_down", "solver_moved_input": false, + "moved_miss_kind": null, "solver_element": 363, "findings": [ [ @@ -1346,6 +1484,7 @@ "reproduced": false, "correctednotes": "no_change", "solver_moved_input": false, + "moved_miss_kind": null, "solver_element": 520, "findings": [ [ @@ -1367,6 +1506,7 @@ "reproduced": false, "correctednotes": "no_change", "solver_moved_input": false, + "moved_miss_kind": null, "solver_element": 150, "findings": [ [ @@ -1393,6 +1533,7 @@ "reproduced": false, "correctednotes": "util_up", "solver_moved_input": true, + "moved_miss_kind": "reset_off_after_step_match", "solver_element": 364, "findings": [ [ @@ -1414,6 +1555,7 @@ "reproduced": false, "correctednotes": "hhsize_down", "solver_moved_input": true, + "moved_miss_kind": "household_size", "solver_element": 150, "findings": [ [ @@ -1435,6 +1577,7 @@ "reproduced": false, "correctednotes": "no_change", "solver_moved_input": false, + "moved_miss_kind": null, "solver_element": 362, "findings": [ [ @@ -1456,6 +1599,7 @@ "reproduced": false, "correctednotes": "hhsize_up", "solver_moved_input": true, + "moved_miss_kind": "household_size", "solver_element": 150, "findings": [ [ @@ -1487,6 +1631,7 @@ "reproduced": false, "correctednotes": "no_change", "solver_moved_input": false, + "moved_miss_kind": null, "solver_element": 362, "findings": [ [ @@ -1508,6 +1653,7 @@ "reproduced": false, "correctednotes": "util_down", "solver_moved_input": true, + "moved_miss_kind": "reset_off_after_step_match", "solver_element": 364, "findings": [ [ @@ -1529,6 +1675,7 @@ "reproduced": false, "correctednotes": "no_change", "solver_moved_input": false, + "moved_miss_kind": null, "solver_element": 362, "findings": [ [ @@ -1556,6 +1703,7 @@ "reproduced": false, "correctednotes": "unearn_error", "solver_moved_input": false, + "moved_miss_kind": null, "solver_element": 350, "findings": [ [ @@ -1582,6 +1730,7 @@ "reproduced": false, "correctednotes": "earn_down", "solver_moved_input": true, + "moved_miss_kind": "steps_stopped_short", "solver_element": 311, "findings": [ [ @@ -1603,6 +1752,7 @@ "reproduced": false, "correctednotes": "no_change", "solver_moved_input": false, + "moved_miss_kind": null, "solver_element": 161, "findings": [ [ @@ -1635,6 +1785,7 @@ "reproduced": false, "correctednotes": "hhsize_down", "solver_moved_input": true, + "moved_miss_kind": "household_size", "solver_element": 150, "findings": [ [ @@ -1656,6 +1807,7 @@ "reproduced": false, "correctednotes": "earn_error", "solver_moved_input": true, + "moved_miss_kind": "steps_stopped_short", "solver_element": 312, "findings": [ [ @@ -1682,6 +1834,7 @@ "reproduced": false, "correctednotes": "rent_error", "solver_moved_input": false, + "moved_miss_kind": null, "solver_element": 363, "findings": [ [ @@ -1708,6 +1861,7 @@ "reproduced": false, "correctednotes": "no_change", "solver_moved_input": false, + "moved_miss_kind": null, "solver_element": 150, "findings": [ [ @@ -1739,6 +1893,7 @@ "reproduced": false, "correctednotes": "no_change", "solver_moved_input": false, + "moved_miss_kind": null, "solver_element": 520, "findings": [ [ @@ -1771,6 +1926,7 @@ "reproduced": false, "correctednotes": "rent_error", "solver_moved_input": true, + "moved_miss_kind": "steps_stopped_short", "solver_element": 363, "findings": [ [ @@ -1797,6 +1953,7 @@ "reproduced": false, "correctednotes": "no_change", "solver_moved_input": false, + "moved_miss_kind": null, "solver_element": 520, "findings": [ [ @@ -1824,6 +1981,7 @@ "reproduced": false, "correctednotes": "earn_error", "solver_moved_input": true, + "moved_miss_kind": "steps_stopped_short", "solver_element": 314, "findings": [ [ @@ -1850,6 +2008,7 @@ "reproduced": false, "correctednotes": "util_up", "solver_moved_input": true, + "moved_miss_kind": "reset_off_after_step_match", "solver_element": 364, "findings": [ [ @@ -1876,6 +2035,7 @@ "reproduced": false, "correctednotes": "earn_down", "solver_moved_input": true, + "moved_miss_kind": "steps_stopped_short", "solver_element": 311, "findings": [ [ @@ -1898,6 +2058,34 @@ "broad_coded_finding_is_computational": false } ], + "utility_reset_rule": { + "candidate_pool_n": 803, + "candidates_by_calendar_year": { + "2023": [ + 0.0, + 91.0, + 560.0 + ], + "2024": [ + 0.0, + 91.0, + 356.0, + 560.0 + ] + }, + "rows_n": 15, + "rule_reproduces_final_amount_n": 15, + "nearest_to_file_util_reproduces_final_amount_n": 8, + "nearest_to_file_util_misses_keys": [ + "202310-40297", + "202402-40603", + "202403-40709", + "202403-40716", + "202407-41181", + "202409-41321", + "202409-41353" + ] + }, "layer2_computational_findings": { "cases": 26, "findings": 29, diff --git a/tests/test_amterr_lab.py b/tests/test_amterr_lab.py index b128459..4879f4c 100644 --- a/tests/test_amterr_lab.py +++ b/tests/test_amterr_lab.py @@ -193,6 +193,77 @@ def test_case_counts_do_not_depend_on_posting(): assert other["layer2"][name]["n"] == metric["n"] +def test_moved_misses_partition_by_what_the_solver_did(): + outcomes = CASES["solver_outcomes"] + moved = outcomes["not_reproduced_solver_moved_input"] + kinds = ("steps_stopped_short", "reset_off_after_step_match", "household_size") + lists = [set(moved[f"{kind}_keys"]) for kind in kinds] + assert sum(len(keys) for keys in lists) == len(set().union(*lists)) + assert moved["n"] == len(set().union(*lists)) + assert moved["n"] == CASES["replay"]["not_reproduced_solver_moved_input_n"] + rows = {row["key"]: row for row in CASES["not_reproduced_cases"]} + for kind, keys in zip(kinds, lists): + for key in keys: + assert rows[key]["solver_moved_input"], key + assert rows[key]["moved_miss_kind"] == kind, key + unmoved = [row for row in rows.values() if not row["solver_moved_input"]] + assert all(row["moved_miss_kind"] is None for row in unmoved) + for key in moved["reset_off_after_step_match_keys"]: + assert rows[key]["correctednotes"] in audit.UTILITY_RESET_NOTES + for key in moved["household_size_keys"]: + assert rows[key]["correctednotes"].startswith("hhsize") + assert ( + moved["reset_off_after_step_match_keys"] + == outcomes["utility_reset"]["reset_off_after_step_match_keys"] + ) + + +def test_solver_outcome_counts_are_consistent(): + outcomes = CASES["solver_outcomes"] + assert ( + outcomes["solver_benefit_reproduces_rawben_recreated_n"] + == (CASES["replay"]["n"]) + ) + household = outcomes["household_size"] + assert len(household["away_from_rawben_keys"]) <= household["n"] + assert household["reproduced_n"] <= household["n"] + missed = {row["key"] for row in CASES["not_reproduced_cases"]} + # A move away from RAWBEN cannot reproduce it. + assert set(household["away_from_rawben_keys"]) <= missed + + utility = outcomes["utility_reset"] + reset_off = utility["reset_off_after_step_match_keys"] + assert len(reset_off) <= utility["steps_within_3_of_rawben_n"] <= utility["n"] + assert utility["reproduced_n"] <= utility["n"] + assert set(reset_off) <= missed + + rule = CASES["utility_reset_rule"] + assert rule["rows_n"] == utility["n"] + assert rule["rule_reproduces_final_amount_n"] == rule["rows_n"] + assert rule["rows_n"] - rule[ + "nearest_to_file_util_reproduces_final_amount_n" + ] == len(rule["nearest_to_file_util_misses_keys"]) + for amounts in rule["candidates_by_calendar_year"].values(): + assert amounts == sorted(set(amounts)) + + at_maximum = outcomes["income_lowered_rawben_at_maximum"] + matched = set(at_maximum["reproduced_keys"]) + assert len(matched) == at_maximum["reproduced_n"] <= at_maximum["n"] + assert matched.isdisjoint(missed) + assert matched.isdisjoint(CASES["replay_inputs_unchanged_keys"]) + assert ( + at_maximum["reproduced_ended_at_zero_income_n"] + + at_maximum["reproduced_ended_at_step_limit_n"] + == at_maximum["reproduced_n"] + ) + assert at_maximum["reproduced_above_threshold_n"] <= at_maximum["reproduced_n"] + # The solver's own at_max flag marks every one of them, and more. + flagged = outcomes["reproduced_at_max_flag"] + assert at_maximum["reproduced_at_max_flag_n"] == at_maximum["reproduced_n"] + assert at_maximum["reproduced_n"] <= flagged["n"] <= CASES["replay"]["reproduced_n"] + assert sum(flagged["by_correctednotes"].values()) == flagged["n"] + + def _millions(value: float, digits: int) -> str: return f"${value / 1e6:.{digits}f}M" @@ -201,6 +272,11 @@ def _percent(value: float, digits: int = 1) -> str: return f"{100 * value:.{digits}f}%" +def _keys(keys: list[str]) -> str: + """Case keys as ANALYSIS.md lists them: "a, b and c".""" + return keys[0] if len(keys) == 1 else f"{', '.join(keys[:-1])} and {keys[-1]}" + + def test_analysis_quotes_the_audit(): """Whole sentences and table rows, so a figure cannot match elsewhere.""" text = " ".join((LAB / "ANALYSIS.md").read_text(encoding="utf-8").split()) @@ -209,6 +285,22 @@ def test_analysis_quotes_the_audit(): counts = CASES["replay"] replay = may["replay"] misses = replay["broad_coded_misses"] + outcomes = CASES["solver_outcomes"] + moved = outcomes["not_reproduced_solver_moved_input"] + reset_off = moved["reset_off_after_step_match_keys"] + household = outcomes["household_size"] + at_maximum = outcomes["income_lowered_rawben_at_maximum"] + rule = CASES["utility_reset_rule"] + broad_reset_off = [ + row["key"] + for row in CASES["broad_coded_misses"]["cases"] + if row["moved_miss_kind"] == "reset_off_after_step_match" + ] + software_reset_off = [ + row["key"] + for row in CASES["software_coded_replayed"]["cases"] + if not row["reproduced"] and row["key"] in reset_off + ] quoted = [ ( f"{counts['reproduced_n']} of {counts['n']} " @@ -221,8 +313,75 @@ def test_analysis_quotes_the_audit(): f"in {counts['reproduced_without_move_n']} nothing moved" ), ( - f"In {counts['not_reproduced_solver_moved_input_n']} the solver moved an " - "input and stopped short" + f"In {moved['n']} the solver moved an input: in " + f"{len(moved['steps_stopped_short_keys'])} its $3 steps stopped short of " + f"RAWBEN; in {len(reset_off)} they reached within $3 of it and the utility " + f"reset then moved the input off that match ({_keys(reset_off)}); and in " + f"{len(moved['household_size_keys'])} the one-person household-size move " + "missed." + ), + ( + f"{at_maximum['reproduced_n']} of the " + f"{counts['reproduced_solver_moved_input_n']} are weakly identified." + ), + ( + "The steps, which cannot pass RAWBEN, ran on to zero income in " + f"{at_maximum['reproduced_ended_at_zero_income_n']} and to the " + f"1,000-step limit in {at_maximum['reproduced_ended_at_step_limit_n']}. " + f"{at_maximum['reproduced_above_threshold_n']} of the " + f"{at_maximum['reproduced_n']} are above the ${audit.THRESHOLD_FY2024} " + "threshold." + ), + ( + f"it is set for {outcomes['reproduced_at_max_flag']['n']} of the " + f"{counts['reproduced_n']} matches, these {at_maximum['reproduced_n']} " + "among them." + ), + ( + f"It does in {len(household['away_from_rawben_keys'])} of the " + f"{household['n']} Colorado household-size moves, " + f"{_keys(household['away_from_rawben_keys'])}." + ), + ( + "Its port of the solver's benefit formula reproduces all " + f"{outcomes['solver_benefit_reproduces_rawben_recreated_n']} recreated " + "benefits, and the reset rule above reproduces the final utility amount " + f"in all {rule['rule_reproduces_final_amount_n']} utility rows." + ), + ( + f"as do the {len(reset_off)} cases the utility reset moved off a match " + "their steps had reached." + ), + ( + f"In {len(broad_reset_off)} of them, {_keys(broad_reset_off)}, the utility " + "steps had reached within $3 of RAWBEN before the reset moved the input " + "off that match." + ), + ( + f"In {_keys(software_reset_off)} its utility steps had reached within $3 " + "of RAWBEN before the reset moved the input off that match." + ), + ( + "that rule gives the final amount in " + f"{rule['nearest_to_file_util_reproduces_final_amount_n']} of the " + f"{rule['rows_n']} utility rows." + ), + ( + "The move takes the benefit away from RAWBEN in " + f"{len(household['away_from_rawben_keys'])} of the {household['n']} " + "Colorado cases." + ), + ( + f'"In {moved["n"]} the solver moved an input and stopped short" was true ' + f"of {len(moved['steps_stopped_short_keys'])} of the {moved['n']}. In " + f"{len(reset_off)} the steps reached within $3 of RAWBEN and the utility " + "reset moved the input off that match; in " + f"{len(moved['household_size_keys'])} the household-size move missed." + ), + ( + f"Added: {at_maximum['reproduced_n']} of the " + f"{counts['reproduced_solver_moved_input_n']} moved matches are weakly " + "identified" ), ( f"They carry {_millions(misses['dollars'], 2)}/yr: " @@ -470,3 +629,85 @@ def test_any_presence_is_monotone_in_the_code_set(frame): broad = audit.class_metric(errors, audit.any_code(errors, audit.BROAD_CODES), total) assert strict["n"] <= broad["n"] assert strict["dollars"] <= broad["dollars"] + 0.01 + + +def test_retired_solver_wording_is_gone(): + """The 2026-10-03 solver wording survives only in the revision history.""" + text = " ".join((LAB / "ANALYSIS.md").read_text(encoding="utf-8").split()) + body = text.split("## Revision history")[0] + for phrase in ( + "moved an input and stopped short", + "every step stops", + "value above (or below) the file's UTIL", + "utility snap", + ): + assert phrase not in body, phrase + assert "value above (or below) the file's UTIL" not in text + + +# --------------------------------------------------------------------------- +# properties of the ported solver benefit, which the weak-identification +# caveat relies on: the benefit never exceeds the maximum allotment, falls as +# income rises and rises with shelter costs and deductions. + + +@st.composite +def solver_units(draw): + """One unit's solver inputs; incomes and costs in whole dollars.""" + dollars = st.integers(min_value=0, max_value=6000) + size = draw(st.integers(min_value=1, max_value=20)) + return { + "rawearn": float(draw(dollars)), + "rawunearn": float(draw(dollars)), + "rawrent": float(draw(dollars)), + "rawutil": float(draw(st.sampled_from((0, 91, 356, 560)))), + "rawmedded": float(draw(st.integers(min_value=0, max_value=1500))), + "rawdepded": float(draw(st.integers(min_value=0, max_value=1500))), + "rawcsded": float(draw(st.integers(min_value=0, max_value=1500))), + "rawstdded": float(draw(st.sampled_from((198, 208, 244, 279)))), + "rawhomeless_ded": float(draw(st.sampled_from((0, 179)))), + "rawbenmax": float(audit.MAX_ALLOTMENT_FY2024[size]), + "rawminimum_ben": 23.0 if size < 3 else 0.0, + "shelter_cap": draw( + st.sampled_from((float(audit.SHELTER_CAP_FY2024), float("inf"))) + ), + } + + +def _benefits(unit: dict, column: str, values: list[float]) -> np.ndarray: + frame = pd.DataFrame([{**unit, column: value} for value in values]) + return audit.solver_benefit(frame).to_numpy() + + +@settings(max_examples=300, deadline=None) +@given( + solver_units(), + st.sampled_from(("rawearn", "rawunearn")), + st.lists(st.integers(min_value=0, max_value=9000), min_size=2, max_size=30), +) +def test_benefit_falls_as_income_rises_and_stays_at_the_maximum_below_break_even( + unit, column, incomes +): + values = sorted(float(v) for v in incomes) + benefits = _benefits(unit, column, values) + assert np.all(np.diff(benefits) <= 0) + assert np.all(benefits <= unit["rawbenmax"]) + assert np.all(benefits >= min(unit["rawminimum_ben"], unit["rawbenmax"])) + # Every income at or below one that yields the maximum also yields it. + at_maximum = benefits == unit["rawbenmax"] + if at_maximum.any(): + last = np.flatnonzero(at_maximum).max() + assert at_maximum[: last + 1].all() + + +@settings(max_examples=300, deadline=None) +@given( + solver_units(), + st.sampled_from(("rawrent", "rawutil", "rawmedded", "rawdepded", "rawcsded")), + st.lists(st.integers(min_value=0, max_value=6000), min_size=2, max_size=30), +) +def test_benefit_rises_with_shelter_costs_and_deductions(unit, column, amounts): + values = sorted(float(v) for v in amounts) + benefits = _benefits(unit, column, values) + assert np.all(np.diff(benefits) >= 0) + assert np.all(benefits <= unit["rawbenmax"]) From 00de6bf94cc952c6c412a166c0e6b8e07fb9b7f8 Mon Sep 17 00:00:00 2001 From: Max Ghenis Date: Sun, 4 Oct 2026 15:28:59 -0400 Subject: [PATCH 2/3] Count every weakly identified replay match and tighten the solver notes An adversarial pass (8 independent agents, each with its own port of the solver) confirmed every count in the previous commit. It also found that the weak-identification caveat understated its own argument. - The caveat now covers every match on a flat stretch of the benefit formula, where the match bounds the moved input on one side only: 60 of the 230 moved matches, 31 of them above the threshold. - 53 are at or within $5 of the maximum allotment. That is the 34 income cuts plus 19 where every further rent, utility or medical amount also reproduces RAWBEN. - 4 are at the $23 minimum benefit. - 3 are rent increases at the shelter-deduction cap. - "Break-even point" usually means the income at which the benefit phases out. It now reads "the point where net income reaches zero": the solver pays the maximum exactly when net income is zero or less, which a new property test checks. - "ran on to zero income" now reads "ran that income down to $0", since other income can remain. - The stop rules now say "at or above the maximum allotment" and add the mirror rule at the minimum benefit. - The correctedamount note now covers household-size moves (always 0) and all 15 utility rows (always the pre-reset change). - The audit check is stated as covering two parts of the account, which is what it does. - The revision note quotes the old reset rule correctly. - New notes: - 202405-40908's reset leaves the benefit farther from RAWBEN than FSBEN. - 2 of the 10 computational misses are reset-off cases. audit_claims.py classifies each match's flat stretch by pushing the moved input without limit in the solver's direction (flat_stretch). The new counts are in claims_audit.json, and test_amterr_lab.py locks them. Co-Authored-By: Claude Opus 5.5 --- paper/snapshot/labs/amterr/ANALYSIS.md | 82 +++++++---- paper/snapshot/labs/amterr/audit_claims.py | 112 +++++++++++++- paper/snapshot/labs/amterr/claims_audit.json | 97 +++++++++++++ tests/test_amterr_lab.py | 145 ++++++++++++++++--- 4 files changed, 384 insertions(+), 52 deletions(-) diff --git a/paper/snapshot/labs/amterr/ANALYSIS.md b/paper/snapshot/labs/amterr/ANALYSIS.md index add8aa0..968e0e1 100644 --- a/paper/snapshot/labs/amterr/ANALYSIS.md +++ b/paper/snapshot/labs/amterr/ANALYSIS.md @@ -110,9 +110,11 @@ finding (ELEMENT1). deduction reaches its cap (when raising) or zero (when lowering); income-raising steps stop when the uncapped benefit falls below zero; and all steps stop when the benefit reaches $0 or after 1,000 steps. Where - RAWBEN is the maximum allotment, which the benefit cannot pass, - income-lowering steps continue until the income reaches zero or the step - limit. + RAWBEN is at or above the maximum allotment, which the benefit cannot + pass, income-lowering steps continue until the income reaches zero or the + step limit. Where it is at or below a one- or two-person unit's minimum + benefit, which the benefit cannot fall below, income-raising steps + continue until the uncapped benefit falls below zero or the step limit. - **The utility reset** then applies to every row the solver labels `util_up` (RAWBEN above FSBEN) or `util_down` (RAWBEN below FSBEN). The candidates are the utility amounts that more than 5 cases report in the @@ -121,7 +123,11 @@ finding (ELEMENT1). row takes the candidate above the file's UTIL that lies nearest the amount its steps reached, and keeps that stepped amount if no candidate lies above. A `util_down` row takes the candidate below the file's UTIL - nearest its stepped amount, and is set to 0 if none lies below. + nearest its stepped amount, and is set to 0 if none lies below. The reset + can also take the benefit away from RAWBEN: in 202405-40908 it leaves the + benefit farther from RAWBEN than FSBEN. With the 2 household-size moves + above, these are the only 3 moved cases that end farther from RAWBEN than + FSBEN. - **Any other first finding** leaves the inputs as they are (`correctednotes == "no_change"`). @@ -129,13 +135,15 @@ The Axiom engine, which reproduces all 856 Colorado FSBEN values at zero tolerance on the file's inputs (axiom-oracles#268), then computes the benefit on the solver's inputs, and the result is compared with RAWBEN at the file's own $5 editing tolerance. Below, "moved" means at least one of -the eight replayed inputs differs from the file value it started from. (The -solver's `correctedamount` column misses two utility rows, 202403-40765 and -202404-40803, because it is recorded before the utility reset.) -`audit_claims.py` checks this account against the solver's output. Its port -of the solver's benefit formula reproduces all 283 recreated benefits, and -the reset rule above reproduces the final utility amount in all 15 utility -rows. +the eight replayed inputs differs from the file value it started from. The +solver's `correctedamount` column is not the final change: household-size +moves never write it, so it is 0 in all 7, and it is recorded before the +utility reset, so it differs from the final change in all 15 utility rows +and is 0 in two of them, 202403-40765 and 202404-40803, which only the reset +moved. `audit_claims.py` checks two parts of this account against the +solver's output: its port of the solver's benefit formula reproduces all 283 +recreated benefits, and the reset rule above, applied to each row's stepped +amount, reproduces all 15 final utility amounts. FSBEN is a constructed variable: the final benefit Mathematica's model calculates from the edited inputs (technical documentation: listed as @@ -151,14 +159,26 @@ certified to receive in the sample month (PDF p. 92). - 246 of 283 (86.9% of cases, 72.0% of the $99.1M replayed error dollars) reproduce the issued benefit within $5. In 230 the solver moved an input; in 16 nothing moved and |RAWBEN − FSBEN| ≤ $5 already. -- 34 of the 230 are weakly identified. In each, RAWBEN is the maximum - allotment and the solver lowered an income. Every value of that income at - or below its break-even point yields the maximum, so the match does not - pin the income down. The steps, which cannot pass RAWBEN, ran on to zero - income in 32 and to the 1,000-step limit in 2. 19 of the 34 are above the - $56 threshold. The solver exports an `at_max` flag to mark such cases (the - uncapped benefit on the replayed inputs is at least the maximum allotment - less $5); it is set for 54 of the 246 matches, these 34 among them. +- 60 of the 230 are weakly identified: the issued benefit sits on a flat + stretch of the benefit formula in the moved input, so the match bounds + that input on one side only. 31 of the 60 are above the $56 threshold. + - 53 are at or within $5 of the maximum allotment, which caps the + benefit. In 34 of them RAWBEN is the maximum and the solver lowered an + income. The solver pays the maximum whenever net income is zero or + less, and net income never falls as income rises, so every value of + that income up to the point where net income reaches zero yields the + maximum. The steps, which cannot pass RAWBEN, ran that income down to + $0 in 32 and to the 1,000-step limit in 2. In the other 19 the solver + moved rent (15), the utility allowance (3) or the medical deduction (1), + and every further amount in the same direction also reproduces RAWBEN. + - 4 are at the $23 minimum benefit of a one- or two-person unit. + - 3 are rent increases that reach the shelter-deduction cap, beyond which + rent no longer changes the benefit. + + The solver exports an `at_max` flag to mark capped cases (the uncapped + benefit on the replayed inputs is at least the maximum allotment less + $5). It is set for 54 of the 246 matches: 52 of the 53, one household-size + match and one match where nothing moved. - 37 do not reproduce. In 20 the solver moved an input: in 9 its $3 steps stopped short of RAWBEN; in 6 they reached within $3 of it and the utility reset then moved the input off that match (202310-40297, @@ -181,7 +201,9 @@ utility reset moved off a match their steps had reached. The replay outcome does not separate the layer-2 computational findings. Of the 26 cases that carry one, 13 reproduce ($3.6M/yr, 3.2% of Colorado error -dollars), 10 do not ($3.5M, 3.1%) and 3 were not replayed. Among the 13 are +dollars), 10 do not ($3.5M, 3.1%) and 3 were not replayed. 2 of the 10, +202310-40297 and 202312-40456, are among the 6 cases the utility reset moved +off a match. Among the 13 are 202404-40794 and 202404-40823, which reproduce after the solver moved the child-support deduction and whose finding is 366/56/17 (computer programming error). @@ -414,20 +436,24 @@ P.L. 119-21 sec. 10105. Rerun instructions and pins: `README.md`. formula and re-applies the utility reset, and the new counts below come from it. Changes: - Utility reset: the solver takes the candidate nearest the amount its - steps reached. The 2026-10-03 text said the candidate nearest the - file's UTIL; that rule gives the final amount in 8 of the 15 utility - rows. The candidates are counted over filtered cases of any review + steps reached. The 2026-10-03 text said the candidate above (or below) + the file's UTIL nearest that UTIL; that rule gives the final amount in + 8 of the 15 utility rows. The candidates are counted over filtered cases of any review status, and a `util_up` row with no candidate above keeps its stepped amount. - Stop rules: rent and utility steps also stop at a zero shelter deduction when lowering, and income-raising steps stop when the - uncapped benefit falls below zero. Where RAWBEN is the maximum - allotment, income-lowering steps run on to zero income or the step - limit. + uncapped benefit falls below zero. Where RAWBEN is at or above the + maximum allotment, income-lowering steps run on to zero income or the + step limit, and at or below the minimum benefit, income-raising steps + run on until the uncapped benefit falls below zero. - Household size: the nature code sets the direction. The move takes the benefit away from RAWBEN in 2 of the 7 Colorado cases. - "In 20 the solver moved an input and stopped short" was true of 9 of the 20. In 6 the steps reached within $3 of RAWBEN and the utility reset moved the input off that match; in 5 the household-size move missed. - - Added: 34 of the 230 moved matches are weakly identified (RAWBEN at the - maximum allotment, income lowered). + - The `correctedamount` note now covers household-size moves and all 15 + utility rows. + - Added: 60 of the 230 moved matches are weakly identified, 34 of them + because RAWBEN is the maximum allotment and the solver lowered an + income. diff --git a/paper/snapshot/labs/amterr/audit_claims.py b/paper/snapshot/labs/amterr/audit_claims.py index 39fab01..29a4e95 100644 --- a/paper/snapshot/labs/amterr/audit_claims.py +++ b/paper/snapshot/labs/amterr/audit_claims.py @@ -162,6 +162,19 @@ # The reset's candidates: amounts more than this many filtered cases report. UTILITY_CANDIDATE_MIN_CASES = 5 UTILITY_RESET_NOTES = ("util_up", "util_down") +# The input each stepped label moves (adjust_* calls in the R script). +SOLVER_NOTE_INPUTS = { + "earn": "rawearn", + "unearn": "rawunearn", + "rent": "rawrent", + "util": "rawutil", + "med": "rawmedded", + "dep": "rawdepded", + "cs": "rawcsded", +} +# Stands in for "without limit" when an input is pushed past the solver's +# value: far beyond every Colorado income, cost and the shelter cap. +SOLVER_PUSH = 100_000 # Labels adjust_income gives a moved income when RAWBEN exceeds the file-side # uncapped benefit, so its steps lower the income (_error: it reached zero). INCOME_LOWERING_NOTES = ("earn_down", "earn_error", "unearn_down", "unearn_error") @@ -523,12 +536,15 @@ def join_replay(replay: pd.DataFrame, posting: pd.DataFrame) -> pd.DataFrame: return classify_solver_moves(joined) -def solver_benefit(inputs: pd.DataFrame) -> pd.Series: - """The solver's benefit: calculate_raw_benefits in reconstruct_co_fy2024.R. +def solver_terms(inputs: pd.DataFrame) -> pd.DataFrame: + """The solver's benefit and its parts: calculate_raw_benefits in + reconstruct_co_fy2024.R. ``inputs`` carries ``SOLVER_BENEFIT_INPUTS`` and ``shelter_cap`` (infinite for a unit with an elderly or disabled member). The operations and their order follow the R function, so the floors land on the same doubles. + Returns the shelter deduction, net income (which may be negative) and the + benefit. """ earn = inputs["rawearn"] net_before_shelter = (earn + inputs["rawunearn"]) - ( @@ -545,9 +561,15 @@ def solver_benefit(inputs: pd.DataFrame) -> pd.Series: ) net = np.floor(net_before_shelter - (shelter + inputs["rawhomeless_ded"])) uncapped = np.floor(inputs["rawbenmax"] - 0.3 * net) - return np.minimum( + benefit = np.minimum( np.maximum(uncapped, inputs["rawminimum_ben"]), inputs["rawbenmax"] ) + return pd.DataFrame({"shelter": shelter, "net": net, "benefit": benefit}) + + +def solver_benefit(inputs: pd.DataFrame) -> pd.Series: + """The solver's benefit (see ``solver_terms``).""" + return solver_terms(inputs)["benefit"] def classify_solver_moves(joined: pd.DataFrame) -> pd.DataFrame: @@ -611,9 +633,57 @@ def classify_solver_moves(joined: pd.DataFrame) -> pd.DataFrame: out["income_lowered_rawben_at_maximum"] = notes.isin( INCOME_LOWERING_NOTES ) & rawben.eq(out["rawbenmax"].astype(float)) + out["flat_stretch"] = flat_stretch(out, inputs, final) + lowered_at_maximum = out["income_lowered_rawben_at_maximum"] & out["reproduced"] + if not out.loc[lowered_at_maximum, "flat_stretch"].eq("maximum_allotment").all(): + raise AssertionError("an at-maximum income match is not on the cap") return out +def flat_stretch( + joined: pd.DataFrame, inputs: pd.DataFrame, final: pd.Series +) -> pd.Series: + """Where a reproduced match sits on a flat stretch of the benefit formula. + + The moved input is pushed without limit in the solver's direction (to $0 + when the solver lowered it, up by ``SOLVER_PUSH`` when it raised it). The + benefit is monotone in each input, so if the pushed benefit still lies + within $5 of RAWBEN, every amount between the solver's and the limit + reproduces RAWBEN too, and the match bounds the input on one side only. + The stretch is named by what makes the formula flat there: the maximum + allotment, the minimum benefit, or (for a raised rent or utility amount) + the shelter-deduction cap. A lowered input whose band merely reaches $0 + is left unnamed. Household-size moves are not stepped and are left out. + """ + notes = joined["correctednotes"] + column = notes.str.split("_").str[0].map(SOLVER_NOTE_INPUTS) + candidate = joined["reproduced"] & joined["solver_moved_input"] & column.notna() + pushed = inputs.copy() + raised = pd.Series(False, index=joined.index) + for name in set(column.dropna()): + rows = candidate & column.eq(name) + start = joined[REPLAY_INPUTS[name]].fillna(0).astype(float) + up = rows & (inputs[name] > start) + raised |= up + pushed.loc[rows, name] = np.where( + up[rows], inputs.loc[rows, name] + SOLVER_PUSH, 0.0 + ) + terms = solver_terms(pushed) + rawben = joined["RAWBEN"].astype(float) + holds = candidate & (rawben - terms["benefit"]).abs().le(REPLAY_TOLERANCE) + shelter_input = column.isin(["rawrent", "rawutil"]) + named = np.select( + [ + holds & terms["benefit"].eq(inputs["rawbenmax"]), + holds & terms["benefit"].eq(inputs["rawminimum_ben"]), + holds & raised & shelter_input & terms["shelter"].eq(inputs["shelter_cap"]), + ], + ["maximum_allotment", "minimum_benefit", "shelter_cap"], + default=None, + ) + return pd.Series(named, index=joined.index, dtype=object) + + def solver_outcomes(joined: pd.DataFrame) -> dict[str, Any]: """Counts behind ANALYSIS.md's account of what the solver's moves did.""" reproduced = joined["reproduced"] @@ -643,6 +713,17 @@ def solver_outcomes(joined: pd.DataFrame) -> dict[str, Any]: "an at-maximum income match ended before zero or the limit" ) flagged = joined["solver_at_max"].astype(bool) & reproduced + above = joined["AMTERR"].gt(THRESHOLD_FY2024) + stretch = joined["flat_stretch"] + weak = stretch.notna() + at_cap = stretch.eq("maximum_allotment") + rawben = joined["RAWBEN"].astype(float) + moved = joined["solver_moved_input"] + farther = moved & ( + (rawben - joined["rawben_recreated"].astype(float)).abs() + > (rawben - joined["FSBEN"].astype(float)).abs() + ) + final_change = joined["rawutil"].astype(float) - joined["UTIL"].astype(float) return { "solver_benefit_reproduces_rawben_recreated_n": len(joined), @@ -651,8 +732,15 @@ def solver_outcomes(joined: pd.DataFrame) -> dict[str, Any]: "reproduced_n": int((household & reproduced).sum()), "away_from_rawben_keys": sorted(keys[away]), }, + "moved_benefit_farther_from_rawben_than_fsben_keys": sorted(keys[farther]), + "household_size_correctedamount_zero_n": int( + (household & joined["correctedamount"].eq(0)).sum() + ), "utility_reset": { "n": int(utility.sum()), + "correctedamount_differs_from_final_change_n": int( + (utility & joined["correctedamount"].ne(final_change)).sum() + ), "steps_within_3_of_rawben_n": int((utility & steps_matched).sum()), "reproduced_n": int((utility & reproduced).sum()), "reset_off_after_step_match_keys": sorted(keys[reset_off]), @@ -675,8 +763,26 @@ def solver_outcomes(joined: pd.DataFrame) -> dict[str, Any]: "reproduced_at_max_flag_n": int((matched & flagged).sum()), "reproduced_keys": sorted(keys[matched]), }, + "weakly_identified": { + "n": int(weak.sum()), + "above_threshold_n": int((weak & above).sum()), + "by_flat_stretch": { + name: int(count) + for name, count in sorted(stretch[weak].value_counts().items()) + }, + "maximum_allotment_by_correctednotes": { + note: int(count) + for note, count in sorted(notes[at_cap].value_counts().items()) + }, + "maximum_allotment_at_max_flag_n": int((at_cap & flagged).sum()), + "keys": { + name: sorted(keys[stretch.eq(name)]) + for name in sorted(set(stretch.dropna())) + }, + }, "reproduced_at_max_flag": { "n": int(flagged.sum()), + "outside_weakly_identified_keys": sorted(keys[flagged & ~weak]), "by_correctednotes": { note: int(count) for note, count in sorted(notes[flagged].value_counts().items()) diff --git a/paper/snapshot/labs/amterr/claims_audit.json b/paper/snapshot/labs/amterr/claims_audit.json index 5a95700..42346c1 100644 --- a/paper/snapshot/labs/amterr/claims_audit.json +++ b/paper/snapshot/labs/amterr/claims_audit.json @@ -100,8 +100,15 @@ "202402-40618" ] }, + "moved_benefit_farther_from_rawben_than_fsben_keys": [ + "202310-40265", + "202402-40618", + "202405-40908" + ], + "household_size_correctedamount_zero_n": 7, "utility_reset": { "n": 15, + "correctedamount_differs_from_final_change_n": 15, "steps_within_3_of_rawben_n": 14, "reproduced_n": 8, "reset_off_after_step_match_keys": [ @@ -186,8 +193,98 @@ "202409-41365" ] }, + "weakly_identified": { + "n": 60, + "above_threshold_n": 31, + "by_flat_stretch": { + "maximum_allotment": 53, + "minimum_benefit": 4, + "shelter_cap": 3 + }, + "maximum_allotment_by_correctednotes": { + "earn_down": 19, + "med_up": 1, + "rent_up": 15, + "unearn_down": 15, + "util_down": 1, + "util_up": 2 + }, + "maximum_allotment_at_max_flag_n": 52, + "keys": { + "maximum_allotment": [ + "202310-40260", + "202310-40281", + "202310-40303", + "202311-40355", + "202311-40362", + "202311-40366", + "202311-40389", + "202312-40448", + "202312-40457", + "202312-40501", + "202312-40506", + "202312-40511", + "202401-40557", + "202401-40561", + "202401-40564", + "202401-40571", + "202401-40578", + "202401-40584", + "202401-40599", + "202402-40635", + "202402-40672", + "202403-40712", + "202403-40744", + "202403-40765", + "202404-40788", + "202404-40803", + "202404-40805", + "202404-40869", + "202405-40879", + "202405-40884", + "202405-40905", + "202405-40919", + "202405-40947", + "202405-40950", + "202405-40960", + "202406-41001", + "202406-41003", + "202406-41005", + "202406-41032", + "202406-41053", + "202406-41059", + "202406-41071", + "202406-41073", + "202407-41112", + "202407-41150", + "202407-41165", + "202407-41175", + "202407-41181", + "202408-41198", + "202408-41240", + "202408-41282", + "202409-41365", + "202409-41373" + ], + "minimum_benefit": [ + "202310-40276", + "202403-40717", + "202403-40756", + "202408-41258" + ], + "shelter_cap": [ + "202402-40643", + "202403-40772", + "202404-40790" + ] + } + }, "reproduced_at_max_flag": { "n": 54, + "outside_weakly_identified_keys": [ + "202311-40381", + "202407-41141" + ], "by_correctednotes": { "earn_down": 19, "hhsize_up": 1, diff --git a/tests/test_amterr_lab.py b/tests/test_amterr_lab.py index 4879f4c..502e701 100644 --- a/tests/test_amterr_lab.py +++ b/tests/test_amterr_lab.py @@ -263,6 +263,40 @@ def test_solver_outcome_counts_are_consistent(): assert at_maximum["reproduced_n"] <= flagged["n"] <= CASES["replay"]["reproduced_n"] assert sum(flagged["by_correctednotes"].values()) == flagged["n"] + weak = outcomes["weakly_identified"] + by_stretch = weak["by_flat_stretch"] + assert sum(by_stretch.values()) == weak["n"] <= CASES["replay"]["n"] + assert weak["above_threshold_n"] <= weak["n"] + keys = {name: set(listed) for name, listed in weak["keys"].items()} + assert {name: len(listed) for name, listed in keys.items()} == by_stretch + assert sum(len(listed) for listed in keys.values()) == len( + set().union(*keys.values()) + ) + weak_keys = set().union(*keys.values()) + # Weakly identified matches are moved matches. + assert weak_keys.isdisjoint(missed) + assert weak_keys.isdisjoint(CASES["replay_inputs_unchanged_keys"]) + assert len(weak_keys) <= CASES["replay"]["reproduced_solver_moved_input_n"] + # The 34 income-lowered matches at the maximum sit on the cap stretch. + assert matched <= keys["maximum_allotment"] + assert ( + sum(weak["maximum_allotment_by_correctednotes"].values()) + == by_stretch["maximum_allotment"] + ) + # The solver's at_max flag overlaps only the cap stretch; the rest of it + # is listed. + assert ( + weak["maximum_allotment_at_max_flag_n"] + + len(flagged["outside_weakly_identified_keys"]) + == (flagged["n"]) + ) + assert set(flagged["outside_weakly_identified_keys"]).isdisjoint(weak_keys) + + farther = set(outcomes["moved_benefit_farther_from_rawben_than_fsben_keys"]) + assert set(household["away_from_rawben_keys"]) <= farther <= missed + assert outcomes["household_size_correctedamount_zero_n"] == household["n"] + assert utility["correctedamount_differs_from_final_change_n"] <= utility["n"] + def _millions(value: float, digits: int) -> str: return f"${value / 1e6:.{digits}f}M" @@ -291,6 +325,21 @@ def test_analysis_quotes_the_audit(): household = outcomes["household_size"] at_maximum = outcomes["income_lowered_rawben_at_maximum"] rule = CASES["utility_reset_rule"] + utility = outcomes["utility_reset"] + weak = outcomes["weakly_identified"] + stretches = weak["by_flat_stretch"] + at_cap = weak["maximum_allotment_by_correctednotes"] + cap_rent = sum(v for k, v in at_cap.items() if k.startswith("rent")) + cap_util = sum(v for k, v in at_cap.items() if k.startswith("util")) + cap_med = sum(v for k, v in at_cap.items() if k.startswith("med")) + cap_other = stretches["maximum_allotment"] - at_maximum["reproduced_n"] + assert cap_rent + cap_util + cap_med == cap_other + flagged = outcomes["reproduced_at_max_flag"] + farther = outcomes["moved_benefit_farther_from_rawben_than_fsben_keys"] + reset_away = sorted(set(farther) - set(household["away_from_rawben_keys"])) + computational_reset_off = sorted( + set(CASES["not_reproduced_with_computational_finding_keys"]) & set(reset_off) + ) broad_reset_off = [ row["key"] for row in CASES["broad_coded_misses"]["cases"] @@ -321,21 +370,38 @@ def test_analysis_quotes_the_audit(): "missed." ), ( - f"{at_maximum['reproduced_n']} of the " - f"{counts['reproduced_solver_moved_input_n']} are weakly identified." + f"{weak['n']} of the {counts['reproduced_solver_moved_input_n']} are " + "weakly identified: the issued benefit sits on a flat stretch of the " + "benefit formula in the moved input, so the match bounds that input on " + f"one side only. {weak['above_threshold_n']} of the {weak['n']} are above " + f"the ${audit.THRESHOLD_FY2024} threshold." ), ( - "The steps, which cannot pass RAWBEN, ran on to zero income in " + f"{stretches['maximum_allotment']} are at or within $5 of the maximum " + f"allotment, which caps the benefit. In {at_maximum['reproduced_n']} of them RAWBEN is the " + "maximum and the solver lowered an income." + ), + ( + "The steps, which cannot pass RAWBEN, ran that income down to $0 in " f"{at_maximum['reproduced_ended_at_zero_income_n']} and to the " f"1,000-step limit in {at_maximum['reproduced_ended_at_step_limit_n']}. " - f"{at_maximum['reproduced_above_threshold_n']} of the " - f"{at_maximum['reproduced_n']} are above the ${audit.THRESHOLD_FY2024} " - "threshold." + f"In the other {cap_other} the solver moved rent ({cap_rent}), the " + f"utility allowance ({cap_util}) or the medical deduction ({cap_med}), " + "and every further amount in the same direction also reproduces RAWBEN." + ), + ( + f"{stretches['minimum_benefit']} are at the $23 minimum benefit of a " + "one- or two-person unit." + ), + ( + f"{stretches['shelter_cap']} are rent increases that reach the " + "shelter-deduction cap" ), ( - f"it is set for {outcomes['reproduced_at_max_flag']['n']} of the " - f"{counts['reproduced_n']} matches, these {at_maximum['reproduced_n']} " - "among them." + f"It is set for {flagged['n']} of the {counts['reproduced_n']} matches: " + f"{weak['maximum_allotment_at_max_flag_n']} of the " + f"{stretches['maximum_allotment']}, one household-size match and one " + "match where nothing moved." ), ( f"It does in {len(household['away_from_rawben_keys'])} of the " @@ -343,10 +409,33 @@ def test_analysis_quotes_the_audit(): f"{_keys(household['away_from_rawben_keys'])}." ), ( - "Its port of the solver's benefit formula reproduces all " + "checks two parts of this account against the solver's output: its port " + "of the solver's benefit formula reproduces all " f"{outcomes['solver_benefit_reproduces_rawben_recreated_n']} recreated " - "benefits, and the reset rule above reproduces the final utility amount " - f"in all {rule['rule_reproduces_final_amount_n']} utility rows." + "benefits, and the reset rule above, applied to each row's stepped " + f"amount, reproduces all {rule['rule_reproduces_final_amount_n']} final " + "utility amounts." + ), + ( + "household-size moves never write it, so it is 0 in all " + f"{outcomes['household_size_correctedamount_zero_n']}, and it is recorded " + "before the utility reset, so it differs from the final change in all " + f"{utility['correctedamount_differs_from_final_change_n']} utility rows " + "and is 0 in two of them, " + f"{_keys(CASES['moved_with_zero_correctedamount_keys'])}, which only the " + "reset moved." + ), + ( + f"in {_keys(reset_away)} it leaves the benefit farther from RAWBEN than " + f"FSBEN. With the {len(household['away_from_rawben_keys'])} " + "household-size moves above, these are the only " + f"{len(farther)} moved cases that end farther from RAWBEN than FSBEN." + ), + ( + f"{len(computational_reset_off)} of the " + f"{replay['not_reproduced_with_computational_finding']['n']}, " + f"{_keys(computational_reset_off)}, are among the {len(reset_off)} cases " + "the utility reset moved off a match." ), ( f"as do the {len(reset_off)} cases the utility reset moved off a match " @@ -362,7 +451,8 @@ def test_analysis_quotes_the_audit(): "of RAWBEN before the reset moved the input off that match." ), ( - "that rule gives the final amount in " + "The 2026-10-03 text said the candidate above (or below) the file's UTIL " + "nearest that UTIL; that rule gives the final amount in " f"{rule['nearest_to_file_util_reproduces_final_amount_n']} of the " f"{rule['rows_n']} utility rows." ), @@ -379,9 +469,9 @@ def test_analysis_quotes_the_audit(): f"{len(moved['household_size_keys'])} the household-size move missed." ), ( - f"Added: {at_maximum['reproduced_n']} of the " - f"{counts['reproduced_solver_moved_input_n']} moved matches are weakly " - "identified" + f"Added: {weak['n']} of the {counts['reproduced_solver_moved_input_n']} " + f"moved matches are weakly identified, {at_maximum['reproduced_n']} of " + "them because RAWBEN is the maximum allotment" ), ( f"They carry {_millions(misses['dollars'], 2)}/yr: " @@ -640,6 +730,8 @@ def test_retired_solver_wording_is_gone(): "every step stops", "value above (or below) the file's UTIL", "utility snap", + "break-even", + "ran on to zero income", ): assert phrase not in body, phrase assert "value above (or below) the file's UTIL" not in text @@ -647,8 +739,9 @@ def test_retired_solver_wording_is_gone(): # --------------------------------------------------------------------------- # properties of the ported solver benefit, which the weak-identification -# caveat relies on: the benefit never exceeds the maximum allotment, falls as -# income rises and rises with shelter costs and deductions. +# caveat relies on: the benefit stays between the minimum benefit and the +# maximum allotment, falls as income rises and rises with shelter costs and +# deductions, so each bound is reached on a flat stretch. @st.composite @@ -679,21 +772,31 @@ def _benefits(unit: dict, column: str, values: list[float]) -> np.ndarray: return audit.solver_benefit(frame).to_numpy() +def _terms(unit: dict, column: str, values: list[float]) -> pd.DataFrame: + frame = pd.DataFrame([{**unit, column: value} for value in values]) + return audit.solver_terms(frame) + + @settings(max_examples=300, deadline=None) @given( solver_units(), st.sampled_from(("rawearn", "rawunearn")), st.lists(st.integers(min_value=0, max_value=9000), min_size=2, max_size=30), ) -def test_benefit_falls_as_income_rises_and_stays_at_the_maximum_below_break_even( +def test_benefit_is_the_maximum_exactly_while_net_income_is_not_positive( unit, column, incomes ): + """The mechanism behind the 34: net income never falls as income rises, + and the benefit is the maximum allotment exactly when net income is zero + or less, so the incomes that yield the maximum run from $0 up to one point.""" values = sorted(float(v) for v in incomes) - benefits = _benefits(unit, column, values) + terms = _terms(unit, column, values) + benefits, net = terms["benefit"].to_numpy(), terms["net"].to_numpy() + assert np.all(np.diff(net) >= 0) assert np.all(np.diff(benefits) <= 0) assert np.all(benefits <= unit["rawbenmax"]) assert np.all(benefits >= min(unit["rawminimum_ben"], unit["rawbenmax"])) - # Every income at or below one that yields the maximum also yields it. + assert np.array_equal(benefits == unit["rawbenmax"], net <= 0) at_maximum = benefits == unit["rawbenmax"] if at_maximum.any(): last = np.flatnonzero(at_maximum).max() From 561f34da4eae8de079b2a6e962ce47367b576f00 Mon Sep 17 00:00:00 2001 From: Max Ghenis Date: Sun, 4 Oct 2026 18:41:58 -0400 Subject: [PATCH 3/3] Address review r1 on the weak-identification caveat The independent Opus review of 00de6bf asked for two required fixes and three recommended ones. All are verified against the solver's formula. - The 3 shelter-cap matches don't reach the cap: the solver stopped $11 to $14 short of $672 on the within-$3 rule. The text now says they "stop within $14 of the shelter-deduction cap", and that past the cap the benefit stays within $5 of RAWBEN. The audit records the gap. - 2 of the 60 (202403-40765 and 202404-40803) reproduce at every utility amount. "On one side only" now reads "on one side at most", and these two are named. The audit pushes each input the opposite way too (opposite_push_holds, bounded_on_neither_side_keys). - The caveat now says "by the solver's formula": the engine was not run at the pushed amounts. The "farther from RAWBEN" list now fails the build unless the engine agrees with the solver's benefit. - Pushes that stay within $5 without reaching a flat stretch are counted and named (unnamed_push_holds_keys: 4 already at $0, 3 whose band runs down to $0). ANALYSIS.md says they are not counted. - More of each lock is derived from the audit: the at_max composition (one household-size match, one unmoved), the $23 minimum, the step limit and the correctedamount "2 of them". Co-Authored-By: Claude Opus 5.5 --- paper/snapshot/labs/amterr/ANALYSIS.md | 25 +++- paper/snapshot/labs/amterr/audit_claims.py | 127 ++++++++++++++----- paper/snapshot/labs/amterr/claims_audit.json | 21 +++ tests/test_amterr_lab.py | 57 +++++++-- 4 files changed, 183 insertions(+), 47 deletions(-) diff --git a/paper/snapshot/labs/amterr/ANALYSIS.md b/paper/snapshot/labs/amterr/ANALYSIS.md index 968e0e1..ea67588 100644 --- a/paper/snapshot/labs/amterr/ANALYSIS.md +++ b/paper/snapshot/labs/amterr/ANALYSIS.md @@ -139,7 +139,7 @@ the eight replayed inputs differs from the file value it started from. The solver's `correctedamount` column is not the final change: household-size moves never write it, so it is 0 in all 7, and it is recorded before the utility reset, so it differs from the final change in all 15 utility rows -and is 0 in two of them, 202403-40765 and 202404-40803, which only the reset +and is 0 in 2 of them, 202403-40765 and 202404-40803, which only the reset moved. `audit_claims.py` checks two parts of this account against the solver's output: its port of the solver's benefit formula reproduces all 283 recreated benefits, and the reset rule above, applied to each row's stepped @@ -159,9 +159,14 @@ certified to receive in the sample month (PDF p. 92). - 246 of 283 (86.9% of cases, 72.0% of the $99.1M replayed error dollars) reproduce the issued benefit within $5. In 230 the solver moved an input; in 16 nothing moved and |RAWBEN − FSBEN| ≤ $5 already. -- 60 of the 230 are weakly identified: the issued benefit sits on a flat - stretch of the benefit formula in the moved input, so the match bounds - that input on one side only. 31 of the 60 are above the $56 threshold. +- 60 of the 230 are weakly identified. In each, the issued benefit lies + within $5 of a stretch where the solver's benefit formula is flat in the + moved input: the maximum allotment, the minimum benefit, or the benefit + past the shelter-deduction cap. By the solver's formula, every amount of + the input on that stretch reproduces the issued benefit, so the match + bounds the input on one side at most. 2 of the 60, 202403-40765 and + 202404-40803, reproduce at every utility amount and bound it on neither + side. 31 of the 60 are above the $56 threshold. - 53 are at or within $5 of the maximum allotment, which caps the benefit. In 34 of them RAWBEN is the maximum and the solver lowered an income. The solver pays the maximum whenever net income is zero or @@ -170,10 +175,16 @@ certified to receive in the sample month (PDF p. 92). maximum. The steps, which cannot pass RAWBEN, ran that income down to $0 in 32 and to the 1,000-step limit in 2. In the other 19 the solver moved rent (15), the utility allowance (3) or the medical deduction (1), - and every further amount in the same direction also reproduces RAWBEN. + and every further amount in the same direction keeps the benefit within + $5 of RAWBEN. - 4 are at the $23 minimum benefit of a one- or two-person unit. - - 3 are rent increases that reach the shelter-deduction cap, beyond which - rent no longer changes the benefit. + - 3 are rent increases that stop within $14 of the shelter-deduction cap. + Past the cap, rent no longer changes the benefit, which stays within $5 + of RAWBEN. + + Not counted: 7 matches whose lowered input reaches no flat stretch, + because the steps took it to $0 (4) or its band of matching amounts runs + down to $0 (3). The solver exports an `at_max` flag to mark capped cases (the uncapped benefit on the replayed inputs is at least the maximum allotment less diff --git a/paper/snapshot/labs/amterr/audit_claims.py b/paper/snapshot/labs/amterr/audit_claims.py index 29a4e95..d73ef58 100644 --- a/paper/snapshot/labs/amterr/audit_claims.py +++ b/paper/snapshot/labs/amterr/audit_claims.py @@ -633,7 +633,8 @@ def classify_solver_moves(joined: pd.DataFrame) -> pd.DataFrame: out["income_lowered_rawben_at_maximum"] = notes.isin( INCOME_LOWERING_NOTES ) & rawben.eq(out["rawbenmax"].astype(float)) - out["flat_stretch"] = flat_stretch(out, inputs, final) + stretch = flat_stretch(out, inputs, final) + out[list(stretch)] = stretch lowered_at_maximum = out["income_lowered_rawben_at_maximum"] & out["reproduced"] if not out.loc[lowered_at_maximum, "flat_stretch"].eq("maximum_allotment").all(): raise AssertionError("an at-maximum income match is not on the cap") @@ -642,46 +643,94 @@ def classify_solver_moves(joined: pd.DataFrame) -> pd.DataFrame: def flat_stretch( joined: pd.DataFrame, inputs: pd.DataFrame, final: pd.Series -) -> pd.Series: - """Where a reproduced match sits on a flat stretch of the benefit formula. - - The moved input is pushed without limit in the solver's direction (to $0 - when the solver lowered it, up by ``SOLVER_PUSH`` when it raised it). The - benefit is monotone in each input, so if the pushed benefit still lies - within $5 of RAWBEN, every amount between the solver's and the limit - reproduces RAWBEN too, and the match bounds the input on one side only. - The stretch is named by what makes the formula flat there: the maximum - allotment, the minimum benefit, or (for a raised rent or utility amount) - the shelter-deduction cap. A lowered input whose band merely reaches $0 - is left unnamed. Household-size moves are not stepped and are left out. +) -> pd.DataFrame: + """Where a reproduced match sits near a flat stretch of the benefit formula. + + All benefits here come from the solver's formula (``solver_terms``); the + engine was not run at the pushed points. The moved input is pushed in the + solver's direction: to $0 when the solver lowered it, up by + ``SOLVER_PUSH`` when it raised it. The benefit is monotone in each input, + so if the pushed benefit still lies within $5 of RAWBEN, every amount + between the solver's and the pushed one does too. Where the formula is + already flat at the pushed amount (the benefit at the maximum allotment + or the minimum benefit, or the shelter deduction at its cap), every + amount beyond it gives the same benefit, so the match bounds the input on + one side at most. The stretch is named by what makes the formula flat: + ``maximum_allotment``, ``minimum_benefit`` or, for a raised rent or + utility amount, ``shelter_cap``. + + Columns returned: + + * ``flat_stretch``: that name, or None; + * ``opposite_push_holds``: for a named row, pushing the other way (up by + ``SOLVER_PUSH`` or to $0) also stays within $5 of RAWBEN, so the match + bounds the input on neither side; + * ``unnamed_push_holds``: the push stays within $5 but no flat stretch is + reached: ``input_already_zero`` (a lowered input the steps took to $0) + or ``band_reaches_zero`` (a lowered input whose band of matching + amounts runs down to $0); + * ``shelter_cap_gap``: for a ``shelter_cap`` row, the cap less the + shelter deduction at the solver's amount. + + Household-size moves are not stepped and are left out. """ notes = joined["correctednotes"] column = notes.str.split("_").str[0].map(SOLVER_NOTE_INPUTS) candidate = joined["reproduced"] & joined["solver_moved_input"] & column.notna() - pushed = inputs.copy() + pushed, opposite = inputs.copy(), inputs.copy() raised = pd.Series(False, index=joined.index) + at_zero = pd.Series(False, index=joined.index) for name in set(column.dropna()): rows = candidate & column.eq(name) start = joined[REPLAY_INPUTS[name]].fillna(0).astype(float) up = rows & (inputs[name] > start) raised |= up - pushed.loc[rows, name] = np.where( - up[rows], inputs.loc[rows, name] + SOLVER_PUSH, 0.0 - ) - terms = solver_terms(pushed) + at_zero |= rows & ~up & inputs[name].eq(0) + far = inputs.loc[rows, name] + SOLVER_PUSH + pushed.loc[rows, name] = np.where(up[rows], far, 0.0) + opposite.loc[rows, name] = np.where(up[rows], 0.0, far) rawben = joined["RAWBEN"].astype(float) + terms = solver_terms(pushed) holds = candidate & (rawben - terms["benefit"]).abs().le(REPLAY_TOLERANCE) + back = solver_terms(opposite)["benefit"] shelter_input = column.isin(["rawrent", "rawutil"]) - named = np.select( - [ - holds & terms["benefit"].eq(inputs["rawbenmax"]), - holds & terms["benefit"].eq(inputs["rawminimum_ben"]), - holds & raised & shelter_input & terms["shelter"].eq(inputs["shelter_cap"]), - ], - ["maximum_allotment", "minimum_benefit", "shelter_cap"], - default=None, + named = pd.Series( + np.select( + [ + holds & terms["benefit"].eq(inputs["rawbenmax"]), + holds & terms["benefit"].eq(inputs["rawminimum_ben"]), + holds + & raised + & shelter_input + & terms["shelter"].eq(inputs["shelter_cap"]), + ], + ["maximum_allotment", "minimum_benefit", "shelter_cap"], + default=None, + ), + index=joined.index, + dtype=object, + ) + unnamed = holds & named.isna() + shelter_at_solver = solver_terms(inputs)["shelter"] + return pd.DataFrame( + { + "flat_stretch": named, + "opposite_push_holds": named.notna() + & (rawben - back).abs().le(REPLAY_TOLERANCE), + "unnamed_push_holds": pd.Series( + np.select( + [unnamed & at_zero, unnamed & ~at_zero], + ["input_already_zero", "band_reaches_zero"], + default=None, + ), + index=joined.index, + dtype=object, + ), + "shelter_cap_gap": (inputs["shelter_cap"] - shelter_at_solver).where( + named.eq("shelter_cap") + ), + } ) - return pd.Series(named, index=joined.index, dtype=object) def solver_outcomes(joined: pd.DataFrame) -> dict[str, Any]: @@ -719,10 +768,15 @@ def solver_outcomes(joined: pd.DataFrame) -> dict[str, Any]: at_cap = stretch.eq("maximum_allotment") rawben = joined["RAWBEN"].astype(float) moved = joined["solver_moved_input"] + from_fsben = (rawben - joined["FSBEN"].astype(float)).abs() farther = moved & ( - (rawben - joined["rawben_recreated"].astype(float)).abs() - > (rawben - joined["FSBEN"].astype(float)).abs() + (rawben - joined["rawben_recreated"].astype(float)).abs() > from_fsben ) + engine_farther = moved & ( + (rawben - joined["engine_on_original"].astype(float)).abs() > from_fsben + ) + if not farther.equals(engine_farther): + raise AssertionError("solver and engine disagree on moves away from RAWBEN") final_change = joined["rawutil"].astype(float) - joined["UTIL"].astype(float) return { @@ -775,6 +829,21 @@ def solver_outcomes(joined: pd.DataFrame) -> dict[str, Any]: for note, count in sorted(notes[at_cap].value_counts().items()) }, "maximum_allotment_at_max_flag_n": int((at_cap & flagged).sum()), + "bounded_on_neither_side_keys": sorted( + keys[weak & joined["opposite_push_holds"]] + ), + "minimum_benefit_rawben_values": sorted( + {float(v) for v in rawben[stretch.eq("minimum_benefit")]} + ), + "shelter_cap_largest_gap_dollars": float( + joined["shelter_cap_gap"].max(skipna=True) + ) + if joined["shelter_cap_gap"].notna().any() + else None, + "unnamed_push_holds_keys": { + reason: sorted(keys[joined["unnamed_push_holds"].eq(reason)]) + for reason in ("input_already_zero", "band_reaches_zero") + }, "keys": { name: sorted(keys[stretch.eq(name)]) for name in sorted(set(stretch.dropna())) diff --git a/paper/snapshot/labs/amterr/claims_audit.json b/paper/snapshot/labs/amterr/claims_audit.json index 42346c1..1339068 100644 --- a/paper/snapshot/labs/amterr/claims_audit.json +++ b/paper/snapshot/labs/amterr/claims_audit.json @@ -210,6 +210,27 @@ "util_up": 2 }, "maximum_allotment_at_max_flag_n": 52, + "bounded_on_neither_side_keys": [ + "202403-40765", + "202404-40803" + ], + "minimum_benefit_rawben_values": [ + 23.0 + ], + "shelter_cap_largest_gap_dollars": 14.0, + "unnamed_push_holds_keys": { + "input_already_zero": [ + "202311-40394", + "202312-40465", + "202402-40675", + "202406-41017" + ], + "band_reaches_zero": [ + "202310-40275", + "202405-40902", + "202406-41007" + ] + }, "keys": { "maximum_allotment": [ "202310-40260", diff --git a/tests/test_amterr_lab.py b/tests/test_amterr_lab.py index 502e701..4d5a107 100644 --- a/tests/test_amterr_lab.py +++ b/tests/test_amterr_lab.py @@ -291,6 +291,23 @@ def test_solver_outcome_counts_are_consistent(): == (flagged["n"]) ) assert set(flagged["outside_weakly_identified_keys"]).isdisjoint(weak_keys) + # "one household-size match and one match where nothing moved" + replay = audit.load_replay().set_index("key") + outside = flagged["outside_weakly_identified_keys"] + unchanged = set(CASES["replay_inputs_unchanged_keys"]) + kinds = sorted( + "unmoved" if key in unchanged else replay.loc[key, "correctednotes"][:6] + for key in outside + ) + assert kinds == ["hhsize", "unmoved"] + # Bounded on neither side, and pushes that hold without a flat stretch. + assert set(weak["bounded_on_neither_side_keys"]) <= weak_keys + for key in weak["bounded_on_neither_side_keys"]: + assert replay.loc[key, "correctednotes"] in audit.UTILITY_RESET_NOTES + unnamed = set().union(*map(set, weak["unnamed_push_holds_keys"].values())) + assert unnamed.isdisjoint(weak_keys) and unnamed.isdisjoint(missed) + assert len(weak["minimum_benefit_rawben_values"]) == 1 + assert 0 < weak["shelter_cap_largest_gap_dollars"] <= audit.SHELTER_CAP_FY2024 farther = set(outcomes["moved_benefit_farther_from_rawben_than_fsben_keys"]) assert set(household["away_from_rawben_keys"]) <= farther <= missed @@ -335,6 +352,9 @@ def test_analysis_quotes_the_audit(): cap_other = stretches["maximum_allotment"] - at_maximum["reproduced_n"] assert cap_rent + cap_util + cap_med == cap_other flagged = outcomes["reproduced_at_max_flag"] + neither = weak["bounded_on_neither_side_keys"] + unnamed = weak["unnamed_push_holds_keys"] + (minimum,) = weak["minimum_benefit_rawben_values"] farther = outcomes["moved_benefit_farther_from_rawben_than_fsben_keys"] reset_away = sorted(set(farther) - set(household["away_from_rawben_keys"])) computational_reset_off = sorted( @@ -371,10 +391,15 @@ def test_analysis_quotes_the_audit(): ), ( f"{weak['n']} of the {counts['reproduced_solver_moved_input_n']} are " - "weakly identified: the issued benefit sits on a flat stretch of the " - "benefit formula in the moved input, so the match bounds that input on " - f"one side only. {weak['above_threshold_n']} of the {weak['n']} are above " - f"the ${audit.THRESHOLD_FY2024} threshold." + "weakly identified. In each, the issued benefit lies within $5 of a " + "stretch where the solver's benefit formula is flat in the moved input: " + "the maximum allotment, the minimum benefit, or the benefit past the " + "shelter-deduction cap. By the solver's formula, every amount of the " + "input on that stretch reproduces the issued benefit, so the match " + f"bounds the input on one side at most. {len(neither)} of the " + f"{weak['n']}, {_keys(neither)}, reproduce at every utility amount and " + f"bound it on neither side. {weak['above_threshold_n']} of the " + f"{weak['n']} are above the ${audit.THRESHOLD_FY2024} threshold." ), ( f"{stretches['maximum_allotment']} are at or within $5 of the maximum " @@ -384,18 +409,28 @@ def test_analysis_quotes_the_audit(): ( "The steps, which cannot pass RAWBEN, ran that income down to $0 in " f"{at_maximum['reproduced_ended_at_zero_income_n']} and to the " - f"1,000-step limit in {at_maximum['reproduced_ended_at_step_limit_n']}. " + f"{audit.SOLVER_MAX_STEPS:,}-step limit in " + f"{at_maximum['reproduced_ended_at_step_limit_n']}. " f"In the other {cap_other} the solver moved rent ({cap_rent}), the " f"utility allowance ({cap_util}) or the medical deduction ({cap_med}), " - "and every further amount in the same direction also reproduces RAWBEN." + "and every further amount in the same direction keeps the benefit within " + f"${audit.REPLAY_TOLERANCE} of RAWBEN." + ), + ( + f"{stretches['minimum_benefit']} are at the ${minimum:.0f} minimum benefit " + "of a one- or two-person unit." ), ( - f"{stretches['minimum_benefit']} are at the $23 minimum benefit of a " - "one- or two-person unit." + f"{stretches['shelter_cap']} are rent increases that stop within " + f"${weak['shelter_cap_largest_gap_dollars']:.0f} of the shelter-deduction " + "cap. Past the cap, rent no longer changes the benefit, which stays " + f"within ${audit.REPLAY_TOLERANCE} of RAWBEN." ), ( - f"{stretches['shelter_cap']} are rent increases that reach the " - "shelter-deduction cap" + f"Not counted: {sum(len(v) for v in unnamed.values())} matches whose " + "lowered input reaches no flat stretch, because the steps took it to $0 " + f"({len(unnamed['input_already_zero'])}) or its band of matching amounts " + f"runs down to $0 ({len(unnamed['band_reaches_zero'])})." ), ( f"It is set for {flagged['n']} of the {counts['reproduced_n']} matches: " @@ -421,7 +456,7 @@ def test_analysis_quotes_the_audit(): f"{outcomes['household_size_correctedamount_zero_n']}, and it is recorded " "before the utility reset, so it differs from the final change in all " f"{utility['correctedamount_differs_from_final_change_n']} utility rows " - "and is 0 in two of them, " + f"and is 0 in {len(CASES['moved_with_zero_correctedamount_keys'])} of them, " f"{_keys(CASES['moved_with_zero_correctedamount_keys'])}, which only the " "reset moved." ),