Skip to content

Add MINT-category group breakdowns for the four blind tests (code only; real-data runs need a #42 registration) - #513

Open
MaxGhenis wants to merge 15 commits into
masterfrom
nasi-group-breakdowns-20261001
Open

MaxGhenis wants to merge 15 commits into
masterfrom
nasi-group-breakdowns-20261001

Conversation

@MaxGhenis

Copy link
Copy Markdown
Contributor

Builds the machinery for the first NASI follow-up from 2026-10-01 (Kathleen Romig, CBPP): report each blind test's results the way SSA publishes MINT results, by group and not only in total. This PR adds code and tests only. It computes no real-data group result. Each real-data run needs its own registration on issue #42 first.

MINT's categories, matched to the page

SSA's MINT8 Table User Guide ("Definitions—Table Rows and Columns", certified 2026-04-01) defines these row groups for its annual beneficiary tables. data/external/mint8_row_categories.json holds the labels verbatim, taken from the row labels of one payroll-tax option table with no data cell read:

Group Rows
Total Total
Sex Female; Male
Race and ethnicity Hispanic or Latino, any race; White, non-Hispanic; Black or African American, non-Hispanic; All other races, non-Hispanic
Country of birth United States; Other countries
Age 60–69; 70–79; 80–89; 90 or older
Marital status Married; Divorced; Widowed; Never married
Highest education level Graduate; Bachelor; Associate; High school; Less than high school
Current-law poverty status Above poverty; In poverty
Current-law household income quintile Highest; Second highest; Middle; Second lowest; Lowest
Current-law benefit type Retired worker only; Widow(er); Spousal; Disabled worker only

MINT8's annual tables have no lifetime earnings quintile. Its birth-cohort tables carry three lifetime measures instead: initial AIME quintile, lifetime payroll tax quintile, and the same shared between spouses for married years. This PR implements those three as the lifetime-earnings dimension.

MINT reports the share of people with a benefit decrease or increase of 1 percent or more, and the 10th, 50th and 90th percentiles of each person's percent change. It reports no means. The tabulator reports those statistics beside each test's own statistic.

What is added

  • Person attributes (data/group_attributes_psid.py, cohorts/group_attributes.py). Label-verified PSID readers for race and Hispanic origin (family files, 1985–2023), completed education (individual file) and country of birth (from 2013), built into a person-keyed side frame.
    • PSID microdata is read only here, in cohort code.
    • People the PSID never asked get an explicit, counted unknown.
    • The existing cohort frames and their pinned digests do not change.
  • Lifetime measures (estimates/lifetime_measures.py). AIME at 62, the present value at 62 of OASDI payroll taxes (own and shared), the Butrica–Uccello average of wage-indexed earnings at ages 22–62, and weighted quintiles with MINT's labels.
  • Tabulation (estimates/group_breakdown.py). Category schemes as data, covering MINT8 and the Butrica–Uccello report rows.
    • Rows without a classification are counted, and enter Total only.
    • Each group floor reuses the full sample's half-split; the subset is never re-split.
    • Cells under SSA's 100-case rule are flagged, never dropped.
  • Adapters (group_breakdowns/). One per test: cola, fra68, uniform_cut, min_benefit.
    • Each re-executes its frozen registered specification by composing the existing modules read-only.
    • Each refuses before loading any group attribute unless every committed cell of its runs/ artifact reproduces exactly.
  • Entry scripts. These refuse without an issue Candidate 2 design: latent-permanent conditioned chained QRF (research memo) #42 pointer, a clean HEAD equal to the registered commit and an absent output. scripts/make_nasi_repro_venv.sh builds the Python environment each committed artifact recorded.

What is not computed

  • Butrica–Uccello education and labor-force-experience rows. The cleared exercise-2 definitions extract leaves both definitions open, so the rows stay unavailable and nothing is invented.
  • Poverty status and household income quintile for exercises 1 and 3. The projection carries no income.
  • Any real-data group cell.

Protected files

This PR edits none of the following:

  • engine/loop.py or engine/steps.py
  • gates.yaml
  • any runs/*.json
  • any file pinned in data/external/track_u2/u1_identity.json
  • the v1 modules under cola_track_a/, fra68_track/ and ss/, or estimates/cola_age_profile.py

Invariants (property-based and example tests)

  • Within each dimension, classified groups plus unclassified equal Total, in both weighted and unweighted counts.
  • Total rows equal what the existing COLA, uniform-cut and Track M tabulators give on the same rows (differential tests).
  • Ratio statistics are invariant to scaling all weights.
  • Percentiles lie within the range of individual changes; decrease, unaffected and increase sum to 100.
  • A group's floor split equals the full split restricted to the group; with fewer than two usable seeds the floor is undefined, not zero.
  • Quintiles are monotone in the value, invariant to weight scaling and row order, and each holds 20 percent of weight up to the largest row's share.
  • Payroll-tax present value is zero for zero earnings, linear below the cap and flat above it; sharing conserves a couple's total.
  • AIME agrees with the ss oracle.
  • Every requested person keeps exactly one attribute row, with labels, code domains and source pins enforced.
  • A mismatched committed cell stops an adapter before any attribute loader is called, and nothing is written.

Tests (targeted, on the host)

  • 806 tests across tests/group_breakdowns, tests/estimates/test_group_breakdown*.py, tests/estimates/test_lifetime_measure*.py, tests/cohorts/test_group_*.py and tests/data/test_group_attributes_psid.py pass.
  • The host-only PSID attribute test passes on the staged files (4 tests, 2 min 42 s, 1.8 GB). It checks structure and coverage only.
  • tests/estimates/test_birth_evidence_artifact.py, tests/track_u2/test_u2_isolation.py and the four tests/test_replication_*.py pin tests pass.
  • Tier manifest: unit 6,199, artifact 3,696, integration_psid 1,346.

axiom: n/a (tabulation and cohort-attribute code; no policy rule changes)

🤖 Generated with Claude Code

MaxGhenis and others added 12 commits October 2, 2026 09:34
estimates/group_breakdown.py tabulates any blind test's person rows by
SSA's MINT8 row groups (sex, race and ethnicity, country of birth, age,
marital status, education, poverty status, household income quintile,
benefit type) and by data-driven alternate schemes. It reports the
tests' own statistics (ratio of scenario means over draws, poverty
change, shares) beside MINT's (percent with a decrease or increase at
the 1 percent threshold, 10th/50th/90th percentiles of individual
percent change, poverty counts). Group floors restrict the full-sample
family split; suppression flags mark cells under SSA's 100-case rule
without dropping them. The MINT8 labels come from captured SSA sources
pinned by SHA-256.

Invariants tested: groups plus unclassified partition Total; ratio
statistics are invariant to weight scaling; Total rows agree with the
existing COLA, uniform-cut and Track M tabulators; percentiles are
bounded; decrease + unaffected + increase = 100.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
estimates/lifetime_measures.py computes, from cohort outputs only:
current-law AIME at age 62 (MINT8's initial AIME), the present value at
62 of OASDI payroll taxes, own and shared between spouses for married
years (MINT8's lifetime payroll tax measures), the Butrica-Uccello
average of wage-indexed earnings at ages 22-62, and weighted quintiles
with MINT's labels, cut per birth cohort or over the whole population.
OASDI tax rates and trust fund effective interest rates are captured
from SSA with SHA-256 pins. Unrecorded conventions are exposed as named
builder defaults for the registration to fix.

Invariants tested: quintile weight bounds and monotonicity; invariance
to weight scaling and row order; zero earnings give zero tax; tax is
linear below the cap and flat above it; sharing conserves a couple's
total; AIME agrees with the ss oracle.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
data/group_attributes_psid.py reads, with exact label checks and
documented code domains, the family-file race and Hispanic-origin
reports of reference persons and spouses, the individual file's
completed education by wave, and country of birth where the PSID asks
it. cohorts/group_attributes.py turns them into a person-keyed side
frame for the 2009 and 2011 projection cohorts, the age-67
observations and the Track M 2023 universe, mapped to MINT8's
categories and to the Butrica-Uccello report rows. Conflicts across
waves resolve deterministically; people never asked (other family-unit
members) get an explicit, counted unknown. The side frame has its own
file audit and SHA-256 seals, so the existing cohort frames and their
pinned digests do not change.

Invariants tested: every requested person keeps exactly one row;
labels, domains and source pins are enforced; resolution is
deterministic; mutation is detected.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…-r2' and 'nasi/g3-group-tabulation' into nasi/g-base
group_breakdowns/cola.py and fra68.py re-execute each frozen registered
specification (projection, benefits, union rows) by composing the
existing modules read-only, and refuse before loading any group
attribute unless every committed cell and per-draw diagnostic matches
runs/replication_urban2010_cola_v1.json or
runs/replication_urban2010_fra68_v1.json. They then join G1's race,
education and nativity side frame and G2's lifetime measures by
person_id and tabulate with G3 for both the tests' own population and
MINT's beneficiaries aged 60 or older in 2030. Poverty status and
household income quintile are not computable in the projection (no
income) and are marked so.

scripts/run_projection_groups_registered.py refuses without an issue
#42 pointer, a clean HEAD equal to the registered commit and an absent
output; scripts/make_nasi_repro_venv.sh builds the environment the
committed artifacts recorded. Invented dry runs only; the large
invented result files are kept outside the repository and the
RESULTS.md summaries are committed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
group_breakdowns/uniform_cut.py re-executes the frozen Track U
specification for every registered row (build_age67_cohort ->
income_rows -> adjusted_incomes -> tabulation_rows ->
tabulate_uniform_cut with the design frame), checks the inputs digest
and every committed cell of runs/replication_boomers2004_uniform_cut_v1
with zero tolerance, and refuses before loading any attribute on a
mismatch. It then adds MINT groups (race and ethnicity, education,
lifetime earnings quintiles own and shared, poverty status, household
income quintile) and the report's own rows, with the adjusted-poverty
statistic, design standard errors and floors, plus MINT's poverty-table
statistics. Every observation must be born 1936-45; the 1946-55 cohort
(the blind U2 target) never enters.

scripts/run_track_u_groups_registered.py refuses without an issue #42
pointer, a clean HEAD equal to the registered commit and an absent
output. Invented data only; the large invented result is kept outside
the repository.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
group_breakdowns/min_benefit.py re-executes Registration 17's
registered computation (cohort, careers, evaluate and the unchanged
tabulate_track_m for every registered row and the d430 sensitivity) and
refuses before loading any attribute unless every committed cell,
n column, floor, design SE and cohort structure of
runs/replication_urban2006_minimum_benefit_v1.json matches. It then
adds MINT groups (marital status in MINT order with unclassified
counted, MINT age bands from 2022 - birth year, sex, race and
ethnicity, education, initial AIME and lifetime payroll tax quintiles,
benefit type) to the share receiving the minimum under each option.

tests/group_breakdowns/test_min_benefit_reproduction_psid.py is a
host-only check that the committed (already published) cells reproduce
exactly; it computes no group cell. Invented data otherwise.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ily from g4a; reconciled in the next commit)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…rarily from g4a; reconciled in the next commit)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
group_breakdowns/common.py is now one shared module for all four
adapters (cola, fra68, uniform_cut, min_benefit): one zero-tolerance
comparator, one registration preflight (issue #42 pointer, HEAD equal
to the registered commit, clean tree including untracked files, a new
runs/ destination), one paired exclusive write with rollback, and one
environment builder. Where the three builders' helpers differed, the
stricter behavior is kept. Every adapter still refuses before any group
attribute loads unless its committed artifact reproduces.

The three captures of SSA's MINT8 Table User Guide collapse to one
canonical capture (certified 2026-04-01) and one label file; tests show
the snapshots' main content and labels are identical.

The Butrica-Uccello report scheme now follows the cleared exercise-2
definitions extract: race rows, own and shared lifetime earnings
(wage-indexed earnings at ages 22-62 over a 41-year divisor, shared in
married years). The extract leaves the education and labor-force-
experience definitions open, so those rows stay unavailable, and the
unrecorded quintile conventions are named choices for the registration.

Adds the ten new src modules to the birth-evidence reducer's post-
review exclusions and the new tests to the tier manifest.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…6, integration_psid 1,346)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@vercel

vercel Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
social-security-model Ready Ready Preview Oct 5, 2026 1:06am UTC

Request Review

MaxGhenis and others added 2 commits October 4, 2026 14:33
…ent years under ALL_AGES

Independent review of cd40245 (REQUEST_CHANGES):
- Major: the projection and minimum-benefit entry scripts wrote the
  caller's --output spelling while preflight checked its absolute
  repository path, so a relative --output from another directory wrote
  elsewhere and left the one-shot guard's file absent. Both now write
  state['output_path'] from preflight. Regression tests run preflight
  from another directory with a relative path and confirm a second
  preflight refuses.
- Minor: under the ALL_AGES divisor, shared report earnings skipped
  married years absent from the person's own career, dropping the
  spouse's half (the reviewer's case returned 487.80 instead of 500).
  Those years now share; the reviewer's case is a test.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…eakdowns PR

Keeps both sides' reducer exclusions. Tier manifest from a full
collection, reconciled with master plus this PR's delta: unit 6,401,
artifact 3,904, integration_psid 1,346.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
CI on 29e6eea failed two pin tests the targeted runs did not cover:
- tests/estimates/test_coordinator.py pins every module under
  estimates/; group_breakdown.py and lifetime_measures.py are
  post-compute re-tabulations outside the registered first-estimates
  surface, declared as their own surface beside Track A's and Track U's
  tabulators (coordinator._ESTIMATOR_SURFACE_SOURCES is unchanged).
- tests/ss/test_statutory_aime.py pins every module that names
  LEGACY_FIXED_35; lifetime_measures.py uses it for the exercise 1 and
  3 AIME conventions, so the projection tests' quintiles use the AIME
  their benefits use.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch was successfully deployed

1 active deployment
Preview — 9227a8f1 Deployed Oct 5, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant