+
+
+
+ An official website of the United States government
+ Here's how you know
+
+
+
+
Official websites use .gov A .gov website belongs to an official government organization in the United States.
+
+
Secure .gov websites use HTTPS A lock (
+
+ ) or https:// means you’ve safely connected to the .gov website. Share sensitive information only on official, secure websites.
+
The first four sets of tables provide results for the analysis years, 2030, 2050, and 2070, while the last two provide results for four birth cohorts: 1960–1969,1980–1989,2000–2009, and 2020–2029.
+
Each set of tables and what they show are discussed below.
The first two columns show the percent of the population with a benefit decrease or increase, and the next three columns show the percent change in individual Social Security benefits at three percentiles.
The first five columns are the same as the Social Security benefits tables except tax changes are shown rather than benefit changes. Columns 6–8 show the dollar amount changes in individual Social Security payroll taxes at the same percentiles as the percentage change columns.
+
Household Income
+
These tables have the same structure and population as the Social Security benefits tables except the effects on household income are shown rather than individual Social Security benefits.
The first two columns show the poverty rate with and without (current law) the proposed option. The next three columns show the number in poverty (expressed in thousands) with and without (current law) the proposed option and the difference between them. The final column shows the percent change in the number in poverty with the proposed change. This percent change is calculated by dividing the change in thousands in poverty (5th column) by the thousands in poverty without the proposal (3rd column).
+
CORRECT interpretation: the number of people in poverty would decline by 2 percent. INCORRECT interpretation: the poverty rate would decline by 2 percent.
+
Benefit/Tax Ratios
+
These tables show the projected changes in the benefit/tax ratios for workers born in four birth cohorts: 1960–1969,1980–1989,2000–2009, and 2020–2029 who have a tax record from which to calculate a benefit/tax ratio. We excluded those who paid zero payroll taxes over their lifetime and, therefore, could not have a benefit/tax ratio calculated.
+
The first five columns are similar to the Social Security benefits and taxes paid tables except the percent of the population with a benefit/tax ratio decrease and increase, and the percent change in the benefit/tax ratio at three percentiles are shown.
+
The last six columns show the distribution of benefit/tax ratios with and without (current law) the proposed option at three percentiles.
+
The benefit/tax ratio is a money's worth measure that is the lifetime present value of benefits divided by the lifetime present value of payroll taxes. The ratio represents how much in benefits an individual received for every dollar of payroll taxes paid.
+
How to interpret the ratios:
+
+
0% means that an individual received no benefits despite paying payroll taxes (or $0.00 in benefits for every $1.00 of taxes).
+
50% means that an individual received half as much benefits as was paid in taxes (or $0.50 in benefits for every $1.00 of taxes).
+
100% means that an individual received the same amount of benefits as was paid in taxes (or $1.00 in benefits for every $1.00 of taxes).
+
1,350% means that an individual received over 13 times the amount of benefits as was paid in taxes (or $13.50 in benefits for every $1.00 of taxes).
+
+
The present value of benefits includes all Social Security benefits the individual received, regardless of earnings record or type of benefit. The present value of payroll taxes includes all the payroll taxes that the individual paid over a lifetime. We use the Social Security Trust Fund interest rate to adjust benefits and taxes to their present values at age 62.
+
Initial Replacement Rates
+
These tables are the same as the benefit/tax ratio tables except initial replacement rates are shown for current-law beneficiaries born in four birth cohorts: 1960–1969,1980–1989,2000–2009, and 2020–2029. Only beneficiaries with both income and benefit records from which to calculate an initial replacement rate are included in the population. Beneficiaries with zero average indexed monthly earnings (AIME) or zero benefit at claiming (due to the earnings test or other fixed-dollar reductions) are excluded because no replacement took place. When no benefit is received, the initial replacement of earnings may take place in a later year or never, if the beneficiary dies before a benefit is paid.
+
The initial replacement rate represents how much of the initial AIME is replaced by the initial total monthly Social Security benefit. It is calculated by dividing the initial monthly benefit by the initial AIME.
+
How to interpret the initial replacement rate values:
+
+
20% means that the initial benefit replaced one-fifth of lifetime earnings (or $0.20 in monthly benefits for every $1.00 of AIME).
+
100% means that the initial benefit replaced all lifetime earnings (or $1.00 in monthly benefits for every $1.00 of AIME).
+
150% means that the initial benefit replaced one and a half times lifetime earnings (or $1.50 in monthly benefits for every $1.00 of AIME).
+
+
The initial monthly benefit is the total individual Social Security benefit received at the person's claiming age, including any spousal, survivor, or disability benefits received. We calculate the replacement rate at claiming age, regardless of what type of benefits the beneficiary claimed.
+
Interpreting the Profile of Beneficiaries by Race & Ethnicity Tables
The tables also include the same information for four race and ethnicity groupings:
+
+
Hispanic or Latino, any race
+
White, non-Hispanic
+
Black or African American, non-Hispanic
+
All other races, non-Hispanic
+
+
Each set of tables and what they show are discussed below.
+
Social Security Benefits
+
These tables show the projected distribution of individual monthly Social Security benefits at three percentiles. The benefit amount is the total monthly benefit an individual would receive, regardless of the type of benefit or earnings record it came from.
+
Poverty Rates and Numbers
+
These tables show the poverty rate and number in poverty under the Official Poverty Measure and the Supplemental Poverty Measure.
+
The first and third columns show the rates and numbers under the Official Poverty Measure while the second and fourth columns show the same information under the Supplemental Poverty Measure. Further information about the Official and Supplemental Poverty Measures and how they relate to the aged, is available in this paper.
The last five columns show the projected mean share of household income from five sources:
+
+
Social Security benefits, which include benefits for the individual, spouse, and any children.
+
Annuitized asset income, which includes income from defined contribution plans (such as 401(k) accounts) and personal savings.
+
Defined benefit pension income, which includes the individual's and any spouse's defined benefit pension income.
+
All earnings, including covered earnings (from which Social Security taxes are withheld) and non-covered earnings (no Social Security taxes withheld) of the individual and his or her spouse.
+
Coresident income, which is the income of non-spousal coresidents in the household.
+
+
Rows may not sum to 100 percent because minor sources of income are excluded.
+
Total Earnings
+
These tables show the projected distribution of annual individual total earnings at three percentiles. Total earnings includes both covered and non-covered wages.
+
Household Wealth
+
These tables show the projected distribution of household wealth at three percentiles. Wealth includes retirement account balances and savings, but excludes household equity.
+
Health Status and Costs
+
These tables show projected health status and expenses. The first two columns cover the percent and number (in thousands) of beneficiaries who are projected to self-report fair or poor health. The third column shows family median annual health insurance premiums, and the fourth column shows family median annual out-of-pocket health expenses.
+
Interpreting the Profile of Taxpayers by Race & Ethnicity Tables
These tables show the projected distribution of annual individual Social Security taxes paid at three percentiles. Both the employer and employee shares of the Social Security portion of FICA (Federal Insurance Contributions Act) taxes are included.
+
Covered Earnings
+
These tables show the projected distribution of annual individual covered earnings at three percentiles. Covered earnings are wages from work that is subject to the Social Security payroll tax. The covered earnings in these tables are not capped at the taxable maximum.
Interpreting the Population Characteristics Tables
+
Each projection table links (at the top) to the corresponding population characteristics table. From there, you can use the tabs on the right side to view each population analyzed across the MINT projections including:
We use 10-year birth cohorts to increase the sample size to a point where the characteristic subgroups could be examined.
+
The three columns in every population characteristics table are:
+
+
Unweighted sample: The number of people in the sample for each row. We use it to verify that we comply with disclosure avoidance policy (see “Sample Size Restrictions”).
+
Population (in thousands): The weighted population in thousands. A value of 71,500 means 71 million, 500 thousand.
+
Share of population: The percentage of the population in that particular characteristic subgroup. The Total row at the top of the table is always 100%. The percentages for the each subgroup underneath should add to 100%. For instance, under Country of Birth, United States may be 84% and Other Countries would be 16%, which adds to 100%.
+
+
CORRECT interpretation: 40% of the population is female, 60% of the population is male INCORRECT interpretation: 40% of females are in this population, 60% of males are in this population
Total: Refers to the total population of the table.
+
Sex: Female or Male.
+
Race/Ethnicity: We list “Hispanic or Latino, any race” first; the rest of the groups (White, Black or African American, and All other races) are non-Hispanic. Additional racial or ethnic identifications are not covered because they are not in the datasets used to build the MINT8 model.
+
Country of Birth: We differentiate between the United States and other countries.
+
Age:
+
+
The beneficiary population includes those aged 60 or older because 60 is the earliest eligibility age for any aged benefits under current law.
+
The taxpayer population includes those aged 31 or older because 31 is the earliest age in MINT for household income and poverty information.
+
+
Marital Status: Refers to the marital status in the year of analysis only. An individual's marital status can change in the future and may have been different in the past.
+
Highest Education Level: Reported number of years of education.
+
+
Graduate means more than 16 years of education,
+
Bachelor means 16 years of education,
+
Associate means 14–15 years of education,
+
High school means 12–13 years of education, and
+
Less than high school means less than 12 years of education.
+
+
Current-Law Poverty Status: Indicates whether the person is in a household that has income above (“above poverty”) or below (“in poverty”) the official poverty line under current law. The household income used for the official poverty measure is the same as the household income used in our results except for how asset income is counted. The official poverty measure of asset income only includes dividend income, interest income, and rental income (non-annuitized) as reported on income tax returns. In contrast, we include annuitized asset income from all household wealth held in defined contribution plans (such as 401(k) accounts) and personal savings in that year. We add the annuitized asset income to account for the expected spend-down of assets in retirement. The asset income value used for official poverty calculations generally produces a substantially lower asset income value than the household income measure.
+
Current-Law Household Income Quintile: Represents an individual's annual household income under current law, including:
+
+
household earnings;
+
asset income (annuitized), which includes income from defined contribution plans (such as 401(k) accounts) and personal savings;
+
defined benefit pensions;
+
means-tested income;
+
non-means-tested income;
+
Social Security;
+
Supplemental Security Income; and
+
non-spousal co-residents' income.
+
+
We calculate the income quintiles for each year (e.g., 2030, 2050, or 2070) for the population analyzed, determine the dollar thresholds for each income quintile, and assign each beneficiary to the appropriate quintile. The dollar ranges are available upon request.
+
Current-Law Benefit Type: Some Social Security benefits are based on one's own work, while others are based on the work of a current, divorced, or deceased spouse. The current-law benefit type refers to one of the following benefit types received in the specified analysis year:
+
+
Retired-worker only: receives only a retired-worker benefit based on his or her earnings record.
+
Widow(er) (includes dually entitled): receives a survivor benefit (may or may not also receive a lower worker benefit from his or her own earnings record, known as dually entitled).
+
Spousal (includes dually entitled): receives a spousal benefit (may or may not also receive a lower worker benefit from his or her own earnings record, known as dually entitled).
+
Disabled-worker only: receives a disabled-worker benefit on his or her earnings record and is under the full retirement age (FRA). Disabled workers convert to retired workers at FRA.
+
+
Our results do not show different benefit types a beneficiary might receive under a policy option/proposal or in a different year under current law.
+
Current-Law Payroll Taxes Quintile: Represents an individual's annual Social Security payroll taxes under current law. We calculate the payroll tax quintiles for each analysis year's population of payroll taxpayers aged 31 or older.
+
Current-Law Initial AIME Quintile: Represents an individual's average indexed monthly earnings (AIME) under current law at age 62, the earliest eligibility age for retired-worker benefits. We calculate the AIME quintiles for each birth cohort. The dollar ranges are available upon request.
+
Lifetime Payroll Tax Quintile: Represents the present value of an individual's current-law payroll taxes at age 62. We calculate the payroll tax quintiles for each birth cohort. The dollar ranges are available upon request.
+
Lifetime Payroll Tax Quintile (Shared): Represents the present value of an individual's current-law payroll taxes at age 62. For married couples, the payroll taxes paid while married are shared equally between them. For never-married individuals, this is the same as the lifetime payroll tax. In any year where an individual is not married, we count only their individual payroll taxes. We calculate the quintiles for each birth cohort. The dollar ranges are available upon request.
+
+
Measures—Column Headings
+
+
Threshold for Categorization in the “Decrease” or “Increase” Groups (“Percent of Population with a—[decrease or increase]” Columns)
+
We categorize individuals as having a “decrease” in the amount being analyzed (benefits, taxes, income, etc.) when a proposal would reduce the analyzed quantity by 1% or more. Individuals are categorized as having an “increase” when a proposal would raise the analyzed quantity by 1% or more. We consider individuals with differences between −1% and 1% to be unaffected.
+
For example, consider two individuals with benefits under an option/proposal that are lower than benefits without the proposal—individual A has a benefit decrease of 0.8% while individual B has a decrease of 1.6%. We would consider individual A unaffected and not categorize him or her in the “decrease” or “increase” columns. However, we would categorize individual B as having a decrease.
+
Thus, some beneficiaries who are technically “affected” by a proposal, but who will still receive essentially the same benefit amount (or pay essentially the same taxes, etc.) are considered to be unaffected in our projections.
+
Percent Change Values (“Percent Change in [item being analyzed] at the—[three percentiles]” Columns)
+
Understanding how to interpret the distribution of percent changes is critical to understanding the results correctly.
+
The formula we use to calculate the percent change for each individual is:
From this distribution of individual percent changes, ranked from high to low, we calculate the 10th percentile, median, and 90th percentile values (see below).
+
CORRECT interpretation: a −5% median indicates that this is the median of the distribution of individual percent changes (half the individuals have a percent change that is higher and half have a percent change that is lower). INCORRECT interpretation: a −5% median indicates that the median amount under the option is 5% less than the median amount under current law.
+
10th, Median, and 90th Percentile Values
+
The percentiles provide a picture of the distribution of policy option effects or beneficiaries/taxpayers financial status distributed from lowest to highest. The example below is for a percent change in benefits, but also applies to distributions of other policy options, dollar amounts, initial replacement rates, household income levels, etc.
+
+
+
Table example
+
+
+
+
+
+
Percent change in Social Security benefits at the—
+
+
+
10th percentile
+
Median
+
90th percentile
+
+
+
+
+
Total
+
2%
+
4%
+
21%
+
+
+
+
+
+
+
+
+
+
10th percentile: “2%” means that 10 percent of the population has a benefit change of less than 2 percent, while 90 percent have a benefit change of more than 2 percent.
+
Median: “4%” means that 50 percent of the population has a benefit change of less than 4 percent, while 50 percent have a benefit change of more than 4 percent.
+
90th percentile: “21%” means that 90 percent of the population has a benefit change of less than 21 percent, while 10 percent have a benefit change of more than 21 percent.
+
CORRECT interpretations include:
+
+
10% of this population has a benefit change of less than 2%.
+
10% of this population has a benefit change of more than 21%.
+
40% of this population has a benefit change of between 2% to 4%.
+
40% of this population has a benefit change of between 4% to 21%.
+
50% of this population has a benefit change of less than 4%.
+
50% of this population has a benefit change of more than 4%.
+
+
+
+
Additional Notes
+
Dollar Amounts
+
All dollar amounts are presented in today's dollars, meaning that they are in real dollars (inflation-adjusted) for the year the table is produced. If a table is run in 2021, the dollars are in 2021 dollars and the table will note that in the column label. Tables run in 2024 will be in 2024 dollars, and so on.
+
Sample Size Restrictions
+
To maintain the privacy of survey respondents, our tables have built-in disclosure avoidance protections that suppress an entire characteristic subgroup if the sample size for any row in that subgroup is less than 100 individuals. This minimum sample size protects privacy so that we can show the 10th and 90th percentiles based on a sample size of at least 10.
+
For example, if there are only 82 widow(er)s in the Widowed row of the Marital Status subgroup, the entire Marital Status subgroup is removed from the table. We remove the entire subgroup to avoid secondary or tertiary disclosure issues. The subgroups below Marital Status in this example would automatically move up the page. A “short” table can reveal at a glance that at least one subgroup has been removed.
+
Columns that show percent of the population with a decrease or increase in whatever value is being shown (e.g. benefits, payroll tax, household income) have disclosure restrictions on the numerator's sample size as well as the denominator. The numerator must have either zero cases or meet a minimum numerator threshold of 10 for the table to display it. The table would suppress any characteristic subgroup that has a percentage based on a numerator of 1–9 cases. (MINT results are weighted, but for this example, everything is presented in unweighted sample sizes.)
+
There is an exception for low numerator situations where the table would show a 0% for any numerator from zero up to and including the minimum numerator threshold. If a particular percentage was based on seven records with a denominator of 30,000, it would produce a percentage of 0.02%. By only showing percentages in single digits, we would display this result as 0%, which would be the same value displayed for any numerator from 0–149.
+
This exception is important because there are policy options where 27,000/30,000 beneficiaries (90%) would receive a benefit increase while 8/30,000 receive a decrease. Without the exception, the very small decrease numerator would suppress a number of subgroups for both the increased and the decreased results, which limits the presentable results more than is necessary for disclosure avoidance.
+
1 All the characteristic subgroups are not in every table. Birth cohort tables (benefit/tax ratios and initial replacement rates) do not have age or marital status breakouts. Characteristic subgroups can also drop out of tables because of sample size restrictions (see “Sample Size Restrictions” for details).
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
\ No newline at end of file
diff --git a/data/external/mint8_table_user_guide.source.provenance.json b/data/external/mint8_table_user_guide.source.provenance.json
new file mode 100644
index 00000000..9cb8a3c6
--- /dev/null
+++ b/data/external/mint8_table_user_guide.source.provenance.json
@@ -0,0 +1,22 @@
+{
+ "schema_version": "external_source_provenance.v1",
+ "committed_source_file": "data/external/mint8_table_user_guide.source.html",
+ "source_url": "https://www.ssa.gov/policy/docs/projections/user-guide.html",
+ "document": "Social Security Administration, Table User Guide - Modeling Income in the Near Term (MINT) 8",
+ "page_title": "Table User Guide - Modeling Income in the Near Term (MINT) 8",
+ "date_certified": "2025-10-01",
+ "date_certified_locator": " in the committed source",
+ "sections_of_record": [
+ "Definitions—Table Rows and Columns > Characteristic Subgroups—Table Rows",
+ "Measures—Column Headings > Threshold for Categorization in the “Decrease” or “Increase” Groups",
+ "Measures—Column Headings > Percent Change Values",
+ "Additional Notes > Sample Size Restrictions"
+ ],
+ "retrieval_date": "2026-10-01",
+ "fetch_method": "Internet Archive raw capture http://web.archive.org/web/20260419231110id_/https://www.ssa.gov/policy/docs/projections/user-guide.html, fetched by the MINT-categories lane (workflow wf_35b496b0-085, orchestrating session 95606380, 2026-10-01) after a direct fetch of ssa.gov returned an Akamai 'Access Denied' page. The capture is identified by its Wayback CDX record: the SHA-1 (base32) of the committed bytes, N2TYAIU2LPKHETR3M5H4E6MUZLL4KFVO, equals the CDX digest of capture 20260419231110 (status 200), checked 2026-10-01 against http://web.archive.org/cdx/search/cdx?url=ssa.gov/policy/docs/projections/user-guide.html.",
+ "source_sha256": "278d5d19c1b50f1d354db1ada515288af563c035a67eb16fb971d25700fb94e9",
+ "source_sha1_base32": "N2TYAIU2LPKHETR3M5H4E6MUZLL4KFVO",
+ "source_length_bytes": 73926,
+ "consumed_by": "src/populace_dynamics/estimates/group_breakdown.py (MINT8 category definitions, the change threshold, the percent-change formula and the disclosure rule; verify_mint8_sources checks every quoted definition against this file's text)",
+ "note": "Methods page only: it defines MINT8's table rows and columns and prints no projection result. It is not a MINT policy-option table."
+}
diff --git a/src/populace_dynamics/estimates/group_breakdown.py b/src/populace_dynamics/estimates/group_breakdown.py
new file mode 100644
index 00000000..cbfc2867
--- /dev/null
+++ b/src/populace_dynamics/estimates/group_breakdown.py
@@ -0,0 +1,3959 @@
+"""Group breakdowns of the blind tests by MINT8 characteristic subgroups.
+
+NASI follow-up package G3 (orchestrating session 95606380, 2026-10-01).
+After the NASI meeting of 2026-10-01 Max asked for each of the four
+finished DYNASIM3 blind tests broken down by SSA's MINT categories. This
+module is the tabulation core of that request. It takes the person rows a
+blind test already produces, with each person's group attributes joined on
+by cohort code, and reduces them to one cell per characteristic subgroup.
+It reads no PSID file, projects nothing, computes no benefit, income or
+poverty status, reads no comparator value, applies no acceptance rule and
+writes no artifact. A real-data breakdown is a post hoc analysis that runs
+only under its own issue #42 registration: ``data_provenance=
+"registered_real"`` requires the registration pointer, output labels and
+the caller's post hoc labels. Development and tests use INVENTED rows.
+
+Category schemes (data, not code paths)
+---------------------------------------
+A :class:`CategoryScheme` is an ordered tuple of :class:`Dimension` (a row
+group such as "Marital status"), each an ordered tuple of
+:class:`Category` (a row such as "Widowed"). A dimension is one of four
+kinds:
+
+* ``total``: one row, every person;
+* ``categorical``: each category lists the attribute codes it holds;
+* ``band``: each category is an inclusive integer interval (age, years of
+ education);
+* ``quintile``: five categories ranked 1 (lowest) to 5 (highest) of a
+ numeric measure, cut by weighted quintile thresholds
+ (:func:`weighted_quintile_ranks`).
+
+:data:`MINT8_SCHEME` holds the row groups and row labels of MINT8's annual
+projected-effects tables for Social Security benefits and household income
+(tables 1-3 and 7-9 of a MINT8 policy-option page; population: current-law
+beneficiaries aged 60 or older), verbatim and in SSA's page order, taken
+from the committed labels-only extract
+``data/external/mint8_row_categories.json``. :data:`MINT8_POVERTY_SCHEME`
+is the official-poverty tables' scheme (tables 10-12), which carries no
+household income quintile rows, and :data:`MINT8_COHORT_SCHEME` the cohort
+tables' scheme (tables 13-20), whose three lifetime measures (initial AIME
+quintile; lifetime payroll tax quintile, own and shared) serve as the
+lifetime-earnings dimension the NASI request names. The two
+``*_WITH_LIFETIME`` schemes append those three measures to the annual
+schemes; they are compositions of this module, not MINT8 table layouts,
+and say so (``composite``). Every definition string a MINT8 dimension
+carries is a verbatim quote of SSA's MINT8 Table User Guide (committed
+capture ``data/external/mint8_table_user_guide.source.html``), and
+:func:`verify_mint8_sources` checks the files' SHA-256 pins, every label
+against the label extract and every quote against the Guide's text.
+
+Alternate schemes (for example the row labels of Butrica and Uccello 2004,
+which ``uniform_cut_tabulation.NOT_COMPUTED_REPORT_ROWS`` already records)
+are added with :func:`register_scheme`, and variants with
+:func:`derive_scheme`; :func:`with_age_bands` swaps a scheme's age bands for
+another set in :data:`AGE_BAND_SETS` (MINT8's beneficiary and taxpayer
+bands, and exercises 1 and 3's registered bands ``50-61`` ... ``80+``,
+taken from ``cola_age_profile.DEFAULT_AGE_GROUPS``).
+
+Group assignment and unclassified rows
+--------------------------------------
+:func:`assign_groups` maps each row of a frame to one category per
+dimension and returns a :class:`GroupAssignment`, whose long frame has one
+row per (row id, dimension) with the label or ``unclassified``. A row
+whose attribute is missing (``None``, ``NaN``, ``pd.NA``), whose code the
+caller declares unclassified (``unclassified_codes``, for example a
+``separated`` marital status SSA does not place), or whose band value lies
+outside every band is unclassified for that dimension: it is counted, by
+reason, and enters Total only. A present code that no category holds, and
+was not declared, is refused, never silently dropped. So within each
+dimension the classified labels partition the classified rows, and
+classified plus unclassified is Total, in counts and weights.
+
+Statistics
+----------
+*Projection benefit change* (exercises 1 and 3;
+:func:`tabulate_projection_breakdown`), per draw and cell, reusing
+``cola_age_profile`` read-only (``_normalize``, ``_membership_masks``,
+``_cell``, ``_draw_summary``, ``_floor_summary``, ``_side_mean``): the
+ratio of scenario means ``100 * (mu_reform / mu_base - 1)`` and the mean
+of individual ratios, with scenario memberships exactly as the
+``ColaAgeProfileConfig`` defines them; and MINT8's benefit statistics over
+the current-law beneficiaries with a positive selected benefit: percent
+with a decrease, percent with an increase (and the unaffected rest), and
+the weighted 10th, 50th and 90th percentiles of each person's percent
+change. Each is reported as the mean over draws with the draws' sample SD.
+
+*Static* cells (exercises 2 and 4, one deterministic run): poverty
+(:func:`tabulate_poverty_breakdown`: official poverty rate under current
+law and with the proposal, the change in points, the number in poverty in
+thousands under each and its change, and the percent change in the number
+in poverty, rates through ``uniform_cut_tabulation._rates``); a weighted
+share (:func:`tabulate_share_breakdown`, exercise 4's statistic); and the
+MINT8 benefit statistics (:func:`tabulate_static_benefit_breakdown`).
+
+Weighted percentile (:data:`PERCENTILE_DEFINITION`)
+---------------------------------------------------
+Zero-weight rows carry no mass and are dropped; values are sorted; with
+``W_k`` the cumulative weight of the k smallest and ``W`` the total, the
+level-p percentile is the smallest ``x_k`` with ``W_k >= p W``, and when
+``W_k = p W`` holds exactly it is the midpoint of ``x_k`` and ``x_(k+1)``.
+Levels are exact rationals and the comparison is exact rational arithmetic
+on the float64 weights (integer numerators over one power-of-two
+denominator), so "exactly" means exactly. This is the rule the G3 brief
+specifies (smallest value whose cumulative weight reaches p, midpoint at
+an exact tie); its first half is the repository's existing rule
+(``uniform_cut_track_u.diagnostics.weighted_quantile``).
+
+Uncertainty
+-----------
+* Projection: the sample SD over the K draws, and the half-sample floor of
+ ``cola_age_profile``: for each floor seed the split units
+ (``config.floor_split_unit``, the opening-wave family unit by default)
+ of **all** input rows are split once with
+ ``harness.panel.split_panel_by_person`` (fraction 0.5), and each group
+ uses that split intersected with its own mask -- never a re-split of the
+ group's subset, which would put family units on different sides
+ (:func:`half_split_masks`). The floor is the summary of
+ ``|side_a - side_b|`` over the usable seeds and is undefined, not zero,
+ with fewer than two.
+* Static: the design-based standard error of
+ ``uniform_cut_tabulation._design_se`` (Taylor linearization of the
+ weighted ratio over the full sample design's strata and clusters, the
+ group as a domain) for rates, changes in rates and shares, and the same
+ half-sample floor with the split computed once on all rows (default
+ split unit: family units linked through shared persons,
+ ``uniform_cut_tabulation.floor_split_units``). No design SE is reported
+ for weighted percentiles, weighted counts or the percent change in the
+ number in poverty; each such cell says why.
+
+Suppression flags (never drops)
+-------------------------------
+Every cell carries its unweighted n and flags: below SSA's disclosure
+minimum of 100 cases, below 30 cases, and, for the percent-with-decrease
+and percent-with-increase statistics, SSA's numerator rule (a numerator of
+1-9 cases). A dimension with any flagged row is flagged
+``ssa_subgroup_suppressed``, because SSA would remove the whole subgroup.
+Nothing is dropped.
+
+Result schema
+-------------
+Every tabulator returns a :class:`GroupBreakdownResult` (frozen
+dataclasses down to each :class:`StatisticCell`); ``as_dict()`` gives a
+JSON-safe mapping (finite numbers only) carrying the labels, the caller's
+post hoc labels, the scheme with its citations, every statistic's
+definition and MINT8 column, and the conventions above.
+"""
+
+from __future__ import annotations
+
+import bisect
+import hashlib
+import itertools
+import json
+import math
+import re
+from collections.abc import Callable, Iterable, Mapping, Sequence
+from dataclasses import dataclass
+from fractions import Fraction
+from html.parser import HTMLParser
+from numbers import Integral, Real
+from pathlib import Path
+from typing import Any
+
+import numpy as np
+import pandas as pd
+
+from populace_dynamics.estimates import cola_age_profile as a7
+from populace_dynamics.estimates import uniform_cut_tabulation as track_u
+from populace_dynamics.harness.panel import split_panel_by_person
+
+__all__ = [
+ "AGE_BAND_SETS",
+ "BAND",
+ "CATEGORICAL",
+ "DATA_PROVENANCES",
+ "DEFAULT_FLOOR_SEEDS",
+ "DIMENSION_KINDS",
+ "FAMILY_UNIT",
+ "FAMILY_UNIT_LINKED_BY_PERSON",
+ "INVENTED",
+ "INVENTED_DATA_LABEL",
+ "MINT8_ANNUAL_WITH_LIFETIME_SCHEME",
+ "MINT8_COHORT_SCHEME",
+ "MINT8_POVERTY_SCHEME",
+ "MINT8_POVERTY_WITH_LIFETIME_SCHEME",
+ "MINT8_SCHEME",
+ "MINT8_SOURCE_SHA256",
+ "MINT_BENEFIT_STATISTICS",
+ "P10",
+ "P50",
+ "P90",
+ "PERCENTILE_DEFINITION",
+ "PERSON",
+ "POVERTY_STATISTICS",
+ "PROJECTION_STATISTICS",
+ "QUINTILE",
+ "QUINTILE_LEVELS",
+ "REGISTERED_REAL",
+ "REGISTRATION_POINTER",
+ "SCHEMA_VERSION",
+ "SHARE_STATISTICS",
+ "SMALL_CELL_MIN_N",
+ "SSA_DISCLOSURE_MIN_N",
+ "SSA_NUMERATOR_MIN",
+ "STATIC_BENEFIT_STATISTICS",
+ "STATISTIC_DEFINITIONS",
+ "TOTAL",
+ "UNCLASSIFIED",
+ "Category",
+ "CategoryScheme",
+ "Dimension",
+ "DimensionResult",
+ "GroupAssignment",
+ "GroupBreakdownError",
+ "GroupBreakdownResult",
+ "GroupCell",
+ "StatisticCell",
+ "assign_groups",
+ "classify_changes",
+ "derive_scheme",
+ "get_scheme",
+ "half_split_masks",
+ "register_scheme",
+ "registered_scheme_ids",
+ "tabulate_poverty_breakdown",
+ "tabulate_projection_breakdown",
+ "tabulate_share_breakdown",
+ "tabulate_static_benefit_breakdown",
+ "verify_mint8_sources",
+ "weighted_percentiles",
+ "weighted_quintile_ranks",
+ "with_age_bands",
+]
+
+SCHEMA_VERSION = "populace_dynamics.group_breakdown.v1"
+
+
+class GroupBreakdownError(ValueError):
+ """The scheme, the rows or the provenance cannot yield the breakdown."""
+
+
+# =========================================================================
+# Constants
+# =========================================================================
+TOTAL = "total"
+CATEGORICAL = "categorical"
+BAND = "band"
+QUINTILE = "quintile"
+DIMENSION_KINDS = (TOTAL, CATEGORICAL, BAND, QUINTILE)
+TOTAL_KEY = "total"
+TOTAL_LABEL = "Total"
+#: The label a row takes, in the long frame, in a dimension it is not
+#: classified in. No category label may equal it.
+UNCLASSIFIED = "unclassified"
+MISSING_REASON = "missing"
+OUTSIDE_BANDS_REASON = "outside_bands"
+
+INVENTED = a7.INVENTED
+REGISTERED_REAL = a7.REGISTERED_REAL
+DATA_PROVENANCES = (INVENTED, REGISTERED_REAL)
+#: The exercise-1 tabulation's invented-data label, reused verbatim.
+INVENTED_DATA_LABEL = a7.INVENTED_DATA_LABEL
+#: ``attrs["provenance_kind"]`` of rows built from staged PSID files (the
+#: Track U and Track M convention).
+PSID_FILES = "psid_files"
+#: An issue #42 comment, the registration issue (the pattern of
+#: ``min_benefit_track_m.tabulation.REGISTRATION_POINTER``, which a test
+#: holds equal).
+REGISTRATION_POINTER = re.compile(
+ r"https://github\.com/PolicyEngine/microcosm-dynamics/issues/42"
+ r"#issuecomment-[0-9]+"
+)
+
+#: SSA's disclosure minimum (MINT8 Table User Guide, "Sample Size
+#: Restrictions"), its numerator threshold, and the repository's small-cell
+#: line.
+SSA_DISCLOSURE_MIN_N = 100
+SSA_NUMERATOR_MIN = 10
+SMALL_CELL_MIN_N = 30
+
+P10 = Fraction(1, 10)
+P50 = Fraction(1, 2)
+P90 = Fraction(9, 10)
+QUINTILE_LEVELS = (
+ Fraction(1, 5),
+ Fraction(2, 5),
+ Fraction(3, 5),
+ Fraction(4, 5),
+)
+
+#: The floor seeds and fraction of every existing tabulator.
+DEFAULT_FLOOR_SEEDS = a7.DEFAULT_FLOOR_SEEDS
+FLOOR_FRACTION = a7.FLOOR_FRACTION
+MIN_FLOOR_SEEDS = a7.MIN_FLOOR_SEEDS
+FAMILY_UNIT_LINKED_BY_PERSON = "family_unit_linked_by_person"
+FAMILY_UNIT = "family_unit_id"
+PERSON = "person_id"
+STATIC_FLOOR_SPLIT_UNITS: dict[str, str] = {
+ FAMILY_UNIT_LINKED_BY_PERSON: (
+ "family units merged through shared persons "
+ "(uniform_cut_tabulation.floor_split_units; with one row per person "
+ "these are the family units), the Track U and Track M floor unit"
+ ),
+ FAMILY_UNIT: "the family_unit_id column as given",
+ PERSON: "the person_id column (not registered by any blind test)",
+}
+
+# ---- statistics -----------------------------------------------------------
+RATIO_OF_SCENARIO_MEANS = a7.RATIO_OF_SCENARIO_MEANS
+MEAN_OF_INDIVIDUAL_RATIOS = a7.MEAN_OF_INDIVIDUAL_RATIOS
+PERCENT_DECREASE = "percent_with_decrease"
+PERCENT_INCREASE = "percent_with_increase"
+PERCENT_UNAFFECTED = "percent_unaffected"
+CHANGE_P10 = "percent_change_p10"
+CHANGE_MEDIAN = "percent_change_median"
+CHANGE_P90 = "percent_change_p90"
+PERCENTILE_STATISTICS: dict[str, Fraction] = {
+ CHANGE_P10: P10,
+ CHANGE_MEDIAN: P50,
+ CHANGE_P90: P90,
+}
+MINT_BENEFIT_STATISTICS = (
+ PERCENT_DECREASE,
+ PERCENT_INCREASE,
+ PERCENT_UNAFFECTED,
+ CHANGE_P10,
+ CHANGE_MEDIAN,
+ CHANGE_P90,
+)
+PROJECTION_STATISTICS = (
+ RATIO_OF_SCENARIO_MEANS,
+ MEAN_OF_INDIVIDUAL_RATIOS,
+ *MINT_BENEFIT_STATISTICS,
+)
+STATIC_BENEFIT_STATISTICS = MINT_BENEFIT_STATISTICS
+
+POVERTY_RATE_CURRENT_LAW = "poverty_rate_current_law"
+POVERTY_RATE_PROPOSAL = "poverty_rate_proposal"
+POVERTY_RATE_CHANGE = "poverty_rate_change_pp"
+NUMBER_POOR_CURRENT_LAW = "number_in_poverty_current_law_thousands"
+NUMBER_POOR_PROPOSAL = "number_in_poverty_proposal_thousands"
+NUMBER_POOR_CHANGE = "number_in_poverty_change_thousands"
+NUMBER_POOR_PERCENT_CHANGE = "percent_change_number_in_poverty"
+POVERTY_STATISTICS = (
+ POVERTY_RATE_CURRENT_LAW,
+ POVERTY_RATE_PROPOSAL,
+ POVERTY_RATE_CHANGE,
+ NUMBER_POOR_CURRENT_LAW,
+ NUMBER_POOR_PROPOSAL,
+ NUMBER_POOR_CHANGE,
+ NUMBER_POOR_PERCENT_CHANGE,
+)
+SHARE = "weighted_share_percent"
+SHARE_STATISTICS = (SHARE,)
+
+PROJECTION_KIND = "projection_benefit_change"
+POVERTY_KIND = "static_poverty"
+SHARE_KIND = "static_share"
+STATIC_BENEFIT_KIND = "static_benefit_change"
+
+_MINT_POPULATION = (
+ "P = rows that are current-law (baseline) recipients with a positive "
+ "current-law amount B_current_law > 0 (MINT8: current-law "
+ "beneficiaries)"
+)
+PERCENTILE_DEFINITION = (
+ "weighted percentile at level p (exact rationals 1/10, 1/2, 9/10; "
+ "quintile thresholds at 1/5, 2/5, 3/5, 4/5): rows with zero weight "
+ "carry no mass and are dropped; the values are sorted ascending; with "
+ "W_k the cumulative weight of the k smallest values and W the total, "
+ "the percentile is the smallest value x_k with W_k >= p * W, except "
+ "that when W_k = p * W holds exactly (exact rational arithmetic on the "
+ "float64 weights) it is the midpoint (x_k + x_(k+1)) / 2 of that value "
+ "and the next one; no other interpolation"
+)
+CHANGE_DEFINITION = (
+ "individual percent change c_i = 100 * (B_option,i / B_current_law,i "
+ "- 1) over P (SSA's formula '(option amount - current-law amount) / "
+ "current-law amount', in percent); a decrease is c_i <= -1 and an "
+ "increase c_i >= +1, decided in exact rational arithmetic on the "
+ "float64 amounts (100 * B_option <= 99 * B_current_law; 100 * B_option "
+ ">= 101 * B_current_law), so a change of exactly 1 percent counts; "
+ "-1 < c_i < 1 is unaffected (MINT8 Table User Guide, 'Threshold for "
+ "Categorization in the Decrease or Increase Groups')"
+)
+QUINTILE_RULE = (
+ "the thresholds t_1..t_4 are the weighted percentiles (the percentile "
+ "rule) at 1/5, 2/5, 3/5 and 4/5 of the measure over the partition's "
+ "rows with a non-missing measure; a row's rank is 1 + the number of "
+ "thresholds strictly below its value, so a value equal to t_k falls in "
+ "the lower quintile; rank 1 is 'Lowest' and 5 'Highest'. Quintiles "
+ "are cut separately within each partition the caller names (for "
+ "example each draw, or each birth cohort, as MINT8's cohort tables "
+ "do). SSA does not publish its tie rule ('The dollar ranges are "
+ "available upon request'), so the tie rule is this module's"
+)
+UNCLASSIFIED_RULE = (
+ "a row whose attribute is missing, whose code the caller declares "
+ "unclassified, or whose band value lies outside every band is "
+ "unclassified in that dimension: counted by reason and entering Total "
+ "only; a present code no category holds and the caller did not "
+ "declare is refused"
+)
+
+#: Each statistic's definition and its MINT8 column (``None`` where it is
+#: not a MINT8 column).
+STATISTIC_DEFINITIONS: dict[str, dict[str, Any]] = {
+ RATIO_OF_SCENARIO_MEANS: {
+ "formula": a7.STATISTIC_DEFINITIONS[RATIO_OF_SCENARIO_MEANS],
+ "membership": (
+ "scenario memberships S_base and S_reform exactly as "
+ "cola_age_profile defines them for the ColaAgeProfileConfig "
+ "passed (recipient_rule, membership_basis)"
+ ),
+ "unit": "percent change of the reform relative to the baseline",
+ "mint8_column": None,
+ "note": "the statistic of exercises 1 and 3 (cola_age_profile)",
+ },
+ MEAN_OF_INDIVIDUAL_RATIOS: {
+ "formula": a7.STATISTIC_DEFINITIONS[MEAN_OF_INDIVIDUAL_RATIOS],
+ "membership": "S_alt as cola_age_profile defines it",
+ "unit": "percent",
+ "mint8_column": None,
+ "note": "cola_age_profile's registered alternative statistic",
+ },
+ PERCENT_DECREASE: {
+ "formula": "100 * sum_{i in P} w_i 1{c_i <= -1} / sum_{i in P} w_i",
+ "population": _MINT_POPULATION,
+ "change": CHANGE_DEFINITION,
+ "unit": "percent of the population",
+ "mint8_column": "Percent of population with a— Benefit decrease",
+ },
+ PERCENT_INCREASE: {
+ "formula": "100 * sum_{i in P} w_i 1{c_i >= 1} / sum_{i in P} w_i",
+ "population": _MINT_POPULATION,
+ "change": CHANGE_DEFINITION,
+ "unit": "percent of the population",
+ "mint8_column": "Percent of population with a— Benefit increase",
+ },
+ PERCENT_UNAFFECTED: {
+ "formula": (
+ "100 * sum_{i in P} w_i 1{-1 < c_i < 1} / sum_{i in P} w_i"
+ ),
+ "population": _MINT_POPULATION,
+ "change": CHANGE_DEFINITION,
+ "unit": "percent of the population",
+ "mint8_column": None,
+ "note": (
+ "not a MINT8 column; reported so that decrease + unaffected + "
+ "increase = 100 can be checked"
+ ),
+ },
+ **{
+ name: {
+ "formula": (
+ f"weighted percentile at level {level} of c_i over P "
+ "(percentile_definition)"
+ ),
+ "population": _MINT_POPULATION,
+ "change": CHANGE_DEFINITION,
+ "unit": "percent change",
+ "mint8_column": (
+ "Percent change in Social Security benefits at the— " + label
+ ),
+ }
+ for name, level, label in (
+ (CHANGE_P10, P10, "10th %ile"),
+ (CHANGE_MEDIAN, P50, "Median"),
+ (CHANGE_P90, P90, "90th %ile"),
+ )
+ },
+ POVERTY_RATE_CURRENT_LAW: {
+ "formula": "100 * sum w 1{poor_baseline} / sum w",
+ "unit": "percent",
+ "mint8_column": "Official poverty rate: Under current law",
+ "computed_by": "uniform_cut_tabulation._rates (baseline_rate)",
+ },
+ POVERTY_RATE_PROPOSAL: {
+ "formula": "100 * sum w 1{poor_reform} / sum w",
+ "unit": "percent",
+ "mint8_column": "Official poverty rate: With proposal",
+ "computed_by": "uniform_cut_tabulation._rates (reform_rate)",
+ },
+ POVERTY_RATE_CHANGE: {
+ "formula": "poverty_rate_proposal - poverty_rate_current_law",
+ "unit": "percentage points",
+ "mint8_column": None,
+ "note": (
+ "not a MINT8 column; the exercise-2 headline "
+ "(uniform_cut_tabulation 'delta')"
+ ),
+ },
+ NUMBER_POOR_CURRENT_LAW: {
+ "formula": "sum w 1{poor_baseline} / 1000",
+ "unit": "thousands of persons (weighted)",
+ "mint8_column": (
+ "Number of population in poverty (in thousands): Under current "
+ "law"
+ ),
+ },
+ NUMBER_POOR_PROPOSAL: {
+ "formula": "sum w 1{poor_reform} / 1000",
+ "unit": "thousands of persons (weighted)",
+ "mint8_column": (
+ "Number of population in poverty (in thousands): With proposal"
+ ),
+ },
+ NUMBER_POOR_CHANGE: {
+ "formula": "(sum w 1{poor_reform} - sum w 1{poor_baseline}) / 1000",
+ "unit": "thousands of persons (weighted)",
+ "mint8_column": (
+ "Number of population in poverty (in thousands): Change"
+ ),
+ },
+ NUMBER_POOR_PERCENT_CHANGE: {
+ "formula": (
+ "100 * (sum w 1{poor_reform} - sum w 1{poor_baseline}) / sum w "
+ "1{poor_baseline}; undefined when nobody is poor under current "
+ "law"
+ ),
+ "unit": "percent change in the number in poverty",
+ "mint8_column": "Percent change in the number in poverty",
+ "source_quote": (
+ "This percent change is calculated by dividing the change in "
+ "thousands in poverty (5th column) by the thousands in poverty "
+ "without the proposal (3rd column)."
+ ),
+ },
+ SHARE: {
+ "formula": "100 * sum w 1{indicator} / sum w",
+ "unit": "percent",
+ "mint8_column": None,
+ "note": (
+ "exercise 4's statistic (min_benefit_track_m.tabulation._share)"
+ ),
+ },
+}
+
+
+# =========================================================================
+# Category schemes
+# =========================================================================
+_KEY_FORM = re.compile(r"[a-z][a-z0-9_]*")
+
+
+def _nonempty_text(value: Any, label: str) -> str:
+ if not isinstance(value, str) or not value.strip():
+ raise GroupBreakdownError(f"{label} must be a non-empty string")
+ return value
+
+
+def _key(value: Any, label: str) -> str:
+ if not isinstance(value, str) or not _KEY_FORM.fullmatch(value):
+ raise GroupBreakdownError(
+ f"{label} must be lower-case letters, digits and '_' starting "
+ f"with a letter; got {value!r}"
+ )
+ return value
+
+
+def _optional_int(value: Any, label: str) -> int | None:
+ if value is None:
+ return None
+ if isinstance(value, bool | np.bool_) or not isinstance(value, Integral):
+ raise GroupBreakdownError(f"{label} must be an integer or None")
+ return int(value)
+
+
+def _string_tuple(values: Any, label: str) -> tuple[str, ...]:
+ if isinstance(values, str):
+ raise GroupBreakdownError(f"{label} must be a sequence of strings")
+ items = tuple(values)
+ for item in items:
+ _nonempty_text(item, label)
+ return items
+
+
+@dataclass(frozen=True)
+class Category:
+ """One table row: its stable ``key``, printed ``label`` and rule.
+
+ ``codes`` (categorical), ``lower``/``upper`` (band; inclusive integer
+ bounds, ``None`` open) or ``rank`` (quintile; 1 lowest, 5 highest)
+ say which rows it holds; the dimension's kind decides which applies.
+ ``definition`` is a verbatim quote of the dimension's source, or empty.
+ """
+
+ key: str
+ label: str
+ codes: tuple[str, ...] = ()
+ lower: int | None = None
+ upper: int | None = None
+ rank: int | None = None
+ definition: str = ""
+
+ def __post_init__(self) -> None:
+ _key(self.key, "category key")
+ _nonempty_text(self.label, f"category {self.key} label")
+ if self.label == UNCLASSIFIED:
+ raise GroupBreakdownError(
+ f"a category label may not be {UNCLASSIFIED!r}"
+ )
+ codes = _string_tuple(self.codes, f"category {self.key} codes")
+ if len(set(codes)) != len(codes):
+ raise GroupBreakdownError(f"category {self.key} repeats a code")
+ object.__setattr__(self, "codes", codes)
+ object.__setattr__(
+ self, "lower", _optional_int(self.lower, f"{self.key} lower")
+ )
+ object.__setattr__(
+ self, "upper", _optional_int(self.upper, f"{self.key} upper")
+ )
+ object.__setattr__(
+ self, "rank", _optional_int(self.rank, f"{self.key} rank")
+ )
+ if not isinstance(self.definition, str):
+ raise GroupBreakdownError(
+ f"category {self.key} definition must be a string"
+ )
+
+ def contains(self, value: int) -> bool:
+ """Whether an integer lies in this band (inclusive bounds)."""
+
+ return (self.lower is None or value >= self.lower) and (
+ self.upper is None or value <= self.upper
+ )
+
+ def as_dict(self) -> dict[str, Any]:
+ return {
+ "key": self.key,
+ "label": self.label,
+ "codes": list(self.codes),
+ "lower": self.lower,
+ "upper": self.upper,
+ "rank": self.rank,
+ "definition": self.definition,
+ }
+
+
+#: The identifier of a quote source that :func:`verify_mint8_sources`
+#: checks verbatim.
+MINT8_GUIDE = "mint8_table_user_guide"
+
+
+@dataclass(frozen=True)
+class Dimension:
+ """One row group (characteristic subgroup) of a table.
+
+ ``definition`` holds verbatim quotes of ``source`` (checked against the
+ committed capture when ``quote_source`` is :data:`MINT8_GUIDE`);
+ ``notes`` hold this module's conventions, never presented as quotes.
+ """
+
+ key: str
+ label: str
+ kind: str
+ categories: tuple[Category, ...]
+ definition: tuple[str, ...] = ()
+ source: str = ""
+ quote_source: str | None = None
+ notes: tuple[str, ...] = ()
+
+ def __post_init__(self) -> None:
+ _key(self.key, "dimension key")
+ _nonempty_text(self.label, f"dimension {self.key} label")
+ if self.kind not in DIMENSION_KINDS:
+ raise GroupBreakdownError(
+ f"dimension {self.key} kind must be one of "
+ f"{list(DIMENSION_KINDS)}"
+ )
+ categories = tuple(self.categories)
+ if not categories or not all(
+ isinstance(c, Category) for c in categories
+ ):
+ raise GroupBreakdownError(
+ f"dimension {self.key} needs a non-empty tuple of Category"
+ )
+ object.__setattr__(self, "categories", categories)
+ keys = [c.key for c in categories]
+ labels = [c.label for c in categories]
+ if len(set(keys)) != len(keys) or len(set(labels)) != len(labels):
+ raise GroupBreakdownError(
+ f"dimension {self.key} repeats a category key or label"
+ )
+ object.__setattr__(
+ self,
+ "definition",
+ _string_tuple(self.definition, f"{self.key} definition"),
+ )
+ object.__setattr__(
+ self, "notes", _string_tuple(self.notes, f"{self.key} notes")
+ )
+ if not isinstance(self.source, str):
+ raise GroupBreakdownError(f"{self.key} source must be a string")
+ if self.quote_source is not None and self.quote_source != (
+ MINT8_GUIDE
+ ):
+ raise GroupBreakdownError(
+ f"{self.key} quote_source must be None or {MINT8_GUIDE!r}"
+ )
+ getattr(self, f"_validate_{self.kind}")(categories)
+
+ def _no_rule(self, category: Category, *, allow: str = "") -> None:
+ fields = {
+ "codes": bool(category.codes),
+ "bounds": category.lower is not None or category.upper is not None,
+ "rank": category.rank is not None,
+ }
+ extra = [
+ name for name, used in fields.items() if used and name != allow
+ ]
+ if extra:
+ raise GroupBreakdownError(
+ f"{self.kind} dimension {self.key}: category "
+ f"{category.key} may not carry {extra}"
+ )
+
+ def _validate_total(self, categories: tuple[Category, ...]) -> None:
+ if self.key != TOTAL_KEY or len(categories) != 1:
+ raise GroupBreakdownError(
+ f"the total dimension has key {TOTAL_KEY!r} and one category"
+ )
+ if categories[0].key != TOTAL_KEY:
+ raise GroupBreakdownError(
+ f"the total category has key {TOTAL_KEY!r}"
+ )
+ self._no_rule(categories[0])
+
+ def _validate_categorical(self, categories: tuple[Category, ...]) -> None:
+ seen: set[str] = set()
+ for category in categories:
+ self._no_rule(category, allow="codes")
+ if not category.codes:
+ raise GroupBreakdownError(
+ f"categorical dimension {self.key}: category "
+ f"{category.key} needs codes"
+ )
+ overlap = seen.intersection(category.codes)
+ if overlap:
+ raise GroupBreakdownError(
+ f"dimension {self.key}: codes {sorted(overlap)} map to "
+ "two categories"
+ )
+ seen.update(category.codes)
+
+ def _validate_band(self, categories: tuple[Category, ...]) -> None:
+ for category in categories:
+ self._no_rule(category, allow="bounds")
+ if category.lower is None and category.upper is None:
+ raise GroupBreakdownError(
+ f"band {category.key} needs at least one bound"
+ )
+ if (
+ category.lower is not None
+ and category.upper is not None
+ and category.upper < category.lower
+ ):
+ raise GroupBreakdownError(
+ f"band {category.key} has upper < lower"
+ )
+ for a, b in itertools.combinations(categories, 2):
+ low = max(
+ -math.inf if a.lower is None else a.lower,
+ -math.inf if b.lower is None else b.lower,
+ )
+ high = min(
+ math.inf if a.upper is None else a.upper,
+ math.inf if b.upper is None else b.upper,
+ )
+ if low <= high:
+ raise GroupBreakdownError(
+ f"dimension {self.key}: bands {a.key} and {b.key} overlap"
+ )
+
+ def _validate_quintile(self, categories: tuple[Category, ...]) -> None:
+ for category in categories:
+ self._no_rule(category, allow="rank")
+ ranks = sorted(c.rank for c in categories if c.rank is not None)
+ if ranks != [1, 2, 3, 4, 5]:
+ raise GroupBreakdownError(
+ f"quintile dimension {self.key} needs ranks 1-5 once each"
+ )
+
+ @property
+ def labels(self) -> tuple[str, ...]:
+ return tuple(c.label for c in self.categories)
+
+ def category(self, key: str) -> Category:
+ for category in self.categories:
+ if category.key == key:
+ return category
+ raise GroupBreakdownError(
+ f"dimension {self.key} has no category {key!r}"
+ )
+
+ def as_dict(self) -> dict[str, Any]:
+ return {
+ "key": self.key,
+ "label": self.label,
+ "kind": self.kind,
+ "categories": [c.as_dict() for c in self.categories],
+ "definition": list(self.definition),
+ "source": self.source,
+ "quote_source": self.quote_source,
+ "notes": list(self.notes),
+ }
+
+
+@dataclass(frozen=True)
+class CategoryScheme:
+ """An ordered set of dimensions with its population and citations.
+
+ The first dimension is the total. ``composite`` marks a scheme this
+ module composed (it is not a published table layout).
+ """
+
+ scheme_id: str
+ title: str
+ population: str
+ dimensions: tuple[Dimension, ...]
+ citations: tuple[str, ...]
+ notes: tuple[str, ...] = ()
+ composite: bool = False
+
+ def __post_init__(self) -> None:
+ _key(self.scheme_id, "scheme_id")
+ _nonempty_text(self.title, "scheme title")
+ _nonempty_text(self.population, "scheme population")
+ dimensions = tuple(self.dimensions)
+ if not dimensions or not all(
+ isinstance(d, Dimension) for d in dimensions
+ ):
+ raise GroupBreakdownError(
+ "a scheme needs a non-empty tuple of Dimension"
+ )
+ if dimensions[0].kind != TOTAL:
+ raise GroupBreakdownError("a scheme's first dimension is Total")
+ if sum(d.kind == TOTAL for d in dimensions) != 1:
+ raise GroupBreakdownError("a scheme has exactly one total")
+ keys = [d.key for d in dimensions]
+ if len(set(keys)) != len(keys):
+ raise GroupBreakdownError("a scheme repeats a dimension key")
+ object.__setattr__(self, "dimensions", dimensions)
+ citations = _string_tuple(self.citations, "scheme citations")
+ if not citations:
+ raise GroupBreakdownError("a scheme needs at least one citation")
+ object.__setattr__(self, "citations", citations)
+ object.__setattr__(
+ self, "notes", _string_tuple(self.notes, "scheme notes")
+ )
+ if not isinstance(self.composite, bool):
+ raise GroupBreakdownError("composite must be a bool")
+
+ @property
+ def keys(self) -> tuple[str, ...]:
+ return tuple(d.key for d in self.dimensions)
+
+ def dimension(self, key: str) -> Dimension:
+ for dimension in self.dimensions:
+ if dimension.key == key:
+ return dimension
+ raise GroupBreakdownError(
+ f"scheme {self.scheme_id} has no dimension {key!r}"
+ )
+
+ def as_dict(self) -> dict[str, Any]:
+ return {
+ "scheme_id": self.scheme_id,
+ "title": self.title,
+ "population": self.population,
+ "composite": self.composite,
+ "citations": list(self.citations),
+ "notes": list(self.notes),
+ "dimensions": [d.as_dict() for d in self.dimensions],
+ }
+
+
+# ---- MINT8 sources ---------------------------------------------------------
+_PROJECT_ROOT = Path(__file__).resolve().parents[3]
+MINT8_GUIDE_FILE = "data/external/mint8_table_user_guide.source.html"
+MINT8_LABELS_FILE = "data/external/mint8_row_categories.json"
+#: SHA-256 pins of the committed MINT8 sources (see each file's
+#: ``.provenance.json``).
+MINT8_SOURCE_SHA256: dict[str, str] = {
+ MINT8_GUIDE_FILE: (
+ "278d5d19c1b50f1d354db1ada515288af563c035a67eb16fb971d25700fb94e9"
+ ),
+ MINT8_LABELS_FILE: (
+ "23fbfbc8dbc14b06144a83bf0583ea0a6c89505638f3f7dc5fd6b34a3a4ac650"
+ ),
+}
+_GUIDE_URL = "https://www.ssa.gov/policy/docs/projections/user-guide.html"
+_OPTION_URL = (
+ "https://www.ssa.gov/policy/docs/projections/policy-options/"
+ "increase-payroll-tax-rate.html"
+)
+MINT8_GUIDE_CITATION = (
+ "Social Security Administration, 'Table User Guide—Modeling Income "
+ f"in the Near Term (MINT) 8', {_GUIDE_URL}, section 'Definitions—"
+ "Table Rows and Columns' > 'Characteristic Subgroups—Table Rows' "
+ "(DCTERMS:dateCertified 2025-10-01); committed capture "
+ f"{MINT8_GUIDE_FILE} (SHA-256 {MINT8_SOURCE_SHA256[MINT8_GUIDE_FILE]}), "
+ "Internet Archive capture 20260419231110"
+)
+
+
+def _row_label_citation(tables: str) -> str:
+ return (
+ "Social Security Administration, MINT8 projected-effects tables on "
+ f"the policy-option page {_OPTION_URL} (DCTERMS:dateCertified "
+ f"2026-04-01), row-group headings and row labels of tables {tables} "
+ "in document order, labels only (raw page SHA-256 "
+ "3cfb6eb636cad6d062eb22734c533fb349a0bc4761014eee8f2ccb0eb6e28d8c, "
+ "Internet Archive capture 20260519190449); committed extract "
+ f"{MINT8_LABELS_FILE} (SHA-256 "
+ f"{MINT8_SOURCE_SHA256[MINT8_LABELS_FILE]})"
+ )
+
+
+_GUIDE_SECTION = (
+ "MINT8 Table User Guide, 'Definitions—Table Rows and Columns' > "
+ "'Characteristic Subgroups—Table Rows'"
+)
+
+# Verbatim quotes of the MINT8 Table User Guide (checked by
+# verify_mint8_sources).
+_Q_TOTAL = "Total: Refers to the total population of the table."
+_Q_SEX = "Sex: Female or Male."
+_Q_RACE = (
+ "Race/Ethnicity: We list “Hispanic or Latino, any race” first; "
+ "the rest of the groups (White, Black or African American, and All "
+ "other races) are non-Hispanic. Additional racial or ethnic "
+ "identifications are not covered because they are not in the datasets "
+ "used to build the MINT8 model."
+)
+_Q_COUNTRY = (
+ "Country of Birth: We differentiate between the United States and "
+ "other countries."
+)
+_Q_AGE_BENEFICIARY = (
+ "The beneficiary population includes those aged 60 or older because 60 "
+ "is the earliest eligibility age for any aged benefits under current "
+ "law."
+)
+_Q_AGE_TAXPAYER = (
+ "The taxpayer population includes those aged 31 or older because 31 is "
+ "the earliest age in MINT for household income and poverty "
+ "information."
+)
+_Q_MARITAL = (
+ "Marital Status: Refers to the marital status in the year of analysis "
+ "only. An individual's marital status can change in the future and may "
+ "have been different in the past."
+)
+_Q_EDUCATION = (
+ "Highest Education Level: Reported number of years of education."
+)
+_Q_POVERTY = (
+ "Current-Law Poverty Status: Indicates whether the person is in a "
+ "household that has income above (“above poverty”) or below "
+ "(“in poverty”) the official poverty line under current law. "
+ "The household income used for the official poverty measure is the "
+ "same as the household income used in our results except for how asset "
+ "income is counted."
+)
+_Q_INCOME = (
+ "Current-Law Household Income Quintile: Represents an individual's "
+ "annual household income under current law, including: household "
+ "earnings; asset income (annuitized), which includes income from "
+ "defined contribution plans (such as 401(k) accounts) and personal "
+ "savings; defined benefit pensions; means-tested income; "
+ "non-means-tested income; Social Security; Supplemental Security "
+ "Income; and non-spousal co-residents' income."
+)
+_Q_INCOME_QUINTILES = (
+ "We calculate the income quintiles for each year (e.g., 2030, 2050, or "
+ "2070) for the population analyzed, determine the dollar thresholds for "
+ "each income quintile, and assign each beneficiary to the appropriate "
+ "quintile. The dollar ranges are available upon request."
+)
+_Q_BENEFIT_TYPE = (
+ "Current-Law Benefit Type: Some Social Security benefits are based on "
+ "one's own work, while others are based on the work of a current, "
+ "divorced, or deceased spouse. The current-law benefit type refers to "
+ "one of the following benefit types received in the specified analysis "
+ "year:"
+)
+_Q_BENEFIT_TYPE_NOTE = (
+ "Our results do not show different benefit types a beneficiary might "
+ "receive under a policy option/proposal or in a different year under "
+ "current law."
+)
+_Q_AIME = (
+ "Current-Law Initial AIME Quintile: Represents an individual's average "
+ "indexed monthly earnings (AIME) under current law at age 62, the "
+ "earliest eligibility age for retired-worker benefits. We calculate the "
+ "AIME quintiles for each birth cohort. The dollar ranges are available "
+ "upon request."
+)
+_Q_PAYROLL = (
+ "Lifetime Payroll Tax Quintile: Represents the present value of an "
+ "individual's current-law payroll taxes at age 62. We calculate the "
+ "payroll tax quintiles for each birth cohort. The dollar ranges are "
+ "available upon request."
+)
+_Q_PAYROLL_SHARED = (
+ "Lifetime Payroll Tax Quintile (Shared): Represents the present value "
+ "of an individual's current-law payroll taxes at age 62. For married "
+ "couples, the payroll taxes paid while married are shared equally "
+ "between them. For never-married individuals, this is the same as the "
+ "lifetime payroll tax. In any year where an individual is not married, "
+ "we count only their individual payroll taxes. We calculate the "
+ "quintiles for each birth cohort. The dollar ranges are available upon "
+ "request."
+)
+_Q_PRESENT_VALUE = (
+ "We use the Social Security Trust Fund interest rate to adjust benefits "
+ "and taxes to their present values at age 62."
+)
+_Q_COHORT_FOOTNOTE = (
+ "Birth cohort tables (benefit/tax ratios and initial replacement rates) "
+ "do not have age or marital status breakouts."
+)
+_Q_DISCLOSURE = (
+ "To maintain the privacy of survey respondents, our tables have "
+ "built-in disclosure avoidance protections that suppress an entire "
+ "characteristic subgroup if the sample size for any row in that "
+ "subgroup is less than 100 individuals."
+)
+_Q_NUMERATOR = (
+ "The numerator must have either zero cases or meet a minimum numerator "
+ "threshold of 10 for the table to display it. The table would suppress "
+ "any characteristic subgroup that has a percentage based on a numerator "
+ "of 1–9 cases."
+)
+_Q_THRESHOLD = (
+ "We categorize individuals as having a “decrease” in the amount "
+ "being analyzed (benefits, taxes, income, etc.) when a proposal would "
+ "reduce the analyzed quantity by 1% or more. Individuals are "
+ "categorized as having an “increase” when a proposal would "
+ "raise the analyzed quantity by 1% or more. We consider individuals "
+ "with differences between −1% and 1% to be unaffected."
+)
+_Q_FORMULA = "(option amount − current-law amount) ÷ current-law amount"
+_Q_POVERTY_PERCENT = STATISTIC_DEFINITIONS[NUMBER_POOR_PERCENT_CHANGE][
+ "source_quote"
+]
+#: Quotes outside the dimensions that this module relies on.
+MINT8_GUIDE_QUOTES: tuple[str, ...] = (
+ _Q_PRESENT_VALUE,
+ _Q_COHORT_FOOTNOTE,
+ _Q_DISCLOSURE,
+ _Q_NUMERATOR,
+ _Q_THRESHOLD,
+ _Q_FORMULA,
+ _Q_POVERTY_PERCENT,
+)
+
+
+def _quintile_categories() -> tuple[Category, ...]:
+ """MINT8's quintile rows in SSA's page order (Highest first)."""
+
+ return (
+ Category("highest", "Highest", rank=5),
+ Category("second_highest", "Second highest", rank=4),
+ Category("middle", "Middle", rank=3),
+ Category("second_lowest", "Second lowest", rank=2),
+ Category("lowest", "Lowest", rank=1),
+ )
+
+
+def _mint8(key: str, label: str, kind: str, categories, *quotes, notes=()):
+ return Dimension(
+ key=key,
+ label=label,
+ kind=kind,
+ categories=tuple(categories),
+ definition=tuple(quotes),
+ source=_GUIDE_SECTION,
+ quote_source=MINT8_GUIDE,
+ notes=tuple(notes),
+ )
+
+
+TOTAL_DIMENSION = _mint8(
+ TOTAL_KEY,
+ TOTAL_LABEL,
+ TOTAL,
+ (Category(TOTAL_KEY, TOTAL_LABEL),),
+ _Q_TOTAL,
+)
+SEX_DIMENSION = _mint8(
+ "sex",
+ "Sex",
+ CATEGORICAL,
+ (
+ Category("female", "Female", codes=("female",)),
+ Category("male", "Male", codes=("male",)),
+ ),
+ _Q_SEX,
+)
+RACE_ETHNICITY_DIMENSION = _mint8(
+ "race_ethnicity",
+ "Race and ethnicity",
+ CATEGORICAL,
+ (
+ Category(
+ "hispanic_any_race",
+ "Hispanic or Latino, any race",
+ codes=("hispanic_any_race",),
+ ),
+ Category(
+ "white_non_hispanic",
+ "White, non-Hispanic",
+ codes=("white_non_hispanic",),
+ ),
+ Category(
+ "black_non_hispanic",
+ "Black or African American, non-Hispanic",
+ codes=("black_non_hispanic",),
+ ),
+ Category(
+ "other_non_hispanic",
+ "All other races, non-Hispanic",
+ codes=("other_non_hispanic",),
+ ),
+ ),
+ _Q_RACE,
+)
+COUNTRY_OF_BIRTH_DIMENSION = _mint8(
+ "country_of_birth",
+ "Country of birth",
+ CATEGORICAL,
+ (
+ Category("united_states", "United States", codes=("united_states",)),
+ Category(
+ "other_countries", "Other countries", codes=("other_countries",)
+ ),
+ ),
+ _Q_COUNTRY,
+)
+_AGE_NOTES = (
+ "age is the integer in the caller's age column; the G3 brief defines "
+ "it as age in the analysis year (the Guide states each population's "
+ "age floor, not the date at which age is measured)",
+ "an age outside every band is unclassified for this dimension and "
+ "still enters Total",
+)
+MINT8_BENEFICIARY_AGE = _mint8(
+ "age",
+ "Age",
+ BAND,
+ (
+ Category("age_60_69", "60–69", lower=60, upper=69),
+ Category("age_70_79", "70–79", lower=70, upper=79),
+ Category("age_80_89", "80–89", lower=80, upper=89),
+ Category("age_90_plus", "90 or older", lower=90),
+ ),
+ _Q_AGE_BENEFICIARY,
+ notes=_AGE_NOTES,
+)
+MINT8_TAXPAYER_AGE = _mint8(
+ "age",
+ "Age",
+ BAND,
+ (
+ Category("age_31_39", "31–39", lower=31, upper=39),
+ Category("age_40_49", "40–49", lower=40, upper=49),
+ Category("age_50_59", "50–59", lower=50, upper=59),
+ Category("age_60_69", "60–69", lower=60, upper=69),
+ Category("age_70_plus", "70 or older", lower=70),
+ ),
+ _Q_AGE_TAXPAYER,
+ notes=_AGE_NOTES,
+)
+
+
+def _a7_band_key(label: str) -> str:
+ return "age_" + label.replace("+", "_plus").replace("-", "_")
+
+
+#: Exercises 1 and 3's registered age groups, from cola_age_profile.
+DYNASIM3_EXERCISES_1_3_AGE = Dimension(
+ key="age",
+ label="Age",
+ kind=BAND,
+ categories=tuple(
+ Category(
+ _a7_band_key(group.label),
+ group.label,
+ lower=group.lower,
+ upper=group.upper,
+ )
+ for group in a7.DEFAULT_AGE_GROUPS
+ ),
+ source=(
+ "estimates/cola_age_profile.py DEFAULT_AGE_GROUPS: the registered "
+ "age groups of exercise 1 (A1 section 9) and exercise 3 (E1), with "
+ "age = 2030 - birth year"
+ ),
+ notes=(
+ "age is the integer in the caller's age column; to reproduce the "
+ "registered groups pass reference_year - birth_year (the "
+ "cola_age_profile age rule)",
+ "an age under 50 is unclassified for this dimension and still "
+ "enters Total",
+ ),
+)
+MARITAL_STATUS_DIMENSION = _mint8(
+ "marital_status",
+ "Marital status",
+ CATEGORICAL,
+ (
+ Category("married", "Married", codes=("married",)),
+ Category("divorced", "Divorced", codes=("divorced",)),
+ Category("widowed", "Widowed", codes=("widowed",)),
+ Category("never_married", "Never married", codes=("never_married",)),
+ ),
+ _Q_MARITAL,
+ notes=(
+ "SSA does not say where separated people go: a 'separated' code is "
+ "refused unless the caller maps it to a category upstream or "
+ "declares it unclassified (cohorts/psid2010.marital_state_at counts "
+ "separated as married by default, separated_is_married=True)",
+ ),
+)
+EDUCATION_DIMENSION = _mint8(
+ "education",
+ "Highest education level",
+ BAND,
+ (
+ Category(
+ "graduate",
+ "Graduate",
+ lower=17,
+ definition="Graduate means more than 16 years of education",
+ ),
+ Category(
+ "bachelor",
+ "Bachelor",
+ lower=16,
+ upper=16,
+ definition="Bachelor means 16 years of education",
+ ),
+ Category(
+ "associate",
+ "Associate",
+ lower=14,
+ upper=15,
+ definition="Associate means 14–15 years of education",
+ ),
+ Category(
+ "high_school",
+ "High school",
+ lower=12,
+ upper=13,
+ definition="High school means 12–13 years of education",
+ ),
+ Category(
+ "less_than_high_school",
+ "Less than high school",
+ upper=11,
+ definition=(
+ "Less than high school means less than 12 years of education"
+ ),
+ ),
+ ),
+ _Q_EDUCATION,
+ notes=(
+ "years of education must be integers; a non-integer value is "
+ "refused",
+ ),
+)
+POVERTY_STATUS_DIMENSION = _mint8(
+ "poverty_status",
+ "Current-law poverty status",
+ CATEGORICAL,
+ (
+ Category("above_poverty", "Above poverty", codes=("above_poverty",)),
+ Category("in_poverty", "In poverty", codes=("in_poverty",)),
+ ),
+ _Q_POVERTY,
+)
+HOUSEHOLD_INCOME_QUINTILE_DIMENSION = _mint8(
+ "household_income_quintile",
+ "Current-law household income quintile",
+ QUINTILE,
+ _quintile_categories(),
+ _Q_INCOME,
+ _Q_INCOME_QUINTILES,
+ notes=(QUINTILE_RULE,),
+)
+BENEFIT_TYPE_DIMENSION = _mint8(
+ "benefit_type",
+ "Current-law benefit type",
+ CATEGORICAL,
+ (
+ Category(
+ "retired_worker_only",
+ "Retired worker only",
+ codes=("retired_worker_only",),
+ definition=(
+ "Retired-worker only: receives only a retired-worker benefit "
+ "based on his or her earnings record."
+ ),
+ ),
+ Category(
+ "widower",
+ "Widow(er) (includes dually entitled)",
+ codes=("widower",),
+ definition=(
+ "Widow(er) (includes dually entitled): receives a survivor "
+ "benefit (may or may not also receive a lower worker benefit "
+ "from his or her own earnings record, known as dually "
+ "entitled)."
+ ),
+ ),
+ Category(
+ "spousal",
+ "Spousal (includes dually entitled)",
+ codes=("spousal",),
+ definition=(
+ "Spousal (includes dually entitled): receives a spousal "
+ "benefit (may or may not also receive a lower worker benefit "
+ "from his or her own earnings record, known as dually "
+ "entitled)."
+ ),
+ ),
+ Category(
+ "disabled_worker_only",
+ "Disabled worker only",
+ codes=("disabled_worker_only",),
+ definition=(
+ "Disabled-worker only: receives a disabled-worker benefit on "
+ "his or her earnings record and is under the full retirement "
+ "age (FRA). Disabled workers convert to retired workers at "
+ "FRA."
+ ),
+ ),
+ ),
+ _Q_BENEFIT_TYPE,
+ _Q_BENEFIT_TYPE_NOTE,
+)
+_LIFETIME_NOTES = (
+ QUINTILE_RULE,
+ "the measure itself (AIME at 62, or the present value at 62 of "
+ "current-law payroll taxes, own or shared) is computed upstream by "
+ "cohort code; this module only cuts its quintiles",
+)
+INITIAL_AIME_QUINTILE_DIMENSION = _mint8(
+ "initial_aime_quintile",
+ "Current-law initial AIME quintile",
+ QUINTILE,
+ _quintile_categories(),
+ _Q_AIME,
+ notes=_LIFETIME_NOTES,
+)
+LIFETIME_PAYROLL_TAX_QUINTILE_DIMENSION = _mint8(
+ "lifetime_payroll_tax_quintile",
+ "Lifetime payroll tax quintile",
+ QUINTILE,
+ _quintile_categories(),
+ _Q_PAYROLL,
+ _Q_PRESENT_VALUE,
+ notes=_LIFETIME_NOTES,
+)
+LIFETIME_PAYROLL_TAX_QUINTILE_SHARED_DIMENSION = _mint8(
+ "lifetime_payroll_tax_quintile_shared",
+ "Lifetime payroll tax quintile (shared)",
+ QUINTILE,
+ _quintile_categories(),
+ _Q_PAYROLL_SHARED,
+ _Q_PRESENT_VALUE,
+ notes=_LIFETIME_NOTES,
+)
+LIFETIME_DIMENSIONS = (
+ INITIAL_AIME_QUINTILE_DIMENSION,
+ LIFETIME_PAYROLL_TAX_QUINTILE_DIMENSION,
+ LIFETIME_PAYROLL_TAX_QUINTILE_SHARED_DIMENSION,
+)
+
+#: Age-band dimensions a scheme's ``age`` dimension can be swapped for.
+AGE_BAND_SETS: dict[str, Dimension] = {
+ "mint8_beneficiary": MINT8_BENEFICIARY_AGE,
+ "mint8_taxpayer": MINT8_TAXPAYER_AGE,
+ "dynasim3_exercises_1_3": DYNASIM3_EXERCISES_1_3_AGE,
+}
+
+_BENEFICIARY_POPULATION = (
+ "MINT8: current-law beneficiaries aged 60 or older in the analysis year "
+ "(2030, 2050, 2070); a blind test's own population is the caller's rows"
+)
+MINT8_SCHEME = CategoryScheme(
+ scheme_id="mint8_beneficiary_annual",
+ title=(
+ "MINT8 projected effects on Social Security benefits and household "
+ "income (tables 1-3 and 7-9)"
+ ),
+ population=_BENEFICIARY_POPULATION,
+ dimensions=(
+ TOTAL_DIMENSION,
+ SEX_DIMENSION,
+ RACE_ETHNICITY_DIMENSION,
+ COUNTRY_OF_BIRTH_DIMENSION,
+ MINT8_BENEFICIARY_AGE,
+ MARITAL_STATUS_DIMENSION,
+ EDUCATION_DIMENSION,
+ POVERTY_STATUS_DIMENSION,
+ HOUSEHOLD_INCOME_QUINTILE_DIMENSION,
+ BENEFIT_TYPE_DIMENSION,
+ ),
+ citations=(MINT8_GUIDE_CITATION, _row_label_citation("1-3 and 7-9")),
+)
+MINT8_POVERTY_SCHEME = CategoryScheme(
+ scheme_id="mint8_beneficiary_poverty",
+ title=(
+ "MINT8 projected effects on the official poverty measure (tables "
+ "10-12)"
+ ),
+ population=_BENEFICIARY_POPULATION,
+ dimensions=tuple(
+ d
+ for d in MINT8_SCHEME.dimensions
+ if d.key != HOUSEHOLD_INCOME_QUINTILE_DIMENSION.key
+ ),
+ citations=(MINT8_GUIDE_CITATION, _row_label_citation("10-12")),
+ notes=(
+ "MINT8's official-poverty tables carry no household income "
+ "quintile rows (row labels of tables 10-12)",
+ ),
+)
+MINT8_COHORT_SCHEME = CategoryScheme(
+ scheme_id="mint8_cohort",
+ title=(
+ "MINT8 projected effects on benefit/tax ratios and initial "
+ "replacement rates (tables 13-20)"
+ ),
+ population=(
+ "MINT8: workers with a benefit/tax ratio, or current-law "
+ "beneficiaries with an initial replacement rate, born 1960–1969, "
+ "1980–1989, 2000–2009 or 2020–2029"
+ ),
+ dimensions=(
+ TOTAL_DIMENSION,
+ SEX_DIMENSION,
+ RACE_ETHNICITY_DIMENSION,
+ COUNTRY_OF_BIRTH_DIMENSION,
+ EDUCATION_DIMENSION,
+ *LIFETIME_DIMENSIONS,
+ ),
+ citations=(MINT8_GUIDE_CITATION, _row_label_citation("13-20")),
+ notes=(_Q_COHORT_FOOTNOTE,),
+)
+_LIFETIME_COMPOSITE_NOTE = (
+ "composite of this module: MINT8's annual tables carry no "
+ "lifetime-earnings dimension, so the three quintile measures of MINT8's "
+ "cohort tables (tables 13-20) are appended as the lifetime-earnings "
+ "dimension the NASI request names (G3 brief, 2026-10-01); MINT8 cuts "
+ "them per birth cohort, and the caller names the partitions"
+)
+MINT8_ANNUAL_WITH_LIFETIME_SCHEME = CategoryScheme(
+ scheme_id="mint8_beneficiary_annual_with_lifetime",
+ title=f"{MINT8_SCHEME.title}, with MINT8's lifetime-earnings quintiles",
+ population=_BENEFICIARY_POPULATION,
+ dimensions=(*MINT8_SCHEME.dimensions, *LIFETIME_DIMENSIONS),
+ citations=(
+ MINT8_GUIDE_CITATION,
+ _row_label_citation("1-3 and 7-9"),
+ _row_label_citation("13-20"),
+ ),
+ notes=(_LIFETIME_COMPOSITE_NOTE,),
+ composite=True,
+)
+MINT8_POVERTY_WITH_LIFETIME_SCHEME = CategoryScheme(
+ scheme_id="mint8_beneficiary_poverty_with_lifetime",
+ title=(
+ f"{MINT8_POVERTY_SCHEME.title}, with MINT8's lifetime-earnings "
+ "quintiles"
+ ),
+ population=_BENEFICIARY_POPULATION,
+ dimensions=(*MINT8_POVERTY_SCHEME.dimensions, *LIFETIME_DIMENSIONS),
+ citations=(
+ MINT8_GUIDE_CITATION,
+ _row_label_citation("10-12"),
+ _row_label_citation("13-20"),
+ ),
+ notes=(*MINT8_POVERTY_SCHEME.notes, _LIFETIME_COMPOSITE_NOTE),
+ composite=True,
+)
+#: Which label-file tables each MINT8 scheme reproduces exactly
+#: (:func:`verify_mint8_sources`).
+MINT8_SCHEME_TABLES: dict[str, tuple[str, ...]] = {
+ MINT8_SCHEME.scheme_id: ("1", "2", "3", "7", "8", "9"),
+ MINT8_POVERTY_SCHEME.scheme_id: ("10", "11", "12"),
+ MINT8_COHORT_SCHEME.scheme_id: tuple(str(n) for n in range(13, 21)),
+}
+#: Composite MINT8 schemes: (base scheme, appended-dimension tables).
+MINT8_COMPOSITES: dict[str, tuple[str, str]] = {
+ MINT8_ANNUAL_WITH_LIFETIME_SCHEME.scheme_id: (
+ MINT8_SCHEME.scheme_id,
+ "13",
+ ),
+ MINT8_POVERTY_WITH_LIFETIME_SCHEME.scheme_id: (
+ MINT8_POVERTY_SCHEME.scheme_id,
+ "13",
+ ),
+}
+
+# ---- registry -------------------------------------------------------------
+_REGISTRY: dict[str, CategoryScheme] = {}
+
+
+def register_scheme(
+ scheme: CategoryScheme, *, replace: bool = False
+) -> CategoryScheme:
+ """Register an alternate scheme under its ``scheme_id``.
+
+ An id already registered is refused unless ``replace`` is true; the
+ MINT8 schemes this module defines can never be replaced.
+ """
+
+ if not isinstance(scheme, CategoryScheme):
+ raise GroupBreakdownError("register_scheme needs a CategoryScheme")
+ existing = _REGISTRY.get(scheme.scheme_id)
+ if existing is not None and existing is not scheme:
+ if not replace or scheme.scheme_id in _BUILTIN_IDS:
+ raise GroupBreakdownError(
+ f"scheme {scheme.scheme_id!r} is already registered"
+ )
+ _REGISTRY[scheme.scheme_id] = scheme
+ return scheme
+
+
+def get_scheme(scheme_id: str) -> CategoryScheme:
+ """The registered scheme ``scheme_id``."""
+
+ try:
+ return _REGISTRY[scheme_id]
+ except KeyError as error:
+ raise GroupBreakdownError(
+ f"no registered scheme {scheme_id!r}; registered: "
+ f"{sorted(_REGISTRY)}"
+ ) from error
+
+
+def registered_scheme_ids() -> tuple[str, ...]:
+ return tuple(sorted(_REGISTRY))
+
+
+def derive_scheme(
+ base: CategoryScheme,
+ *,
+ scheme_id: str,
+ title: str,
+ replace: Mapping[str, Dimension] | None = None,
+ append: Sequence[Dimension] = (),
+ drop: Sequence[str] = (),
+ population: str | None = None,
+ citations: Sequence[str] = (),
+ notes: Sequence[str] = (),
+) -> CategoryScheme:
+ """A composite variant of ``base`` (not registered).
+
+ ``replace`` maps an existing dimension key to its replacement (which
+ must keep the key), ``drop`` removes dimensions (never the total) and
+ ``append`` adds dimensions at the end. Citations and notes are
+ appended to the base's; the result is marked ``composite``.
+ """
+
+ if not isinstance(base, CategoryScheme):
+ raise GroupBreakdownError("derive_scheme needs a CategoryScheme")
+ replace = dict(replace or {})
+ for key, dimension in replace.items():
+ base.dimension(key)
+ if not isinstance(dimension, Dimension) or dimension.key != key:
+ raise GroupBreakdownError(
+ f"the replacement for {key!r} must be a Dimension with that "
+ "key"
+ )
+ drop = tuple(drop)
+ for key in drop:
+ if key == TOTAL_KEY:
+ raise GroupBreakdownError("the total dimension cannot be dropped")
+ base.dimension(key)
+ dimensions = [
+ replace.get(d.key, d) for d in base.dimensions if d.key not in drop
+ ]
+ dimensions.extend(append)
+ return CategoryScheme(
+ scheme_id=scheme_id,
+ title=title,
+ population=base.population if population is None else population,
+ dimensions=tuple(dimensions),
+ citations=(*base.citations, *_string_tuple(citations, "citations")),
+ notes=(*base.notes, *_string_tuple(notes, "notes")),
+ composite=True,
+ )
+
+
+def with_age_bands(
+ base: CategoryScheme,
+ band_set: str,
+ *,
+ scheme_id: str | None = None,
+) -> CategoryScheme:
+ """``base`` with its ``age`` dimension replaced by an age-band set."""
+
+ if band_set not in AGE_BAND_SETS:
+ raise GroupBreakdownError(
+ f"band_set must be one of {sorted(AGE_BAND_SETS)}"
+ )
+ bands = AGE_BAND_SETS[band_set]
+ return derive_scheme(
+ base,
+ scheme_id=scheme_id or f"{base.scheme_id}__{band_set}_age",
+ title=f"{base.title}, age bands {list(bands.labels)}",
+ replace={"age": bands},
+ citations=(bands.source,) if bands.quote_source is None else (),
+ notes=(
+ f"age bands replaced by the {band_set!r} set: "
+ f"{list(bands.labels)}",
+ ),
+ )
+
+
+_BUILTIN_SCHEMES = (
+ MINT8_SCHEME,
+ MINT8_POVERTY_SCHEME,
+ MINT8_COHORT_SCHEME,
+ MINT8_ANNUAL_WITH_LIFETIME_SCHEME,
+ MINT8_POVERTY_WITH_LIFETIME_SCHEME,
+)
+_BUILTIN_IDS = frozenset(s.scheme_id for s in _BUILTIN_SCHEMES)
+for _scheme in _BUILTIN_SCHEMES:
+ register_scheme(_scheme)
+del _scheme
+
+
+# =========================================================================
+# MINT8 source verification
+# =========================================================================
+class _GuideText(HTMLParser):
+ """The visible text of the committed Guide (scripts and styles left
+ out), whitespace-normalized by :func:`_guide_text`."""
+
+ _BLOCKS = {"p", "li", "h1", "h2", "h3", "h4", "br", "div", "tr", "td"}
+
+ def __init__(self) -> None:
+ super().__init__(convert_charrefs=True)
+ self.parts: list[str] = []
+ self._skip = 0
+
+ def handle_starttag(self, tag: str, attrs: Any) -> None:
+ if tag in ("script", "style"):
+ self._skip += 1
+ if tag in self._BLOCKS:
+ self.parts.append(" ")
+
+ def handle_endtag(self, tag: str) -> None:
+ if tag in ("script", "style"):
+ self._skip = max(0, self._skip - 1)
+
+ def handle_data(self, data: str) -> None:
+ if not self._skip:
+ self.parts.append(data)
+
+
+def _normalize_whitespace(text: str) -> str:
+ return " ".join(text.split())
+
+
+def _guide_text(raw: bytes) -> str:
+ parser = _GuideText()
+ parser.feed(raw.decode("utf-8"))
+ return _normalize_whitespace("".join(parser.parts))
+
+
+def _sha256(path: Path) -> str:
+ return hashlib.sha256(path.read_bytes()).hexdigest()
+
+
+def _dimension_quotes(scheme: CategoryScheme) -> list[str]:
+ quotes = []
+ for dimension in scheme.dimensions:
+ if dimension.quote_source != MINT8_GUIDE:
+ continue
+ quotes.extend(dimension.definition)
+ quotes.extend(
+ c.definition for c in dimension.categories if c.definition
+ )
+ return quotes
+
+
+def _table_groups(table: Mapping[str, Any]) -> list[tuple[str, list[str]]]:
+ return [(g["group"], list(g["labels"])) for g in table["groups"]]
+
+
+def _scheme_groups(
+ dimensions: Iterable[Dimension],
+) -> list[tuple[str, list[str]]]:
+ return [
+ (d.label, [] if d.kind == TOTAL else list(d.labels))
+ for d in dimensions
+ ]
+
+
+def verify_mint8_sources(root: Path | str | None = None) -> dict[str, Any]:
+ """Check the MINT8 schemes against the committed sources.
+
+ Verifies both files' SHA-256 pins; that each MINT8 scheme's row-group
+ headings and row labels equal, verbatim and in order, those of every
+ label-file table it stands for (:data:`MINT8_SCHEME_TABLES`); that each
+ composite scheme is its base plus dimensions of the cohort tables; and
+ that every MINT8 definition quote appears verbatim (whitespace
+ normalized) in the Guide's text. Raises :class:`GroupBreakdownError`
+ on the first mismatch; returns a record of what was checked.
+ """
+
+ base = _PROJECT_ROOT if root is None else Path(root)
+ hashes = {}
+ for relative, expected in MINT8_SOURCE_SHA256.items():
+ actual = _sha256(base / relative)
+ if actual != expected:
+ raise GroupBreakdownError(
+ f"{relative} SHA-256 {actual} differs from the pin {expected}"
+ )
+ hashes[relative] = actual
+ labels = json.loads((base / MINT8_LABELS_FILE).read_text("utf-8"))
+ tables = labels["tables"]
+ checked: dict[str, list[str]] = {}
+ for scheme_id, numbers in MINT8_SCHEME_TABLES.items():
+ expected_groups = _scheme_groups(get_scheme(scheme_id).dimensions)
+ for number in numbers:
+ if _table_groups(tables[number]) != expected_groups:
+ raise GroupBreakdownError(
+ f"scheme {scheme_id} differs from label-file table "
+ f"{number}"
+ )
+ checked[scheme_id] = list(numbers)
+ for scheme_id, (base_id, extra_table) in MINT8_COMPOSITES.items():
+ scheme = get_scheme(scheme_id)
+ base_scheme = get_scheme(base_id)
+ n_base = len(base_scheme.dimensions)
+ if scheme.dimensions[:n_base] != base_scheme.dimensions:
+ raise GroupBreakdownError(
+ f"composite {scheme_id} does not start with {base_id}"
+ )
+ cohort_groups = _table_groups(tables[extra_table])
+ for group in _scheme_groups(scheme.dimensions[n_base:]):
+ if group not in cohort_groups:
+ raise GroupBreakdownError(
+ f"composite {scheme_id} appends {group[0]!r}, which "
+ f"label-file table {extra_table} does not carry"
+ )
+ checked[scheme_id] = [f"{base_id} + table {extra_table}"]
+ text = _guide_text((base / MINT8_GUIDE_FILE).read_bytes())
+ quotes = list(MINT8_GUIDE_QUOTES)
+ for scheme in _BUILTIN_SCHEMES:
+ quotes.extend(_dimension_quotes(scheme))
+ for band in AGE_BAND_SETS.values():
+ if band.quote_source == MINT8_GUIDE:
+ quotes.extend(band.definition)
+ unique = list(dict.fromkeys(quotes))
+ for quote in unique:
+ if _normalize_whitespace(quote) not in text:
+ raise GroupBreakdownError(
+ f"quote not found verbatim in the MINT8 Guide: {quote!r}"
+ )
+ return {
+ "files_sha256": hashes,
+ "schemes_checked": checked,
+ "n_quotes_checked": len(unique),
+ }
+
+
+# =========================================================================
+# Weighted percentiles, quintiles and change classes
+# =========================================================================
+def _exact_numerators(weights: np.ndarray) -> np.ndarray:
+ """Each float64 weight as an exact integer over one common
+ power-of-two denominator (an object array of Python ints)."""
+
+ ratios = [float(w).as_integer_ratio() for w in weights.tolist()]
+ out = np.empty(len(ratios), dtype=object)
+ if ratios:
+ common = max(den for _, den in ratios)
+ out[:] = [num * (common // den) for num, den in ratios]
+ return out
+
+
+def _midpoint(a: float, b: float) -> float:
+ if a == b:
+ return a
+ middle = (a + b) / 2.0
+ if not math.isfinite(middle):
+ middle = a / 2.0 + b / 2.0
+ return middle
+
+
+def _percentiles_of(
+ values: np.ndarray,
+ numerators: np.ndarray,
+ levels: Sequence[Fraction],
+) -> list[float]:
+ """The percentile rule on positive-numerator rows (none empty)."""
+
+ order = np.argsort(values, kind="stable")
+ ordered = values[order]
+ cumulative = list(itertools.accumulate(numerators[order].tolist()))
+ total = cumulative[-1]
+ out = []
+ for level in levels:
+ target = level.numerator * total
+ threshold = -(-target // level.denominator)
+ k = bisect.bisect_left(cumulative, threshold)
+ value = float(ordered[k])
+ tie = cumulative[k] * level.denominator == target
+ if tie and k + 1 < len(cumulative):
+ value = _midpoint(value, float(ordered[k + 1]))
+ out.append(value)
+ return out
+
+
+def _levels(levels: Iterable[Any]) -> tuple[Fraction, ...]:
+ out = tuple(levels)
+ if not out:
+ raise GroupBreakdownError("percentile levels must not be empty")
+ for level in out:
+ if not isinstance(level, Fraction) or not 0 < level < 1:
+ raise GroupBreakdownError(
+ "percentile levels must be fractions.Fraction in (0, 1), "
+ f"not {level!r} (exact levels keep exact ties exact)"
+ )
+ return out
+
+
+def _real_array(values: Any, label: str) -> np.ndarray:
+ array = np.asarray(values)
+ if array.dtype == bool or array.ndim != 1:
+ raise GroupBreakdownError(f"{label} must be a 1-D numeric sequence")
+ try:
+ array = array.astype(np.float64)
+ except (TypeError, ValueError) as error:
+ raise GroupBreakdownError(f"{label} must be numeric") from error
+ if not np.all(np.isfinite(array)):
+ raise GroupBreakdownError(f"{label} must be finite")
+ return array
+
+
+def weighted_percentiles(
+ values: Any, weights: Any, levels: Sequence[Fraction]
+) -> list[float]:
+ """The weighted percentiles of ``values`` at ``levels``.
+
+ The rule is :data:`PERCENTILE_DEFINITION`. ``weights`` must be finite,
+ non-negative and not all zero; ``levels`` are
+ :class:`fractions.Fraction` in (0, 1).
+ """
+
+ values = _real_array(values, "values")
+ weights = _real_array(weights, "weights")
+ if values.size == 0 or values.size != weights.size:
+ raise GroupBreakdownError(
+ "values and weights must be non-empty and the same length"
+ )
+ if (weights < 0).any():
+ raise GroupBreakdownError("weights must be non-negative")
+ keep = weights > 0
+ if not keep.any():
+ raise GroupBreakdownError("weights must not all be zero")
+ return _percentiles_of(
+ values[keep], _exact_numerators(weights[keep]), _levels(levels)
+ )
+
+
+def weighted_quintile_ranks(
+ values: Any, weights: Any
+) -> tuple[np.ndarray, list[float]]:
+ """Each value's quintile rank (1 lowest, 5 highest) and the thresholds.
+
+ The rule is :data:`QUINTILE_RULE`: thresholds are the weighted
+ percentiles at 1/5 ... 4/5 and a value equal to a threshold falls in
+ the lower quintile. Zero-weight rows are ranked but do not move the
+ thresholds.
+ """
+
+ array = _real_array(values, "values")
+ thresholds = weighted_percentiles(array, weights, QUINTILE_LEVELS)
+ ranks = 1 + (array[:, None] > np.asarray(thresholds)[None, :]).sum(axis=1)
+ return ranks.astype(np.int64), thresholds
+
+
+#: A change this close to a +-1 percent threshold is classified exactly.
+_CHANGE_MARGIN = 1e-9
+
+
+def classify_changes(
+ base: Any, reform: Any
+) -> tuple[np.ndarray, np.ndarray, np.ndarray, np.ndarray]:
+ """Individual percent changes and MINT8's change classes.
+
+ Returns ``(population, change, decrease, increase)``: ``population`` is
+ ``base > 0``; ``change`` is ``100 * (reform / base - 1)`` there and NaN
+ elsewhere; ``decrease`` and ``increase`` follow
+ :data:`CHANGE_DEFINITION`, decided exactly (rational arithmetic on the
+ float64 amounts) for every change within ``1e-9`` of a threshold.
+ """
+
+ base = _real_array(base, "base amounts")
+ reform = _real_array(reform, "reform amounts")
+ if base.size != reform.size:
+ raise GroupBreakdownError("base and reform amounts differ in length")
+ if (base < 0).any() or (reform < 0).any():
+ raise GroupBreakdownError("benefit amounts must be non-negative")
+ population = base > 0
+ change = np.full(base.size, np.nan)
+ with np.errstate(over="ignore", invalid="ignore"):
+ change[population] = 100.0 * (
+ reform[population] / base[population] - 1.0
+ )
+ if not np.all(np.isfinite(change[population])):
+ raise GroupBreakdownError("a percent change overflows float64")
+ decrease = population & (change <= -1.0)
+ increase = population & (change >= 1.0)
+ near = population & (
+ (np.abs(change + 1.0) <= _CHANGE_MARGIN)
+ | (np.abs(change - 1.0) <= _CHANGE_MARGIN)
+ )
+ for index in np.flatnonzero(near):
+ b_num, b_den = float(base[index]).as_integer_ratio()
+ r_num, r_den = float(reform[index]).as_integer_ratio()
+ scaled_reform = 100 * r_num * b_den
+ decrease[index] = scaled_reform <= 99 * b_num * r_den
+ increase[index] = scaled_reform >= 101 * b_num * r_den
+ return population, change, decrease, increase
+
+
+# =========================================================================
+# Group assignment
+# =========================================================================
+def _is_missing(value: Any) -> bool:
+ if value is None or value is pd.NA or value is pd.NaT:
+ return True
+ if isinstance(value, float | np.floating):
+ return math.isnan(value)
+ return False
+
+
+def _plain(value: Any) -> Any:
+ if isinstance(value, np.generic):
+ return value.item()
+ return value
+
+
+def _key_digest(frame: pd.DataFrame, key_columns: tuple[str, ...]) -> str:
+ keys = []
+ for row in frame[list(key_columns)].itertuples(index=False, name=None):
+ if any(_is_missing(v) for v in row):
+ raise GroupBreakdownError(
+ f"key columns {list(key_columns)} must not be missing"
+ )
+ keys.append(tuple(_plain(v) for v in row))
+ if len(set(keys)) != len(keys):
+ raise GroupBreakdownError(
+ f"key columns {list(key_columns)} must identify each row"
+ )
+ return hashlib.sha256(repr(keys).encode("utf-8")).hexdigest()
+
+
+def _band_value(value: Any, label: str) -> int:
+ if isinstance(value, bool | np.bool_):
+ raise GroupBreakdownError(f"{label}: a boolean is not a band value")
+ if isinstance(value, Integral):
+ return int(value)
+ if isinstance(value, Real) and math.isfinite(float(value)):
+ if float(value).is_integer():
+ return int(value)
+ raise GroupBreakdownError(
+ f"{label}: band values must be integers; got {value!r}"
+ )
+
+
+def _measure_value(value: Any, label: str) -> float:
+ if isinstance(value, bool | np.bool_) or not isinstance(value, Real):
+ raise GroupBreakdownError(f"{label}: measures must be real numbers")
+ number = float(value)
+ if not math.isfinite(number):
+ raise GroupBreakdownError(f"{label}: measures must be finite")
+ return number
+
+
+@dataclass(frozen=True, eq=False)
+class GroupAssignment:
+ """Each row's category per dimension of a scheme.
+
+ ``long`` has one row per (``row_id``, ``dimension``) with ``label``
+ (the category's printed label, or ``unclassified``), ``category`` (its
+ key, or ``unclassified``) and ``classified``. ``row_id`` is the
+ row's position in the frame passed to :func:`assign_groups`, and
+ ``key_digest`` fingerprints its ``key_columns`` so a tabulator can
+ refuse changed row identities or order. Classifications are a fixed
+ snapshot: recreate the assignment to change attributes, quintile
+ partitions or quintile thresholds. Tabulators can use the snapshot
+ with alternative outcome columns or analysis weights for those same
+ identified rows; they do not recompute classifications.
+ """
+
+ scheme: CategoryScheme
+ n_rows: int
+ key_columns: tuple[str, ...]
+ key_digest: str
+ columns: Mapping[str, str]
+ weight_column: str | None
+ quintile_partition: tuple[str, ...]
+ quintile_partitions: Mapping[str, tuple[str, ...]]
+ long: pd.DataFrame
+ dimension_summaries: tuple[Mapping[str, Any], ...]
+ _category_index: Mapping[str, np.ndarray]
+
+ def _codes(self, dimension: str) -> np.ndarray:
+ try:
+ return self._category_index[dimension]
+ except KeyError as error:
+ raise GroupBreakdownError(
+ f"scheme {self.scheme.scheme_id} has no dimension "
+ f"{dimension!r}"
+ ) from error
+
+ def mask(self, dimension: str, category: str) -> np.ndarray:
+ """The rows of ``category`` (a key) in ``dimension``."""
+
+ self.scheme.dimension(dimension).category(category)
+ return self._codes(dimension) == category
+
+ def unclassified_mask(self, dimension: str) -> np.ndarray:
+ return self._codes(dimension) == UNCLASSIFIED
+
+ def as_dict(self) -> dict[str, Any]:
+ return _json_safe(
+ {
+ "scheme_id": self.scheme.scheme_id,
+ "n_rows": self.n_rows,
+ "key_columns": list(self.key_columns),
+ "key_digest": self.key_digest,
+ "columns": dict(self.columns),
+ "weight_column": self.weight_column,
+ "quintile_partition": list(self.quintile_partition),
+ "quintile_partitions": {
+ key: list(columns)
+ for key, columns in self.quintile_partitions.items()
+ },
+ "unclassified_rule": UNCLASSIFIED_RULE,
+ "dimensions": [dict(s) for s in self.dimension_summaries],
+ }
+ )
+
+
+def _categorical_codes(
+ dimension: Dimension,
+ values: list[Any],
+ extra: frozenset[str],
+ column: str,
+) -> tuple[np.ndarray, dict[str, int]]:
+ lookup = {
+ code: category.key
+ for category in dimension.categories
+ for code in category.codes
+ }
+ out = np.empty(len(values), dtype=object)
+ reasons: dict[str, int] = {}
+ for index, value in enumerate(values):
+ if _is_missing(value):
+ reason = MISSING_REASON
+ elif not isinstance(value, str):
+ raise GroupBreakdownError(
+ f"column {column!r} ({dimension.key}) row {index}: codes "
+ f"must be strings or missing; got {value!r}"
+ )
+ elif value in lookup:
+ out[index] = lookup[value]
+ continue
+ elif value in extra:
+ reason = f"code:{value}"
+ else:
+ raise GroupBreakdownError(
+ f"column {column!r} ({dimension.key}) row {index}: code "
+ f"{value!r} is in no category and was not declared "
+ "unclassified"
+ )
+ out[index] = UNCLASSIFIED
+ reasons[reason] = reasons.get(reason, 0) + 1
+ return out, reasons
+
+
+def _band_codes(
+ dimension: Dimension, values: list[Any], column: str
+) -> tuple[np.ndarray, dict[str, int]]:
+ out = np.empty(len(values), dtype=object)
+ reasons: dict[str, int] = {}
+ for index, value in enumerate(values):
+ reason = None
+ if _is_missing(value):
+ reason = MISSING_REASON
+ else:
+ number = _band_value(
+ value, f"column {column!r} ({dimension.key}) row {index}"
+ )
+ hits = [c.key for c in dimension.categories if c.contains(number)]
+ if hits:
+ out[index] = hits[0]
+ else:
+ reason = OUTSIDE_BANDS_REASON
+ if reason is not None:
+ out[index] = UNCLASSIFIED
+ reasons[reason] = reasons.get(reason, 0) + 1
+ return out, reasons
+
+
+def _quintile_codes(
+ dimension: Dimension,
+ frame: pd.DataFrame,
+ column: str,
+ weights: np.ndarray,
+ partition: tuple[str, ...],
+) -> tuple[np.ndarray, dict[str, int], list[dict[str, Any]]]:
+ values = frame[column].tolist()
+ present = np.zeros(len(values), dtype=bool)
+ numbers = np.zeros(len(values), dtype=np.float64)
+ for index, value in enumerate(values):
+ if _is_missing(value):
+ continue
+ numbers[index] = _measure_value(
+ value, f"column {column!r} ({dimension.key}) row {index}"
+ )
+ present[index] = True
+ if partition:
+ keys = list(frame[list(partition)].itertuples(index=False, name=None))
+ else:
+ keys = [()] * len(values)
+ groups: dict[tuple, list[int]] = {}
+ for index, key in enumerate(keys):
+ if present[index]:
+ plain = tuple(_plain(v) for v in key)
+ groups.setdefault(plain, []).append(index)
+ by_rank = {c.rank: c.key for c in dimension.categories}
+ out = np.empty(len(values), dtype=object)
+ out[:] = UNCLASSIFIED
+ records = []
+ for key, members in groups.items():
+ index = np.asarray(members, dtype=np.int64)
+ member_weights = weights[index]
+ if not (member_weights > 0).any():
+ raise GroupBreakdownError(
+ f"{dimension.key}: partition "
+ f"{dict(zip(partition, key, strict=True))} has no positive "
+ "weight, so its quintile thresholds are undefined"
+ )
+ ranks, thresholds = weighted_quintile_ranks(
+ numbers[index], member_weights
+ )
+ for row, rank in zip(index.tolist(), ranks.tolist(), strict=True):
+ out[row] = by_rank[rank]
+ records.append(
+ {
+ "partition": dict(zip(partition, key, strict=True)),
+ "n_rows": int(index.size),
+ "n_positive_weight": int((member_weights > 0).sum()),
+ "weighted_n": math.fsum(member_weights.tolist()),
+ "thresholds": thresholds,
+ }
+ )
+ n_missing = int((~present).sum())
+ reasons = {MISSING_REASON: n_missing} if n_missing else {}
+ return out, reasons, records
+
+
+def assign_groups(
+ frame: pd.DataFrame,
+ scheme: CategoryScheme,
+ columns_map: Mapping[str, str],
+ *,
+ key_columns: Sequence[str],
+ weight_column: str | None = "weight",
+ quintile_partition: Sequence[str] = (),
+ quintile_partitions: Mapping[str, Sequence[str]] | None = None,
+ unclassified_codes: Mapping[str, Sequence[str]] | None = None,
+) -> GroupAssignment:
+ """Assign every row of ``frame`` to one category per dimension.
+
+ ``columns_map`` maps each non-total dimension key of ``scheme`` to the
+ column holding its attribute: codes (categorical), integers (band) or
+ a numeric measure (quintile). ``key_columns`` identify each row (for
+ example ``("draw", "person_id")`` for projection rows);
+ ``weight_column`` gives the weights quintile thresholds use (required
+ only when the scheme has a quintile dimension); ``quintile_partition``
+ names the columns within whose values quintiles are cut separately;
+ ``quintile_partitions`` overrides those columns per dimension, for
+ example household income within each draw/year and lifetime payroll
+ taxes within each draw/birth cohort in a composite scheme;
+ ``unclassified_codes`` declares, per categorical dimension, present
+ codes that are unclassified. See :data:`UNCLASSIFIED_RULE`.
+ """
+
+ if not isinstance(frame, pd.DataFrame):
+ raise GroupBreakdownError("frame must be a DataFrame")
+ if not frame.columns.is_unique:
+ raise GroupBreakdownError("frame has repeated column names")
+ if frame.empty:
+ raise GroupBreakdownError("frame has no rows")
+ if not isinstance(scheme, CategoryScheme):
+ raise GroupBreakdownError("scheme must be a CategoryScheme")
+ if not isinstance(columns_map, Mapping):
+ raise GroupBreakdownError("columns_map must be a mapping")
+ needed = {d.key for d in scheme.dimensions if d.kind != TOTAL}
+ unknown = sorted(set(columns_map) - needed)
+ missing = sorted(needed - set(columns_map))
+ if unknown or missing:
+ raise GroupBreakdownError(
+ f"columns_map must name exactly the scheme's non-total "
+ f"dimensions; unknown {unknown}, missing {missing}"
+ )
+ for dimension, column in columns_map.items():
+ if column not in frame.columns:
+ raise GroupBreakdownError(
+ f"column {column!r} for {dimension} is not in the frame"
+ )
+ key_columns = _string_tuple(key_columns, "key_columns")
+ if not key_columns:
+ raise GroupBreakdownError("key_columns must not be empty")
+ absent = [c for c in key_columns if c not in frame.columns]
+ if absent:
+ raise GroupBreakdownError(f"key columns {absent} are not in frame")
+ frame = frame.reset_index(drop=True)
+ digest = _key_digest(frame, key_columns)
+ partition = _string_tuple(quintile_partition, "quintile_partition")
+ absent = [c for c in partition if c not in frame.columns]
+ if absent:
+ raise GroupBreakdownError(f"partition columns {absent} are absent")
+ if partition and frame[list(partition)].isna().any(axis=None):
+ raise GroupBreakdownError("partition columns must not be missing")
+ if quintile_partitions is not None and not isinstance(
+ quintile_partitions, Mapping
+ ):
+ raise GroupBreakdownError("quintile_partitions must be a mapping")
+ partitions = {
+ d.key: partition for d in scheme.dimensions if d.kind == QUINTILE
+ }
+ for key, columns in (quintile_partitions or {}).items():
+ dimension = scheme.dimension(key)
+ if dimension.kind != QUINTILE:
+ raise GroupBreakdownError(
+ f"quintile_partitions applies to quintile dimensions; "
+ f"{key} is {dimension.kind}"
+ )
+ columns = _string_tuple(columns, f"{key} partition columns")
+ absent = [c for c in columns if c not in frame.columns]
+ if absent:
+ raise GroupBreakdownError(f"partition columns {absent} are absent")
+ if columns and frame[list(columns)].isna().any(axis=None):
+ raise GroupBreakdownError("partition columns must not be missing")
+ partitions[key] = columns
+ declared = dict(unclassified_codes or {})
+ for key, codes in declared.items():
+ dimension = scheme.dimension(key)
+ if dimension.kind != CATEGORICAL:
+ raise GroupBreakdownError(
+ f"unclassified_codes applies to categorical dimensions; "
+ f"{key} is {dimension.kind}"
+ )
+ codes = frozenset(_string_tuple(codes, f"{key} unclassified codes"))
+ held = codes.intersection(
+ code for c in dimension.categories for code in c.codes
+ )
+ if held:
+ raise GroupBreakdownError(
+ f"{key}: codes {sorted(held)} are a category's and cannot "
+ "be declared unclassified"
+ )
+ declared[key] = codes
+ has_quintile = any(d.kind == QUINTILE for d in scheme.dimensions)
+ weights = None
+ if has_quintile:
+ if weight_column is None or weight_column not in frame.columns:
+ raise GroupBreakdownError(
+ "a quintile dimension needs the weight column"
+ )
+ weights = _real_array(frame[weight_column].to_numpy(), "weights")
+ if (weights < 0).any():
+ raise GroupBreakdownError("weights must be non-negative")
+
+ n = len(frame)
+ index: dict[str, np.ndarray] = {}
+ summaries = []
+ long_parts = []
+ for dimension in scheme.dimensions:
+ records: list[dict[str, Any]] = []
+ column = columns_map.get(dimension.key)
+ if dimension.kind == TOTAL:
+ codes = np.empty(n, dtype=object)
+ codes[:] = TOTAL_KEY
+ reasons: dict[str, int] = {}
+ elif dimension.kind == CATEGORICAL:
+ codes, reasons = _categorical_codes(
+ dimension,
+ frame[column].tolist(),
+ declared.get(dimension.key, frozenset()),
+ column,
+ )
+ elif dimension.kind == BAND:
+ codes, reasons = _band_codes(
+ dimension, frame[column].tolist(), column
+ )
+ else:
+ codes, reasons, records = _quintile_codes(
+ dimension, frame, column, weights, partitions[dimension.key]
+ )
+ index[dimension.key] = codes
+ label_of = {c.key: c.label for c in dimension.categories}
+ label_of[UNCLASSIFIED] = UNCLASSIFIED
+ classified = codes != UNCLASSIFIED
+ long_parts.append(
+ pd.DataFrame(
+ {
+ "row_id": np.arange(n, dtype=np.int64),
+ "dimension": dimension.key,
+ "label": [label_of[c] for c in codes.tolist()],
+ "category": codes.tolist(),
+ "classified": classified,
+ }
+ )
+ )
+ summary = {
+ "key": dimension.key,
+ "label": dimension.label,
+ "kind": dimension.kind,
+ "column": column,
+ "n_rows": n,
+ "n_classified": int(classified.sum()),
+ "n_unclassified": int((~classified).sum()),
+ "unclassified_reasons": dict(sorted(reasons.items())),
+ "n_by_category": {
+ c.key: int((codes == c.key).sum())
+ for c in dimension.categories
+ },
+ }
+ if dimension.kind == QUINTILE:
+ summary["quintile_rule"] = QUINTILE_RULE
+ summary["quintile_partitions"] = records
+ summaries.append(summary)
+ return GroupAssignment(
+ scheme=scheme,
+ n_rows=n,
+ key_columns=key_columns,
+ key_digest=digest,
+ columns=dict(columns_map),
+ weight_column=weight_column,
+ quintile_partition=partition,
+ quintile_partitions=partitions,
+ long=pd.concat(long_parts, ignore_index=True),
+ dimension_summaries=tuple(summaries),
+ _category_index=index,
+ )
+
+
+def _check_assignment(
+ assignment: Any,
+ rows: pd.DataFrame,
+ key_columns: tuple[str, ...],
+) -> None:
+ if not isinstance(assignment, GroupAssignment):
+ raise GroupBreakdownError("assignment must be a GroupAssignment")
+ if assignment.key_columns != key_columns:
+ raise GroupBreakdownError(
+ f"the assignment was keyed on {list(assignment.key_columns)}, "
+ f"not {list(key_columns)}"
+ )
+ if assignment.n_rows != len(rows):
+ raise GroupBreakdownError(
+ f"the assignment covers {assignment.n_rows} rows, not "
+ f"{len(rows)}"
+ )
+ missing = [c for c in key_columns if c not in rows.columns]
+ if missing:
+ raise GroupBreakdownError(f"rows lack key columns {missing}")
+ digest = _key_digest(rows.reset_index(drop=True), key_columns)
+ if digest != assignment.key_digest:
+ raise GroupBreakdownError(
+ "the rows are not the rows the assignment was built from (key "
+ "digest differs): same rows, same order"
+ )
+
+
+# =========================================================================
+# Shared tabulation pieces
+# =========================================================================
+def half_split_masks(
+ unit_values: Sequence[Any], seeds: Sequence[int]
+) -> dict[int, np.ndarray]:
+ """Side-A membership of every row, per floor seed.
+
+ The split is computed once over the split units of all rows passed, as
+ ``cola_age_profile`` and ``uniform_cut_tabulation`` compute it
+ (``split_panel_by_person`` with fraction 0.5 over the sorted unique
+ units); a group's halves are this split intersected with its mask.
+ """
+
+ seeds = tuple(_floor_seeds(seeds))
+ frame = pd.DataFrame({"split_unit": list(unit_values)})
+ out = {}
+ for seed in seeds:
+ side_a, _ = split_panel_by_person(
+ frame, "split_unit", fraction=FLOOR_FRACTION, seed=seed
+ )
+ in_a = np.zeros(len(frame), dtype=bool)
+ in_a[side_a.index.to_numpy()] = True
+ out[seed] = in_a
+ return out
+
+
+def _floor_seeds(seeds: Iterable[Any]) -> list[int]:
+ out = []
+ for seed in seeds:
+ if isinstance(seed, bool | np.bool_) or not isinstance(seed, Integral):
+ raise GroupBreakdownError("floor seeds must be integers")
+ out.append(int(seed))
+ if not out or len(set(out)) != len(out) or min(out) < 0:
+ raise GroupBreakdownError(
+ "floor seeds must be a non-empty set of distinct integers >= 0"
+ )
+ return out
+
+
+def _labels(labels: Any, label: str) -> tuple[str, ...]:
+ if isinstance(labels, str):
+ raise GroupBreakdownError(f"{label} must be a sequence of strings")
+ out = tuple(labels)
+ if not all(isinstance(item, str) and item.strip() for item in out):
+ raise GroupBreakdownError(f"{label} must be non-empty strings")
+ return out
+
+
+def _validate_provenance(
+ rows: pd.DataFrame,
+ data_provenance: Any,
+ registration_pointer: Any,
+ labels: Any,
+ post_hoc_labels: Any,
+) -> tuple[list[str], tuple[str, ...]]:
+ if data_provenance not in DATA_PROVENANCES:
+ raise GroupBreakdownError(
+ f"data_provenance must be one of {list(DATA_PROVENANCES)}"
+ )
+ labels = _labels(labels, "labels")
+ post_hoc = _labels(post_hoc_labels, "post_hoc_labels")
+ kind = rows.attrs.get("provenance_kind")
+ if data_provenance == REGISTERED_REAL:
+ if not isinstance(
+ registration_pointer, str
+ ) or not REGISTRATION_POINTER.fullmatch(registration_pointer):
+ raise GroupBreakdownError(
+ "a real-data breakdown requires its issue #42 registration "
+ "pointer (https://github.com/PolicyEngine/microcosm-"
+ "dynamics/issues/42#issuecomment-), which must exist "
+ "before the run"
+ )
+ if not labels:
+ raise GroupBreakdownError("a real-data breakdown needs labels")
+ if not post_hoc:
+ raise GroupBreakdownError(
+ "a real-data breakdown needs the caller's post hoc labels "
+ "(for example 'registered, one-shot, post hoc, not blind')"
+ )
+ if INVENTED_DATA_LABEL in labels:
+ raise GroupBreakdownError(
+ "a registered_real result cannot carry the invented label"
+ )
+ if kind == INVENTED:
+ raise GroupBreakdownError(
+ "rows marked invented cannot be tabulated as registered_real"
+ )
+ else:
+ if registration_pointer is not None and not isinstance(
+ registration_pointer, str
+ ):
+ raise GroupBreakdownError(
+ "registration_pointer must be a string or None"
+ )
+ if kind == PSID_FILES:
+ raise GroupBreakdownError(
+ "rows built from staged PSID files cannot be tabulated as "
+ "invented data"
+ )
+ output = [label for label in labels if label != INVENTED_DATA_LABEL]
+ if data_provenance == INVENTED:
+ output.insert(0, INVENTED_DATA_LABEL)
+ return output, post_hoc
+
+
+def _fsum(values: np.ndarray) -> float:
+ try:
+ total = math.fsum(values.tolist())
+ except OverflowError as error:
+ raise GroupBreakdownError(
+ "a weighted sum overflows float64"
+ ) from error
+ if not math.isfinite(total):
+ raise GroupBreakdownError("a weighted sum is not finite")
+ return total
+
+
+def _n_flags(n: int) -> dict[str, bool]:
+ return {
+ "ssa_disclosure_below_100": n < SSA_DISCLOSURE_MIN_N,
+ "below_30": n < SMALL_CELL_MIN_N,
+ }
+
+
+def _numerator_flag(numerators: Sequence[int]) -> dict[str, Any]:
+ low = [int(n) for n in numerators if 1 <= n < SSA_NUMERATOR_MIN]
+ return {"ssa_numerator_1_to_9": bool(low)}
+
+
+SUPPRESSION_RULES: dict[str, Any] = {
+ "applied": (
+ "flags only: no cell or dimension is dropped; a reader applies "
+ "SSA's suppression by the flags"
+ ),
+ "ssa_disclosure_below_100": {
+ "threshold": SSA_DISCLOSURE_MIN_N,
+ "rule": "the cell's unweighted n is below 100",
+ "dimension_flag": (
+ "ssa_subgroup_suppressed: any category of the dimension is "
+ "flagged, since SSA removes the whole subgroup"
+ ),
+ "source_quote": _Q_DISCLOSURE,
+ },
+ "below_30": {
+ "threshold": SMALL_CELL_MIN_N,
+ "rule": "the cell's unweighted n is below 30",
+ },
+ "ssa_numerator_1_to_9": {
+ "threshold": SSA_NUMERATOR_MIN,
+ "rule": (
+ "percent with a decrease / increase only: an unweighted "
+ "numerator of 1-9 cases (in any draw, for projections); SSA's "
+ "low-numerator display exception is not modelled"
+ ),
+ "source_quote": _Q_NUMERATOR,
+ },
+}
+
+
+def _json_safe(value: Any) -> Any:
+ if isinstance(value, Mapping):
+ out = {}
+ for key, item in value.items():
+ if not isinstance(key, str):
+ raise GroupBreakdownError(f"non-string key {key!r}")
+ out[key] = _json_safe(item)
+ return out
+ if isinstance(value, list | tuple):
+ return [_json_safe(item) for item in value]
+ if isinstance(value, np.generic):
+ value = value.item()
+ if isinstance(value, float):
+ if not math.isfinite(value):
+ raise GroupBreakdownError("a result number is not finite")
+ return value
+ if value is None or isinstance(value, str | bool | int):
+ return value
+ raise GroupBreakdownError(f"value {value!r} is not JSON-safe")
+
+
+# =========================================================================
+# Result schema
+# =========================================================================
+@dataclass(frozen=True)
+class StatisticCell:
+ """One statistic of one group cell.
+
+ ``value`` is the reported value (the mean over draws for projections),
+ ``None`` when undefined with ``undefined_reason``; ``unweighted_n`` and
+ ``weighted_n`` are the statistic's population in the cell (projections:
+ the minimum over draws and the mean over draws); ``uncertainty`` holds
+ the draw SD, design SE and floor as applicable; ``flags`` the
+ suppression flags.
+ """
+
+ statistic: str
+ value: float | None
+ defined: bool
+ undefined_reason: str | None
+ unweighted_n: int
+ weighted_n: float
+ uncertainty: Mapping[str, Any]
+ flags: Mapping[str, Any]
+
+ def as_dict(self) -> dict[str, Any]:
+ return {
+ "statistic": self.statistic,
+ "value": self.value,
+ "defined": self.defined,
+ "undefined_reason": self.undefined_reason,
+ "unweighted_n": self.unweighted_n,
+ "weighted_n": self.weighted_n,
+ "uncertainty": dict(self.uncertainty),
+ "flags": dict(self.flags),
+ }
+
+
+@dataclass(frozen=True)
+class GroupCell:
+ """One row of the breakdown: a category of a dimension."""
+
+ dimension: str
+ dimension_label: str
+ category: str
+ label: str
+ counts: Mapping[str, Any]
+ statistics: tuple[StatisticCell, ...]
+ flags: Mapping[str, Any]
+
+ def statistic(self, name: str) -> StatisticCell:
+ for cell in self.statistics:
+ if cell.statistic == name:
+ return cell
+ raise GroupBreakdownError(f"no statistic {name!r}")
+
+ def as_dict(self) -> dict[str, Any]:
+ return {
+ "dimension": self.dimension,
+ "dimension_label": self.dimension_label,
+ "category": self.category,
+ "label": self.label,
+ "counts": dict(self.counts),
+ "flags": dict(self.flags),
+ "statistics": [s.as_dict() for s in self.statistics],
+ }
+
+
+@dataclass(frozen=True)
+class DimensionResult:
+ """A dimension's cells, its unclassified rows and its SSA flag."""
+
+ key: str
+ label: str
+ kind: str
+ cells: tuple[GroupCell, ...]
+ unclassified: Mapping[str, Any]
+ ssa_subgroup_suppressed: bool
+ suppressed_by: tuple[str, ...]
+
+ def cell(self, category: str) -> GroupCell:
+ for cell in self.cells:
+ if category in (cell.category, cell.label):
+ return cell
+ raise GroupBreakdownError(f"{self.key} has no category {category!r}")
+
+ def as_dict(self) -> dict[str, Any]:
+ return {
+ "key": self.key,
+ "label": self.label,
+ "kind": self.kind,
+ "ssa_subgroup_suppressed": self.ssa_subgroup_suppressed,
+ "suppressed_by": list(self.suppressed_by),
+ "unclassified": dict(self.unclassified),
+ "cells": [c.as_dict() for c in self.cells],
+ }
+
+
+@dataclass(frozen=True)
+class GroupBreakdownResult:
+ """A complete breakdown; ``as_dict()`` is its JSON-safe form."""
+
+ kind: str
+ statistic_id: str
+ data_provenance: str
+ registration_pointer: str | None
+ labels: tuple[str, ...]
+ post_hoc_labels: tuple[str, ...]
+ scheme: CategoryScheme
+ statistics: tuple[str, ...]
+ conventions: Mapping[str, Any]
+ dimensions: tuple[DimensionResult, ...]
+ assignment: Mapping[str, Any]
+ input_summary: Mapping[str, Any]
+ upstream: Mapping[str, Any]
+ schema_version: str = SCHEMA_VERSION
+
+ def dimension(self, key: str) -> DimensionResult:
+ for dimension in self.dimensions:
+ if dimension.key == key:
+ return dimension
+ raise GroupBreakdownError(f"no dimension {key!r}")
+
+ def cell(self, dimension: str, category: str) -> GroupCell:
+ return self.dimension(dimension).cell(category)
+
+ def as_dict(self) -> dict[str, Any]:
+ return _json_safe(
+ {
+ "schema_version": self.schema_version,
+ "kind": self.kind,
+ "statistic_id": self.statistic_id,
+ "data_provenance": self.data_provenance,
+ "registration_pointer": self.registration_pointer,
+ "labels": list(self.labels),
+ "post_hoc_labels": list(self.post_hoc_labels),
+ "scheme": self.scheme.as_dict(),
+ "statistics": list(self.statistics),
+ "statistic_definitions": {
+ name: STATISTIC_DEFINITIONS[name]
+ for name in self.statistics
+ },
+ "conventions": dict(self.conventions),
+ "assignment": dict(self.assignment),
+ "input_summary": dict(self.input_summary),
+ "upstream": dict(self.upstream),
+ "dimensions": [d.as_dict() for d in self.dimensions],
+ }
+ )
+
+
+def _cell_flags(
+ n: int, statistics: Sequence[StatisticCell]
+) -> dict[str, bool]:
+ flags = _n_flags(n)
+ for name in flags:
+ flags[name] = flags[name] or any(
+ s.flags.get(name, False) for s in statistics
+ )
+ if any(s.flags.get("ssa_numerator_1_to_9", False) for s in statistics):
+ flags["ssa_numerator_1_to_9"] = True
+ return flags
+
+
+def _dimension_result(
+ dimension: Dimension,
+ cells: list[GroupCell],
+ unclassified: Mapping[str, Any],
+) -> DimensionResult:
+ flagged = tuple(
+ cell.label
+ for cell in cells
+ if cell.flags.get("ssa_disclosure_below_100")
+ or cell.flags.get("ssa_numerator_1_to_9")
+ )
+ return DimensionResult(
+ key=dimension.key,
+ label=dimension.label,
+ kind=dimension.kind,
+ cells=tuple(cells),
+ unclassified=dict(unclassified),
+ ssa_subgroup_suppressed=bool(flagged),
+ suppressed_by=flagged,
+ )
+
+
+def _upstream(upstream: Mapping[str, Any] | None) -> dict[str, Any]:
+ try:
+ return a7._json_scalar_mapping(upstream, "upstream")
+ except a7.ColaTabulationError as error:
+ raise GroupBreakdownError(str(error)) from error
+
+
+def _statistic_id(value: Any) -> str:
+ return _nonempty_text(value, "statistic_id")
+
+
+def _base_conventions() -> dict[str, Any]:
+ return {
+ "unclassified_rule": UNCLASSIFIED_RULE,
+ "percentile_definition": PERCENTILE_DEFINITION,
+ "quintile_rule": QUINTILE_RULE,
+ "suppression": SUPPRESSION_RULES,
+ "acceptance_rule": None,
+ "acceptance_rule_note": (
+ "none: a breakdown reports values only; it scores nothing"
+ ),
+ }
+
+
+# =========================================================================
+# Projection breakdown (exercises 1 and 3)
+# =========================================================================
+_PROJECTION_POPULATIONS = {
+ RATIO_OF_SCENARIO_MEANS: ("n_base", "weight_base"),
+ MEAN_OF_INDIVIDUAL_RATIOS: ("n_alternative", "weight_alternative"),
+ **{
+ name: ("n_current_law", "weight_current_law")
+ for name in MINT_BENEFIT_STATISTICS
+ },
+}
+_NUMERATORS = {PERCENT_DECREASE: "n_decrease", PERCENT_INCREASE: "n_increase"}
+
+
+@dataclass(frozen=True, eq=False)
+class _Projection:
+ rows: Any # cola_age_profile._Rows
+ config: a7.ColaAgeProfileConfig
+ masks: tuple[np.ndarray, np.ndarray, np.ndarray]
+ by_draw: Mapping[int, np.ndarray]
+ population: np.ndarray
+ change: np.ndarray
+ decrease: np.ndarray
+ increase: np.ndarray
+ numerators: np.ndarray
+ positive: np.ndarray
+
+
+def _mint_values(
+ weights: np.ndarray,
+ change: np.ndarray,
+ decrease: np.ndarray,
+ increase: np.ndarray,
+ numerators: np.ndarray,
+ positive: np.ndarray,
+ in_population: np.ndarray,
+) -> tuple[dict[str, Any], dict[str, str]]:
+ """MINT8's benefit statistics over ``in_population`` (P in the cell)."""
+
+ n = int(np.count_nonzero(in_population))
+ total = _fsum(weights[in_population])
+ counts = {
+ "n_current_law": n,
+ "weight_current_law": total,
+ "n_decrease": int(np.count_nonzero(in_population & decrease)),
+ "n_increase": int(np.count_nonzero(in_population & increase)),
+ }
+ if n == 0 or total <= 0.0:
+ reason = (
+ "empty_current_law_population"
+ if n == 0
+ else "zero_current_law_weight"
+ )
+ values = {name: None for name in MINT_BENEFIT_STATISTICS}
+ return {**values, **counts}, {
+ name: reason for name in MINT_BENEFIT_STATISTICS
+ }
+ unaffected = in_population & ~decrease & ~increase
+ with_mass = in_population & positive
+ p10, p50, p90 = _percentiles_of(
+ change[with_mass], numerators[with_mass], (P10, P50, P90)
+ )
+ values = {
+ PERCENT_DECREASE: 100.0
+ * _fsum(weights[in_population & decrease])
+ / total,
+ PERCENT_INCREASE: 100.0
+ * _fsum(weights[in_population & increase])
+ / total,
+ PERCENT_UNAFFECTED: 100.0 * _fsum(weights[unaffected]) / total,
+ CHANGE_P10: p10,
+ CHANGE_MEDIAN: p50,
+ CHANGE_P90: p90,
+ }
+ return {**values, **counts}, {}
+
+
+def _projection_draw(
+ context: _Projection, subset: np.ndarray, draw: int
+) -> dict[str, Any]:
+ try:
+ cell = a7._cell(context.rows, subset, context.masks, draw)
+ except a7.ColaTabulationError as error:
+ raise GroupBreakdownError(str(error)) from error
+ mint, reasons = _mint_values(
+ context.rows.weight,
+ context.change,
+ context.decrease,
+ context.increase,
+ context.numerators,
+ context.positive,
+ subset & context.population,
+ )
+ merged = dict(cell["undefined_reasons"])
+ merged.update(reasons)
+ return {**cell, **mint, "undefined_reasons": merged}
+
+
+def _projection_side(
+ context: _Projection, subset: np.ndarray
+) -> dict[str, float | None]:
+ """The reported values recomputed on one half (cola_age_profile's
+ ``_side_values``: the mean over all draws, null if any is undefined)."""
+
+ cells = [
+ _projection_draw(context, subset & context.by_draw[d], d)
+ for d in context.config.draw_indices
+ ]
+ out: dict[str, float | None] = {}
+ for statistic in PROJECTION_STATISTICS:
+ values = [cell[statistic] for cell in cells]
+ if any(v is None for v in values):
+ out[statistic] = None
+ else:
+ try:
+ out[statistic] = a7._side_mean(values)
+ except a7.ColaTabulationError as error:
+ raise GroupBreakdownError(str(error)) from error
+ return out
+
+
+def _projection_cell(
+ context: _Projection,
+ split: Mapping[int, np.ndarray],
+ dimension: Dimension,
+ category: Category,
+ mask: np.ndarray,
+) -> GroupCell:
+ draws = context.config.draw_indices
+ cells = [
+ _projection_draw(context, context.by_draw[d] & mask, d) for d in draws
+ ]
+ sides = {
+ seed: (
+ _projection_side(context, mask & in_a),
+ _projection_side(context, mask & ~in_a),
+ )
+ for seed, in_a in split.items()
+ }
+ statistics = []
+ for statistic in PROJECTION_STATISTICS:
+ summary = a7._draw_summary(cells, statistic)
+ gaps, dropped = [], []
+ for seed, (side_a, side_b) in sides.items():
+ a, b = side_a[statistic], side_b[statistic]
+ if a is None or b is None:
+ dropped.append(seed)
+ else:
+ gaps.append(abs(a - b))
+ try:
+ floor = {**a7._floor_summary(gaps), "dropped_seeds": dropped}
+ except a7.ColaTabulationError as error:
+ raise GroupBreakdownError(str(error)) from error
+ n_key, w_key = _PROJECTION_POPULATIONS[statistic]
+ per_draw_n = [int(c[n_key]) for c in cells]
+ if statistic == RATIO_OF_SCENARIO_MEANS:
+ per_draw_n = [min(c["n_base"], c["n_reform"]) for c in cells]
+ per_draw_w = [float(c[w_key]) for c in cells]
+ flags: dict[str, Any] = _n_flags(min(per_draw_n))
+ if statistic in _NUMERATORS:
+ flags.update(
+ _numerator_flag([c[_NUMERATORS[statistic]] for c in cells])
+ )
+ undefined = summary["undefined_draws"]
+ statistics.append(
+ StatisticCell(
+ statistic=statistic,
+ value=summary["mean"],
+ defined=summary["defined"],
+ undefined_reason=(
+ None
+ if summary["defined"]
+ else (
+ f"undefined in {len(undefined)} of "
+ f"{summary['n_draws']} draws: "
+ f"{sorted({u['reason'] for u in undefined})}"
+ )
+ ),
+ unweighted_n=min(per_draw_n),
+ weighted_n=math.fsum(per_draw_w) / len(per_draw_w),
+ uncertainty={
+ "method": "draws_and_half_split_floor",
+ "per_draw": summary["per_draw"],
+ "sample_sd": summary["sample_sd"],
+ "n_draws": summary["n_draws"],
+ "n_defined_draws": summary["n_defined_draws"],
+ "undefined_draws": undefined,
+ "unweighted_n_per_draw": per_draw_n,
+ "weighted_n_per_draw": per_draw_w,
+ "floor": floor,
+ },
+ flags=flags,
+ )
+ )
+ n_current = [int(c["n_current_law"]) for c in cells]
+ counts = {
+ "n_rows": int(np.count_nonzero(mask)),
+ "n_persons": int(len(set(context.rows.person_id[mask].tolist()))),
+ "unweighted_n_current_law_per_draw": n_current,
+ "weighted_n_current_law_per_draw": [
+ float(c["weight_current_law"]) for c in cells
+ ],
+ "unweighted_n_base_per_draw": [int(c["n_base"]) for c in cells],
+ "weighted_n_base_per_draw": [float(c["weight_base"]) for c in cells],
+ "unweighted_n_reform_per_draw": [int(c["n_reform"]) for c in cells],
+ "weighted_n_reform_per_draw": [
+ float(c["weight_reform"]) for c in cells
+ ],
+ "unweighted_n": min(n_current),
+ "unweighted_n_rule": (
+ "minimum over draws of the cell's current-law beneficiaries "
+ "with a positive selected benefit"
+ ),
+ }
+ return GroupCell(
+ dimension=dimension.key,
+ dimension_label=dimension.label,
+ category=category.key,
+ label=category.label,
+ counts=counts,
+ statistics=tuple(statistics),
+ flags=_cell_flags(min(n_current), statistics),
+ )
+
+
+def _projection_unclassified(
+ context: _Projection, mask: np.ndarray, summary: Mapping[str, Any]
+) -> dict[str, Any]:
+ per_draw = []
+ for d in context.config.draw_indices:
+ subset = mask & context.by_draw[d]
+ base = subset & context.masks[0]
+ current = subset & context.population
+ per_draw.append(
+ {
+ "draw": int(d),
+ "n_rows": int(np.count_nonzero(subset)),
+ "n_current_law": int(np.count_nonzero(current)),
+ "weight_current_law": _fsum(context.rows.weight[current]),
+ "n_base": int(np.count_nonzero(base)),
+ "weight_base": _fsum(context.rows.weight[base]),
+ }
+ )
+ return {
+ "n_rows": int(np.count_nonzero(mask)),
+ "reasons": dict(summary["unclassified_reasons"]),
+ "per_draw": per_draw,
+ "statistics": None,
+ "note": "unclassified rows enter Total only; counts are reported",
+ }
+
+
+PROJECTION_STATISTIC_ID = "group_breakdown_projection_benefit_change"
+_PROJECTION_KEYS = ("draw", "person_id")
+
+
+def tabulate_projection_breakdown(
+ rows: pd.DataFrame,
+ assignment: GroupAssignment,
+ *,
+ data_provenance: str,
+ config: a7.ColaAgeProfileConfig | None = None,
+ registration_pointer: str | None = None,
+ labels: Sequence[str] = (),
+ post_hoc_labels: Sequence[str] = (),
+ statistic_id: str = PROJECTION_STATISTIC_ID,
+ upstream: Mapping[str, Any] | None = None,
+) -> GroupBreakdownResult:
+ """Break a projection's per-person-per-draw rows down by group.
+
+ ``rows`` are ``cola_age_profile`` input rows (its
+ ``REQUIRED_COLUMNS``; attribute columns are ignored here) and
+ ``assignment`` is :func:`assign_groups` of the same frame keyed on
+ ``("draw", "person_id")``. ``config`` is the
+ ``ColaAgeProfileConfig`` of the run (exercise 3 passes
+ ``allow_membership_difference=True``); scenario memberships, draws,
+ floor seeds and the floor's split unit all come from it, and rows
+ whose memberships differ are refused unless it allows them, as
+ ``tabulate_cola_age_profile`` does. Every cell carries
+ :data:`PROJECTION_STATISTICS`.
+ """
+
+ config = a7.ColaAgeProfileConfig() if config is None else config
+ if not isinstance(config, a7.ColaAgeProfileConfig):
+ raise GroupBreakdownError("config must be a ColaAgeProfileConfig")
+ if not isinstance(rows, pd.DataFrame):
+ raise GroupBreakdownError("rows must be a DataFrame")
+ output_labels, post_hoc = _validate_provenance(
+ rows, data_provenance, registration_pointer, labels, post_hoc_labels
+ )
+ statistic_id = _statistic_id(statistic_id)
+ recorded_upstream = _upstream(upstream)
+ _check_assignment(assignment, rows, _PROJECTION_KEYS)
+ try:
+ normalized = a7._normalize(rows.reset_index(drop=True), config)
+ except a7.ColaTabulationError as error:
+ raise GroupBreakdownError(str(error)) from error
+ differs = normalized.recipient_base != normalized.recipient_reform
+ if differs.any() and not config.allow_membership_difference:
+ raise GroupBreakdownError(
+ f"{int(np.count_nonzero(differs))} rows are recipients in only "
+ "one scenario; pass a config with allow_membership_difference "
+ "to tabulate them under its membership_basis"
+ )
+ population_base = normalized.recipient_base & (
+ normalized.selected_base > 0.0
+ )
+ _, change, decrease, increase = classify_changes(
+ np.where(population_base, normalized.selected_base, 0.0),
+ normalized.selected_reform,
+ )
+ context = _Projection(
+ rows=normalized,
+ config=config,
+ masks=a7._membership_masks(normalized, config.membership_basis),
+ by_draw={d: normalized.draw == d for d in config.draw_indices},
+ population=population_base,
+ change=change,
+ decrease=decrease & population_base,
+ increase=increase & population_base,
+ numerators=_exact_numerators(normalized.weight),
+ positive=normalized.weight > 0.0,
+ )
+ unit_values = (
+ normalized.family_unit_id
+ if config.floor_split_unit == a7.FAMILY_UNIT
+ else normalized.person_id
+ )
+ split = half_split_masks(unit_values.tolist(), config.floor_seeds)
+ summaries = {s["key"]: s for s in assignment.dimension_summaries}
+ dimensions = []
+ for dimension in assignment.scheme.dimensions:
+ cells = [
+ _projection_cell(
+ context,
+ split,
+ dimension,
+ category,
+ assignment.mask(dimension.key, category.key),
+ )
+ for category in dimension.categories
+ ]
+ unclassified = _projection_unclassified(
+ context,
+ assignment.unclassified_mask(dimension.key),
+ summaries[dimension.key],
+ )
+ dimensions.append(_dimension_result(dimension, cells, unclassified))
+ conventions = {
+ **_base_conventions(),
+ "change_definition": CHANGE_DEFINITION,
+ "cola_age_profile_config": config.as_dict(),
+ "membership": {
+ "recipient_rule": config.recipient_rule,
+ "recipient_rule_definition": a7.RECIPIENT_RULE_DEFINITIONS[
+ config.recipient_rule
+ ],
+ "membership_basis": config.membership_basis,
+ "membership_basis_definition": a7.MEMBERSHIP_BASIS_DEFINITIONS[
+ config.membership_basis
+ ],
+ "allow_membership_difference": (
+ config.allow_membership_difference
+ ),
+ "mint_population": _MINT_POPULATION,
+ },
+ "uncertainty": {
+ "draws": (
+ "the reported value is the mean over the configured draws "
+ "of the per-draw statistic, with the ddof=1 sample SD "
+ "(cola_age_profile._draw_summary); a statistic undefined in "
+ "any draw has no mean"
+ ),
+ "floor": (
+ "for each floor seed the split units "
+ f"({config.floor_split_unit}) of all input rows are split "
+ "once (half_split_masks); each group's halves are that "
+ "split intersected with the group's mask, never a re-split "
+ "of the group; the side value is the mean over draws; the "
+ "floor summarizes |side_a - side_b| over usable seeds "
+ "(cola_age_profile._floor_summary) and is undefined, not "
+ "zero, with fewer than two"
+ ),
+ "floor_seeds": list(config.floor_seeds),
+ "floor_split_unit": config.floor_split_unit,
+ "scale": "half sample; not rescaled",
+ },
+ "unweighted_n": (
+ "per statistic: the minimum over draws of the per-draw count "
+ "of the statistic's population in the cell (the smaller of "
+ "S_base and S_reform for the ratio of means, S_alt for the "
+ "mean of ratios, P for the MINT8 "
+ "statistics); the cell's flags include any statistic's "
+ "small-population or numerator flag"
+ ),
+ "weighted_n": (
+ "per statistic: mean over draws of population weight; the "
+ "ratio of means reports baseline weight, with both scenarios' "
+ "counts and weights also retained in each cell's counts"
+ ),
+ }
+ input_summary = {
+ "n_rows": normalized.n,
+ "n_persons": int(len(set(normalized.person_id.tolist()))),
+ "n_family_units": int(len(set(normalized.family_unit_id.tolist()))),
+ "draws": list(config.draw_indices),
+ "n_rows_per_draw": {
+ str(d): int(np.count_nonzero(mask))
+ for d, mask in context.by_draw.items()
+ },
+ "n_rows_membership_differs": int(np.count_nonzero(differs)),
+ "n_current_law_rows": int(np.count_nonzero(population_base)),
+ "extra_columns_ignored": list(normalized.extra_columns),
+ }
+ return GroupBreakdownResult(
+ kind=PROJECTION_KIND,
+ statistic_id=statistic_id,
+ data_provenance=data_provenance,
+ registration_pointer=registration_pointer,
+ labels=tuple(output_labels),
+ post_hoc_labels=post_hoc,
+ scheme=assignment.scheme,
+ statistics=PROJECTION_STATISTICS,
+ conventions=conventions,
+ dimensions=tuple(dimensions),
+ assignment=assignment.as_dict(),
+ input_summary=input_summary,
+ upstream=recorded_upstream,
+ )
+
+
+# =========================================================================
+# Static breakdowns (exercises 2 and 4)
+# =========================================================================
+_STATIC_BASE_COLUMNS = (
+ "person_id",
+ "family_unit_id",
+ "weight",
+ "stratum",
+ "cluster",
+)
+
+
+def _design_identifiers(frame: pd.DataFrame) -> pd.DataFrame:
+ """Preserve integer survey identifiers; refuse lossy coercions.
+
+ Track U normalizes design identifiers to int64. Check the values
+ before calling it so a fractional cluster cannot silently merge with
+ another cluster, and a boolean cannot be mistaken for an identifier.
+ """
+
+ out = frame.copy()
+ limits = np.iinfo(np.int64)
+ for column in ("stratum", "cluster"):
+ values = [
+ _band_value(value, f"{column} identifier")
+ for value in frame[column].tolist()
+ ]
+ if any(value < limits.min or value > limits.max for value in values):
+ raise GroupBreakdownError(f"{column} identifiers exceed int64")
+ out[column] = np.asarray(values, dtype=np.int64)
+ return out
+
+
+def _static_frame(
+ rows: Any,
+ *,
+ id_column: str,
+ bool_columns: Sequence[str] = (),
+ amount_columns: Sequence[str] = (),
+) -> pd.DataFrame:
+ if not isinstance(rows, pd.DataFrame):
+ raise GroupBreakdownError("rows must be a DataFrame")
+ if not rows.columns.is_unique:
+ raise GroupBreakdownError("rows have repeated column names")
+ required = list(
+ dict.fromkeys(
+ (id_column, *_STATIC_BASE_COLUMNS, *bool_columns, *amount_columns)
+ )
+ )
+ missing = [c for c in required if c not in rows.columns]
+ if missing:
+ raise GroupBreakdownError(f"rows lack columns {missing}")
+ out = rows[required].reset_index(drop=True).copy()
+ if out.empty:
+ raise GroupBreakdownError("no rows")
+ if out[id_column].isna().any() or out[id_column].duplicated().any():
+ raise GroupBreakdownError(f"{id_column} must be present and unique")
+ identifiers = ["person_id", "family_unit_id", "stratum", "cluster"]
+ if out[identifiers].isna().any(axis=None):
+ raise GroupBreakdownError(
+ "person_id, family_unit_id, stratum and cluster must be present"
+ )
+ out["weight"] = _real_array(out["weight"].to_numpy(), "weights")
+ if (out["weight"] < 0).any():
+ raise GroupBreakdownError("weights must be non-negative")
+ for column in bool_columns:
+ values = out[column]
+ if not values.map(lambda v: isinstance(v, bool | np.bool_)).all():
+ raise GroupBreakdownError(f"{column} must be boolean")
+ out[column] = values.astype(bool)
+ for column in amount_columns:
+ amounts = _real_array(out[column].to_numpy(), column)
+ if (amounts < 0).any():
+ raise GroupBreakdownError(f"{column} must be non-negative")
+ out[column] = amounts
+ return _design_identifiers(out)
+
+
+def _design(
+ normalized: pd.DataFrame, design: Any
+) -> tuple[pd.MultiIndex, dict[str, Any]]:
+ try:
+ if isinstance(design, pd.DataFrame) and all(
+ c in design.columns for c in ("stratum", "cluster")
+ ):
+ design = _design_identifiers(design)
+ clusters_all = track_u._design_clusters(design)
+ except track_u.UniformCutTabulationError as error:
+ raise GroupBreakdownError(str(error)) from error
+ if clusters_all is None:
+ raise GroupBreakdownError(
+ "the design-based standard error needs the design frame (every "
+ "stratum and cluster of the sample's positive-weight persons)"
+ )
+ present = pd.MultiIndex.from_frame(normalized[["stratum", "cluster"]])
+ outside = present[~present.isin(clusters_all)].unique()
+ if len(outside):
+ raise GroupBreakdownError(
+ f"{len(outside)} (stratum, cluster) pairs of the rows are not in "
+ "the design frame"
+ )
+ sizes = pd.Series(1, index=clusters_all).groupby(level=0).sum()
+ return clusters_all, {
+ "method": "taylor_linearization",
+ "estimator": "uniform_cut_tabulation._design_se",
+ "domain": "full_sample_design",
+ "n_strata": int(len(sizes)),
+ "n_clusters": int(sizes.sum()),
+ "singleton_strata": [int(s) for s in sizes[sizes < 2].index],
+ "unit": "percentage points",
+ }
+
+
+def _static_split(
+ normalized: pd.DataFrame, floor_split_unit: str, seeds: Sequence[int]
+) -> dict[int, np.ndarray]:
+ if floor_split_unit not in STATIC_FLOOR_SPLIT_UNITS:
+ raise GroupBreakdownError(
+ f"floor_split_unit must be one of {sorted(STATIC_FLOOR_SPLIT_UNITS)}"
+ )
+ if floor_split_unit == FAMILY_UNIT_LINKED_BY_PERSON:
+ units = track_u.floor_split_units(normalized)
+ else:
+ units = normalized[floor_split_unit].to_numpy()
+ return half_split_masks(list(units), seeds)
+
+
+def _design_se(
+ frame: pd.DataFrame,
+ mask: np.ndarray,
+ indicator: np.ndarray,
+ clusters_all: pd.MultiIndex,
+) -> dict[str, Any]:
+ try:
+ return track_u._design_se(frame, mask, indicator, clusters_all)
+ except track_u.UniformCutTabulationError as error:
+ raise GroupBreakdownError(str(error)) from error
+
+
+def _static_floor(gaps: list[float], dropped: list[int]) -> dict[str, Any]:
+ try:
+ return {**track_u._floor_summary(gaps), "dropped_seeds": dropped}
+ except track_u.UniformCutTabulationError as error:
+ raise GroupBreakdownError(str(error)) from error
+
+
+@dataclass(frozen=True, eq=False)
+class _StaticSpec:
+ """How one static kind computes its statistics on a row mask."""
+
+ statistics: tuple[str, ...]
+ #: mask -> (values, undefined reasons, counts with unweighted_n,
+ #: weighted_n and numerators).
+ compute: Callable[[np.ndarray], tuple[dict, dict, dict]]
+ #: (mask, statistic) -> (design SE record or None, note or None).
+ design_se: Callable[[np.ndarray, str], tuple[dict | None, str | None]]
+ floor_notes: Mapping[str, str]
+ population: np.ndarray
+
+
+def _static_cell(
+ spec: _StaticSpec,
+ split: Mapping[int, np.ndarray],
+ dimension: Dimension,
+ category: Category,
+ mask: np.ndarray,
+) -> GroupCell:
+ values, reasons, counts = spec.compute(mask)
+ sides = {
+ seed: (spec.compute(mask & in_a)[0], spec.compute(mask & ~in_a)[0])
+ for seed, in_a in split.items()
+ }
+ statistics = []
+ n = counts["unweighted_n"]
+ for statistic in spec.statistics:
+ value = values[statistic]
+ defined = value is not None
+ floor_note = spec.floor_notes.get(statistic)
+ gaps, dropped = [], []
+ for seed, (side_a, side_b) in sides.items():
+ a, b = side_a[statistic], side_b[statistic]
+ if a is None or b is None:
+ dropped.append(seed)
+ else:
+ gaps.append(abs(a - b))
+ floor = _static_floor(gaps, dropped)
+ if defined:
+ se, se_note = spec.design_se(mask, statistic)
+ else:
+ se, se_note = None, "statistic undefined"
+ flags: dict[str, Any] = _n_flags(n)
+ if statistic in _NUMERATORS:
+ flags.update(
+ _numerator_flag([counts["numerators"][_NUMERATORS[statistic]]])
+ )
+ statistics.append(
+ StatisticCell(
+ statistic=statistic,
+ value=value,
+ defined=defined,
+ undefined_reason=None if defined else reasons[statistic],
+ unweighted_n=n,
+ weighted_n=counts["weighted_n"],
+ uncertainty={
+ "method": "design_se_and_half_split_floor",
+ "design_se": se,
+ "design_se_note": se_note,
+ "floor": floor,
+ "floor_note": floor_note,
+ },
+ flags=flags,
+ )
+ )
+ return GroupCell(
+ dimension=dimension.key,
+ dimension_label=dimension.label,
+ category=category.key,
+ label=category.label,
+ counts={
+ "n_rows": int(np.count_nonzero(mask)),
+ **counts,
+ },
+ statistics=tuple(statistics),
+ flags=_cell_flags(n, statistics),
+ )
+
+
+def _static_result(
+ *,
+ kind: str,
+ rows: pd.DataFrame,
+ assignment: GroupAssignment,
+ normalized: pd.DataFrame,
+ spec: _StaticSpec,
+ id_column: str,
+ design_summary: Mapping[str, Any],
+ floor_seeds: Sequence[int],
+ floor_split_unit: str,
+ data_provenance: str,
+ registration_pointer: str | None,
+ output_labels: list[str],
+ post_hoc: tuple[str, ...],
+ statistic_id: str,
+ upstream: Mapping[str, Any],
+ extra_conventions: Mapping[str, Any],
+) -> GroupBreakdownResult:
+ split = _static_split(normalized, floor_split_unit, floor_seeds)
+ summaries = {s["key"]: s for s in assignment.dimension_summaries}
+ weights = normalized["weight"].to_numpy()
+ dimensions = []
+ for dimension in assignment.scheme.dimensions:
+ cells = [
+ _static_cell(
+ spec,
+ split,
+ dimension,
+ category,
+ assignment.mask(dimension.key, category.key),
+ )
+ for category in dimension.categories
+ ]
+ unclassified = assignment.unclassified_mask(dimension.key)
+ in_population = unclassified & spec.population
+ dimensions.append(
+ _dimension_result(
+ dimension,
+ cells,
+ {
+ "n_rows": int(np.count_nonzero(unclassified)),
+ "unweighted_n": int(np.count_nonzero(in_population)),
+ "weighted_n": _fsum(weights[in_population]),
+ "reasons": dict(
+ summaries[dimension.key]["unclassified_reasons"]
+ ),
+ "statistics": None,
+ "note": (
+ "unclassified rows enter Total only; counts are "
+ "reported over the statistics' population"
+ ),
+ },
+ )
+ )
+ conventions = {
+ **_base_conventions(),
+ **extra_conventions,
+ "uncertainty": {
+ "draws": 1,
+ "design_se": dict(design_summary),
+ "floor": (
+ "for each floor seed the split units "
+ f"({floor_split_unit}: "
+ f"{STATIC_FLOOR_SPLIT_UNITS[floor_split_unit]}) of all rows "
+ "are split once (half_split_masks); each group's halves are "
+ "that split intersected with the group's mask; the floor "
+ "summarizes |side_a - side_b| over usable seeds "
+ "(uniform_cut_tabulation._floor_summary) and is undefined, "
+ "not zero, with fewer than two"
+ ),
+ "floor_seeds": list(_floor_seeds(floor_seeds)),
+ "floor_split_unit": floor_split_unit,
+ "scale": "half sample; not rescaled",
+ },
+ "unweighted_n": (
+ "the number of rows of the statistics' population in the cell"
+ ),
+ "id_column": id_column,
+ }
+ input_summary = {
+ "n_rows": int(len(normalized)),
+ "n_persons": int(normalized["person_id"].nunique()),
+ "n_family_units": int(normalized["family_unit_id"].nunique()),
+ "n_zero_weight": int((normalized["weight"] == 0).sum()),
+ "n_population_rows": int(np.count_nonzero(spec.population)),
+ "extra_columns_ignored": sorted(
+ str(c) for c in rows.columns if c not in normalized.columns
+ ),
+ }
+ return GroupBreakdownResult(
+ kind=kind,
+ statistic_id=statistic_id,
+ data_provenance=data_provenance,
+ registration_pointer=registration_pointer,
+ labels=tuple(output_labels),
+ post_hoc_labels=post_hoc,
+ scheme=assignment.scheme,
+ statistics=spec.statistics,
+ conventions=conventions,
+ dimensions=tuple(dimensions),
+ assignment=assignment.as_dict(),
+ input_summary=input_summary,
+ upstream=upstream,
+ )
+
+
+def _static_prelude(
+ rows: Any,
+ assignment: Any,
+ *,
+ id_column: str,
+ data_provenance: Any,
+ registration_pointer: Any,
+ labels: Any,
+ post_hoc_labels: Any,
+ statistic_id: Any,
+ upstream: Any,
+ bool_columns: Sequence[str] = (),
+ amount_columns: Sequence[str] = (),
+) -> tuple[pd.DataFrame, list[str], tuple[str, ...], str, dict[str, Any]]:
+ if not isinstance(rows, pd.DataFrame):
+ raise GroupBreakdownError("rows must be a DataFrame")
+ _nonempty_text(id_column, "id_column")
+ output_labels, post_hoc = _validate_provenance(
+ rows, data_provenance, registration_pointer, labels, post_hoc_labels
+ )
+ statistic_id = _statistic_id(statistic_id)
+ recorded_upstream = _upstream(upstream)
+ _check_assignment(assignment, rows, (id_column,))
+ normalized = _static_frame(
+ rows,
+ id_column=id_column,
+ bool_columns=bool_columns,
+ amount_columns=amount_columns,
+ )
+ return normalized, output_labels, post_hoc, statistic_id, recorded_upstream
+
+
+POVERTY_STATISTIC_ID = "group_breakdown_static_poverty"
+_NOT_A_RATIO = (
+ "not computed: uniform_cut_tabulation._design_se linearizes a weighted "
+ "ratio, and this statistic is a weighted count or a ratio of two "
+ "weighted counts"
+)
+_COUNT_FLOOR = (
+ "raw half-to-half gap in thousands, not rescaled: each half's weighted "
+ "count is on half the full sample's scale; this describes the split "
+ "diagnostic and is not a full-sample count standard error"
+)
+
+
+def tabulate_poverty_breakdown(
+ rows: pd.DataFrame,
+ assignment: GroupAssignment,
+ *,
+ design: pd.DataFrame,
+ data_provenance: str,
+ id_column: str = "person_id",
+ floor_seeds: Sequence[int] = DEFAULT_FLOOR_SEEDS,
+ floor_split_unit: str = FAMILY_UNIT_LINKED_BY_PERSON,
+ registration_pointer: str | None = None,
+ labels: Sequence[str] = (),
+ post_hoc_labels: Sequence[str] = (),
+ statistic_id: str = POVERTY_STATISTIC_ID,
+ upstream: Mapping[str, Any] | None = None,
+) -> GroupBreakdownResult:
+ """Break static poverty flags down by group (exercise 2).
+
+ ``rows`` carry ``id_column`` (unique), ``person_id``,
+ ``family_unit_id``, ``weight``, ``stratum``, ``cluster`` and the
+ booleans ``poor_baseline`` and ``poor_reform`` (for example
+ ``uniform_cut_tabulation.tabulation_rows`` output, ``id_column=
+ "observation_id"``); ``assignment`` is keyed on ``(id_column,)``.
+ Rates and their change come from ``uniform_cut_tabulation._rates`` and
+ their design SEs from its ``_design_se`` on the full design frame
+ (``design``), so the Total cell equals ``tabulate_uniform_cut``'s
+ ``all`` cell.
+ """
+
+ normalized, output_labels, post_hoc, statistic_id, recorded_upstream = (
+ _static_prelude(
+ rows,
+ assignment,
+ id_column=id_column,
+ data_provenance=data_provenance,
+ registration_pointer=registration_pointer,
+ labels=labels,
+ post_hoc_labels=post_hoc_labels,
+ statistic_id=statistic_id,
+ upstream=upstream,
+ bool_columns=("poor_baseline", "poor_reform"),
+ )
+ )
+ clusters_all, design_summary = _design(normalized, design)
+ weights = normalized["weight"].to_numpy()
+ poor_base = normalized["poor_baseline"].to_numpy()
+ poor_reform = normalized["poor_reform"].to_numpy()
+ base_f = poor_base.astype(np.float64)
+ reform_f = poor_reform.astype(np.float64)
+ indicators = {
+ POVERTY_RATE_CURRENT_LAW: base_f,
+ POVERTY_RATE_PROPOSAL: reform_f,
+ POVERTY_RATE_CHANGE: reform_f - base_f,
+ }
+
+ def compute(mask: np.ndarray) -> tuple[dict, dict, dict]:
+ try:
+ rates = track_u._rates(normalized, mask)
+ except track_u.UniformCutTabulationError as error:
+ raise GroupBreakdownError(str(error)) from error
+ counts = {
+ "unweighted_n": int(np.count_nonzero(mask)),
+ "weighted_n": _fsum(weights[mask]),
+ "numerators": {
+ "n_poor_current_law": int(np.count_nonzero(mask & poor_base)),
+ "n_poor_proposal": int(np.count_nonzero(mask & poor_reform)),
+ },
+ }
+ if not rates["defined"]:
+ reason = rates["undefined_reason"]
+ return (
+ {s: None for s in POVERTY_STATISTICS},
+ {s: reason for s in POVERTY_STATISTICS},
+ counts,
+ )
+ number_base = _fsum(weights[mask & poor_base])
+ number_reform = _fsum(weights[mask & poor_reform])
+ values = {
+ POVERTY_RATE_CURRENT_LAW: rates["baseline_rate"],
+ POVERTY_RATE_PROPOSAL: rates["reform_rate"],
+ POVERTY_RATE_CHANGE: rates["delta"],
+ NUMBER_POOR_CURRENT_LAW: number_base / 1000.0,
+ NUMBER_POOR_PROPOSAL: number_reform / 1000.0,
+ NUMBER_POOR_CHANGE: (number_reform - number_base) / 1000.0,
+ NUMBER_POOR_PERCENT_CHANGE: (
+ 100.0 * (number_reform - number_base) / number_base
+ if number_base > 0.0
+ else None
+ ),
+ }
+ reasons = {}
+ if values[NUMBER_POOR_PERCENT_CHANGE] is None:
+ reasons[NUMBER_POOR_PERCENT_CHANGE] = (
+ "nobody in the cell is poor under current law"
+ )
+ return values, reasons, counts
+
+ def design_se(mask: np.ndarray, statistic: str):
+ if statistic in indicators:
+ return (
+ _design_se(
+ normalized, mask, indicators[statistic], clusters_all
+ ),
+ None,
+ )
+ return None, _NOT_A_RATIO
+
+ spec = _StaticSpec(
+ statistics=POVERTY_STATISTICS,
+ compute=compute,
+ design_se=design_se,
+ floor_notes={
+ NUMBER_POOR_CURRENT_LAW: _COUNT_FLOOR,
+ NUMBER_POOR_PROPOSAL: _COUNT_FLOOR,
+ NUMBER_POOR_CHANGE: _COUNT_FLOOR,
+ },
+ population=np.ones(len(normalized), dtype=bool),
+ )
+ return _static_result(
+ kind=POVERTY_KIND,
+ rows=rows,
+ assignment=assignment,
+ normalized=normalized,
+ spec=spec,
+ id_column=id_column,
+ design_summary=design_summary,
+ floor_seeds=floor_seeds,
+ floor_split_unit=floor_split_unit,
+ data_provenance=data_provenance,
+ registration_pointer=registration_pointer,
+ output_labels=output_labels,
+ post_hoc=post_hoc,
+ statistic_id=statistic_id,
+ upstream=recorded_upstream,
+ extra_conventions={"population": "every row"},
+ )
+
+
+SHARE_STATISTIC_ID = "group_breakdown_static_share"
+
+
+def tabulate_share_breakdown(
+ rows: pd.DataFrame,
+ assignment: GroupAssignment,
+ *,
+ indicator_column: str,
+ design: pd.DataFrame,
+ data_provenance: str,
+ id_column: str = "person_id",
+ floor_seeds: Sequence[int] = DEFAULT_FLOOR_SEEDS,
+ floor_split_unit: str = FAMILY_UNIT_LINKED_BY_PERSON,
+ registration_pointer: str | None = None,
+ labels: Sequence[str] = (),
+ post_hoc_labels: Sequence[str] = (),
+ statistic_id: str = SHARE_STATISTIC_ID,
+ upstream: Mapping[str, Any] | None = None,
+) -> GroupBreakdownResult:
+ """Break a static weighted share down by group (exercise 4).
+
+ The share is ``100 * sum w 1{indicator} / sum w``, the formula of
+ ``min_benefit_track_m.tabulation._share``, with the design SE of
+ ``uniform_cut_tabulation._design_se`` and the half-sample floor; the
+ Total cell equals ``tabulate_track_m``'s ``all`` cell for the same
+ indicator.
+ """
+
+ _nonempty_text(indicator_column, "indicator_column")
+ normalized, output_labels, post_hoc, statistic_id, recorded_upstream = (
+ _static_prelude(
+ rows,
+ assignment,
+ id_column=id_column,
+ data_provenance=data_provenance,
+ registration_pointer=registration_pointer,
+ labels=labels,
+ post_hoc_labels=post_hoc_labels,
+ statistic_id=statistic_id,
+ upstream=upstream,
+ bool_columns=(indicator_column,),
+ )
+ )
+ clusters_all, design_summary = _design(normalized, design)
+ weights = normalized["weight"].to_numpy()
+ receiving = normalized[indicator_column].to_numpy()
+
+ def compute(mask: np.ndarray) -> tuple[dict, dict, dict]:
+ total = _fsum(weights[mask])
+ counts = {
+ "unweighted_n": int(np.count_nonzero(mask)),
+ "weighted_n": total,
+ "numerators": {
+ "n_indicator": int(np.count_nonzero(mask & receiving))
+ },
+ }
+ if not mask.any():
+ return {SHARE: None}, {SHARE: "empty cell"}, counts
+ if total <= 0:
+ return {SHARE: None}, {SHARE: "zero total weight"}, counts
+ numerator = _fsum(weights[mask & receiving])
+ counts["weighted_indicator"] = numerator
+ return {SHARE: 100.0 * numerator / total}, {}, counts
+
+ def design_se(mask: np.ndarray, statistic: str):
+ return (
+ _design_se(
+ normalized, mask, receiving.astype(np.float64), clusters_all
+ ),
+ None,
+ )
+
+ spec = _StaticSpec(
+ statistics=SHARE_STATISTICS,
+ compute=compute,
+ design_se=design_se,
+ floor_notes={},
+ population=np.ones(len(normalized), dtype=bool),
+ )
+ return _static_result(
+ kind=SHARE_KIND,
+ rows=rows,
+ assignment=assignment,
+ normalized=normalized,
+ spec=spec,
+ id_column=id_column,
+ design_summary=design_summary,
+ floor_seeds=floor_seeds,
+ floor_split_unit=floor_split_unit,
+ data_provenance=data_provenance,
+ registration_pointer=registration_pointer,
+ output_labels=output_labels,
+ post_hoc=post_hoc,
+ statistic_id=statistic_id,
+ upstream=recorded_upstream,
+ extra_conventions={
+ "population": "every row",
+ "indicator_column": indicator_column,
+ },
+ )
+
+
+STATIC_BENEFIT_STATISTIC_ID = "group_breakdown_static_benefit_change"
+_NO_PERCENTILE_SE = (
+ "not computed: a design-based standard error of a weighted percentile "
+ "needs a density estimate or replicate weights, which this module does "
+ "not implement"
+)
+
+
+def tabulate_static_benefit_breakdown(
+ rows: pd.DataFrame,
+ assignment: GroupAssignment,
+ *,
+ design: pd.DataFrame,
+ data_provenance: str,
+ base_column: str = "benefit_base",
+ reform_column: str = "benefit_reform",
+ id_column: str = "person_id",
+ floor_seeds: Sequence[int] = DEFAULT_FLOOR_SEEDS,
+ floor_split_unit: str = FAMILY_UNIT_LINKED_BY_PERSON,
+ registration_pointer: str | None = None,
+ labels: Sequence[str] = (),
+ post_hoc_labels: Sequence[str] = (),
+ statistic_id: str = STATIC_BENEFIT_STATISTIC_ID,
+ upstream: Mapping[str, Any] | None = None,
+) -> GroupBreakdownResult:
+ """MINT8's benefit statistics for a static test, by group.
+
+ Over the rows with a positive current-law amount (``base_column``):
+ percent with a decrease / increase (and the unaffected rest), each
+ with the design SE of ``uniform_cut_tabulation._design_se`` (the group
+ and population as the domain), and the weighted 10th, 50th and 90th
+ percentiles of the individual percent change, which carry no design
+ SE (each cell says why). Every statistic has the half-sample floor.
+ """
+
+ normalized, output_labels, post_hoc, statistic_id, recorded_upstream = (
+ _static_prelude(
+ rows,
+ assignment,
+ id_column=id_column,
+ data_provenance=data_provenance,
+ registration_pointer=registration_pointer,
+ labels=labels,
+ post_hoc_labels=post_hoc_labels,
+ statistic_id=statistic_id,
+ upstream=upstream,
+ amount_columns=(base_column, reform_column),
+ )
+ )
+ clusters_all, design_summary = _design(normalized, design)
+ weights = normalized["weight"].to_numpy()
+ population, change, decrease, increase = classify_changes(
+ normalized[base_column].to_numpy(),
+ normalized[reform_column].to_numpy(),
+ )
+ numerators = _exact_numerators(weights)
+ positive = weights > 0.0
+ unaffected = population & ~decrease & ~increase
+ indicators = {
+ PERCENT_DECREASE: decrease.astype(np.float64),
+ PERCENT_INCREASE: increase.astype(np.float64),
+ PERCENT_UNAFFECTED: unaffected.astype(np.float64),
+ }
+
+ def compute(mask: np.ndarray) -> tuple[dict, dict, dict]:
+ values, reasons = _mint_values(
+ weights,
+ change,
+ decrease,
+ increase,
+ numerators,
+ positive,
+ mask & population,
+ )
+ counts = {
+ "unweighted_n": values.pop("n_current_law"),
+ "weighted_n": values.pop("weight_current_law"),
+ "numerators": {
+ "n_decrease": values.pop("n_decrease"),
+ "n_increase": values.pop("n_increase"),
+ },
+ }
+ return values, reasons, counts
+
+ def design_se(mask: np.ndarray, statistic: str):
+ if statistic in indicators:
+ return (
+ _design_se(
+ normalized,
+ mask & population,
+ indicators[statistic],
+ clusters_all,
+ ),
+ None,
+ )
+ return None, _NO_PERCENTILE_SE
+
+ spec = _StaticSpec(
+ statistics=STATIC_BENEFIT_STATISTICS,
+ compute=compute,
+ design_se=design_se,
+ floor_notes={},
+ population=population,
+ )
+ return _static_result(
+ kind=STATIC_BENEFIT_KIND,
+ rows=rows,
+ assignment=assignment,
+ normalized=normalized,
+ spec=spec,
+ id_column=id_column,
+ design_summary=design_summary,
+ floor_seeds=floor_seeds,
+ floor_split_unit=floor_split_unit,
+ data_provenance=data_provenance,
+ registration_pointer=registration_pointer,
+ output_labels=output_labels,
+ post_hoc=post_hoc,
+ statistic_id=statistic_id,
+ upstream=recorded_upstream,
+ extra_conventions={
+ "population": (
+ f"rows with {base_column} > 0 (current-law beneficiaries)"
+ ),
+ "change_definition": CHANGE_DEFINITION,
+ "base_column": base_column,
+ "reform_column": reform_column,
+ },
+ )
diff --git a/tests/estimates/test_group_breakdown.py b/tests/estimates/test_group_breakdown.py
new file mode 100644
index 00000000..91cf560b
--- /dev/null
+++ b/tests/estimates/test_group_breakdown.py
@@ -0,0 +1,1773 @@
+"""Group-breakdown tabulation core (NASI package G3), on INVENTED rows only.
+
+Every row in this module is invented for testing. No PSID record, model
+projection, oracle benefit or comparator value is read or produced, and no
+number here is a model result. Expected values are hand-computed in the
+comments next to each assertion, or come from the existing tabulators
+(``cola_age_profile``, ``uniform_cut_tabulation``, Track M's
+``tabulation``) run on the same invented rows, or from an independent
+plain-Python recomputation written in this file.
+
+Invariants exercised with hypothesis:
+
+* within each dimension, classified categories plus unclassified rows
+ partition Total, in unweighted and weighted counts (per draw);
+* the ratio of scenario means and every MINT8 statistic are invariant to
+ scaling all weights (exactly, for powers of two);
+* Total and age-group cells equal ``tabulate_cola_age_profile``; Total and
+ sex/marital cells equal ``tabulate_uniform_cut``; Total and sex cells
+ equal ``tabulate_track_m`` (differential tests);
+* the percentile rule equals an independent exact-Fraction reference and
+ agrees with ``uniform_cut_track_u.diagnostics.weighted_quantile`` away
+ from exact ties; percentiles lie within [min, max] of the changes;
+* percent decrease + unaffected + increase = 100;
+* floors are undefined, never zero, with fewer than two usable seeds;
+* a group's halves are the full-sample split restricted to the group.
+"""
+
+from __future__ import annotations
+
+import dataclasses
+import json
+import math
+from fractions import Fraction
+
+import numpy as np
+import pandas as pd
+import pytest
+from hypothesis import HealthCheck, given, settings
+from hypothesis import strategies as st
+
+from populace_dynamics.estimates import cola_age_profile as a7
+from populace_dynamics.estimates import group_breakdown as gb
+from populace_dynamics.estimates import uniform_cut_tabulation as ut
+from populace_dynamics.min_benefit_track_m import tabulation as track_m
+from populace_dynamics.uniform_cut_track_u.diagnostics import (
+ weighted_quantile,
+)
+
+SLOW = settings(
+ max_examples=40,
+ deadline=None,
+ suppress_health_check=[HealthCheck.too_slow],
+)
+POINTER = (
+ "https://github.com/PolicyEngine/microcosm-dynamics/issues/42"
+ "#issuecomment-123"
+)
+
+
+# =========================================================================
+# Invented-row builders
+# =========================================================================
+def _projection_row(
+ draw,
+ person_id,
+ *,
+ birth_year,
+ base,
+ reform,
+ weight=1.0,
+ family_unit_id=None,
+ **attributes,
+):
+ """One INVENTED cola_age_profile row; one component carries it all."""
+
+ components = (
+ {"retired_worker": {"base": base, "reform": reform}}
+ if base > 0 or reform > 0
+ else {}
+ )
+ return {
+ "draw": draw,
+ "person_id": person_id,
+ "family_unit_id": (
+ person_id if family_unit_id is None else family_unit_id
+ ),
+ "weight": weight,
+ "birth_year": birth_year,
+ "beneficiary_base": base > 0,
+ "beneficiary_reform": reform > 0,
+ "benefit_base": base,
+ "benefit_reform": reform,
+ "benefit_components": components,
+ "age": 2030 - birth_year,
+ **attributes,
+ }
+
+
+def _sex_age_scheme(scheme_id="sex_age", age=None):
+ """MINT8's Total, Sex and Age dimensions (age bands swappable)."""
+
+ base = gb.get_scheme("mint8_beneficiary_annual")
+ keep = {"total", "sex", "age"}
+ return gb.derive_scheme(
+ base,
+ scheme_id=scheme_id,
+ title="test: total, sex, age",
+ drop=[d.key for d in base.dimensions if d.key not in keep],
+ replace={} if age is None else {"age": age},
+ )
+
+
+def _sex_marital_scheme(scheme_id="sex_marital"):
+ base = gb.get_scheme("mint8_beneficiary_annual")
+ keep = {"total", "sex", "marital_status"}
+ return gb.derive_scheme(
+ base,
+ scheme_id=scheme_id,
+ title="test: total, sex, marital status",
+ drop=[d.key for d in base.dimensions if d.key not in keep],
+ )
+
+
+def _project(frame, scheme, config, **kwargs):
+ assignment = gb.assign_groups(
+ frame,
+ scheme,
+ {"sex": "sex", "age": "age"},
+ key_columns=("draw", "person_id"),
+ unclassified_codes=kwargs.pop("unclassified_codes", None),
+ )
+ return gb.tabulate_projection_breakdown(
+ frame,
+ assignment,
+ data_provenance="invented",
+ config=config,
+ **kwargs,
+ )
+
+
+@st.composite
+def projection_frames(draw, allow_difference=False):
+ """INVENTED per-person-per-draw rows with sex and age attributes."""
+
+ n_persons = draw(st.integers(1, 9))
+ n_draws = draw(st.integers(1, 3))
+ persons = [
+ {
+ "family": draw(st.integers(0, 4)),
+ "birth_year": draw(st.integers(1935, 1985)),
+ "weight": draw(st.sampled_from([0.0, 0.5, 1.0, 2.0, 3.0, 7.25])),
+ "sex": draw(st.sampled_from(["female", "male", None])),
+ }
+ for _ in range(n_persons)
+ ]
+ bases = [0.0, 100.0, 495.0, 500.0, 777.0, 1234.5]
+ factors = [0.9, 0.99, 1.0, 1.01, 1.05]
+ rows = []
+ for d in range(n_draws):
+ for p, info in enumerate(persons):
+ base = draw(st.sampled_from(bases))
+ reform = base * draw(st.sampled_from(factors))
+ if allow_difference:
+ kind = draw(st.sampled_from(["same", "lost", "gained"]))
+ if kind == "lost":
+ reform = 0.0
+ elif kind == "gained" and base == 0.0:
+ reform = 250.0
+ rows.append(
+ _projection_row(
+ d,
+ p,
+ birth_year=info["birth_year"],
+ base=base,
+ reform=reform,
+ weight=info["weight"],
+ family_unit_id=info["family"],
+ sex=info["sex"],
+ )
+ )
+ frame = pd.DataFrame(rows)
+ config = a7.ColaAgeProfileConfig(
+ draw_indices=tuple(range(n_draws)),
+ allow_membership_difference=allow_difference,
+ )
+ return frame, config
+
+
+def _poverty_frame(rng, n):
+ """INVENTED Track U-style observations (not PSID)."""
+
+ statuses = ["married", "widowed", "divorced", "never_married"]
+ return pd.DataFrame(
+ {
+ "observation_id": [f"o{i}" for i in range(n)],
+ "person_id": np.arange(n) + 1,
+ "family_unit_id": rng.integers(0, max(1, n // 2), n),
+ "weight": rng.choice([0.5, 1.0, 2.0, 4.0], n),
+ "sex": rng.choice(["female", "male"], n),
+ "marital_status_4": rng.choice([*statuses, ut.UNCLASSIFIED], n),
+ "birth_year": rng.choice([1937, 1939, 1941], n),
+ "stratum": rng.integers(1, 4, n),
+ "cluster": rng.integers(1, 3, n),
+ "poor_baseline": rng.random(n) < 0.3,
+ "poor_reform": rng.random(n) < 0.5,
+ }
+ )
+
+
+def _design(rows, *extra):
+ pairs = rows[["stratum", "cluster"]].drop_duplicates()
+ if extra:
+ pairs = pd.concat(
+ [pairs, pd.DataFrame(extra, columns=["stratum", "cluster"])],
+ ignore_index=True,
+ )
+ return pairs.reset_index(drop=True)
+
+
+def _track_m_frame(rng, n):
+ """INVENTED Track M-style evaluation rows (not PSID)."""
+
+ frame = pd.DataFrame(
+ {
+ "person_id": [str(i) for i in range(n)],
+ "family_unit_id": rng.integers(0, max(1, n // 2), n),
+ "weight": rng.choice([0.5, 1.0, 3.0], n),
+ "sex": rng.choice(["female", "male", "unknown"], n),
+ "stratum": rng.integers(1, 4, n),
+ "cluster": rng.integers(1, 3, n),
+ **{f"receives_{k}": rng.random(n) < 0.4 for k in (2, 3, 4, 5)},
+ }
+ )
+ frame.attrs["provenance_kind"] = track_m.INVENTED
+ return frame
+
+
+# =========================================================================
+# Category schemes
+# =========================================================================
+def test_mint8_scheme_rows_follow_the_brief_in_page_order():
+ scheme = gb.MINT8_SCHEME
+ assert [(d.label, list(d.labels)) for d in scheme.dimensions] == [
+ ("Total", ["Total"]),
+ ("Sex", ["Female", "Male"]),
+ (
+ "Race and ethnicity",
+ [
+ "Hispanic or Latino, any race",
+ "White, non-Hispanic",
+ "Black or African American, non-Hispanic",
+ "All other races, non-Hispanic",
+ ],
+ ),
+ ("Country of birth", ["United States", "Other countries"]),
+ ("Age", ["60–69", "70–79", "80–89", "90 or older"]),
+ (
+ "Marital status",
+ ["Married", "Divorced", "Widowed", "Never married"],
+ ),
+ (
+ "Highest education level",
+ [
+ "Graduate",
+ "Bachelor",
+ "Associate",
+ "High school",
+ "Less than high school",
+ ],
+ ),
+ ("Current-law poverty status", ["Above poverty", "In poverty"]),
+ (
+ "Current-law household income quintile",
+ ["Highest", "Second highest", "Middle", "Second lowest", "Lowest"],
+ ),
+ (
+ "Current-law benefit type",
+ [
+ "Retired worker only",
+ "Widow(er) (includes dually entitled)",
+ "Spousal (includes dually entitled)",
+ "Disabled worker only",
+ ],
+ ),
+ ]
+ assert not scheme.composite
+ assert any("user-guide.html" in c for c in scheme.citations)
+ assert any("increase-payroll-tax-rate" in c for c in scheme.citations)
+
+
+def test_education_bands_follow_ssa_year_definitions():
+ education = gb.MINT8_SCHEME.dimension("education")
+ # Graduate > 16, Bachelor 16, Associate 14-15, High school 12-13,
+ # Less than high school < 12 (MINT8 Table User Guide).
+ expected = {
+ 0: "less_than_high_school",
+ 11: "less_than_high_school",
+ 12: "high_school",
+ 13: "high_school",
+ 14: "associate",
+ 15: "associate",
+ 16: "bachelor",
+ 17: "graduate",
+ 20: "graduate",
+ }
+ for years, key in expected.items():
+ hits = [c.key for c in education.categories if c.contains(years)]
+ assert hits == [key], years
+
+
+def test_poverty_and_cohort_schemes():
+ poverty = gb.MINT8_POVERTY_SCHEME
+ assert "household_income_quintile" not in poverty.keys
+ assert poverty.keys == tuple(
+ k for k in gb.MINT8_SCHEME.keys if k != "household_income_quintile"
+ )
+ cohort = gb.MINT8_COHORT_SCHEME
+ assert cohort.keys == (
+ "total",
+ "sex",
+ "race_ethnicity",
+ "country_of_birth",
+ "education",
+ "initial_aime_quintile",
+ "lifetime_payroll_tax_quintile",
+ "lifetime_payroll_tax_quintile_shared",
+ )
+ for scheme in (
+ gb.MINT8_ANNUAL_WITH_LIFETIME_SCHEME,
+ gb.MINT8_POVERTY_WITH_LIFETIME_SCHEME,
+ ):
+ assert scheme.composite
+ assert scheme.keys[-3:] == cohort.keys[-3:]
+ assert any("composite" in note for note in scheme.notes)
+
+
+def test_quintile_rows_are_ranked_highest_first():
+ quintile = gb.MINT8_SCHEME.dimension("household_income_quintile")
+ assert [(c.label, c.rank) for c in quintile.categories] == [
+ ("Highest", 5),
+ ("Second highest", 4),
+ ("Middle", 3),
+ ("Second lowest", 2),
+ ("Lowest", 1),
+ ]
+
+
+def test_age_band_sets_include_the_exercise_bands():
+ bands = gb.AGE_BAND_SETS["dynasim3_exercises_1_3"]
+ assert bands.labels == tuple(g.label for g in a7.DEFAULT_AGE_GROUPS)
+ assert [(c.lower, c.upper) for c in bands.categories] == [
+ (g.lower, g.upper) for g in a7.DEFAULT_AGE_GROUPS
+ ]
+ derived = gb.with_age_bands(gb.MINT8_SCHEME, "dynasim3_exercises_1_3")
+ assert derived.dimension("age") is bands
+ assert derived.composite
+ assert derived.keys == gb.MINT8_SCHEME.keys
+ taxpayer = gb.AGE_BAND_SETS["mint8_taxpayer"]
+ assert taxpayer.labels[-1] == "70 or older"
+ with pytest.raises(gb.GroupBreakdownError, match="band_set"):
+ gb.with_age_bands(gb.MINT8_SCHEME, "nope")
+ with pytest.raises(gb.GroupBreakdownError, match="no dimension"):
+ gb.with_age_bands(gb.MINT8_COHORT_SCHEME, "mint8_beneficiary")
+
+
+def test_registry_takes_an_alternate_scheme_from_existing_row_labels():
+ # Butrica and Uccello (2004) rows, as uniform_cut_tabulation records
+ # them (labels only; the race codes are this test's own).
+ race_rows = ut.NOT_COMPUTED_REPORT_ROWS["race_ethnicity"]["rows"]
+ race = gb.Dimension(
+ key="race_ethnicity",
+ label=ut.NOT_COMPUTED_REPORT_ROWS["race_ethnicity"]["section"],
+ kind=gb.CATEGORICAL,
+ categories=tuple(
+ gb.Category(f"race_{i}", label, codes=(f"r{i}",))
+ for i, label in enumerate(race_rows)
+ ),
+ source="uniform_cut_tabulation.NOT_COMPUTED_REPORT_ROWS",
+ )
+ scheme = gb.CategoryScheme(
+ scheme_id="test_boomers2004_rows",
+ title="test alternate scheme",
+ population="test population",
+ dimensions=(
+ gb.MINT8_SCHEME.dimension("total"),
+ race,
+ ),
+ citations=("Butrica and Uccello (2004), labels via Track U",),
+ )
+ try:
+ assert gb.register_scheme(scheme) is scheme
+ assert gb.get_scheme("test_boomers2004_rows") is scheme
+ assert gb.register_scheme(scheme) is scheme # idempotent
+ twin = gb.CategoryScheme(
+ scheme_id="test_boomers2004_rows",
+ title="another",
+ population="p",
+ dimensions=scheme.dimensions,
+ citations=scheme.citations,
+ )
+ with pytest.raises(gb.GroupBreakdownError, match="already"):
+ gb.register_scheme(twin)
+ assert gb.register_scheme(twin, replace=True) is twin
+ finally:
+ gb._REGISTRY.pop("test_boomers2004_rows", None)
+ with pytest.raises(gb.GroupBreakdownError, match="already"):
+ gb.register_scheme(
+ gb.derive_scheme(
+ gb.MINT8_SCHEME,
+ scheme_id="mint8_beneficiary_annual",
+ title="x",
+ ),
+ replace=True,
+ )
+ with pytest.raises(gb.GroupBreakdownError, match="no registered"):
+ gb.get_scheme("missing_scheme")
+
+
+@pytest.mark.parametrize(
+ ("build", "match"),
+ [
+ (
+ lambda: gb.Dimension(
+ "x",
+ "X",
+ gb.BAND,
+ (
+ gb.Category("a", "A", lower=0, upper=10),
+ gb.Category("b", "B", lower=10, upper=20),
+ ),
+ ),
+ "overlap",
+ ),
+ (
+ lambda: gb.Dimension(
+ "x",
+ "X",
+ gb.CATEGORICAL,
+ (
+ gb.Category("a", "A", codes=("c",)),
+ gb.Category("b", "B", codes=("c",)),
+ ),
+ ),
+ "two categories",
+ ),
+ (
+ lambda: gb.Dimension(
+ "x",
+ "X",
+ gb.QUINTILE,
+ tuple(gb.Category(f"q{i}", f"Q{i}", rank=1) for i in range(5)),
+ ),
+ "ranks 1-5",
+ ),
+ (
+ lambda: gb.Dimension(
+ "x", "X", gb.CATEGORICAL, (gb.Category("a", "A", rank=1),)
+ ),
+ "may not carry",
+ ),
+ (lambda: gb.Category("a", gb.UNCLASSIFIED), "unclassified"),
+ (lambda: gb.Category("Bad Key", "A"), "lower-case"),
+ (
+ lambda: gb.CategoryScheme(
+ "s",
+ "t",
+ "p",
+ (gb.SEX_DIMENSION,),
+ ("c",),
+ ),
+ "first dimension is Total",
+ ),
+ (
+ lambda: gb.CategoryScheme(
+ "s", "t", "p", (gb.TOTAL_DIMENSION,), ()
+ ),
+ "citation",
+ ),
+ (
+ lambda: gb.derive_scheme(
+ gb.MINT8_SCHEME, scheme_id="s", title="t", drop=("total",)
+ ),
+ "cannot be dropped",
+ ),
+ (
+ lambda: gb.derive_scheme(
+ gb.MINT8_SCHEME,
+ scheme_id="s",
+ title="t",
+ replace={"age": gb.SEX_DIMENSION},
+ ),
+ "with that key",
+ ),
+ ],
+)
+def test_scheme_validation_refuses(build, match):
+ with pytest.raises(gb.GroupBreakdownError, match=match):
+ build()
+
+
+# =========================================================================
+# Group assignment
+# =========================================================================
+def _assignment_frame():
+ """INVENTED: six persons with every unclassified route."""
+
+ return pd.DataFrame(
+ {
+ "person_id": [1, 2, 3, 4, 5, 6],
+ "weight": [1.0, 2.0, 3.0, 4.0, 0.0, 5.0],
+ "sex": ["female", "male", None, "female", "male", np.nan],
+ "marital": [
+ "married",
+ "separated",
+ "widowed",
+ None,
+ "never_married",
+ "divorced",
+ ],
+ "age": [59, 65, 75.0, 92, np.nan, 85],
+ "income": [10.0, 20.0, 30.0, 40.0, 50.0, None],
+ }
+ )
+
+
+def _assignment_scheme():
+ base = gb.MINT8_SCHEME
+ keep = {
+ "total",
+ "sex",
+ "marital_status",
+ "age",
+ "household_income_quintile",
+ }
+ return gb.derive_scheme(
+ base,
+ scheme_id="assignment_test",
+ title="t",
+ drop=[d.key for d in base.dimensions if d.key not in keep],
+ )
+
+
+def _assign(frame=None, **kwargs):
+ frame = _assignment_frame() if frame is None else frame
+ options = {
+ "key_columns": ("person_id",),
+ "unclassified_codes": {"marital_status": ("separated",)},
+ **kwargs,
+ }
+ return gb.assign_groups(
+ frame,
+ _assignment_scheme(),
+ {
+ "sex": "sex",
+ "marital_status": "marital",
+ "age": "age",
+ "household_income_quintile": "income",
+ },
+ **options,
+ )
+
+
+def test_assignment_counts_every_unclassified_route():
+ assignment = _assign()
+ summaries = {s["key"]: s for s in assignment.dimension_summaries}
+ assert summaries["sex"]["unclassified_reasons"] == {"missing": 2}
+ assert summaries["marital_status"]["unclassified_reasons"] == {
+ "code:separated": 1,
+ "missing": 1,
+ }
+ # age 59 is below 60-69; NaN is missing
+ assert summaries["age"]["unclassified_reasons"] == {
+ "missing": 1,
+ "outside_bands": 1,
+ }
+ assert summaries["age"]["n_by_category"] == {
+ "age_60_69": 1,
+ "age_70_79": 1,
+ "age_80_89": 1,
+ "age_90_plus": 1,
+ }
+ assert list(assignment.mask("sex", "female")) == [
+ True,
+ False,
+ False,
+ True,
+ False,
+ False,
+ ]
+ long = assignment.long
+ assert len(long) == 6 * len(assignment.scheme.dimensions)
+ assert set(long.columns) == {
+ "row_id",
+ "dimension",
+ "label",
+ "category",
+ "classified",
+ }
+ row3 = long[(long.row_id == 2) & (long.dimension == "sex")].iloc[0]
+ assert (row3.label, row3.category, bool(row3.classified)) == (
+ gb.UNCLASSIFIED,
+ gb.UNCLASSIFIED,
+ False,
+ )
+ total = long[long.dimension == "total"]
+ assert set(total.label) == {"Total"} and total.classified.all()
+ json.dumps(assignment.as_dict(), allow_nan=False)
+
+
+def test_quintiles_by_hand():
+ # Weights 1, 2, 3, 4, 0 over incomes 10..50 (the sixth is missing):
+ # W = 10; cumulative 1, 3, 6, 10 (50 has no mass). t1: 2 -> 20;
+ # t2: 4 -> 30; t3: 6 hits exactly at 30 -> midpoint (30 + 40) / 2 =
+ # 35; t4: 8 -> 40. Ranks: 10 -> 1, 20 -> 1 (equal to t1 is lower),
+ # 30 -> 2, 40 -> 4, 50 -> 5.
+ assignment = _assign()
+ (summary,) = [
+ s
+ for s in assignment.dimension_summaries
+ if s["key"] == "household_income_quintile"
+ ]
+ (record,) = summary["quintile_partitions"]
+ assert record["thresholds"] == [20.0, 30.0, 35.0, 40.0]
+ assert record["n_rows"] == 5 and record["n_positive_weight"] == 4
+ codes = [
+ assignment._category_index["household_income_quintile"][i]
+ for i in range(6)
+ ]
+ assert codes == [
+ "lowest",
+ "lowest",
+ "second_lowest",
+ "second_highest",
+ "highest",
+ gb.UNCLASSIFIED,
+ ]
+
+
+def test_quintiles_are_cut_within_partitions():
+ frame = pd.DataFrame(
+ {
+ "draw": [0, 0, 0, 0, 0, 1, 1, 1, 1, 1],
+ "person_id": [1, 2, 3, 4, 5] * 2,
+ "weight": [1.0] * 10,
+ "income": [1.0, 2.0, 3.0, 4.0, 5.0, 10.0, 20.0, 30.0, 40.0, 50.0],
+ }
+ )
+ scheme = gb.derive_scheme(
+ gb.MINT8_SCHEME,
+ scheme_id="partition_test",
+ title="t",
+ drop=[
+ d.key
+ for d in gb.MINT8_SCHEME.dimensions
+ if d.key not in {"total", "household_income_quintile"}
+ ],
+ )
+ assignment = gb.assign_groups(
+ frame,
+ scheme,
+ {"household_income_quintile": "income"},
+ key_columns=("draw", "person_id"),
+ quintile_partition=("draw",),
+ )
+ codes = list(assignment._category_index["household_income_quintile"])
+ order = ["lowest", "second_lowest", "middle", "second_highest", "highest"]
+ assert codes == order + order
+
+
+@pytest.mark.parametrize(
+ ("change", "match"),
+ [
+ ({"marital": "cohabiting"}, "no category"),
+ ({"age": 65.5}, "integers"),
+ ({"age": True}, "boolean"),
+ ({"sex": 1}, "strings"),
+ ({"income": np.inf}, "finite"),
+ ({"person_id": 2}, "identify each row"),
+ ],
+)
+def test_assignment_refuses_bad_values(change, match):
+ frame = _assignment_frame()
+ for column, value in change.items():
+ frame[column] = frame[column].astype(object)
+ frame.loc[0, column] = value
+ with pytest.raises(gb.GroupBreakdownError, match=match):
+ _assign(frame)
+
+
+def test_assignment_refuses_bad_maps():
+ frame = _assignment_frame()
+ with pytest.raises(gb.GroupBreakdownError, match="missing"):
+ gb.assign_groups(
+ frame,
+ _assignment_scheme(),
+ {"sex": "sex"},
+ key_columns=("person_id",),
+ )
+ with pytest.raises(gb.GroupBreakdownError, match="a category's"):
+ _assign(unclassified_codes={"marital_status": ("married",)})
+ with pytest.raises(gb.GroupBreakdownError, match="categorical"):
+ _assign(unclassified_codes={"age": ("x",)})
+ zero = frame.assign(weight=0.0)
+ with pytest.raises(gb.GroupBreakdownError, match="no positive weight"):
+ _assign(zero)
+
+
+@st.composite
+def attribute_frames(draw):
+ n = draw(st.integers(1, 30))
+ return pd.DataFrame(
+ {
+ "person_id": list(range(n)),
+ "weight": draw(
+ st.lists(
+ st.sampled_from([0.0, 0.5, 1.0, 2.0, 9.0]),
+ min_size=n,
+ max_size=n,
+ ).filter(lambda ws: any(w > 0 for w in ws))
+ ),
+ "sex": draw(
+ st.lists(
+ st.sampled_from(["female", "male", None]),
+ min_size=n,
+ max_size=n,
+ )
+ ),
+ "marital": draw(
+ st.lists(
+ st.sampled_from(
+ [
+ "married",
+ "divorced",
+ "widowed",
+ "never_married",
+ "separated",
+ None,
+ ]
+ ),
+ min_size=n,
+ max_size=n,
+ )
+ ),
+ "age": draw(
+ st.lists(
+ st.one_of(st.integers(40, 105), st.none()),
+ min_size=n,
+ max_size=n,
+ )
+ ),
+ "income": draw(
+ st.lists(
+ st.one_of(
+ st.sampled_from([0.0, 1.0, 2.5, 7.0, 100.0]),
+ st.none(),
+ ),
+ min_size=n,
+ max_size=n,
+ )
+ ),
+ }
+ )
+
+
+@SLOW
+@given(attribute_frames())
+def test_assignment_partitions_every_dimension(frame):
+ if (
+ frame["income"].isna().all()
+ or not (frame.loc[frame["income"].notna(), "weight"] > 0).any()
+ ):
+ return # no quintile population: refused, tested above
+ assignment = _assign(frame)
+ n = len(frame)
+ for dimension in assignment.scheme.dimensions:
+ masks = [
+ assignment.mask(dimension.key, c.key) for c in dimension.categories
+ ]
+ unclassified = assignment.unclassified_mask(dimension.key)
+ stacked = np.vstack([*masks, unclassified]).astype(int)
+ # every row is in exactly one category or unclassified
+ assert (stacked.sum(axis=0) == 1).all()
+ summary = [
+ s
+ for s in assignment.dimension_summaries
+ if s["key"] == dimension.key
+ ][0]
+ assert summary["n_classified"] + summary["n_unclassified"] == n
+ assert sum(summary["unclassified_reasons"].values()) == int(
+ unclassified.sum()
+ )
+ assert not assignment.unclassified_mask("total").any()
+
+
+# =========================================================================
+# Weighted percentiles
+# =========================================================================
+def _reference_percentile(values, weights, level):
+ """Independent exact reference: Fractions, no numpy."""
+
+ pairs = sorted(
+ (float(v), Fraction(float(w)))
+ for v, w in zip(values, weights, strict=True)
+ if w > 0
+ )
+ total = sum(w for _, w in pairs)
+ running = Fraction(0)
+ for index, (value, weight) in enumerate(pairs):
+ running += weight
+ if running >= level * total:
+ if running == level * total and index + 1 < len(pairs):
+ return (value + pairs[index + 1][0]) / 2
+ return value
+ raise AssertionError("unreachable")
+
+
+def test_percentiles_by_hand():
+ # values 10, 20, 30, 40, weights 1 each: W = 4. p10: 0.4 -> 10;
+ # median: 2 is hit exactly at 20 -> (20 + 30) / 2 = 25; p90: 3.6 -> 40.
+ assert gb.weighted_percentiles(
+ [40, 10, 30, 20], [1, 1, 1, 1], (gb.P10, gb.P50, gb.P90)
+ ) == [10.0, 25.0, 40.0]
+ # weights 1, 1, 2 over 1, 2, 3: W = 4; median 2 hit exactly at 2 ->
+ # (2 + 3) / 2 = 2.5; p10 0.4 -> 1; p90 3.6 -> 3
+ assert gb.weighted_percentiles(
+ [1, 2, 3], [1, 1, 2], (gb.P10, gb.P50, gb.P90)
+ ) == [1.0, 2.5, 3.0]
+ # a zero-weight value carries no mass and is never the "next" value
+ assert gb.weighted_percentiles([1, 2, 3], [1, 0, 1], (gb.P50,)) == [2.0]
+ # equal values: the midpoint of a value with itself is the value
+ assert gb.weighted_percentiles([5, 5], [1, 1], (gb.P50,)) == [5.0]
+
+
+@pytest.mark.parametrize(
+ ("values", "weights", "levels", "match"),
+ [
+ ([1.0], [1.0], (0.5,), "Fraction"),
+ ([1.0], [1.0], (Fraction(1),), "Fraction"),
+ ([1.0], [0.0], (gb.P50,), "all be zero"),
+ ([1.0], [-1.0], (gb.P50,), "non-negative"),
+ ([np.nan], [1.0], (gb.P50,), "finite"),
+ ([], [], (gb.P50,), "non-empty"),
+ ],
+)
+def test_percentile_refusals(values, weights, levels, match):
+ with pytest.raises(gb.GroupBreakdownError, match=match):
+ gb.weighted_percentiles(values, weights, levels)
+
+
+_VALUES = st.lists(
+ st.sampled_from([-100.0, -5.0, -1.0, 0.0, 0.5, 1.0, 2.0, 37.5]),
+ min_size=1,
+ max_size=25,
+)
+
+
+@st.composite
+def weighted_samples(draw, integer_weights=False):
+ values = draw(_VALUES)
+ weight = (
+ st.integers(0, 6).map(float)
+ if integer_weights
+ else st.sampled_from([0.0, 0.1, 0.25, 1.0, 1.5, 3.0, 1e-3])
+ )
+ weights = draw(
+ st.lists(weight, min_size=len(values), max_size=len(values)).filter(
+ lambda ws: any(w > 0 for w in ws)
+ )
+ )
+ return values, weights
+
+
+_LEVELS = (gb.P10, gb.P50, gb.P90, *gb.QUINTILE_LEVELS)
+
+
+@settings(max_examples=300, deadline=None)
+@given(weighted_samples())
+def test_percentiles_equal_the_exact_reference(sample):
+ values, weights = sample
+ got = gb.weighted_percentiles(values, weights, _LEVELS)
+ expected = [_reference_percentile(values, weights, p) for p in _LEVELS]
+ assert got == expected
+ positive = [v for v, w in zip(values, weights, strict=True) if w > 0]
+ assert all(min(positive) <= g <= max(positive) for g in got)
+ p10, p50, p90 = got[:3]
+ assert p10 <= p50 <= p90
+
+
+@settings(max_examples=300, deadline=None)
+@given(weighted_samples(integer_weights=True))
+def test_percentiles_agree_with_track_u_away_from_exact_ties(sample):
+ values, weights = sample
+ total = sum(Fraction(w) for w in weights)
+ for level in _LEVELS:
+ running, tie = Fraction(0), False
+ for _value, weight in sorted(zip(values, weights, strict=True)):
+ if weight == 0:
+ continue
+ running += Fraction(weight)
+ tie = tie or running == level * total
+ got = gb.weighted_percentiles(values, weights, (level,))[0]
+ theirs = weighted_quantile(values, weights, float(level))
+ if not tie:
+ assert got == theirs
+ else:
+ assert got >= theirs # the midpoint lies above the lower value
+
+
+@settings(max_examples=200, deadline=None)
+@given(weighted_samples(), st.integers(-8, 8), st.randoms())
+def test_percentiles_ignore_weight_scale_and_row_order(sample, power, rnd):
+ values, weights = sample
+ scaled = [w * 2.0**power for w in weights]
+ order = list(range(len(values)))
+ rnd.shuffle(order)
+ base = gb.weighted_percentiles(values, weights, _LEVELS)
+ assert gb.weighted_percentiles(values, scaled, _LEVELS) == base
+ assert (
+ gb.weighted_percentiles(
+ [values[i] for i in order], [weights[i] for i in order], _LEVELS
+ )
+ == base
+ )
+
+
+@settings(max_examples=300, deadline=None)
+@given(weighted_samples())
+def test_quintile_thresholds_split_the_weight_exactly(sample):
+ values, weights = sample
+ ranks, thresholds = gb.weighted_quintile_ranks(values, weights)
+ assert set(ranks.tolist()) <= {1, 2, 3, 4, 5}
+ total = sum(Fraction(w) for w in weights)
+ for k, t in enumerate(thresholds, start=1):
+ below = sum(
+ Fraction(w) for v, w in zip(values, weights, strict=True) if v < t
+ )
+ at_or_below = sum(
+ Fraction(w) for v, w in zip(values, weights, strict=True) if v <= t
+ )
+ assert below <= Fraction(k, 5) * total <= at_or_below
+ # rank is monotone in value
+ pairs = sorted(zip(values, ranks.tolist(), strict=True))
+ assert [r for _, r in pairs] == sorted(r for _, r in pairs)
+
+
+# =========================================================================
+# Change classes
+# =========================================================================
+def test_change_classes_are_exact_at_the_one_percent_threshold():
+ base = np.array([500.0, 500.0, 777.0, 100.0, 100.0, 0.0, 200.0])
+ reform = np.array([495.0, 505.0, 769.23, 99.5, 101.0, 50.0, 0.0])
+ population, change, decrease, increase = gb.classify_changes(base, reform)
+ assert population.tolist() == [True] * 5 + [False, True]
+ # 495/500 is exactly -1 percent: a decrease; 505/500 exactly +1: an
+ # increase. 769.23 (as a float, 769.2300000000000182...) is above
+ # 0.99 * 777 = 769.23, so the change is just above -1 percent and
+ # unaffected, although the float change rounds to -1.0000000000000009.
+ assert change[2] <= -1.0
+ assert decrease.tolist() == [True, False, False, False, False, False, True]
+ assert increase.tolist() == [False, True, False, False, True, False, False]
+ assert math.isnan(change[5])
+ assert change[6] == -100.0
+
+
+@settings(max_examples=300, deadline=None)
+@given(
+ st.lists(
+ st.tuples(
+ st.floats(0.01, 1e6, allow_nan=False),
+ st.floats(0.0, 2e6, allow_nan=False),
+ ),
+ min_size=1,
+ max_size=20,
+ )
+)
+def test_change_classes_equal_exact_rational_arithmetic(pairs):
+ base = np.array([b for b, _ in pairs])
+ reform = np.array([r for _, r in pairs])
+ _, _, decrease, increase = gb.classify_changes(base, reform)
+ for i, (b, r) in enumerate(pairs):
+ assert decrease[i] == (100 * Fraction(r) <= 99 * Fraction(b))
+ assert increase[i] == (100 * Fraction(r) >= 101 * Fraction(b))
+
+
+# =========================================================================
+# Projection breakdowns (exercises 1 and 3)
+# =========================================================================
+def _a7_group(result, label):
+ (group,) = [g for g in result["groups"] if g["label"] == label]
+ return group
+
+
+def _assert_matches_a7(cell, a7_group):
+ for statistic in a7.STATISTICS:
+ ours = cell.statistic(statistic)
+ theirs = a7_group[statistic]
+ assert ours.defined == theirs["defined"]
+ assert ours.value == theirs["mean"]
+ assert ours.uncertainty["sample_sd"] == theirs["sample_sd"]
+ assert ours.uncertainty["per_draw"] == theirs["per_draw"]
+ assert ours.uncertainty["floor"] == theirs["floor"]
+ assert ours.uncertainty["n_defined_draws"] == (
+ theirs["n_defined_draws"]
+ )
+
+
+@SLOW
+@given(st.booleans().flatmap(lambda b: projection_frames(b)))
+def test_total_and_age_cells_equal_cola_age_profile(sample):
+ frame, config = sample
+ scheme = _sex_age_scheme(age=gb.AGE_BAND_SETS["dynasim3_exercises_1_3"])
+ result = _project(frame, scheme, config)
+ by_age = a7.tabulate_cola_age_profile(
+ frame, data_provenance="invented", config=config
+ )
+ for group in a7.DEFAULT_AGE_GROUPS:
+ _assert_matches_a7(
+ result.cell("age", group.label), _a7_group(by_age, group.label)
+ )
+ everyone = dataclasses.replace(config, age_groups=(a7.AgeGroup("all", 0),))
+ total = a7.tabulate_cola_age_profile(
+ frame, data_provenance="invented", config=everyone
+ )
+ _assert_matches_a7(result.cell("total", "total"), _a7_group(total, "all"))
+
+
+@SLOW
+@given(projection_frames(allow_difference=True))
+def test_scenario_specific_membership_matches_cola_age_profile(sample):
+ frame, config = sample
+ scheme = _sex_age_scheme(age=gb.AGE_BAND_SETS["dynasim3_exercises_1_3"])
+ result = _project(frame, scheme, config)
+ by_age = a7.tabulate_cola_age_profile(
+ frame, data_provenance="invented", config=config
+ )
+ for group in a7.DEFAULT_AGE_GROUPS:
+ _assert_matches_a7(
+ result.cell("age", group.label), _a7_group(by_age, group.label)
+ )
+
+
+@SLOW
+@given(projection_frames(allow_difference=True))
+def test_projection_partitions_and_mint_identities(sample):
+ frame, config = sample
+ result = _project(frame, _sex_age_scheme(), config)
+ total = result.cell("total", "total")
+ k = len(config.draw_indices)
+ for key in ("sex", "age"):
+ dimension = result.dimension(key)
+ for d in range(k):
+ n = sum(
+ c.counts["unweighted_n_current_law_per_draw"][d]
+ for c in dimension.cells
+ )
+ w = math.fsum(
+ c.counts["weighted_n_current_law_per_draw"][d]
+ for c in dimension.cells
+ )
+ extra = dimension.unclassified["per_draw"][d]
+ assert n + extra["n_current_law"] == (
+ total.counts["unweighted_n_current_law_per_draw"][d]
+ )
+ assert math.isclose(
+ w + extra["weight_current_law"],
+ total.counts["weighted_n_current_law_per_draw"][d],
+ rel_tol=1e-12,
+ abs_tol=1e-9,
+ )
+ for dimension in result.dimensions:
+ for cell in dimension.cells:
+ stats = {s.statistic: s for s in cell.statistics}
+ per_draw = {
+ name: stats[name].uncertainty["per_draw"]
+ for name in gb.MINT_BENEFIT_STATISTICS
+ }
+ for d in range(k):
+ shares = [
+ per_draw[name][d]
+ for name in (
+ gb.PERCENT_DECREASE,
+ gb.PERCENT_UNAFFECTED,
+ gb.PERCENT_INCREASE,
+ )
+ ]
+ if shares[0] is None:
+ assert all(s is None for s in shares)
+ continue
+ assert math.isclose(sum(shares), 100.0, rel_tol=1e-12)
+ p10, p50, p90 = (
+ per_draw[name][d]
+ for name in (
+ gb.CHANGE_P10,
+ gb.CHANGE_MEDIAN,
+ gb.CHANGE_P90,
+ )
+ )
+ assert p10 <= p50 <= p90
+ json.dumps(result.as_dict(), allow_nan=False)
+
+
+@SLOW
+@given(projection_frames(), st.integers(-6, 6))
+def test_projection_statistics_ignore_weight_scale(sample, power):
+ frame, config = sample
+ scaled = frame.assign(weight=frame["weight"] * 2.0**power)
+ a = _project(frame, _sex_age_scheme(), config)
+ b = _project(scaled, _sex_age_scheme(), config)
+ for dim_a, dim_b in zip(a.dimensions, b.dimensions, strict=True):
+ for cell_a, cell_b in zip(dim_a.cells, dim_b.cells, strict=True):
+ for s_a, s_b in zip(
+ cell_a.statistics, cell_b.statistics, strict=True
+ ):
+ assert s_a.value == s_b.value
+ assert s_a.uncertainty["floor"] == s_b.uncertainty["floor"]
+
+
+@SLOW
+@given(projection_frames(), st.sampled_from([0.3, 3.7, 1e3]))
+def test_ratio_of_means_ignores_any_weight_scale(sample, factor):
+ frame, config = sample
+ scaled = frame.assign(weight=frame["weight"] * factor)
+ a = _project(frame, _sex_age_scheme(), config)
+ b = _project(scaled, _sex_age_scheme(), config)
+ for dim_a, dim_b in zip(a.dimensions, b.dimensions, strict=True):
+ for cell_a, cell_b in zip(dim_a.cells, dim_b.cells, strict=True):
+ s_a = cell_a.statistic(gb.RATIO_OF_SCENARIO_MEANS)
+ s_b = cell_b.statistic(gb.RATIO_OF_SCENARIO_MEANS)
+ assert s_a.defined == s_b.defined
+ if s_a.defined:
+ assert math.isclose(
+ s_a.value, s_b.value, rel_tol=1e-9, abs_tol=1e-9
+ )
+
+
+def _manual_side_ratio(frame, side_rows, draws):
+ """Plain-Python mean over draws of the ratio of weighted means."""
+
+ values = []
+ for d in draws:
+ part = frame[side_rows & (frame["draw"] == d).to_numpy()]
+ base = part[part["benefit_base"] > 0]
+ reform = part[part["benefit_reform"] > 0]
+ if base["weight"].sum() <= 0 or reform["weight"].sum() <= 0:
+ return None
+ mu_b = math.fsum(base["weight"] * base["benefit_base"]) / math.fsum(
+ base["weight"]
+ )
+ mu_r = math.fsum(
+ reform["weight"] * reform["benefit_reform"]
+ ) / math.fsum(reform["weight"])
+ values.append(100 * (mu_r / mu_b - 1))
+ return math.fsum(values) / len(values)
+
+
+@SLOW
+@given(projection_frames())
+def test_group_floor_uses_the_full_split_restricted_to_the_group(sample):
+ frame, config = sample
+ result = _project(frame, _sex_age_scheme(), config)
+ split = gb.half_split_masks(
+ frame["family_unit_id"].tolist(), config.floor_seeds
+ )
+ for sex in ("female", "male"):
+ in_group = (frame["sex"] == sex).to_numpy()
+ gaps = []
+ for in_a in split.values():
+ a = _manual_side_ratio(frame, in_group & in_a, config.draw_indices)
+ b = _manual_side_ratio(
+ frame, in_group & ~in_a, config.draw_indices
+ )
+ if a is not None and b is not None:
+ gaps.append(abs(a - b))
+ floor = (
+ result.cell("sex", sex)
+ .statistic(gb.RATIO_OF_SCENARIO_MEANS)
+ .uncertainty["floor"]
+ )
+ assert floor["n_seeds"] == len(gaps)
+ for ours, manual in zip(floor["values"], gaps, strict=True):
+ assert math.isclose(ours, manual, rel_tol=1e-9, abs_tol=1e-9)
+
+
+def test_a_resplit_of_the_group_would_differ():
+ # Ten one-person family units 0-9; women are the even-numbered units.
+ # Use interleaved IDs: a prefix would reuse the RNG stream's prefix.
+ # The full
+ # split restricted to women differs from re-splitting the five women's
+ # units alone (seed 0), which is why the group floor never re-splits.
+ units = list(range(10))
+ full = gb.half_split_masks(units, (0,))[0]
+ women = np.array([u % 2 == 0 for u in units])
+ resplit = gb.half_split_masks(units[::2], (0,))[0]
+ assert full[women].tolist() != resplit.tolist()
+
+
+def test_floor_is_undefined_not_zero_with_one_seed():
+ rows = [
+ _projection_row(0, p, birth_year=1960, base=1000.0, reform=990.0)
+ for p in range(6)
+ ]
+ frame = pd.DataFrame([{**r, "sex": "female"} for r in rows])
+ config = a7.ColaAgeProfileConfig(draw_indices=(0,), floor_seeds=(0,))
+ result = _project(frame, _sex_age_scheme(), config)
+ for dimension in result.dimensions:
+ for cell in dimension.cells:
+ for statistic in cell.statistics:
+ floor = statistic.uncertainty["floor"]
+ assert floor["defined"] is False
+ assert floor["mean"] is None and floor["sd"] is None
+ assert floor["n_seeds"] <= 1
+
+
+def test_projection_by_hand():
+ # Draw 0 only. Women: p0 w1 1000 -> 990 (-1%: decrease), p1 w3
+ # 2000 -> 2100 (+5%: increase). mu_base = (1000 + 6000) / 4 = 1750;
+ # mu_reform = (990 + 6300) / 4 = 1822.5; ratio 100 * (1822.5 / 1750 -
+ # 1) = 4.142857... Decrease share 25%, increase 75%; changes -1 (w1)
+ # and +5 (w3): p10 -> -1, median: 2 of 4 reached at +5 -> 5, p90 5.
+ frame = pd.DataFrame(
+ [
+ _projection_row(
+ 0, 0, birth_year=1960, base=1000.0, reform=990.0, sex="female"
+ ),
+ _projection_row(
+ 0,
+ 1,
+ birth_year=1960,
+ base=2000.0,
+ reform=2100.0,
+ weight=3.0,
+ sex="female",
+ ),
+ _projection_row(
+ 0, 2, birth_year=1960, base=500.0, reform=500.0, sex=None
+ ),
+ ]
+ )
+ config = a7.ColaAgeProfileConfig(draw_indices=(0,))
+ result = _project(frame, _sex_age_scheme(), config)
+ women = result.cell("sex", "Female")
+ stats = {s.statistic: s.value for s in women.statistics}
+ assert stats[gb.RATIO_OF_SCENARIO_MEANS] == pytest.approx(
+ 100 * (1822.5 / 1750 - 1)
+ )
+ assert stats[gb.PERCENT_DECREASE] == 25.0
+ assert stats[gb.PERCENT_INCREASE] == 75.0
+ assert stats[gb.PERCENT_UNAFFECTED] == 0.0
+ assert stats[gb.CHANGE_P10] == pytest.approx(-1.0)
+ assert stats[gb.CHANGE_MEDIAN] == pytest.approx(5.0)
+ assert stats[gb.CHANGE_P90] == pytest.approx(5.0)
+ men = result.cell("sex", "male").statistic(gb.RATIO_OF_SCENARIO_MEANS)
+ assert not men.defined and men.value is None
+ assert "empty_baseline_membership" in men.undefined_reason
+ sex = result.dimension("sex")
+ assert sex.unclassified["reasons"] == {"missing": 1}
+ assert sex.ssa_subgroup_suppressed
+ assert set(sex.suppressed_by) == {"Female", "Male"}
+ assert women.flags == {
+ "ssa_disclosure_below_100": True,
+ "below_30": True,
+ "ssa_numerator_1_to_9": True,
+ }
+ decrease = women.statistic(gb.PERCENT_DECREASE)
+ assert decrease.flags["ssa_numerator_1_to_9"] is True
+ assert decrease.unweighted_n == 2 and decrease.weighted_n == 4.0
+
+
+def test_membership_difference_is_refused_unless_configured():
+ frame = pd.DataFrame(
+ [
+ _projection_row(
+ 0, 0, birth_year=1960, base=1000.0, reform=0.0, sex="female"
+ ),
+ _projection_row(
+ 0, 1, birth_year=1960, base=1000.0, reform=990.0, sex="male"
+ ),
+ ]
+ )
+ with pytest.raises(gb.GroupBreakdownError, match="only one scenario"):
+ _project(
+ frame,
+ _sex_age_scheme(),
+ a7.ColaAgeProfileConfig(draw_indices=(0,)),
+ )
+
+
+def test_projection_refuses_rows_the_assignment_was_not_built_from():
+ frame, config = _small_projection()
+ assignment = gb.assign_groups(
+ frame,
+ _sex_age_scheme(),
+ {"sex": "sex", "age": "age"},
+ key_columns=("draw", "person_id"),
+ )
+ reordered = frame.iloc[::-1]
+ with pytest.raises(gb.GroupBreakdownError, match="key digest"):
+ gb.tabulate_projection_breakdown(
+ reordered, assignment, data_provenance="invented", config=config
+ )
+ static_keys = gb.assign_groups(
+ frame.assign(row=range(len(frame))),
+ _sex_age_scheme(),
+ {"sex": "sex", "age": "age"},
+ key_columns=("row",),
+ )
+ with pytest.raises(gb.GroupBreakdownError, match="keyed on"):
+ gb.tabulate_projection_breakdown(
+ frame, static_keys, data_provenance="invented", config=config
+ )
+
+
+def _small_projection():
+ frame = pd.DataFrame(
+ [
+ _projection_row(
+ d,
+ p,
+ birth_year=1950 + p,
+ base=1000.0 + p,
+ reform=990.0 + p,
+ sex="female" if p % 2 else "male",
+ )
+ for d in (0, 1)
+ for p in range(4)
+ ]
+ )
+ return frame, a7.ColaAgeProfileConfig(draw_indices=(0, 1))
+
+
+# =========================================================================
+# Static breakdowns (exercises 2 and 4)
+# =========================================================================
+def _poverty(rows, **kwargs):
+ assignment = gb.assign_groups(
+ rows,
+ _sex_marital_scheme(),
+ {"sex": "sex", "marital_status": "marital_status_4"},
+ key_columns=("observation_id",),
+ unclassified_codes={"marital_status": (ut.UNCLASSIFIED,)},
+ )
+ return gb.tabulate_poverty_breakdown(
+ rows,
+ assignment,
+ design=kwargs.pop("design", _design(rows, (9, 1))),
+ data_provenance="invented",
+ id_column="observation_id",
+ **kwargs,
+ )
+
+
+_TRACK_U_CELLS = {
+ ("total", "total"): "all",
+ ("sex", "female"): "women",
+ ("sex", "male"): "men",
+ ("marital_status", "married"): "married",
+ ("marital_status", "widowed"): "widowed",
+ ("marital_status", "divorced"): "divorced",
+ ("marital_status", "never_married"): "never_married",
+}
+_TRACK_U_STATISTICS = {
+ gb.POVERTY_RATE_CURRENT_LAW: "baseline_rate",
+ gb.POVERTY_RATE_PROPOSAL: "reform_rate",
+ gb.POVERTY_RATE_CHANGE: "delta",
+}
+
+
+@SLOW
+@given(st.integers(0, 2**32 - 1), st.integers(1, 40))
+def test_poverty_cells_equal_uniform_cut_tabulation(seed, n):
+ rows = _poverty_frame(np.random.default_rng(seed), n)
+ design = _design(rows, (9, 1))
+ ours = _poverty(rows, design=design)
+ theirs = {
+ cell["cell"]: cell
+ for cell in ut.tabulate_uniform_cut(
+ rows, data_provenance="invented", design=design
+ )["cells"]
+ }
+ for (dimension, category), name in _TRACK_U_CELLS.items():
+ cell = ours.cell(dimension, category)
+ reference = theirs[name]
+ assert cell.counts["unweighted_n"] == reference["n_observations"]
+ assert cell.counts["weighted_n"] == reference["weight_total"]
+ for statistic, key in _TRACK_U_STATISTICS.items():
+ value = cell.statistic(statistic)
+ assert value.defined == reference["defined"]
+ if reference["defined"]:
+ assert value.value == reference[key]
+ assert value.uncertainty["design_se"] == (
+ reference["design_se"][key]
+ )
+ assert value.uncertainty["floor"] == reference["floor"][key]
+ # partition of the weighted total within each dimension
+ total = ours.cell("total", "total").counts["weighted_n"]
+ for key in ("sex", "marital_status"):
+ dimension = ours.dimension(key)
+ parts = math.fsum(c.counts["weighted_n"] for c in dimension.cells)
+ assert math.isclose(
+ parts + dimension.unclassified["weighted_n"],
+ total,
+ rel_tol=1e-12,
+ abs_tol=1e-12,
+ )
+
+
+def test_poverty_numbers_by_hand():
+ rows = pd.DataFrame(
+ {
+ "observation_id": ["o1", "o2", "o3", "o4"],
+ "person_id": [1, 2, 3, 4],
+ "family_unit_id": [1, 1, 2, 3],
+ "weight": [1000.0, 3000.0, 2000.0, 4000.0],
+ "sex": ["female", "female", "male", "male"],
+ "marital_status_4": ["married", "widowed", "married", "divorced"],
+ "birth_year": [1941] * 4,
+ "stratum": [1, 1, 2, 2],
+ "cluster": [1, 2, 1, 2],
+ "poor_baseline": [True, False, False, False],
+ "poor_reform": [True, True, False, True],
+ }
+ )
+ result = _poverty(rows)
+ total = result.cell("total", "total")
+ values = {s.statistic: s.value for s in total.statistics}
+ # W = 10000; poor under current law 1000 (10%); with the proposal
+ # 1000 + 3000 + 4000 = 8000 (80%); change 70 points; in thousands 1
+ # and 8, change 7; percent change in the number 700%.
+ assert values == {
+ gb.POVERTY_RATE_CURRENT_LAW: 10.0,
+ gb.POVERTY_RATE_PROPOSAL: 80.0,
+ gb.POVERTY_RATE_CHANGE: 70.0,
+ gb.NUMBER_POOR_CURRENT_LAW: 1.0,
+ gb.NUMBER_POOR_PROPOSAL: 8.0,
+ gb.NUMBER_POOR_CHANGE: 7.0,
+ gb.NUMBER_POOR_PERCENT_CHANGE: 700.0,
+ }
+ men = result.cell("sex", "male").statistic(gb.NUMBER_POOR_PERCENT_CHANGE)
+ assert not men.defined
+ assert (
+ men.undefined_reason == "nobody in the cell is poor under current law"
+ )
+ count = total.statistic(gb.NUMBER_POOR_CURRENT_LAW)
+ assert count.uncertainty["design_se"] is None
+ assert "weighted count" in count.uncertainty["design_se_note"]
+ assert count.uncertainty["floor"] is not None
+ assert "half the full" in count.uncertainty["floor_note"]
+ json.dumps(result.as_dict(), allow_nan=False)
+
+
+def _shares(rows, **kwargs):
+ base = gb.MINT8_SCHEME
+ scheme = gb.derive_scheme(
+ base,
+ scheme_id="share_test",
+ title="t",
+ drop=[d.key for d in base.dimensions if d.key not in {"total", "sex"}],
+ )
+ assignment = gb.assign_groups(
+ rows,
+ scheme,
+ {"sex": "sex"},
+ key_columns=("person_id",),
+ unclassified_codes={"sex": ("unknown",)},
+ )
+ return gb.tabulate_share_breakdown(
+ rows,
+ assignment,
+ indicator_column=kwargs.pop("indicator_column", "receives_2"),
+ design=kwargs.pop("design", _design(rows)),
+ data_provenance="invented",
+ **kwargs,
+ )
+
+
+@SLOW
+@given(st.integers(0, 2**32 - 1), st.integers(1, 40))
+def test_share_cells_equal_track_m(seed, n):
+ rows = _track_m_frame(np.random.default_rng(seed), n)
+ design = _design(rows, (7, 1))
+ theirs = track_m.tabulate_track_m(
+ rows, row_id="MS0", data_provenance="invented", design=design
+ )
+ for option in track_m.TABLE6_OPTIONS:
+ ours = _shares(
+ rows, indicator_column=f"receives_{option}", design=design
+ )
+ for row, key in (
+ ("all", "total"),
+ ("men", "male"),
+ ("women", "female"),
+ ):
+ (reference,) = [
+ c
+ for c in theirs["cells"]
+ if c["option"] == option and c["row"] == row
+ ]
+ dimension = "total" if row == "all" else "sex"
+ share = ours.cell(dimension, key).statistic(gb.SHARE)
+ assert share.defined == reference["defined"]
+ if reference["defined"]:
+ assert share.value == reference["share_percent"]
+ assert share.weighted_n == reference["weighted_n"]
+ assert share.unweighted_n == reference["unweighted_n"]
+ assert share.uncertainty["design_se"] == (
+ reference["design_se"]
+ )
+ assert share.uncertainty["floor"] == reference["floor"]
+
+
+def test_suppression_flags_never_drop_cells():
+ # INVENTED: 100 women, 29 men, 1 unknown sex.
+ n = 130
+ rows = pd.DataFrame(
+ {
+ "person_id": [str(i) for i in range(n)],
+ "family_unit_id": list(range(n)),
+ "weight": [1.0] * n,
+ "sex": ["female"] * 100 + ["male"] * 29 + ["unknown"],
+ "stratum": [1, 2] * (n // 2),
+ "cluster": [1] * 65 + [2] * 65,
+ "receives_2": [i % 3 == 0 for i in range(n)],
+ }
+ )
+ result = _shares(rows)
+ sex = result.dimension("sex")
+ assert [c.label for c in sex.cells] == ["Female", "Male"]
+ assert sex.cell("female").flags == {
+ "ssa_disclosure_below_100": False,
+ "below_30": False,
+ }
+ assert sex.cell("male").flags == {
+ "ssa_disclosure_below_100": True,
+ "below_30": True,
+ }
+ assert sex.ssa_subgroup_suppressed and sex.suppressed_by == ("Male",)
+ assert sex.unclassified["unweighted_n"] == 1
+ assert sex.unclassified["reasons"] == {"code:unknown": 1}
+ total = result.dimension("total")
+ assert not total.ssa_subgroup_suppressed
+ assert sex.cell("male").statistic(gb.SHARE).value is not None
+
+
+def _static_benefits(rows, **kwargs):
+ base = gb.MINT8_SCHEME
+ scheme = gb.derive_scheme(
+ base,
+ scheme_id="static_benefit_test",
+ title="t",
+ drop=[d.key for d in base.dimensions if d.key not in {"total", "sex"}],
+ )
+ assignment = gb.assign_groups(
+ rows, scheme, {"sex": "sex"}, key_columns=("person_id",)
+ )
+ return gb.tabulate_static_benefit_breakdown(
+ rows,
+ assignment,
+ design=_design(rows),
+ data_provenance="invented",
+ **kwargs,
+ )
+
+
+def _static_benefit_frame(rng, n):
+ base = rng.choice([0.0, 100.0, 500.0, 777.0], n)
+ return pd.DataFrame(
+ {
+ "person_id": np.arange(n),
+ "family_unit_id": rng.integers(0, max(1, n // 2), n),
+ "weight": rng.choice([0.0, 1.0, 2.5], n),
+ "sex": rng.choice(["female", "male"], n),
+ "stratum": rng.integers(1, 3, n),
+ "cluster": rng.integers(1, 3, n),
+ "benefit_base": base,
+ "benefit_reform": base * rng.choice([0.87, 0.99, 1.0, 1.01], n),
+ }
+ )
+
+
+def test_static_benefit_by_hand():
+ # Uniform 13% cut for three beneficiaries; a fourth has no benefit.
+ rows = pd.DataFrame(
+ {
+ "person_id": [1, 2, 3, 4],
+ "family_unit_id": [1, 2, 3, 4],
+ "weight": [1.0, 1.0, 2.0, 5.0],
+ "sex": ["female", "male", "female", "male"],
+ "stratum": [1, 1, 2, 2],
+ "cluster": [1, 2, 1, 2],
+ "benefit_base": [1000.0, 2000.0, 1500.0, 0.0],
+ "benefit_reform": [870.0, 1740.0, 1305.0, 0.0],
+ }
+ )
+ result = _static_benefits(rows)
+ total = result.cell("total", "total")
+ values = {s.statistic: s.value for s in total.statistics}
+ assert values[gb.PERCENT_DECREASE] == 100.0
+ assert values[gb.PERCENT_INCREASE] == 0.0
+ for name in (gb.CHANGE_P10, gb.CHANGE_MEDIAN, gb.CHANGE_P90):
+ assert values[name] == pytest.approx(-13.0)
+ assert total.statistic(gb.PERCENT_DECREASE).unweighted_n == 3
+ assert total.statistic(gb.PERCENT_DECREASE).weighted_n == 4.0
+ median = total.statistic(gb.CHANGE_MEDIAN)
+ assert median.uncertainty["design_se"] is None
+ assert "percentile" in median.uncertainty["design_se_note"]
+ assert total.statistic(gb.PERCENT_DECREASE).uncertainty["design_se"]
+ json.dumps(result.as_dict(), allow_nan=False)
+
+
+@SLOW
+@given(st.integers(0, 2**32 - 1), st.integers(1, 30))
+def test_static_benefit_identities(seed, n):
+ rows = _static_benefit_frame(np.random.default_rng(seed), n)
+ result = _static_benefits(rows)
+ population = rows["benefit_base"].to_numpy() > 0
+ weights = rows["weight"].to_numpy()
+ total_w = math.fsum(weights[population].tolist())
+ sex = result.dimension("sex")
+ parts = math.fsum(
+ c.statistic(gb.PERCENT_DECREASE).weighted_n for c in sex.cells
+ )
+ assert math.isclose(
+ parts + sex.unclassified["weighted_n"], total_w, abs_tol=1e-12
+ )
+ for dimension in result.dimensions:
+ for cell in dimension.cells:
+ stats = {s.statistic: s for s in cell.statistics}
+ if not stats[gb.PERCENT_DECREASE].defined:
+ continue
+ assert math.isclose(
+ stats[gb.PERCENT_DECREASE].value
+ + stats[gb.PERCENT_UNAFFECTED].value
+ + stats[gb.PERCENT_INCREASE].value,
+ 100.0,
+ rel_tol=1e-12,
+ )
+ mask = population & (
+ np.ones(n, dtype=bool)
+ if dimension.key == "total"
+ else (rows["sex"] == cell.category).to_numpy()
+ )
+ mask &= weights > 0
+ changes = 100 * (
+ rows["benefit_reform"].to_numpy()[mask]
+ / rows["benefit_base"].to_numpy()[mask]
+ - 1
+ )
+ for name in (gb.CHANGE_P10, gb.CHANGE_MEDIAN, gb.CHANGE_P90):
+ assert changes.min() <= stats[name].value <= changes.max()
+
+
+# =========================================================================
+# Provenance and the result schema
+# =========================================================================
+def test_registration_pointer_pattern_equals_track_m():
+ assert gb.REGISTRATION_POINTER.pattern == (
+ track_m.REGISTRATION_POINTER.pattern
+ )
+
+
+def test_invented_results_carry_labels_scheme_and_definitions():
+ frame, config = _small_projection()
+ result = _project(
+ frame,
+ _sex_age_scheme(),
+ config,
+ labels=("PSID-seeded closed cohort",),
+ post_hoc_labels=("post hoc, not blind",),
+ upstream={"row": "R0"},
+ )
+ out = result.as_dict()
+ assert out["labels"] == [
+ gb.INVENTED_DATA_LABEL,
+ "PSID-seeded closed cohort",
+ ]
+ assert out["post_hoc_labels"] == ["post hoc, not blind"]
+ assert out["scheme"]["citations"]
+ assert out["statistic_definitions"][gb.CHANGE_MEDIAN]["mint8_column"] == (
+ "Percent change in Social Security benefits at the— Median"
+ )
+ assert out["conventions"]["percentile_definition"] == (
+ gb.PERCENTILE_DEFINITION
+ )
+ assert out["upstream"] == {"row": "R0"}
+ assert out["schema_version"] == gb.SCHEMA_VERSION
+ again = _project(
+ frame,
+ _sex_age_scheme(),
+ config,
+ labels=("PSID-seeded closed cohort",),
+ post_hoc_labels=("post hoc, not blind",),
+ upstream={"row": "R0"},
+ )
+ assert json.dumps(again.as_dict(), sort_keys=True) == json.dumps(
+ out, sort_keys=True
+ )
+
+
+@pytest.mark.parametrize(
+ ("overrides", "match"),
+ [
+ ({"registration_pointer": None}, "registration pointer"),
+ (
+ {"registration_pointer": "https://example.org/42"},
+ "registration pointer",
+ ),
+ ({"labels": ()}, "needs labels"),
+ ({"post_hoc_labels": ()}, "post hoc labels"),
+ ({"labels": (gb.INVENTED_DATA_LABEL,)}, "invented label"),
+ ({"labels": "one string"}, "sequence"),
+ ],
+)
+def test_registered_real_refusals(overrides, match):
+ frame, config = _small_projection()
+ assignment = gb.assign_groups(
+ frame,
+ _sex_age_scheme(),
+ {"sex": "sex", "age": "age"},
+ key_columns=("draw", "person_id"),
+ )
+ options = {
+ "registration_pointer": POINTER,
+ "labels": ("label",),
+ "post_hoc_labels": ("registered, one-shot, post hoc, not blind",),
+ **overrides,
+ }
+ with pytest.raises(gb.GroupBreakdownError, match=match):
+ gb.tabulate_projection_breakdown(
+ frame,
+ assignment,
+ data_provenance="registered_real",
+ config=config,
+ **options,
+ )
+
+
+def test_provenance_kind_guards():
+ rows = _track_m_frame(np.random.default_rng(3), 12)
+ rows.attrs["provenance_kind"] = gb.PSID_FILES
+ with pytest.raises(gb.GroupBreakdownError, match="staged PSID"):
+ _shares(rows)
+ rows.attrs["provenance_kind"] = track_m.INVENTED
+ base = gb.MINT8_SCHEME
+ scheme = gb.derive_scheme(
+ base,
+ scheme_id="guard_test",
+ title="t",
+ drop=[d.key for d in base.dimensions if d.key != "total"],
+ )
+ assignment = gb.assign_groups(rows, scheme, {}, key_columns=("person_id",))
+ with pytest.raises(gb.GroupBreakdownError, match="marked invented"):
+ gb.tabulate_share_breakdown(
+ rows,
+ assignment,
+ indicator_column="receives_2",
+ design=_design(rows),
+ data_provenance="registered_real",
+ registration_pointer=POINTER,
+ labels=("label",),
+ post_hoc_labels=("post hoc",),
+ )
+
+
+def test_static_inputs_are_refused_when_malformed():
+ rows = _track_m_frame(np.random.default_rng(5), 10)
+ with pytest.raises(gb.GroupBreakdownError, match="not in the design"):
+ _shares(rows, design=pd.DataFrame({"stratum": [99], "cluster": [1]}))
+ with pytest.raises(gb.GroupBreakdownError, match="boolean"):
+ _shares(rows.assign(receives_2=1))
+ with pytest.raises(gb.GroupBreakdownError, match="floor_split_unit"):
+ _shares(rows, floor_split_unit="household")
+ with pytest.raises(gb.GroupBreakdownError, match="floor seeds"):
+ _shares(rows, floor_seeds=(0, 0))
diff --git a/tests/estimates/test_group_breakdown_invariants.py b/tests/estimates/test_group_breakdown_invariants.py
new file mode 100644
index 00000000..3034ca17
--- /dev/null
+++ b/tests/estimates/test_group_breakdown_invariants.py
@@ -0,0 +1,337 @@
+"""Additional G3 invariants and adversarial inputs on INVENTED rows.
+
+These unit tests open no microdata or outcome artifacts. All rows and
+amounts below are INVENTED; reference statistics are recomputed from those
+same rows. The tests cover per-dimension quintile populations, static
+full-sample splits, disclosure propagation and explicit design refusals.
+"""
+
+from __future__ import annotations
+
+import math
+
+import numpy as np
+import pandas as pd
+import pytest
+from hypothesis import given, settings
+from hypothesis import strategies as st
+
+from populace_dynamics.estimates import cola_age_profile as a7
+from populace_dynamics.estimates import group_breakdown as gb
+from populace_dynamics.estimates import uniform_cut_tabulation as ut
+
+
+def _scheme(base=gb.MINT8_SCHEME, keep=("total", "sex")):
+ """An INVENTED test composition of captured category dimensions."""
+ return gb.derive_scheme(
+ base,
+ scheme_id="invented_adversarial_test",
+ title="INVENTED test composition",
+ drop=[d.key for d in base.dimensions if d.key not in keep],
+ )
+
+
+def _rows(n=12, weights=None):
+ """INVENTED static rows, with both sexes in each family unit."""
+ rows = pd.DataFrame(
+ {
+ "person_id": np.arange(n),
+ "family_unit_id": np.arange(n) // 2,
+ "weight": np.ones(n) if weights is None else weights,
+ "sex": ["female" if i % 2 == 0 else "male" for i in range(n)],
+ "stratum": np.ones(n, dtype=int),
+ "cluster": 1 + np.arange(n) % 2,
+ "indicator": [i % 3 == 0 for i in range(n)],
+ "poor_baseline": [i % 4 == 0 for i in range(n)],
+ "poor_reform": [i % 3 == 0 for i in range(n)],
+ "benefit_base": np.full(n, 100.0),
+ "benefit_reform": [
+ 90.0 if i % 3 == 0 else 100.0 for i in range(n)
+ ],
+ }
+ )
+ rows.attrs["provenance_kind"] = gb.INVENTED
+ return rows
+
+
+def _assignment(rows):
+ return gb.assign_groups(
+ rows, _scheme(), {"sex": "sex"}, key_columns=("person_id",)
+ )
+
+
+def _static(rows, kind, **kwargs):
+ options = {
+ "design": pd.DataFrame({"stratum": [1, 1], "cluster": [1, 2]}),
+ "data_provenance": gb.INVENTED,
+ "floor_split_unit": gb.FAMILY_UNIT,
+ **kwargs,
+ }
+ assignment = _assignment(rows)
+ if kind == "share":
+ return gb.tabulate_share_breakdown(
+ rows, assignment, indicator_column="indicator", **options
+ )
+ if kind == "poverty":
+ return gb.tabulate_poverty_breakdown(rows, assignment, **options)
+ return gb.tabulate_static_benefit_breakdown(rows, assignment, **options)
+
+
+@pytest.mark.parametrize("kind", ["share", "poverty", "benefit"])
+@pytest.mark.parametrize("column", ["stratum", "cluster"])
+def test_static_design_identifiers_refuse_fractional_values(kind, column):
+ rows = _rows()
+ rows[column] = rows[column].astype(float)
+ rows.loc[0, column] = 1.5
+ with pytest.raises(gb.GroupBreakdownError, match="integer|integral"):
+ _static(rows, kind)
+
+
+@pytest.mark.parametrize("column", ["stratum", "cluster"])
+def test_full_design_frame_refuses_fractional_identifiers(column):
+ design = pd.DataFrame({"stratum": [1.0, 1.0], "cluster": [1.0, 2.0]})
+ design.loc[0, column] += 0.5
+ with pytest.raises(gb.GroupBreakdownError, match="integer|integral"):
+ _static(_rows(), "share", design=design)
+
+
+@pytest.mark.parametrize("kind", ["share", "poverty", "benefit"])
+def test_static_floor_requires_two_usable_seeds(kind):
+ result = _static(_rows(), kind, floor_seeds=(0,))
+ statistic = {
+ "share": gb.SHARE,
+ "poverty": gb.POVERTY_RATE_CHANGE,
+ "benefit": gb.PERCENT_DECREASE,
+ }[kind]
+ floor = (
+ result.cell("total", "total").statistic(statistic).uncertainty["floor"]
+ )
+ assert floor["n_seeds"] == 1
+ assert not floor["defined"]
+ assert floor["mean"] is None
+ assert floor["sd"] is None
+
+
+@settings(max_examples=25, deadline=None)
+@given(
+ st.lists(st.integers(1, 100), min_size=8, max_size=24).filter(
+ lambda values: len(values) % 2 == 0
+ )
+)
+def test_static_group_floor_restricts_the_full_family_split(weights):
+ rows = _rows(len(weights), np.asarray(weights, dtype=float))
+ result = _static(rows, "share")
+ full = gb.half_split_masks(rows.family_unit_id.tolist(), (0, 1, 2, 3, 4))
+ group = (rows.sex == "female").to_numpy()
+ w = rows.weight.to_numpy()
+ receiving = rows.indicator.to_numpy()
+
+ def share(mask):
+ denominator = math.fsum(w[mask].tolist())
+ if denominator == 0:
+ return None
+ return 100.0 * math.fsum(w[mask & receiving].tolist()) / denominator
+
+ gaps, dropped = [], []
+ for seed, side in full.items():
+ first, second = share(group & side), share(group & ~side)
+ if first is None or second is None:
+ dropped.append(seed)
+ else:
+ gaps.append(abs(first - second))
+ floor = (
+ result.cell("sex", "female").statistic(gb.SHARE).uncertainty["floor"]
+ )
+ assert floor == {**ut._floor_summary(gaps), "dropped_seeds": dropped}
+
+
+def test_poverty_count_floors_use_the_same_raw_half_scale():
+ rows = _rows(weights=np.arange(1.0, 13.0) * 1000)
+ result = _static(rows, "poverty")
+ split = gb.half_split_masks(rows.family_unit_id.tolist(), (0, 1, 2, 3, 4))
+ w = rows.weight.to_numpy()
+ poor = rows.poor_baseline.to_numpy()
+ gaps = [
+ abs(
+ math.fsum(w[side & poor].tolist())
+ - math.fsum(w[~side & poor].tolist())
+ )
+ / 1000.0
+ for side in split.values()
+ ]
+ floor = (
+ result.cell("total", "total")
+ .statistic(gb.NUMBER_POOR_CURRENT_LAW)
+ .uncertainty["floor"]
+ )
+ assert floor == {**ut._floor_summary(gaps), "dropped_seeds": []}
+ assert result.conventions["uncertainty"]["scale"] == (
+ "half sample; not rescaled"
+ )
+
+
+def test_numerator_disclosure_flag_suppresses_the_whole_subgroup():
+ # INVENTED: 110 cases in each sex category, three decreases in Female.
+ rows = _rows(220)
+ rows["benefit_reform"] = 100.0
+ rows.loc[[0, 2, 4], "benefit_reform"] = 50.0
+ result = _static(rows, "benefit")
+ sex = result.dimension("sex")
+ decrease = sex.cell("female").statistic(gb.PERCENT_DECREASE)
+ assert decrease.unweighted_n == 110
+ assert not decrease.flags["ssa_disclosure_below_100"]
+ assert decrease.flags["ssa_numerator_1_to_9"]
+ assert sex.ssa_subgroup_suppressed
+ assert sex.suppressed_by == ("Female",)
+ assert len(sex.cells) == 2
+
+
+def test_projection_ratio_flags_the_smaller_scenario_membership():
+ # INVENTED: 120 current-law recipients, 20 proposal recipients.
+ rows = pd.DataFrame(
+ [
+ {
+ "draw": 0,
+ "person_id": person,
+ "family_unit_id": person // 2,
+ "weight": 1.0,
+ "birth_year": 1960,
+ "beneficiary_base": True,
+ "beneficiary_reform": person < 20,
+ "benefit_base": 100.0,
+ "benefit_reform": 110.0 if person < 20 else 0.0,
+ "benefit_components": {
+ "retired_worker": {
+ "base": 100.0,
+ "reform": 110.0 if person < 20 else 0.0,
+ }
+ },
+ }
+ for person in range(120)
+ ]
+ )
+ rows.attrs["provenance_kind"] = gb.INVENTED
+ assignment = gb.assign_groups(
+ rows,
+ _scheme(keep=("total",)),
+ {},
+ key_columns=("draw", "person_id"),
+ )
+ result = gb.tabulate_projection_breakdown(
+ rows,
+ assignment,
+ data_provenance=gb.INVENTED,
+ config=a7.ColaAgeProfileConfig(
+ draw_indices=(0,),
+ floor_seeds=(0, 1),
+ allow_membership_difference=True,
+ ),
+ )
+ cell = result.cell("total", "total")
+ ratio = cell.statistic(gb.RATIO_OF_SCENARIO_MEANS)
+ assert ratio.value == pytest.approx(10.0)
+ assert cell.counts["unweighted_n_base_per_draw"] == [120]
+ assert cell.counts["unweighted_n_reform_per_draw"] == [20]
+ assert ratio.unweighted_n == 20
+ assert ratio.flags["ssa_disclosure_below_100"]
+ assert ratio.flags["below_30"]
+ assert cell.flags["ssa_disclosure_below_100"]
+ assert cell.flags["below_30"]
+ assert result.dimension("total").ssa_subgroup_suppressed
+
+
+def _quintile_rows(lower=0, upper=100):
+ """INVENTED two draws, two birth cohorts, five distinct values each."""
+ rows = pd.DataFrame(
+ [
+ {
+ "draw": draw,
+ "person_id": 10 * cohort + rank,
+ "birth_cohort": cohort,
+ "weight": 1.0,
+ "household_income": offset + rank,
+ "initial_aime": offset + rank,
+ "expected_rank": rank,
+ }
+ for draw in (0, 1)
+ for cohort, offset in enumerate((lower, upper))
+ for rank in range(1, 6)
+ ]
+ )
+ rows.attrs["provenance_kind"] = gb.INVENTED
+ return rows
+
+
+def _quintile_scheme():
+ return _scheme(
+ gb.MINT8_ANNUAL_WITH_LIFETIME_SCHEME,
+ ("total", "household_income_quintile", "initial_aime_quintile"),
+ )
+
+
+def _assign_quintiles(rows, overrides):
+ return gb.assign_groups(
+ rows,
+ _quintile_scheme(),
+ {
+ "household_income_quintile": "household_income",
+ "initial_aime_quintile": "initial_aime",
+ },
+ key_columns=("draw", "person_id"),
+ quintile_partition=("draw",),
+ quintile_partitions=overrides,
+ )
+
+
+@settings(max_examples=20, deadline=None)
+@given(st.integers(0, 100), st.integers(201, 1000))
+def test_each_quintile_dimension_uses_its_own_population(lower, upper):
+ rows = _quintile_rows(lower, upper)
+ assignment = _assign_quintiles(
+ rows, {"initial_aime_quintile": ("draw", "birth_cohort")}
+ )
+ for dimension in ("household_income_quintile", "initial_aime_quintile"):
+ rank_by_category = {
+ category.key: category.rank
+ for category in assignment.scheme.dimension(dimension).categories
+ }
+ assigned = assignment.long[
+ assignment.long.dimension == dimension
+ ].sort_values("row_id")
+ ranks = assigned.category.map(rank_by_category).to_numpy()
+ if dimension == "initial_aime_quintile":
+ np.testing.assert_array_equal(ranks, rows.expected_rank)
+ else:
+ for draw in (0, 1):
+ mask = rows.draw == draw
+ expected, _ = gb.weighted_quintile_ranks(
+ rows.household_income[mask], rows.weight[mask]
+ )
+ np.testing.assert_array_equal(ranks[mask], expected)
+ assert assignment.as_dict()["quintile_partitions"] == {
+ "household_income_quintile": ["draw"],
+ "initial_aime_quintile": ["draw", "birth_cohort"],
+ }
+
+
+@pytest.mark.parametrize(
+ "overrides",
+ [
+ {"total": ("draw",)},
+ {"unknown_dimension": ("draw",)},
+ {"initial_aime_quintile": ("missing_column",)},
+ {"initial_aime_quintile": "draw"},
+ ],
+)
+def test_invalid_quintile_partition_overrides_are_refused(overrides):
+ with pytest.raises(gb.GroupBreakdownError):
+ _assign_quintiles(_quintile_rows(), overrides)
+
+
+def test_missing_per_dimension_partition_values_are_refused():
+ rows = _quintile_rows()
+ rows.loc[0, "birth_cohort"] = np.nan
+ with pytest.raises(gb.GroupBreakdownError, match="missing"):
+ _assign_quintiles(
+ rows, {"initial_aime_quintile": ("draw", "birth_cohort")}
+ )
diff --git a/tests/estimates/test_group_breakdown_sources.py b/tests/estimates/test_group_breakdown_sources.py
new file mode 100644
index 00000000..d7945c70
--- /dev/null
+++ b/tests/estimates/test_group_breakdown_sources.py
@@ -0,0 +1,122 @@
+"""G3 source artifacts: captured MINT8 definitions and labels only.
+
+These artifact-tier checks read only the committed source captures under
+``data/external`` and their provenance records. They never fetch a page,
+open microdata or read model outcomes or comparator values.
+"""
+
+from __future__ import annotations
+
+import base64
+import hashlib
+import json
+from pathlib import Path
+
+import pytest
+
+from populace_dynamics.estimates import group_breakdown as gb
+
+_ROOT = Path(__file__).resolve().parents[2]
+_EXTERNAL = _ROOT / "data" / "external"
+_CAPTURES = (
+ (
+ "mint8_row_categories.json",
+ "mint8_row_categories.provenance.json",
+ "https://www.ssa.gov/policy/docs/projections/policy-options/"
+ "increase-payroll-tax-rate.html",
+ "2026-04-01",
+ ),
+ (
+ "mint8_table_user_guide.source.html",
+ "mint8_table_user_guide.source.provenance.json",
+ "https://www.ssa.gov/policy/docs/projections/user-guide.html",
+ "2025-10-01",
+ ),
+)
+
+
+@pytest.mark.parametrize(
+ "filename,provenance_filename,source_url,date_certified", _CAPTURES
+)
+def test_capture_bytes_match_source_pins_and_provenance(
+ filename, provenance_filename, source_url, date_certified
+):
+ raw = (_EXTERNAL / filename).read_bytes()
+ relative = f"data/external/{filename}"
+ provenance = json.loads(
+ (_EXTERNAL / provenance_filename).read_text(encoding="utf-8")
+ )
+ digest = hashlib.sha256(raw).hexdigest()
+ assert digest == gb.MINT8_SOURCE_SHA256[relative]
+ assert digest == provenance["source_sha256"]
+ assert len(raw) == provenance["source_length_bytes"]
+ assert provenance["schema_version"] == "external_source_provenance.v1"
+ assert provenance["committed_source_file"] == relative
+ assert provenance["source_url"] == source_url
+ assert provenance["date_certified"] == date_certified
+ assert provenance["retrieval_date"] == "2026-10-01"
+
+
+def test_label_capture_metadata_matches_its_provenance():
+ labels = json.loads(
+ (_EXTERNAL / "mint8_row_categories.json").read_text(encoding="utf-8")
+ )
+ provenance = json.loads(
+ (_EXTERNAL / "mint8_row_categories.provenance.json").read_text(
+ encoding="utf-8"
+ )
+ )
+ assert labels["source_url"] == provenance["source_url"]
+ assert labels["dateCertified"] == provenance["date_certified"]
+ assert labels["raw_sha256"] == provenance["raw_page_sha256"]
+ assert "LABELS ONLY" in labels["note"]
+ # Only captions, headings and labels are retained from the option page.
+ for table in labels["tables"].values():
+ assert set(table) == {"caption", "columns", "groups"}
+ for group in table["groups"]:
+ assert set(group) == {"group", "labels"}
+
+
+def test_guide_capture_matches_recorded_archive_digest_and_date_locator():
+ raw = (_EXTERNAL / "mint8_table_user_guide.source.html").read_bytes()
+ provenance = json.loads(
+ (
+ _EXTERNAL / "mint8_table_user_guide.source.provenance.json"
+ ).read_text(encoding="utf-8")
+ )
+ digest = base64.b32encode(hashlib.sha1(raw).digest()).decode("ascii")
+ assert digest == provenance["source_sha1_base32"]
+ assert provenance["date_certified_locator"].split(" in ")[0] in (
+ raw.decode("utf-8")
+ )
+
+
+def test_verification_checks_all_published_layouts_and_composites():
+ verification = gb.verify_mint8_sources()
+ assert verification["files_sha256"] == gb.MINT8_SOURCE_SHA256
+ assert verification["schemes_checked"] == {
+ "mint8_beneficiary_annual": ["1", "2", "3", "7", "8", "9"],
+ "mint8_beneficiary_poverty": ["10", "11", "12"],
+ "mint8_cohort": [str(number) for number in range(13, 21)],
+ "mint8_beneficiary_annual_with_lifetime": [
+ "mint8_beneficiary_annual + table 13"
+ ],
+ "mint8_beneficiary_poverty_with_lifetime": [
+ "mint8_beneficiary_poverty + table 13"
+ ],
+ }
+ assert verification["n_quotes_checked"] > 0
+
+
+@pytest.mark.parametrize("relative", tuple(gb.MINT8_SOURCE_SHA256))
+def test_verification_refuses_tampered_capture_bytes(tmp_path, relative):
+ # Temporary copies contain captured definitions/labels only. Do not
+ # modify the committed source files to exercise their hash guard.
+ for source in gb.MINT8_SOURCE_SHA256:
+ target = tmp_path / source
+ target.parent.mkdir(parents=True, exist_ok=True)
+ target.write_bytes((_ROOT / source).read_bytes())
+ target = tmp_path / relative
+ target.write_bytes(target.read_bytes() + b"\nINVENTED tampering\n")
+ with pytest.raises(gb.GroupBreakdownError, match="SHA-256.*pin"):
+ gb.verify_mint8_sources(tmp_path)
From 1e107f26b335dcaae360bc4ff8563b27b3240058 Mon Sep 17 00:00:00 2001
From: Max Ghenis
Date: Fri, 2 Oct 2026 09:35:19 -0400
Subject: [PATCH 02/10] WIP: partial lifetime-measures work from two
interrupted builders (to be reviewed)
---
.../lifetime_measure_sources.provenance.json | 52 +
.../mint8_lifetime_quintile_definitions.json | 68 +
...ctive_interest_rates_1940_1979.source.html | 374 ++++
...ctive_interest_rates_1980_2025.source.html | 618 ++++++
.../ssa_mint8_user_guide_2026.source.html | 635 ++++++
data/external/ssa_oasdi_tax_rates_2026.json | 572 +++++
.../ssa_oasdi_tax_rates_2026.source.html | 733 ++++++
.../ssa_trust_fund_interest_rates_2026.json | 562 +++++
...trust_fund_interest_rates_2026.source.html | 836 +++++++
scripts/extract_lifetime_measure_sources.py | 778 +++++++
.../estimates/lifetime_measures.py | 1966 +++++++++++++++++
.../test_lifetime_measure_sources.py | 266 +++
tests/estimates/test_lifetime_measures.py | 1198 ++++++++++
13 files changed, 8658 insertions(+)
create mode 100644 data/external/lifetime_measure_sources.provenance.json
create mode 100644 data/external/mint8_lifetime_quintile_definitions.json
create mode 100644 data/external/ssa_effective_interest_rates_1940_1979.source.html
create mode 100644 data/external/ssa_effective_interest_rates_1980_2025.source.html
create mode 100644 data/external/ssa_mint8_user_guide_2026.source.html
create mode 100644 data/external/ssa_oasdi_tax_rates_2026.json
create mode 100644 data/external/ssa_oasdi_tax_rates_2026.source.html
create mode 100644 data/external/ssa_trust_fund_interest_rates_2026.json
create mode 100644 data/external/ssa_trust_fund_interest_rates_2026.source.html
create mode 100644 scripts/extract_lifetime_measure_sources.py
create mode 100644 src/populace_dynamics/estimates/lifetime_measures.py
create mode 100644 tests/estimates/test_lifetime_measure_sources.py
create mode 100644 tests/estimates/test_lifetime_measures.py
diff --git a/data/external/lifetime_measure_sources.provenance.json b/data/external/lifetime_measure_sources.provenance.json
new file mode 100644
index 00000000..d222d4a3
--- /dev/null
+++ b/data/external/lifetime_measure_sources.provenance.json
@@ -0,0 +1,52 @@
+{
+ "schema_version": "external_source_provenance_set.v1",
+ "source": "Social Security Administration",
+ "retrieval_date": "2026-10-01T22:31:33Z",
+ "fetch_method": "Direct HTTPS GET of the live ssa.gov page with curl -A 'Wget/1.21.4' on 2026-10-01 (HTTP 200, content-type text/html; charset=UTF-8; response Date header Thu, 01 Oct 2026 22:31:33-34 GMT). The body is committed byte for byte. A per-request Akamai mPulse script in changes between fetches; an earlier fetch the same day parsed to the identical tables.",
+ "acquired_by": "Claude Code (Opus 5.5), NASI follow-up package G2 (branch nasi/g2-lifetime-measures)",
+ "sources": {
+ "trust_fund_interest_rates": {
+ "committed_source_file": "data/external/ssa_trust_fund_interest_rates_2026.source.html",
+ "source_url": "https://www.ssa.gov/oact/ProgData/annualinterestrates.html",
+ "document": "Average and Effective Interest Rates",
+ "source_sha256": "7c11df685faaf0602dd77c5e4e124d64709e8cffd58e463d938f35e7254a0d18",
+ "source_length_bytes": 41498
+ },
+ "effective_rates_1980_on": {
+ "committed_source_file": "data/external/ssa_effective_interest_rates_1980_2025.source.html",
+ "source_url": "https://www.ssa.gov/oact/ProgData/effectiveRates.html",
+ "document": "Effective Interest Rates",
+ "source_sha256": "eaa8870898da2907bbc39da55d6a4c14f349e38c518aefb0ff87be2ddae09322",
+ "source_length_bytes": 39756
+ },
+ "effective_rates_1940_1979": {
+ "committed_source_file": "data/external/ssa_effective_interest_rates_1940_1979.source.html",
+ "source_url": "https://www.ssa.gov/oact/ProgData/effectiveRts1940-79.html",
+ "document": "Effective Interest Rates, 1940-79 (historical document)",
+ "source_sha256": "fb16fdaff44b3427eec2216740734792aab3b94e2851acde57c26e4057094e3d",
+ "source_length_bytes": 19037
+ },
+ "oasdi_tax_rates": {
+ "committed_source_file": "data/external/ssa_oasdi_tax_rates_2026.source.html",
+ "source_url": "https://www.ssa.gov/oact/ProgData/oasdiRates.html",
+ "document": "Social Security Tax Rates",
+ "source_sha256": "7aad3e899ab2ee4e2dfd91c2dc9208c79e40bf4cb7f3315867a11ba7efd8a9a4",
+ "source_length_bytes": 41837
+ },
+ "mint8_user_guide": {
+ "committed_source_file": "data/external/ssa_mint8_user_guide_2026.source.html",
+ "source_url": "https://www.ssa.gov/policy/docs/projections/user-guide.html",
+ "document": "MINT8 Table User Guide",
+ "source_sha256": "8d5bc3f0de17831c2ed07383003257a63bb657df179d0c54868fda052f23f100",
+ "source_length_bytes": 73967
+ }
+ },
+ "consumed_by": {
+ "scripts/extract_lifetime_measure_sources.py": [
+ "data/external/ssa_trust_fund_interest_rates_2026.json",
+ "data/external/ssa_oasdi_tax_rates_2026.json",
+ "data/external/mint8_lifetime_quintile_definitions.json"
+ ]
+ },
+ "reproducible": "parsed deterministically from the committed, sha256-verified HTML bodies; no network or wall-clock timestamp is used"
+}
diff --git a/data/external/mint8_lifetime_quintile_definitions.json b/data/external/mint8_lifetime_quintile_definitions.json
new file mode 100644
index 00000000..bab32025
--- /dev/null
+++ b/data/external/mint8_lifetime_quintile_definitions.json
@@ -0,0 +1,68 @@
+{
+ "schema_version": "mint8_lifetime_quintile_definitions.v1",
+ "document": "MINT8 Table User Guide",
+ "date_certified": "2026-04-01",
+ "definitions": {
+ "initial_aime_quintile": {
+ "element_id": "AIME",
+ "text": "Current-Law Initial AIME Quintile: Represents an individual's average indexed monthly earnings (AIME) under current law at age 62, the earliest eligibility age for retired-worker benefits. We calculate the AIME quintiles for each birth cohort. The dollar ranges are available upon request."
+ },
+ "lifetime_payroll_tax_quintile": {
+ "element_id": "lifetime-tax",
+ "text": "Lifetime Payroll Tax Quintile: Represents the present value of an individual's current-law payroll taxes at age 62. We calculate the payroll tax quintiles for each birth cohort. The dollar ranges are available upon request."
+ },
+ "lifetime_payroll_tax_quintile_shared": {
+ "element_id": "lifetime-tax-shared",
+ "text": "Lifetime Payroll Tax Quintile (Shared): Represents the present value of an individual's current-law payroll taxes at age 62. For married couples, the payroll taxes paid while married are shared equally between them. For never-married individuals, this is the same as the lifetime payroll tax. In any year where an individual is not married, we count only their individual payroll taxes. We calculate the quintiles for each birth cohort. The dollar ranges are available upon request."
+ },
+ "current_law_payroll_taxes_quintile": {
+ "element_id": "taxes",
+ "text": "Current-Law Payroll Taxes Quintile: Represents an individual's annual Social Security payroll taxes under current law. We calculate the payroll tax quintiles for each analysis year's population of payroll taxpayers aged 31 or older."
+ },
+ "present_value_convention": {
+ "contains": "We use the Social Security Trust Fund interest rate to adjust benefits and taxes to their present values at age 62.",
+ "text": "The present value of benefits includes all Social Security benefits the individual received, regardless of earnings record or type of benefit. The present value of payroll taxes includes all the payroll taxes that the individual paid over a lifetime. We use the Social Security Trust Fund interest rate to adjust benefits and taxes to their present values at age 62."
+ },
+ "ten_year_birth_cohorts": {
+ "contains": "We use 10-year birth cohorts to increase the sample size",
+ "text": "We use 10-year birth cohorts to increase the sample size to a point where the characteristic subgroups could be examined."
+ },
+ "household_income_quintile_assignment": {
+ "contains": "determine the dollar thresholds for each income quintile",
+ "text": "We calculate the income quintiles for each year (e.g., 2030, 2050, or 2070) for the population analyzed, determine the dollar thresholds for each income quintile, and assign each beneficiary to the appropriate quintile. The dollar ranges are available upon request."
+ },
+ "sample_size_restriction": {
+ "contains": "suppress an entire characteristic subgroup",
+ "text": "To maintain the privacy of survey respondents, our tables have built-in disclosure avoidance protections that suppress an entire characteristic subgroup if the sample size for any row in that subgroup is less than 100 individuals. This minimum sample size protects privacy so that we can show the 10th and 90th percentiles based on a sample size of at least 10."
+ }
+ },
+ "cohort_table_row_groups": {
+ "initial_aime_quintile": "Current-law initial AIME quintile",
+ "lifetime_payroll_tax_quintile": "Lifetime payroll tax quintile",
+ "lifetime_payroll_tax_quintile_shared": "Lifetime payroll tax quintile (shared)"
+ },
+ "quintile_labels_high_to_low": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ],
+ "cohort_table_label_provenance": {
+ "source_url": "https://www.ssa.gov/policy/docs/projections/policy-options/increase-payroll-tax-rate.html",
+ "retrieved_via": "http://web.archive.org/web/20260519190449id_/https://www.ssa.gov/policy/docs/projections/policy-options/increase-payroll-tax-rate.html",
+ "raw_sha256": "3cfb6eb636cad6d062eb22734c533fb349a0bc4761014eee8f2ccb0eb6e28d8c",
+ "date_certified": "2026-04-01",
+ "label_file": "mint8_row_categories.json (MINT-categories lane)",
+ "label_file_sha256": "23fbfbc8dbc14b06144a83bf0583ea0a6c89505638f3f7dc5fd6b34a3a4ac650",
+ "tables": "13-20 (benefit/tax ratios and initial replacement rates)",
+ "note": "Labels only. The lane's parser emitted th, caption and heading text and no data cell; the option table itself is not committed and supplies no input to this package."
+ },
+ "build": {
+ "built_by": "scripts/extract_lifetime_measure_sources.py",
+ "sources": [
+ "mint8_user_guide"
+ ],
+ "provenance_file": "data/external/lifetime_measure_sources.provenance.json"
+ }
+}
diff --git a/data/external/ssa_effective_interest_rates_1940_1979.source.html b/data/external/ssa_effective_interest_rates_1940_1979.source.html
new file mode 100644
index 00000000..fc0837e7
--- /dev/null
+++ b/data/external/ssa_effective_interest_rates_1940_1979.source.html
@@ -0,0 +1,374 @@
+
+
+
+
+Effective Interest Rates
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
Estimated effective interest rates earned by the
+ Old-Age and Survivors Insurance (OASI) and
+ Disability Insurance (DI) Trust Funds are shown below
+ for 1940-1979. (OASDI refers to the two funds combined.)
+
+
+
+
+
+
Estimated Effective Interest Rates Earned By the
+Assets of the OASI and DI Trust Funds, 1940-79 [Percent]
+
+
+
+
Calendar year
+
OASI
+
DI
+
OASDI
+
+
Calendar year
+
OASI
+
DI
+
OASDI
+
+
+
+
+
1940............
+
2.4
+
--
+
2.4
+
+
1960............
+
2.6
+
2.6
+
2.6
+
+
1941............
+
2.4
+
--
+
2.4
+
+
1961............
+
2.7
+
2.9
+
2.8
+
+
1942............
+
2.3
+
--
+
2.3
+
+
1962............
+
2.8
+
2.9
+
2.8
+
+
1943............
+
2.1
+
--
+
2.1
+
+
1963............
+
2.9
+
3.0
+
2.9
+
+
1944............
+
2.0
+
--
+
2.0
+
+
1964............
+
3.1
+
3.2
+
3.1
+
+
1945............
+
2.1
+
--
+
2.1
+
+
1965............
+
3.2
+
3.4
+
3.2
+
+
1946............
+
2.0
+
--
+
2.0
+
+
1966............
+
3.5
+
3.7
+
3.5
+
+
1947............
+
1.9
+
--
+
1.9
+
+
1967............
+
3.7
+
4.1
+
3.8
+
+
1948............
+
2.8
+
--
+
2.8
+
+
1968............
+
3.9
+
4.4
+
4.0
+
+
1949............
+
1.3
+
--
+
1.3
+
+
1969............
+
4.4
+
5.1
+
4.4
+
+
1950............
+
2.0
+
--
+
2.0
+
+
1970............
+
5.0
+
5.8
+
5.1
+
+
1951............
+
2.9
+
--
+
2.9
+
+
1971............
+
5.2
+
6.0
+
5.3
+
+
1952............
+
2.2
+
--
+
2.2
+
+
1972............
+
5.3
+
6.0
+
5.4
+
+
1953............
+
2.3
+
--
+
2.3
+
+
1973............
+
5.7
+
6.2
+
5.8
+
+
1954............
+
2.3
+
--
+
2.3
+
+
1974............
+
6.2
+
6.5
+
6.2
+
+
1955............
+
2.2
+
--
+
2.2
+
+
1975............
+
6.6
+
6.7
+
6.6
+
+
1956............
+
2.4
+
--
+
2.4
+
+
1976............
+
6.7
+
6.8
+
6.7
+
+
1957............
+
2.5
+
2.3
+
2.5
+
+
1977............
+
6.9
+
7.1
+
7.0
+
+
1958............
+
2.5
+
2.4
+
2.5
+
+
1978............
+
7.2
+
7.5
+
7.2
+
+
1959............
+
2.6
+
2.5
+
2.6
+
+
1979............
+
7.4
+
8.0
+
7.5
+
+
+
+
+
+OASI refers to the Old-Age and Survivors Insurance Trust Fund, DI
+refers to the Disability Insurance Trust Fund, and OASDI refers to the two
+funds combined
+
+
+
+
+ An official website of the United States government
+ Here's how you know
+
+
+
+
Official websites use .gov A .gov website belongs to an official government organization in the United States.
+
+
Secure .gov websites use HTTPS A lock (
+
+ ) or https:// means you’ve safely connected to the .gov website. Share sensitive information only on official, secure websites.
+
+
+
+
+ An official website of the United States government
+ Here's how you know
+
+
+
+
Official websites use .gov A .gov website belongs to an official government organization in the United States.
+
+
Secure .gov websites use HTTPS A lock (
+
+ ) or https:// means you’ve safely connected to the .gov website. Share sensitive information only on official, secure websites.
+
The first four sets of tables provide results for the analysis years, 2030, 2050, and 2070, while the last two provide results for four birth cohorts: 1960–1969,1980–1989,2000–2009, and 2020–2029.
+
Each set of tables and what they show are discussed below.
The first two columns show the percent of the population with a benefit decrease or increase, and the next three columns show the percent change in individual Social Security benefits at three percentiles.
The first five columns are the same as the Social Security benefits tables except tax changes are shown rather than benefit changes. Columns 6–8 show the dollar amount changes in individual Social Security payroll taxes at the same percentiles as the percentage change columns.
+
Household Income
+
These tables have the same structure and population as the Social Security benefits tables except the effects on household income are shown rather than individual Social Security benefits.
The first two columns show the poverty rate with and without (current law) the proposed option. The next three columns show the number in poverty (expressed in thousands) with and without (current law) the proposed option and the difference between them. The final column shows the percent change in the number in poverty with the proposed change. This percent change is calculated by dividing the change in thousands in poverty (5th column) by the thousands in poverty without the proposal (3rd column).
+
CORRECT interpretation: the number of people in poverty would decline by 2 percent. INCORRECT interpretation: the poverty rate would decline by 2 percent.
+
Benefit/Tax Ratios
+
These tables show the projected changes in the benefit/tax ratios for workers born in four birth cohorts: 1960–1969,1980–1989,2000–2009, and 2020–2029 who have a tax record from which to calculate a benefit/tax ratio. We excluded those who paid zero payroll taxes over their lifetime and, therefore, could not have a benefit/tax ratio calculated.
+
The first five columns are similar to the Social Security benefits and taxes paid tables except the percent of the population with a benefit/tax ratio decrease and increase, and the percent change in the benefit/tax ratio at three percentiles are shown.
+
The last six columns show the distribution of benefit/tax ratios with and without (current law) the proposed option at three percentiles.
+
The benefit/tax ratio is a money's worth measure that is the lifetime present value of benefits divided by the lifetime present value of payroll taxes. The ratio represents how much in benefits an individual received for every dollar of payroll taxes paid.
+
How to interpret the ratios:
+
+
0% means that an individual received no benefits despite paying payroll taxes (or $0.00 in benefits for every $1.00 of taxes).
+
50% means that an individual received half as much benefits as was paid in taxes (or $0.50 in benefits for every $1.00 of taxes).
+
100% means that an individual received the same amount of benefits as was paid in taxes (or $1.00 in benefits for every $1.00 of taxes).
+
1,350% means that an individual received over 13 times the amount of benefits as was paid in taxes (or $13.50 in benefits for every $1.00 of taxes).
+
+
The present value of benefits includes all Social Security benefits the individual received, regardless of earnings record or type of benefit. The present value of payroll taxes includes all the payroll taxes that the individual paid over a lifetime. We use the Social Security Trust Fund interest rate to adjust benefits and taxes to their present values at age 62.
+
Initial Replacement Rates
+
These tables are the same as the benefit/tax ratio tables except initial replacement rates are shown for current-law beneficiaries born in four birth cohorts: 1960–1969,1980–1989,2000–2009, and 2020–2029. Only beneficiaries with both income and benefit records from which to calculate an initial replacement rate are included in the population. Beneficiaries with zero average indexed monthly earnings (AIME) or zero benefit at claiming (due to the earnings test or other fixed-dollar reductions) are excluded because no replacement took place. When no benefit is received, the initial replacement of earnings may take place in a later year or never, if the beneficiary dies before a benefit is paid.
+
The initial replacement rate represents how much of the initial AIME is replaced by the initial total monthly Social Security benefit. It is calculated by dividing the initial monthly benefit by the initial AIME.
+
How to interpret the initial replacement rate values:
+
+
20% means that the initial benefit replaced one-fifth of lifetime earnings (or $0.20 in monthly benefits for every $1.00 of AIME).
+
100% means that the initial benefit replaced all lifetime earnings (or $1.00 in monthly benefits for every $1.00 of AIME).
+
150% means that the initial benefit replaced one and a half times lifetime earnings (or $1.50 in monthly benefits for every $1.00 of AIME).
+
+
The initial monthly benefit is the total individual Social Security benefit received at the person's claiming age, including any spousal, survivor, or disability benefits received. We calculate the replacement rate at claiming age, regardless of what type of benefits the beneficiary claimed.
+
Interpreting the Profile of Beneficiaries by Race & Ethnicity Tables
The tables also include the same information for four race and ethnicity groupings:
+
+
Hispanic or Latino, any race
+
White, non-Hispanic
+
Black or African American, non-Hispanic
+
All other races, non-Hispanic
+
+
Each set of tables and what they show are discussed below.
+
Social Security Benefits
+
These tables show the projected distribution of individual monthly Social Security benefits at three percentiles. The benefit amount is the total monthly benefit an individual would receive, regardless of the type of benefit or earnings record it came from.
+
Poverty Rates and Numbers
+
These tables show the poverty rate and number in poverty under the Official Poverty Measure and the Supplemental Poverty Measure.
+
The first and third columns show the rates and numbers under the Official Poverty Measure while the second and fourth columns show the same information under the Supplemental Poverty Measure. Further information about the Official and Supplemental Poverty Measures and how they relate to the aged, is available in this paper.
The last five columns show the projected mean share of household income from five sources:
+
+
Social Security benefits, which include benefits for the individual, spouse, and any children.
+
Annuitized asset income, which includes income from defined contribution plans (such as 401(k) accounts) and personal savings.
+
Defined benefit pension income, which includes the individual's and any spouse's defined benefit pension income.
+
All earnings, including covered earnings (from which Social Security taxes are withheld) and non-covered earnings (no Social Security taxes withheld) of the individual and his or her spouse.
+
Coresident income, which is the income of non-spousal coresidents in the household.
+
+
Rows may not sum to 100 percent because minor sources of income are excluded.
+
Total Earnings
+
These tables show the projected distribution of annual individual total earnings at three percentiles. Total earnings includes both covered and non-covered wages.
+
Household Wealth
+
These tables show the projected distribution of household wealth at three percentiles. Wealth includes retirement account balances and savings, but excludes household equity.
+
Health Status and Costs
+
These tables show projected health status and expenses. The first two columns cover the percent and number (in thousands) of beneficiaries who are projected to self-report fair or poor health. The third column shows family median annual health insurance premiums, and the fourth column shows family median annual out-of-pocket health expenses.
+
Interpreting the Profile of Taxpayers by Race & Ethnicity Tables
These tables show the projected distribution of annual individual Social Security taxes paid at three percentiles. Both the employer and employee shares of the Social Security portion of FICA (Federal Insurance Contributions Act) taxes are included.
+
Covered Earnings
+
These tables show the projected distribution of annual individual covered earnings at three percentiles. Covered earnings are wages from work that is subject to the Social Security payroll tax. The covered earnings in these tables are not capped at the taxable maximum.
Interpreting the Population Characteristics Tables
+
Each projection table links (at the top) to the corresponding population characteristics table. From there, you can use the tabs on the right side to view each population analyzed across the MINT projections including:
We use 10-year birth cohorts to increase the sample size to a point where the characteristic subgroups could be examined.
+
The three columns in every population characteristics table are:
+
+
Unweighted sample: The number of people in the sample for each row. We use it to verify that we comply with disclosure avoidance policy (see “Sample Size Restrictions”).
+
Population (in thousands): The weighted population in thousands. A value of 71,500 means 71 million, 500 thousand.
+
Share of population: The percentage of the population in that particular characteristic subgroup. The Total row at the top of the table is always 100%. The percentages for the each subgroup underneath should add to 100%. For instance, under Country of Birth, United States may be 84% and Other Countries would be 16%, which adds to 100%.
+
+
CORRECT interpretation: 40% of the population is female, 60% of the population is male INCORRECT interpretation: 40% of females are in this population, 60% of males are in this population
Total: Refers to the total population of the table.
+
Sex: Female or Male.
+
Race/Ethnicity: We list “Hispanic or Latino, any race” first; the rest of the groups (White, Black or African American, and All other races) are non-Hispanic. Additional racial or ethnic identifications are not covered because they are not in the datasets used to build the MINT8 model.
+
Country of Birth: We differentiate between the United States and other countries.
+
Age:
+
+
The beneficiary population includes those aged 60 or older because 60 is the earliest eligibility age for any aged benefits under current law.
+
The taxpayer population includes those aged 31 or older because 31 is the earliest age in MINT for household income and poverty information.
+
+
Marital Status: Refers to the marital status in the year of analysis only. An individual's marital status can change in the future and may have been different in the past.
+
Highest Education Level: Reported number of years of education.
+
+
Graduate means more than 16 years of education,
+
Bachelor means 16 years of education,
+
Associate means 14–15 years of education,
+
High school means 12–13 years of education, and
+
Less than high school means less than 12 years of education.
+
+
Current-Law Poverty Status: Indicates whether the person is in a household that has income above (“above poverty”) or below (“in poverty”) the official poverty line under current law. The household income used for the official poverty measure is the same as the household income used in our results except for how asset income is counted. The official poverty measure of asset income only includes dividend income, interest income, and rental income (non-annuitized) as reported on income tax returns. In contrast, we include annuitized asset income from all household wealth held in defined contribution plans (such as 401(k) accounts) and personal savings in that year. We add the annuitized asset income to account for the expected spend-down of assets in retirement. The asset income value used for official poverty calculations generally produces a substantially lower asset income value than the household income measure.
+
Current-Law Household Income Quintile: Represents an individual's annual household income under current law, including:
+
+
household earnings;
+
asset income (annuitized), which includes income from defined contribution plans (such as 401(k) accounts) and personal savings;
+
defined benefit pensions;
+
means-tested income;
+
non-means-tested income;
+
Social Security;
+
Supplemental Security Income; and
+
non-spousal co-residents' income.
+
+
We calculate the income quintiles for each year (e.g., 2030, 2050, or 2070) for the population analyzed, determine the dollar thresholds for each income quintile, and assign each beneficiary to the appropriate quintile. The dollar ranges are available upon request.
+
Current-Law Benefit Type: Some Social Security benefits are based on one's own work, while others are based on the work of a current, divorced, or deceased spouse. The current-law benefit type refers to one of the following benefit types received in the specified analysis year:
+
+
Retired-worker only: receives only a retired-worker benefit based on his or her earnings record.
+
Widow(er) (includes dually entitled): receives a survivor benefit (may or may not also receive a lower worker benefit from his or her own earnings record, known as dually entitled).
+
Spousal (includes dually entitled): receives a spousal benefit (may or may not also receive a lower worker benefit from his or her own earnings record, known as dually entitled).
+
Disabled-worker only: receives a disabled-worker benefit on his or her earnings record and is under the full retirement age (FRA). Disabled workers convert to retired workers at FRA.
+
+
Our results do not show different benefit types a beneficiary might receive under a policy option/proposal or in a different year under current law.
+
Current-Law Payroll Taxes Quintile: Represents an individual's annual Social Security payroll taxes under current law. We calculate the payroll tax quintiles for each analysis year's population of payroll taxpayers aged 31 or older.
+
Current-Law Initial AIME Quintile: Represents an individual's average indexed monthly earnings (AIME) under current law at age 62, the earliest eligibility age for retired-worker benefits. We calculate the AIME quintiles for each birth cohort. The dollar ranges are available upon request.
+
Lifetime Payroll Tax Quintile: Represents the present value of an individual's current-law payroll taxes at age 62. We calculate the payroll tax quintiles for each birth cohort. The dollar ranges are available upon request.
+
Lifetime Payroll Tax Quintile (Shared): Represents the present value of an individual's current-law payroll taxes at age 62. For married couples, the payroll taxes paid while married are shared equally between them. For never-married individuals, this is the same as the lifetime payroll tax. In any year where an individual is not married, we count only their individual payroll taxes. We calculate the quintiles for each birth cohort. The dollar ranges are available upon request.
+
+
Measures—Column Headings
+
+
Threshold for Categorization in the “Decrease” or “Increase” Groups (“Percent of Population with a—[decrease or increase]” Columns)
+
We categorize individuals as having a “decrease” in the amount being analyzed (benefits, taxes, income, etc.) when a proposal would reduce the analyzed quantity by 1% or more. Individuals are categorized as having an “increase” when a proposal would raise the analyzed quantity by 1% or more. We consider individuals with differences between −1% and 1% to be unaffected.
+
For example, consider two individuals with benefits under an option/proposal that are lower than benefits without the proposal—individual A has a benefit decrease of 0.8% while individual B has a decrease of 1.6%. We would consider individual A unaffected and not categorize him or her in the “decrease” or “increase” columns. However, we would categorize individual B as having a decrease.
+
Thus, some beneficiaries who are technically “affected” by a proposal, but who will still receive essentially the same benefit amount (or pay essentially the same taxes, etc.) are considered to be unaffected in our projections.
+
Percent Change Values (“Percent Change in [item being analyzed] at the—[three percentiles]” Columns)
+
Understanding how to interpret the distribution of percent changes is critical to understanding the results correctly.
+
The formula we use to calculate the percent change for each individual is:
From this distribution of individual percent changes, ranked from high to low, we calculate the 10th percentile, median, and 90th percentile values (see below).
+
CORRECT interpretation: a −5% median indicates that this is the median of the distribution of individual percent changes (half the individuals have a percent change that is higher and half have a percent change that is lower). INCORRECT interpretation: a −5% median indicates that the median amount under the option is 5% less than the median amount under current law.
+
10th, Median, and 90th Percentile Values
+
The percentiles provide a picture of the distribution of policy option effects or beneficiaries/taxpayers financial status distributed from lowest to highest. The example below is for a percent change in benefits, but also applies to distributions of other policy options, dollar amounts, initial replacement rates, household income levels, etc.
+
+
+
Table example
+
+
+
+
+
+
Percent change in Social Security benefits at the—
+
+
+
10th percentile
+
Median
+
90th percentile
+
+
+
+
+
Total
+
2%
+
4%
+
21%
+
+
+
+
+
+
+
+
+
+
10th percentile: “2%” means that 10 percent of the population has a benefit change of less than 2 percent, while 90 percent have a benefit change of more than 2 percent.
+
Median: “4%” means that 50 percent of the population has a benefit change of less than 4 percent, while 50 percent have a benefit change of more than 4 percent.
+
90th percentile: “21%” means that 90 percent of the population has a benefit change of less than 21 percent, while 10 percent have a benefit change of more than 21 percent.
+
CORRECT interpretations include:
+
+
10% of this population has a benefit change of less than 2%.
+
10% of this population has a benefit change of more than 21%.
+
40% of this population has a benefit change of between 2% to 4%.
+
40% of this population has a benefit change of between 4% to 21%.
+
50% of this population has a benefit change of less than 4%.
+
50% of this population has a benefit change of more than 4%.
+
+
+
+
Additional Notes
+
Dollar Amounts
+
All dollar amounts are presented in today's dollars, meaning that they are in real dollars (inflation-adjusted) for the year the table is produced. If a table is run in 2021, the dollars are in 2021 dollars and the table will note that in the column label. Tables run in 2024 will be in 2024 dollars, and so on.
+
Sample Size Restrictions
+
To maintain the privacy of survey respondents, our tables have built-in disclosure avoidance protections that suppress an entire characteristic subgroup if the sample size for any row in that subgroup is less than 100 individuals. This minimum sample size protects privacy so that we can show the 10th and 90th percentiles based on a sample size of at least 10.
+
For example, if there are only 82 widow(er)s in the Widowed row of the Marital Status subgroup, the entire Marital Status subgroup is removed from the table. We remove the entire subgroup to avoid secondary or tertiary disclosure issues. The subgroups below Marital Status in this example would automatically move up the page. A “short” table can reveal at a glance that at least one subgroup has been removed.
+
Columns that show percent of the population with a decrease or increase in whatever value is being shown (e.g. benefits, payroll tax, household income) have disclosure restrictions on the numerator's sample size as well as the denominator. The numerator must have either zero cases or meet a minimum numerator threshold of 10 for the table to display it. The table would suppress any characteristic subgroup that has a percentage based on a numerator of 1–9 cases. (MINT results are weighted, but for this example, everything is presented in unweighted sample sizes.)
+
There is an exception for low numerator situations where the table would show a 0% for any numerator from zero up to and including the minimum numerator threshold. If a particular percentage was based on seven records with a denominator of 30,000, it would produce a percentage of 0.02%. By only showing percentages in single digits, we would display this result as 0%, which would be the same value displayed for any numerator from 0–149.
+
This exception is important because there are policy options where 27,000/30,000 beneficiaries (90%) would receive a benefit increase while 8/30,000 receive a decrease. Without the exception, the very small decrease numerator would suppress a number of subgroups for both the increased and the decreased results, which limits the presentable results more than is necessary for disclosure avoidance.
+
1 All the characteristic subgroups are not in every table. Birth cohort tables (benefit/tax ratios and initial replacement rates) do not have age or marital status breakouts. Characteristic subgroups can also drop out of tables because of sample size restrictions (see “Sample Size Restrictions” for details).
+
+
+
+ An official website of the United States government
+ Here's how you know
+
+
+
+
Official websites use .gov A .gov website belongs to an official government organization in the United States.
+
+
Secure .gov websites use HTTPS A lock (
+
+ ) or https:// means you’ve safely connected to the .gov website. Share sensitive information only on official, secure websites.
+
+ The rates shown reflect the amounts received by the trust funds.
+ In certain years, the effective rate paid by employees, employers, and/or
+ self-employed workers was less than the rate received by the trust funds, with
+ the difference covered by general revenue. See the footnotes for details.
+
a
+ In 1984 only, an immediate credit of 0.3 percent of taxable wages was
+ allowed against the OASDI taxes paid by employees, resulting in an
+ effective employee tax rate of 5.4 percent. The OASI and DI trust funds, however,
+ received general revenue equivalent to 0.3 percent of taxable wages for
+ 1984. Similar credits of 2.7 percent, 2.3 percent, and 2.0 percent were allowed
+ against the combined OASDI and HI taxes on net earnings from self-employment
+ in 1984, 1985, and 1986-89, respectively.
+ b
+ Beginning in 1990, self-employed workers are allowed a deduction, for purposes
+ of computing their net earnings, equal to half of the combined OASDI and HI
+ contributions that would be payable without regard to the contribution and
+ benefit base. The OASDI contribution rate is then applied to net earnings after
+ this deduction, but subject to the OASDI base.
+ c
+ For 2010, most employers were exempt from paying the employer share of OASDI
+ tax on wages paid to certain qualified individuals hired after February 3. For
+ 2011 and 2012, the OASDI tax rate is reduced by 2 percentage points for employees and
+ for self-employed workers, resulting in a 4.2 percent effective tax rate for
+ employees and a 10.4 percent effective tax rate for self-employed workers. These
+ reductions in tax revenue due to lower tax rates are being made up
+ by transfers from the general fund of the Treasury to the OASI and DI trust
+ funds.
+ d
+ Public Law 114-74, the Bipartisan Budget Act of 2015, temporarily re-allocated a portion
+ of the OASI tax rate to DI for calendar years 2016 through 2018. Beginning in 2019, the
+ tax rates for each fund revert to the rates in effect from 2000 through 2015.
+
+
+
+
+ An official website of the United States government
+ Here's how you know
+
+
+
+
Official websites use .gov A .gov website belongs to an official government organization in the United States.
+
+
Secure .gov websites use HTTPS A lock (
+
+ ) or https:// means you’ve safely connected to the .gov website. Share sensitive information only on official, secure websites.
+
An effective interest rate for a calendar year is the interest earned in
+ that year divided by the average level of assets held during the year. This rate reflects
+ the entire portfolio of securites held by the Social Security trust funds
+ (OASI and DI).
+ Effective rates for the trust funds on a combined basis are shown below; rates for
+ each trust fund, separately, are also available.
+
+
+
+
+
+
+
+
+
+
+
Average annual special-issue interest rates on new issues and effective
+ annual interest rates (percent)
+
+
+
+
+
Year
+
Average
+
Effective
+
+
+
1940
+
2.5
+
2.4
+
+
+
1941
+
2.5
+
2.4
+
+
+
1942
+
2.2
+
2.3
+
+
+
1943
+
1.9
+
2.1
+
+
+
1944
+
1.9
+
2.0
+
+
+
1945
+
1.9
+
2.1
+
+
+
1946
+
1.9
+
2.0
+
+
+
1947
+
2.0
+
1.9
+
+
+
1948
+
2.1
+
2.8
+
+
+
1949
+
2.1
+
1.3
+
+
+
1950
+
2.1
+
2.0
+
+
+
1951
+
2.2
+
2.9
+
+
+
1952
+
2.3
+
2.2
+
+
+
1953
+
2.4
+
2.3
+
+
+
1954
+
2.3
+
2.3
+
+
+
1955
+
2.3
+
2.2
+
+
+
1956
+
2.5
+
2.4
+
+
+
1957
+
2.5
+
2.5
+
+
+
1958
+
2.6
+
2.5
+
+
+
1959
+
2.6
+
2.6
+
+
+
1960
+
2.9
+
2.6
+
+
+
1961
+
3.8
+
2.8
+
+
+
1962
+
3.9
+
2.8
+
+
+
1963
+
3.9
+
2.9
+
+
+
1964
+
4.1
+
3.1
+
+
+
1965
+
4.2
+
3.2
+
+
+
1966
+
4.9
+
3.5
+
+
+
1967
+
5.0
+
3.8
+
+
+
1968
+
5.5
+
4.0
+
+
+
1969
+
6.6
+
4.4
+
+
+
+
+
+
+
Year
+
Average
+
Effective
+
+
+
1970
+
7.3
+
5.1
+
+
+
1971
+
6.0
+
5.3
+
+
+
1972
+
5.9
+
5.4
+
+
+
1973
+
6.6
+
5.8
+
+
+
1974
+
7.5
+
6.2
+
+
+
1975
+
7.4
+
6.6
+
+
+
1976
+
7.1
+
6.7
+
+
+
1977
+
7.1
+
7.0
+
+
+
1978
+
8.2
+
7.2
+
+
+
1979
+
9.1
+
7.5
+
+
+
1980
+
11.0
+
8.6
+
+
+
1981
+
13.3
+
9.9
+
+
+
1982
+
12.8
+
11.2
+
+
+
1983
+
11.0
+
10.8
+
+
+
1984
+
12.4
+
11.6
+
+
+
1985
+
10.8
+
11.2
+
+
+
1986
+
8.0
+
11.1
+
+
+
1987
+
8.4
+
10.1
+
+
+
1988
+
8.8
+
9.8
+
+
+
1989
+
8.7
+
9.6
+
+
+
1990
+
8.6
+
9.3
+
+
+
1991
+
8.0
+
9.1
+
+
+
1992
+
7.1
+
8.7
+
+
+
1993
+
6.1
+
8.3
+
+
+
1994
+
7.1
+
8.0
+
+
+
1995
+
6.9
+
7.8
+
+
+
1996
+
6.6
+
7.6
+
+
+
1997
+
6.6
+
7.5
+
+
+
1998
+
5.6
+
7.2
+
+
+
1999
+
5.9
+
6.9
+
+
+
+
+
+
+
Year
+
Average
+
Effective
+
+
+
2000
+
6.2
+
6.9
+
+
+
2001
+
5.2
+
6.6
+
+
+
2002
+
4.9
+
6.4
+
+
+
2003
+
4.1
+
6.0
+
+
+
2004
+
4.3
+
5.7
+
+
+
2005
+
4.3
+
5.5
+
+
+
2006
+
4.8
+
5.3
+
+
+
2007
+
4.7
+
5.3
+
+
+
2008
+
3.6
+
5.1
+
+
+
2009
+
2.9
+
4.9
+
+
+
2010
+
2.8
+
4.6
+
+
+
2011
+
2.4
+
4.4
+
+
+
2012
+
1.5
+
4.1
+
+
+
2013
+
1.9
+
3.8
+
+
+
2014
+
2.3
+
3.6
+
+
+
2015
+
2.0
+
3.4
+
+
+
2016
+
1.8
+
3.2
+
+
+
2017
+
2.3
+
3.0
+
+
+
2018
+
2.9
+
2.9
+
+
+
2019
+
2.2
+
2.8
+
+
+
2020
+
1.0
+
2.6
+
+
+
2021
+
1.4
+
2.5
+
+
+
2022
+
3.0
+
2.4
+
+
+
2023
+
4.1
+
2.4
+
+
+
2024
+
4.3
+
2.5
+
+
+
2025
+
4.3
+
2.6
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
diff --git a/scripts/extract_lifetime_measure_sources.py b/scripts/extract_lifetime_measure_sources.py
new file mode 100644
index 00000000..7bf41ed1
--- /dev/null
+++ b/scripts/extract_lifetime_measure_sources.py
@@ -0,0 +1,778 @@
+"""Extract the SSA sources behind the G2 lifetime-earnings measures.
+
+Package G2 of the NASI follow-ups (group breakdowns of the four blind
+tests) needs three external inputs that the repository did not hold:
+
+* the OASDI payroll tax rate schedule from 1937 (SSA OACT, "Social
+ Security Tax Rates", ``oasdiRates.html``);
+* the trust fund interest rates from 1940, both the effective rate
+ earned by the combined OASI and DI trust funds and the average annual
+ special-issue rate on new issues (SSA OACT, "Average and Effective
+ Interest Rates", ``annualinterestrates.html``), cross-checked against
+ the two per-fund effective-rate pages (``effectiveRates.html`` for
+ 1980 on and ``effectiveRts1940-79.html``);
+* the verbatim MINT8 definitions of the lifetime-earnings quintile rows
+ (SSA, "MINT8 Table User Guide", ``user-guide.html``).
+
+Every source is the exact HTTP 200 response body fetched from ssa.gov on
+2026-10-01 with ``curl -A 'Wget/1.21.4'`` (ssa.gov refuses some default
+clients). The pages carry a per-request Akamai mPulse script in
+````, so a second fetch has different bytes; two fetches taken
+minutes apart on 2026-10-01 parsed to identical tables. The committed
+bytes are pinned by SHA-256 below and verified before parsing.
+
+``ssa_effective_interest_rates_2014.json`` (the M7 vintage file, 1980-2013)
+is not edited. The new rate file is a 2026 vintage: it must not be used
+by any code path bound to the M7 2014 information boundary
+(``engine.refit.validate_external_vintage`` rejects vintage 2026).
+
+The MINT8 table row-group labels ("Current-law initial AIME quintile",
+...) are not on the user guide page. They are transcribed from the
+label-only extraction of SSA's payroll-tax option table made by the
+MINT-categories lane (no data cell was extracted); its provenance is
+recorded in :data:`MINT8_TABLE_LABEL_PROVENANCE`.
+
+Run from the repository root::
+
+ .venv/bin/python scripts/extract_lifetime_measure_sources.py
+"""
+
+from __future__ import annotations
+
+import hashlib
+import json
+import re
+from html.parser import HTMLParser
+from pathlib import Path
+from typing import Any
+
+ROOT = Path(__file__).resolve().parents[1]
+EXTERNAL = ROOT / "data" / "external"
+
+RETRIEVAL_DATE = "2026-10-01T22:31:33Z"
+FETCH_METHOD = (
+ "Direct HTTPS GET of the live ssa.gov page with curl -A "
+ "'Wget/1.21.4' on 2026-10-01 (HTTP 200, content-type text/html; "
+ "charset=UTF-8; response Date header Thu, 01 Oct 2026 22:31:33-34 "
+ "GMT). The body is committed byte for byte. A per-request Akamai "
+ "mPulse script in changes between fetches; an earlier fetch "
+ "the same day parsed to the identical tables."
+)
+ACQUIRED_BY = (
+ "Claude Code (Opus 5.5), NASI follow-up package G2 "
+ "(branch nasi/g2-lifetime-measures)"
+)
+
+#: Committed source bodies, each pinned by SHA-256 and length.
+SOURCES: dict[str, dict[str, Any]] = {
+ "trust_fund_interest_rates": {
+ "file": "ssa_trust_fund_interest_rates_2026.source.html",
+ "url": "https://www.ssa.gov/oact/ProgData/annualinterestrates.html",
+ "title": "Average and Effective Interest Rates",
+ "sha256": (
+ "7c11df685faaf0602dd77c5e4e124d64709e8cffd58e463d938f35e7254a0d18"
+ ),
+ "bytes": 41498,
+ },
+ "effective_rates_1980_on": {
+ "file": "ssa_effective_interest_rates_1980_2025.source.html",
+ "url": "https://www.ssa.gov/oact/ProgData/effectiveRates.html",
+ "title": "Effective Interest Rates",
+ "sha256": (
+ "eaa8870898da2907bbc39da55d6a4c14f349e38c518aefb0ff87be2ddae09322"
+ ),
+ "bytes": 39756,
+ },
+ "effective_rates_1940_1979": {
+ "file": "ssa_effective_interest_rates_1940_1979.source.html",
+ "url": "https://www.ssa.gov/oact/ProgData/effectiveRts1940-79.html",
+ "title": "Effective Interest Rates, 1940-79 (historical document)",
+ "sha256": (
+ "fb16fdaff44b3427eec2216740734792aab3b94e2851acde57c26e4057094e3d"
+ ),
+ "bytes": 19037,
+ },
+ "oasdi_tax_rates": {
+ "file": "ssa_oasdi_tax_rates_2026.source.html",
+ "url": "https://www.ssa.gov/oact/ProgData/oasdiRates.html",
+ "title": "Social Security Tax Rates",
+ "sha256": (
+ "7aad3e899ab2ee4e2dfd91c2dc9208c79e40bf4cb7f3315867a11ba7efd8a9a4"
+ ),
+ "bytes": 41837,
+ },
+ "mint8_user_guide": {
+ "file": "ssa_mint8_user_guide_2026.source.html",
+ "url": "https://www.ssa.gov/policy/docs/projections/user-guide.html",
+ "title": "MINT8 Table User Guide",
+ "sha256": (
+ "8d5bc3f0de17831c2ed07383003257a63bb657df179d0c54868fda052f23f100"
+ ),
+ "bytes": 73967,
+ },
+}
+
+INTEREST_OUT = EXTERNAL / "ssa_trust_fund_interest_rates_2026.json"
+TAX_OUT = EXTERNAL / "ssa_oasdi_tax_rates_2026.json"
+MINT8_OUT = EXTERNAL / "mint8_lifetime_quintile_definitions.json"
+PROVENANCE_OUT = EXTERNAL / "lifetime_measure_sources.provenance.json"
+
+INTEREST_SCHEMA = "ssa_trust_fund_interest_rates.v1"
+TAX_SCHEMA = "ssa_oasdi_tax_rates.v1"
+MINT8_SCHEMA = "mint8_lifetime_quintile_definitions.v1"
+PROVENANCE_SCHEMA = "external_source_provenance_set.v1"
+
+VINTAGE_YEAR = 2026
+FIRST_INTEREST_YEAR = 1940
+LATEST_INTEREST_YEAR = 2025
+INTEREST_CAPTION = (
+ "Average annual special-issue interest rates on new issues and "
+ "effective annual interest rates (percent)"
+)
+INTEREST_HEADERS = ("Year", "Average", "Effective")
+EFFECTIVE_1980_CAPTION = (
+ "Effective Interest Rates Earned By the Invested Assets of the OASI "
+ "and DI Trust Funds [Percent]"
+)
+EFFECTIVE_1940_CAPTION = (
+ "Estimated Effective Interest Rates Earned By the Assets of the OASI "
+ "and DI Trust Funds, 1940-79[Percent]"
+)
+EFFECTIVE_HEADERS = ("Calendar year", "OASI", "DI", "OASDI")
+EFFECTIVE_1940_HEADERS = ("Calendaryear", "OASI", "DI", "OASDI")
+INTEREST_DEFINITIONS = {
+ "average_new_issue": (
+ "The average special-issue interest rate for a calendar year is "
+ "the average of the 12 monthly interest rates on new issues "
+ "during the year."
+ ),
+ "effective_oasdi": (
+ "An effective interest rate for a calendar year is the interest "
+ "earned in that year divided by the average level of assets held "
+ "during the year. This rate reflects the entire portfolio of "
+ "securites held by the Social Security trust funds"
+ ),
+}
+
+TAX_SUMMARY = "Tax rate table for Social Security trust funds"
+TAX_GROUP_HEADERS = (
+ "Tax rate for employees and employers, each",
+ "Tax rate for self-employed workers",
+)
+TAX_PREAMBLE_SENTENCE = (
+ "The rates shown reflect the amounts received by the trust funds."
+)
+FIRST_TAX_YEAR = 1937
+OPEN_ENDED_TAX_YEAR = 2019
+
+#: Footnote sentences that set an effective employee rate below the
+#: trust-fund rate for every covered employee in a year (the only footnote
+#: provisions that apply to wage earners as a class). Each value is
+#: parsed from its quoted sentence, and the quote is asserted verbatim.
+PAID_RATE_QUOTES: dict[str, dict[str, Any]] = {
+ "a": {
+ "years": (1984,),
+ "quote": (
+ "resulting in an effective employee tax rate of 5.4 percent"
+ ),
+ "pattern": r"effective employee tax rate of (\d+\.\d) percent",
+ },
+ "c": {
+ "years": (2011, 2012),
+ "quote": (
+ "resulting in a 4.2 percent effective tax rate for employees"
+ ),
+ "pattern": (
+ r"resulting in a (\d+\.\d) percent effective tax rate for "
+ r"employees"
+ ),
+ },
+}
+#: Footnote provisions recorded but not encoded: none applies to every
+#: covered wage earner, or they concern self-employment, which the careers
+#: frames do not separate from wages.
+TAX_NOT_ENCODED = {
+ "a": (
+ "the 1984-89 credits against the combined OASDI and HI taxes on "
+ "net earnings from self-employment (self-employment income is not "
+ "separated in the careers frames)"
+ ),
+ "b": (
+ "the self-employment deduction from 1990 (self-employment income "
+ "is not separated in the careers frames)"
+ ),
+ "c": (
+ "the 2010 employer exemption, which applied only to wages paid to "
+ "certain qualified individuals hired after February 3"
+ ),
+ "d": (
+ "the 2016-18 reallocation between OASI and DI, which leaves the "
+ "OASDI total unchanged"
+ ),
+}
+
+#: Paragraphs of the MINT8 user guide that G2 uses, located by element id
+#: or by a sentence that only that paragraph contains.
+MINT8_PARAGRAPHS: dict[str, dict[str, str]] = {
+ "initial_aime_quintile": {"element_id": "AIME"},
+ "lifetime_payroll_tax_quintile": {"element_id": "lifetime-tax"},
+ "lifetime_payroll_tax_quintile_shared": {
+ "element_id": "lifetime-tax-shared"
+ },
+ "current_law_payroll_taxes_quintile": {"element_id": "taxes"},
+ "present_value_convention": {
+ "contains": (
+ "We use the Social Security Trust Fund interest rate to adjust "
+ "benefits and taxes to their present values at age 62."
+ )
+ },
+ "ten_year_birth_cohorts": {
+ "contains": "We use 10-year birth cohorts to increase the sample size"
+ },
+ "household_income_quintile_assignment": {
+ "contains": "determine the dollar thresholds for each income quintile"
+ },
+ "sample_size_restriction": {
+ "contains": "suppress an entire characteristic subgroup"
+ },
+}
+
+#: Provenance of the MINT8 cohort-table row-group labels, transcribed from
+#: the label-only extraction (tables 13-20 of the payroll-tax option page).
+MINT8_TABLE_LABEL_PROVENANCE = {
+ "source_url": (
+ "https://www.ssa.gov/policy/docs/projections/policy-options/"
+ "increase-payroll-tax-rate.html"
+ ),
+ "retrieved_via": (
+ "http://web.archive.org/web/20260519190449id_/https://www.ssa.gov/"
+ "policy/docs/projections/policy-options/"
+ "increase-payroll-tax-rate.html"
+ ),
+ "raw_sha256": (
+ "3cfb6eb636cad6d062eb22734c533fb349a0bc4761014eee8f2ccb0eb6e28d8c"
+ ),
+ "date_certified": "2026-04-01",
+ "label_file": "mint8_row_categories.json (MINT-categories lane)",
+ "label_file_sha256": (
+ "23fbfbc8dbc14b06144a83bf0583ea0a6c89505638f3f7dc5fd6b34a3a4ac650"
+ ),
+ "tables": "13-20 (benefit/tax ratios and initial replacement rates)",
+ "note": (
+ "Labels only. The lane's parser emitted th, caption and heading "
+ "text and no data cell; the option table itself is not committed "
+ "and supplies no input to this package."
+ ),
+}
+MINT8_COHORT_ROW_GROUPS = {
+ "initial_aime_quintile": "Current-law initial AIME quintile",
+ "lifetime_payroll_tax_quintile": "Lifetime payroll tax quintile",
+ "lifetime_payroll_tax_quintile_shared": (
+ "Lifetime payroll tax quintile (shared)"
+ ),
+}
+MINT8_QUINTILE_LABELS = (
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest",
+)
+
+
+def _normalize(text: str) -> str:
+ """Collapse whitespace and unify dashes; keep every other character."""
+ for dash in "‐‑‒–—−":
+ text = text.replace(dash, "-")
+ return " ".join(text.replace("\xa0", " ").split())
+
+
+def _page_text(source: str) -> str:
+ """Tag-stripped, whitespace-normalized page text."""
+ return _normalize(re.sub(r"<[^>]+>", " ", source))
+
+
+class _PageParser(HTMLParser):
+ """Collect table rows, captions, summaries, paragraphs and footnotes."""
+
+ def __init__(self) -> None:
+ super().__init__(convert_charrefs=True)
+ self.rows: list[list[str]] = []
+ self.captions: list[str] = []
+ self.summaries: list[str] = []
+ self.paragraphs: list[tuple[str | None, str]] = []
+ self.footnotes: dict[str, list[str]] = {}
+ self._row_stack: list[dict[str, Any]] = []
+ self._caption: list[str] | None = None
+ self._paragraph: tuple[str | None, list[str]] | None = None
+ self._footnote: str | None = None
+
+ def handle_starttag(
+ self, tag: str, attrs: list[tuple[str, str | None]]
+ ) -> None:
+ attributes = dict(attrs)
+ if tag == "table" and attributes.get("summary"):
+ self.summaries.append(_normalize(attributes["summary"] or ""))
+ if tag == "tr":
+ self._row_stack.append({"cells": [], "cell": None})
+ elif tag in {"th", "td"} and self._row_stack:
+ self._row_stack[-1]["cell"] = []
+ elif tag == "caption":
+ self._caption = []
+ elif tag == "p":
+ self._paragraph = (attributes.get("id"), [])
+ elif tag == "a":
+ match = re.fullmatch(r"fn([a-z])", attributes.get("name") or "")
+ if match:
+ self._footnote = match.group(1)
+ self.footnotes[self._footnote] = []
+ elif tag == "br" and self._footnote is not None:
+ self.footnotes[self._footnote].append(" ")
+
+ def handle_data(self, data: str) -> None:
+ if self._row_stack and self._row_stack[-1]["cell"] is not None:
+ self._row_stack[-1]["cell"].append(data)
+ if self._caption is not None:
+ self._caption.append(data)
+ if self._paragraph is not None:
+ self._paragraph[1].append(data)
+ if self._footnote is not None:
+ self.footnotes[self._footnote].append(data)
+
+ def handle_endtag(self, tag: str) -> None:
+ if tag in {"th", "td"} and self._row_stack:
+ context = self._row_stack[-1]
+ if context["cell"] is not None:
+ context["cells"].append(_normalize("".join(context["cell"])))
+ context["cell"] = None
+ self._footnote = None
+ elif tag == "tr" and self._row_stack:
+ self.rows.append(self._row_stack.pop()["cells"])
+ elif tag == "caption" and self._caption is not None:
+ self.captions.append(_normalize("".join(self._caption)))
+ self._caption = None
+ elif tag == "p" and self._paragraph is not None:
+ element_id, parts = self._paragraph
+ self.paragraphs.append((element_id, _normalize("".join(parts))))
+ self._paragraph = None
+
+
+def read_source(key: str) -> str:
+ """Return a committed source body after checking its SHA-256."""
+ spec = SOURCES[key]
+ raw = (EXTERNAL / spec["file"]).read_bytes()
+ digest = hashlib.sha256(raw).hexdigest()
+ if digest != spec["sha256"] or len(raw) != spec["bytes"]:
+ raise ValueError(
+ f"{spec['file']} sha256 {digest} ({len(raw)} bytes) != pinned "
+ f"{spec['sha256']} ({spec['bytes']} bytes); re-verify the "
+ "source before rebuilding"
+ )
+ return raw.decode("utf-8")
+
+
+def _parse(key: str) -> _PageParser:
+ parser = _PageParser()
+ parser.feed(read_source(key))
+ parser.close()
+ return parser
+
+
+def _percent(text: str) -> float | None:
+ """A printed percent cell; ``--`` (no such tax or fund) is None."""
+ if text == "--":
+ return None
+ if not re.fullmatch(r"\d+\.\d+", text):
+ raise ValueError(f"non-numeric rate cell {text!r}")
+ return float(text)
+
+
+# ---------------------------------------------------------------------------
+# Trust fund interest rates
+# ---------------------------------------------------------------------------
+def parse_interest_rates() -> dict[int, dict[str, float | None]]:
+ """Parse the combined table and cross-check the per-fund pages."""
+ combined = _parse("trust_fund_interest_rates")
+ if INTEREST_CAPTION not in combined.captions:
+ raise ValueError(f"caption {INTEREST_CAPTION!r} not found")
+ headers = sum(tuple(row) == INTEREST_HEADERS for row in combined.rows)
+ if headers != 3:
+ raise ValueError(f"expected three {INTEREST_HEADERS!r} panels")
+ rates: dict[int, dict[str, float | None]] = {}
+ for cells in combined.rows:
+ if len(cells) != 3 or not re.fullmatch(r"\d{4}", cells[0]):
+ continue
+ year = int(cells[0])
+ if year in rates:
+ raise ValueError(f"duplicate year {year}")
+ rates[year] = {
+ "average_new_issue": _percent(cells[1]),
+ "effective_oasdi": _percent(cells[2]),
+ }
+ expected = list(range(FIRST_INTEREST_YEAR, LATEST_INTEREST_YEAR + 1))
+ if sorted(rates) != expected:
+ raise ValueError(f"years {sorted(rates)} != {expected}")
+
+ per_fund: dict[int, list[float | None]] = {}
+ recent = _parse("effective_rates_1980_on")
+ if EFFECTIVE_1980_CAPTION not in recent.captions:
+ raise ValueError(f"caption {EFFECTIVE_1980_CAPTION!r} not found")
+ if sum(tuple(row) == EFFECTIVE_HEADERS for row in recent.rows) != 2:
+ raise ValueError("expected two per-fund header rows (1980 on)")
+ for cells in recent.rows:
+ if len(cells) == 4 and re.fullmatch(r"\d{4}", cells[0]):
+ per_fund[int(cells[0])] = [_percent(cell) for cell in cells[1:]]
+ early = _parse("effective_rates_1940_1979")
+ if EFFECTIVE_1940_CAPTION not in early.captions:
+ raise ValueError(f"caption {EFFECTIVE_1940_CAPTION!r} not found")
+ header = (*EFFECTIVE_1940_HEADERS, "", *EFFECTIVE_1940_HEADERS)
+ if sum(tuple(row) == header for row in early.rows) != 1:
+ raise ValueError("expected one two-panel header row (1940-79)")
+ for cells in early.rows:
+ for offset in (0, 5):
+ if len(cells) < offset + 4:
+ continue
+ match = re.fullmatch(r"(\d{4})\.*", cells[offset])
+ if match:
+ year = int(match.group(1))
+ if year in per_fund:
+ raise ValueError(f"duplicate per-fund year {year}")
+ per_fund[year] = [
+ _percent(cell) for cell in cells[offset + 1 : offset + 4]
+ ]
+ if sorted(per_fund) != expected:
+ raise ValueError("per-fund pages do not cover 1940-2025 exactly")
+ for year in expected:
+ oasi, di, oasdi = per_fund[year]
+ if oasdi != rates[year]["effective_oasdi"]:
+ raise ValueError(
+ f"{year}: combined effective "
+ f"{rates[year]['effective_oasdi']} != per-fund OASDI {oasdi}"
+ )
+ rates[year]["effective_oasi"] = oasi
+ rates[year]["effective_di"] = di
+ return rates
+
+
+def build_interest() -> dict[str, Any]:
+ """The trust fund interest-rate artifact (percent, as printed)."""
+ rates = parse_interest_rates()
+ text = _page_text(read_source("trust_fund_interest_rates"))
+ for key, sentence in INTEREST_DEFINITIONS.items():
+ if sentence not in text:
+ raise ValueError(f"definition sentence for {key} not found")
+ return {
+ "schema_version": INTEREST_SCHEMA,
+ "table": INTEREST_CAPTION,
+ "unit": "percent",
+ "vintage_year": VINTAGE_YEAR,
+ "latest_observation_year": LATEST_INTEREST_YEAR,
+ "series": {
+ "effective_oasdi": {
+ "label": "Effective (combined OASI and DI trust funds)",
+ "definition": INTEREST_DEFINITIONS["effective_oasdi"],
+ "source": "trust_fund_interest_rates",
+ },
+ "average_new_issue": {
+ "label": "Average annual special-issue rate on new issues",
+ "definition": INTEREST_DEFINITIONS["average_new_issue"],
+ "source": "trust_fund_interest_rates",
+ },
+ "effective_oasi": {
+ "label": "Effective (OASI trust fund)",
+ "source": (
+ "effective_rates_1940_1979, effective_rates_1980_on"
+ ),
+ },
+ "effective_di": {
+ "label": "Effective (DI trust fund; none before 1957)",
+ "source": (
+ "effective_rates_1940_1979, effective_rates_1980_on"
+ ),
+ },
+ },
+ "vintage_note": (
+ "Fetched 2026-10-01; covers 1940-2025. A post-2014 vintage, so "
+ "not for code bound to the M7 2014 information boundary."
+ ),
+ "validation": {
+ "first_observation_year": FIRST_INTEREST_YEAR,
+ "latest_observation_year": LATEST_INTEREST_YEAR,
+ "n_observations": len(rates),
+ "continuous_calendar_years": True,
+ "combined_effective_equals_per_fund_oasdi": True,
+ },
+ "build": {
+ "built_by": "scripts/extract_lifetime_measure_sources.py",
+ "sources": [
+ "trust_fund_interest_rates",
+ "effective_rates_1980_on",
+ "effective_rates_1940_1979",
+ ],
+ "provenance_file": (
+ "data/external/lifetime_measure_sources.provenance.json"
+ ),
+ },
+ "data": {str(year): rates[year] for year in sorted(rates)},
+ }
+
+
+# ---------------------------------------------------------------------------
+# OASDI tax rates
+# ---------------------------------------------------------------------------
+def _period(label: str) -> tuple[int, int | None, list[str]]:
+ """``"1937-49"`` -> (1937, 1949, []); ``"2019 and later b"`` -> open."""
+ match = re.fullmatch(
+ r"(\d{4})(?:-(\d{2}))?( and later)?((?: ?,? ?[a-d])*)", label
+ )
+ if not match:
+ raise ValueError(f"unparsed period label {label!r}")
+ first = int(match.group(1))
+ last: int | None
+ if match.group(3):
+ last = None
+ elif match.group(2):
+ last = (first // 100) * 100 + int(match.group(2))
+ if last < first:
+ raise ValueError(f"period {label!r} runs backwards")
+ else:
+ last = first
+ notes = re.findall(r"[a-d]", match.group(4) or "")
+ return first, last, notes
+
+
+def parse_tax_rates() -> tuple[list[dict[str, Any]], dict[str, str]]:
+ """Parse the rate rows and footnotes; check labels and coverage."""
+ page = _parse("oasdi_tax_rates")
+ if TAX_SUMMARY not in page.summaries:
+ raise ValueError(f"table summary {TAX_SUMMARY!r} not found")
+ flat = [cell for row in page.rows for cell in row]
+ for label in TAX_GROUP_HEADERS:
+ if label not in flat:
+ raise ValueError(f"header {label!r} not found")
+ if ["OASI", "DI", "Total", "OASI", "DI", "Total"] not in page.rows:
+ raise ValueError("OASI/DI/Total header row not found")
+ rows: list[dict[str, Any]] = []
+ for cells in page.rows:
+ if len(cells) != 7 or not re.match(r"\d{4}", cells[0]):
+ continue
+ first, last, notes = _period(cells[0])
+ values = [_percent(cell) for cell in cells[1:]]
+ each = dict(zip(("oasi", "di", "total"), values[:3], strict=True))
+ self_employed = dict(
+ zip(("oasi", "di", "total"), values[3:], strict=True)
+ )
+ if each["total"] is None:
+ raise ValueError(f"{cells[0]}: no employee/employer total")
+ parts = [part for part in (each["oasi"], each["di"]) if part]
+ if abs(sum(parts) - each["total"]) > 1e-9:
+ raise ValueError(f"{cells[0]}: OASI + DI != total")
+ rows.append(
+ {
+ "period": cells[0],
+ "first_year": first,
+ "last_year": last,
+ "footnotes": notes,
+ "employee_employer_each": each,
+ "self_employed": self_employed,
+ }
+ )
+ expected_first = FIRST_TAX_YEAR
+ for index, row in enumerate(rows):
+ if row["first_year"] != expected_first:
+ raise ValueError(
+ f"tax periods not contiguous at {row['period']!r}"
+ )
+ if row["last_year"] is None:
+ if index != len(rows) - 1:
+ raise ValueError("open-ended period is not the last row")
+ break
+ expected_first = row["last_year"] + 1
+ last_row = rows[-1]
+ if (
+ last_row["first_year"] != OPEN_ENDED_TAX_YEAR
+ or last_row["last_year"] is not None
+ ):
+ raise ValueError("the last row must be '2019 and later'")
+ footnotes = {}
+ for key, parts in sorted(page.footnotes.items()):
+ text = _normalize("".join(parts))
+ # The anchor's own superscript marker opens each footnote.
+ if not text.startswith(f"{key} "):
+ raise ValueError(f"footnote {key} does not open with its marker")
+ footnotes[key] = text[len(key) + 1 :]
+ if sorted(footnotes) != ["a", "b", "c", "d"]:
+ raise ValueError(f"footnotes {sorted(footnotes)} != a-d")
+ return rows, footnotes
+
+
+def _paid_adjustments(
+ rows: list[dict[str, Any]], footnotes: dict[str, str]
+) -> list[dict[str, Any]]:
+ adjustments = []
+ for key, spec in PAID_RATE_QUOTES.items():
+ text = footnotes[key]
+ if spec["quote"] not in text:
+ raise ValueError(f"footnote {key} quote not found")
+ match = re.search(spec["pattern"], text)
+ if not match:
+ raise ValueError(f"footnote {key} rate not parsed")
+ employee = float(match.group(1))
+ for year in spec["years"]:
+ row = next(
+ row
+ for row in rows
+ if row["first_year"] <= year
+ and (row["last_year"] is None or year <= row["last_year"])
+ )
+ if key not in row["footnotes"]:
+ raise ValueError(f"{year}: row lacks footnote {key}")
+ adjustments.append(
+ {
+ "year": year,
+ "footnote": key,
+ "quote": spec["quote"],
+ "employee_effective_rate": employee,
+ "employer_rate": row["employee_employer_each"]["total"],
+ }
+ )
+ return adjustments
+
+
+def build_tax() -> dict[str, Any]:
+ """The OASDI tax rate artifact (percent of taxable earnings)."""
+ rows, footnotes = parse_tax_rates()
+ if TAX_PREAMBLE_SENTENCE not in _page_text(read_source("oasdi_tax_rates")):
+ raise ValueError("the trust-fund preamble sentence is missing")
+ return {
+ "schema_version": TAX_SCHEMA,
+ "table": TAX_SUMMARY,
+ "unit": "percent of taxable earnings",
+ "vintage_year": VINTAGE_YEAR,
+ "rates_reflect": TAX_PREAMBLE_SENTENCE,
+ "rows": rows,
+ "footnotes": footnotes,
+ "paid_rate_adjustments": _paid_adjustments(rows, footnotes),
+ "not_encoded": TAX_NOT_ENCODED,
+ "validation": {
+ "first_year": FIRST_TAX_YEAR,
+ "open_ended_from": OPEN_ENDED_TAX_YEAR,
+ "n_rows": len(rows),
+ "contiguous_periods": True,
+ "oasi_plus_di_equals_total": True,
+ },
+ "build": {
+ "built_by": "scripts/extract_lifetime_measure_sources.py",
+ "sources": ["oasdi_tax_rates"],
+ "provenance_file": (
+ "data/external/lifetime_measure_sources.provenance.json"
+ ),
+ },
+ }
+
+
+# ---------------------------------------------------------------------------
+# MINT8 definitions
+# ---------------------------------------------------------------------------
+def build_mint8() -> dict[str, Any]:
+ """Verbatim MINT8 definitions used by the lifetime-earnings rows."""
+ page = _parse("mint8_user_guide")
+ certified = re.search(
+ r' dict[str, Any]:
+ """One provenance record for the five committed source bodies."""
+ for key in SOURCES:
+ read_source(key)
+ return {
+ "schema_version": PROVENANCE_SCHEMA,
+ "source": "Social Security Administration",
+ "retrieval_date": RETRIEVAL_DATE,
+ "fetch_method": FETCH_METHOD,
+ "acquired_by": ACQUIRED_BY,
+ "sources": {
+ key: {
+ "committed_source_file": f"data/external/{spec['file']}",
+ "source_url": spec["url"],
+ "document": spec["title"],
+ "source_sha256": spec["sha256"],
+ "source_length_bytes": spec["bytes"],
+ }
+ for key, spec in SOURCES.items()
+ },
+ "consumed_by": {
+ "scripts/extract_lifetime_measure_sources.py": [
+ "data/external/ssa_trust_fund_interest_rates_2026.json",
+ "data/external/ssa_oasdi_tax_rates_2026.json",
+ "data/external/mint8_lifetime_quintile_definitions.json",
+ ]
+ },
+ "reproducible": (
+ "parsed deterministically from the committed, sha256-verified "
+ "HTML bodies; no network or wall-clock timestamp is used"
+ ),
+ }
+
+
+def render(document: dict[str, Any]) -> str:
+ """The committed JSON text of an artifact."""
+ return json.dumps(document, indent=2, ensure_ascii=False) + "\n"
+
+
+def build_all() -> dict[Path, dict[str, Any]]:
+ """Every artifact this script writes, keyed by output path."""
+ return {
+ PROVENANCE_OUT: build_provenance(),
+ INTEREST_OUT: build_interest(),
+ TAX_OUT: build_tax(),
+ MINT8_OUT: build_mint8(),
+ }
+
+
+def main() -> None:
+ for path, document in build_all().items():
+ path.write_text(render(document), encoding="utf-8")
+ digest = hashlib.sha256(path.read_bytes()).hexdigest()
+ print(f"wrote {path.relative_to(ROOT)} sha256 {digest}")
+
+
+if __name__ == "__main__":
+ main()
diff --git a/src/populace_dynamics/estimates/lifetime_measures.py b/src/populace_dynamics/estimates/lifetime_measures.py
new file mode 100644
index 00000000..5739e3fa
--- /dev/null
+++ b/src/populace_dynamics/estimates/lifetime_measures.py
@@ -0,0 +1,1966 @@
+"""Lifetime-earnings measures and quintiles for the group breakdowns (G2).
+
+Package G2 of the NASI follow-ups (2026-10-01): the lifetime-earnings
+dimension of the group breakdowns of the four blind tests. Every function
+here is pure over cohort outputs and never opens a PSID file:
+
+* ``careers``: one row per person-year, ``person_id``, ``year``,
+ ``earnings`` (nominal labor earnings) and optionally ``provenance``
+ (the cohort builders' career frame, e.g. ``psid2010._careers``);
+* ``persons``: ``person_id`` and ``birth_year`` (one row per person);
+* ``marriage_episodes``: the already-loaded marriage history in the shape
+ of :func:`populace_dynamics.data.marriage.marriage_episodes`, optionally
+ with ``separation_year`` alongside (as ``psid2010.marital_state_at``
+ reads it);
+* :class:`~populace_dynamics.ss.params.SSAParameters` (NAWI and the
+ contribution and benefit base).
+
+MINT8 definitions (SSA, MINT8 Table User Guide, dateCertified 2026-04-01,
+captured 2026-10-01; verbatim in
+``data/external/mint8_lifetime_quintile_definitions.json``). MINT8's
+annual beneficiary tables have no lifetime-earnings row group; its cohort
+tables (benefit/tax ratios, initial replacement rates) have three, which
+this module implements as the lifetime-earnings dimension:
+
+* "Current-law initial AIME quintile": "Represents an individual's average
+ indexed monthly earnings (AIME) under current law at age 62 ... We
+ calculate the AIME quintiles for each birth cohort." ->
+ :func:`initial_aime_at_62`;
+* "Lifetime payroll tax quintile": "Represents the present value of an
+ individual's current-law payroll taxes at age 62." ->
+ :func:`lifetime_payroll_tax_pv_at_62` with ``shared=False``;
+* "Lifetime payroll tax quintile (shared)": "For married couples, the
+ payroll taxes paid while married are shared equally between them. For
+ never-married individuals, this is the same as the lifetime payroll
+ tax. In any year where an individual is not married, we count only
+ their individual payroll taxes." -> ``shared=True``.
+
+MINT8 says "We use the Social Security Trust Fund interest rate to adjust
+benefits and taxes to their present values at age 62" and "The present
+value of payroll taxes includes all the payroll taxes that the individual
+paid over a lifetime." It does not say which trust fund rate, which tax
+rate, how a year is timed, how ties fall at quintile boundaries, or where
+separated people go. Every such choice below is a **registered builder
+default**: an explicit, named parameter recorded in each result's
+provenance, never a silent fallback.
+
+The report scheme (Butrica and Uccello 2004, exercise 2). The repository
+records only that both of that Report's lifetime-earnings measures
+"average wage-indexed earnings at ages 22-62, and the own measure includes
+uncovered earnings and earnings above the taxable maximum"
+(``docs/design/boomers2004_uniform_cut_comparison.md``, named omissions;
+``estimates/uniform_cut_tabulation.py`` ``NOT_COMPUTED_REPORT_ROWS``).
+:func:`report_average_indexed_earnings_22_62` implements exactly that and
+names every unrecorded convention in :data:`REPORT_EARNINGS_BUILDER_DEFAULTS`.
+That scheme stays unregistered (:func:`boomers2004_scheme`): the Report's
+quintile order ("1st Quintile" lowest or highest) and quintile population
+are not recorded, so it refuses to label.
+
+Stated invariants (property-tested in ``tests/estimates/
+test_lifetime_measures.py``):
+
+1. Quintiles: within a cell with no tied values every quintile holds
+ 20 percent of the weight to within the largest row's weight share;
+ labels are invariant to exact positive weight scaling and to row
+ order; a higher value never gets a lower quintile in the same cell;
+ missing values get :data:`NOT_COMPUTED` and nothing else does.
+2. Payroll tax present value: zero earnings give zero; doubling earnings
+ that stay below the contribution and benefit base doubles it;
+ earnings above the base do not change it; it is linear in the tax rate.
+3. Sharing: for a couple married to each other in every year either
+ spouse paid tax, with the same birth year, the two shared present
+ values sum to the two own present values; year by year the shared taxes
+ of the two sum exactly to their own taxes.
+4. AIME: equals the repository oracle (``ss.statutory_aime.oracle_aime``)
+ on the same truncated history, and under the exercise-4 convention it
+ equals Track M's ``rules.history_pia`` AIME for an entitlement at 62.
+"""
+
+from __future__ import annotations
+
+import hashlib
+import json
+import math
+from collections.abc import Callable, Collection, Mapping
+from dataclasses import dataclass, field
+from enum import Enum
+from fractions import Fraction
+from pathlib import Path
+from typing import Any, Protocol
+
+import numpy as np
+import pandas as pd
+
+from populace_dynamics.ss import statutory_aime
+from populace_dynamics.ss.params import SSAParameters
+from populace_dynamics.ss.statutory_aime import ComputationYears
+
+__all__ = [
+ "AIME_CONVENTIONS",
+ "COMPUTED",
+ "MINT8_DEFINITIONS_PATH",
+ "MINT8_DIMENSION_MEASURES",
+ "MINT8_DEFINITIONS_SHA256",
+ "NOT_COMPUTED",
+ "OASDI_TAX_RATES_PATH",
+ "OASDI_TAX_RATES_SHA256",
+ "PAYROLL_TAX_BUILDER_DEFAULTS",
+ "QUINTILE_LABELS",
+ "REPORT_EARNINGS_BUILDER_DEFAULTS",
+ "REPORT_EARNINGS_RECORDED",
+ "SCHEMA_VERSION",
+ "SCHEMES",
+ "TRUST_FUND_INTEREST_RATES_PATH",
+ "TRUST_FUND_INTEREST_RATES_SHA256",
+ "AimeBasis",
+ "AimeConvention",
+ "AverageDivisor",
+ "InterestSeries",
+ "LifetimeEarningsScheme",
+ "MeasureResult",
+ "MissingRatePolicy",
+ "MissingSpousePolicy",
+ "OASDITaxRates",
+ "QuintileDimension",
+ "QuintileScope",
+ "ReportEarningsConventions",
+ "TaxRateBasis",
+ "TrustFundInterestRates",
+ "accumulation_factor",
+ "annual_payroll_taxes",
+ "boomers2004_scheme",
+ "initial_aime_at_62",
+ "lifetime_payroll_tax_pv_at_62",
+ "load_mint8_definitions",
+ "load_mint8_scheme",
+ "load_oasdi_tax_rates",
+ "load_trust_fund_interest_rates",
+ "quintile_cells",
+ "quintile_summary",
+ "report_average_indexed_earnings_22_62",
+ "ten_year_birth_cohort",
+ "weighted_quintiles",
+]
+
+SCHEMA_VERSION = "lifetime_measures.v1"
+
+_PROJECT_ROOT = Path(__file__).resolve().parents[3]
+_EXTERNAL = _PROJECT_ROOT / "data" / "external"
+#: SSA OACT "Social Security Tax Rates" (oasdiRates.html), captured
+#: 2026-10-01 and extracted by scripts/extract_lifetime_measure_sources.py.
+OASDI_TAX_RATES_PATH = _EXTERNAL / "ssa_oasdi_tax_rates_2026.json"
+OASDI_TAX_RATES_SHA256 = (
+ "63267c7e7596439224d0bb088290e5c0be92a1187091b448fdf4074aabcc1615"
+)
+#: SSA OACT "Average and Effective Interest Rates" (annualinterestrates
+#: .html), 1940-2025, cross-checked against the per-fund pages.
+TRUST_FUND_INTEREST_RATES_PATH = (
+ _EXTERNAL / "ssa_trust_fund_interest_rates_2026.json"
+)
+TRUST_FUND_INTEREST_RATES_SHA256 = (
+ "a55c4722caa80caf26a0442492492cebfc52b412535fff584617a35188812766"
+)
+#: Verbatim MINT8 Table User Guide definitions and the cohort-table labels.
+MINT8_DEFINITIONS_PATH = _EXTERNAL / "mint8_lifetime_quintile_definitions.json"
+MINT8_DEFINITIONS_SHA256 = (
+ "0df7298e0655eecdd4aca5b04b41ec9c520228baa7329080e7208950e7769608"
+)
+
+#: MINT8's quintile labels in its table order (always these five).
+QUINTILE_LABELS: tuple[str, ...] = (
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest",
+)
+#: The label of a person with no value (status ``not computed``).
+NOT_COMPUTED = "not computed"
+#: The status of a person whose measure was computed.
+COMPUTED = "computed"
+
+_N_QUINTILES = 5
+_RETIREMENT_AGE = 62
+#: The age at which the career frames' coverage starts at the earliest
+#: (estimates/career.py: max(1968, birth year + 22)); a later first row
+#: flags a left-censored history.
+_FIRST_AGE = 22
+_AIME_INDEXING_AGE = 60
+_FIRST_COMPUTATION_BASE_YEAR = 1951
+_EN_DASH = "–"
+_EPISODE_COLUMNS = (
+ "person_id",
+ "marriage_order",
+ "start_year",
+ "episode_end_year",
+ "how_ended",
+ "spouse_person_id",
+)
+
+
+# ---------------------------------------------------------------------------
+# Conventions
+# ---------------------------------------------------------------------------
+class TaxRateBasis(str, Enum):
+ """Which OASDI rate a year's payroll tax uses.
+
+ ``TRUST_FUND_RECEIVED`` (registered builder default): twice the
+ "Tax rate for employees and employers, each" total of SSA's table,
+ which "reflect[s] the amounts received by the trust funds".
+ ``EMPLOYEE_EMPLOYER_PAID``: the same, except the employee share in
+ 1984 (5.4 percent, footnote a) and 2011-2012 (4.2 percent, footnote c),
+ when general revenue covered the difference.
+ """
+
+ TRUST_FUND_RECEIVED = "trust_fund_received"
+ EMPLOYEE_EMPLOYER_PAID = "employee_employer_paid"
+
+
+class InterestSeries(str, Enum):
+ """Which SSA trust fund interest rate discounts and accumulates.
+
+ ``EFFECTIVE_OASDI`` (registered builder default): "the interest earned
+ in that year divided by the average level of assets held during the
+ year" for the combined OASI and DI trust funds.
+ ``AVERAGE_NEW_ISSUE``: "the average of the 12 monthly interest rates on
+ new issues during the year".
+ """
+
+ EFFECTIVE_OASDI = "effective_oasdi"
+ AVERAGE_NEW_ISSUE = "average_new_issue"
+
+
+class MissingSpousePolicy(str, Enum):
+ """A married year whose spouse has no career in the careers frame.
+
+ ``OWN_ONLY`` (registered builder default): the year counts the
+ person's own amount, and the year is counted in the result.
+ ``NOT_COMPUTED``: the person's shared measure is not computed.
+ """
+
+ OWN_ONLY = "own_only"
+ NOT_COMPUTED = "not_computed"
+
+
+class MissingRatePolicy(str, Enum):
+ """A year whose interest rate the series does not cover.
+
+ ``REFUSE`` (default): raise, naming the years. ``NOT_COMPUTED``: mark
+ the affected persons not computed. Coverage can be extended only by
+ :meth:`TrustFundInterestRates.extended`, with a named source.
+ """
+
+ REFUSE = "refuse"
+ NOT_COMPUTED = "not_computed"
+
+
+class AimeBasis(str, Enum):
+ """``AGE_62``: the person attains 62 by the analysis year.
+ ``PROVISIONAL_THROUGH_LAST_OBSERVED``: younger than 62 in the analysis
+ year; the AIME uses the history through the analysis year (flagged).
+ """
+
+ AGE_62 = "age_62"
+ PROVISIONAL_THROUGH_LAST_OBSERVED = "provisional_through_last_observed"
+
+
+class AverageDivisor(str, Enum):
+ """The divisor of the report's average indexed earnings.
+
+ ``COVERED_AGES`` (registered builder default): the ages 22-62 at which
+ the careers frame has a row. ``ALL_AGES``: all 41 ages (a missing
+ age counts as zero earnings).
+ """
+
+ COVERED_AGES = "covered_ages"
+ ALL_AGES = "all_ages"
+
+
+class QuintileScope(str, Enum):
+ """The population over which a quintile is cut.
+
+ MINT8: "We calculate the AIME quintiles for each birth cohort", and it
+ uses "10-year birth cohorts" (1960-1969, ...). For an annual
+ population both cuts are exposed.
+ """
+
+ TEN_YEAR_BIRTH_COHORT = "ten_year_birth_cohort"
+ WHOLE_POPULATION = "whole_population"
+
+
+@dataclass(frozen=True)
+class AimeConvention:
+ """How an initial AIME is computed: the oracle's computation-year
+ convention and the last age whose earnings enter.
+
+ ``last_earnings_age`` ``None`` passes the history as supplied (up to
+ the analysis year), as Track A's ``_Calculator._level`` does.
+ """
+
+ name: str
+ computation_years: ComputationYears
+ last_earnings_age: int | None
+ source: str
+
+
+#: Named AIME conventions, one per way the repository computes an
+#: old-age AIME. Exercise 2 (Track U) computes no AIME: it reads
+#: reported Social Security amounts.
+AIME_CONVENTIONS: dict[str, AimeConvention] = {
+ "mint8_initial_aime": AimeConvention(
+ name="mint8_initial_aime",
+ computation_years=ComputationYears.STATUTORY,
+ last_earnings_age=61,
+ source=(
+ "MINT8: 'AIME under current law at age 62'. Statutory 415(b)(2) "
+ "computation years; earnings through the year before the year "
+ "of attaining 62 (registered builder default: the history an "
+ "entitlement beginning at 62 uses under Track M's window-year-"
+ "less-one cut, min_benefit_track_m/rules.py history_pia)."
+ ),
+ ),
+ "exercise_1_cola": AimeConvention(
+ name="exercise_1_cola",
+ computation_years=ComputationYears.LEGACY_FIXED_35,
+ last_earnings_age=None,
+ source=(
+ "cola_track_a/benefits.py TRACK_A_COMPUTATION_YEARS "
+ "(LEGACY_FIXED_35, Registration 13); _Calculator._level passes "
+ "the whole opening-year career to eligibility_pia_for_clock."
+ ),
+ ),
+ "exercise_3_fra68": AimeConvention(
+ name="exercise_3_fra68",
+ computation_years=ComputationYears.LEGACY_FIXED_35,
+ last_earnings_age=None,
+ source=(
+ "fra68_track/config.py MAX_RULINGS['benefit_computation_years'] "
+ "(LEGACY_FIXED_35, d188 item (a): run exercise 3 exactly like "
+ "Track A); benefits through the Track A calculator."
+ ),
+ ),
+ "exercise_4_min_benefit": AimeConvention(
+ name="exercise_4_min_benefit",
+ computation_years=ComputationYears.STATUTORY,
+ last_earnings_age=61,
+ source=(
+ "min_benefit_track_m/rules.py history_pia (old_age): "
+ "ss.statutory_aime.aime over the history through the window "
+ "year less one (record_years: computation base years end "
+ "before the year of first entitlement); for an entitlement in "
+ "the year of attaining 62 that is the year of attaining 61."
+ ),
+ ),
+}
+
+
+@dataclass(frozen=True)
+class ReportEarningsConventions:
+ """Conventions of :func:`report_average_indexed_earnings_22_62`.
+
+ ``first_age``, ``last_age`` and, for the own measure,
+ ``cap_at_taxable_maximum=False`` are recorded
+ (:data:`REPORT_EARNINGS_RECORDED`); the rest are registered builder
+ defaults (:data:`REPORT_EARNINGS_BUILDER_DEFAULTS`).
+ """
+
+ first_age: int = 22
+ last_age: int = 62
+ cap_at_taxable_maximum: bool = False
+ index_age: int = 60
+ divisor: AverageDivisor = AverageDivisor.COVERED_AGES
+ separated_is_married: bool = True
+ missing_spouse: MissingSpousePolicy = MissingSpousePolicy.OWN_ONLY
+
+
+_REPORT_RECORD = (
+ "docs/design/boomers2004_uniform_cut_comparison.md (named omissions); "
+ "estimates/uniform_cut_tabulation.py NOT_COMPUTED_REPORT_ROWS"
+)
+#: What the repository records about the Report's measure, with locators.
+REPORT_EARNINGS_RECORDED: dict[str, str] = {
+ "measure": (
+ "both measures 'average wage-indexed earnings at ages 22-62' "
+ f"({_REPORT_RECORD})"
+ ),
+ "own_includes_above_taxable_maximum": (
+ "'the own measure includes uncovered earnings and earnings above "
+ f"the taxable maximum' ({_REPORT_RECORD}); so the own measure is "
+ "not capped. PSID labor earnings do not separate covered from "
+ "uncovered work, so 'includes uncovered earnings' holds trivially."
+ ),
+ "rows": (
+ "Lifetime Earnings (Own) and (Shared), five rows each, '1st "
+ "Quintile' to '5th Quintile' (uniform_cut_tabulation.py)"
+ ),
+}
+#: Every convention the record leaves open, with this builder's default.
+REPORT_EARNINGS_BUILDER_DEFAULTS: dict[str, str] = {
+ "age_in_year": "age = calendar year - birth year (the oracle's rule)",
+ "age_range_inclusive": "ages 22 through 62 inclusive (41 ages)",
+ "wage_index": (
+ "the SSA national average wage index (SSAParameters.nawi), each "
+ "year's earnings times NAWI(index year) / NAWI(year), years at or "
+ "after the index year nominal (ss.benefits.indexed_history's rule)"
+ ),
+ "index_age": "60, the AIME indexing age",
+ "divisor": (
+ "the ages 22-62 with a career row (COVERED_AGES); the PSID does "
+ "not observe every age, so missing ages are not read as zeros"
+ ),
+ "shared_rule": (
+ "in a year the person is married at year end, the mean of the two "
+ "spouses' indexed earnings (both indexed to the person's index "
+ "year); otherwise own earnings (MINT8's sharing rule, since the "
+ "Report's is not recorded)"
+ ),
+ "shared_cap": "uncapped, as recorded for the own measure",
+ "marital_state": (
+ "psid2010.marital_state_at at the end of each year; separated "
+ "counts as married (the repository's default)"
+ ),
+ "missing_spouse": "own earnings for that year, counted (OWN_ONLY)",
+ "quintile_order_and_population": (
+ "not registered: whether '1st Quintile' is the lowest and over "
+ "which population the quintiles are cut are unrecorded, so the "
+ "scheme refuses to label (boomers2004_scheme)"
+ ),
+}
+
+#: Registered builder defaults of the MINT8 payroll-tax measures.
+PAYROLL_TAX_BUILDER_DEFAULTS: dict[str, str] = {
+ "tax_rate": (
+ "combined employee and employer OASDI rate by year (twice SSA's "
+ "'each' total, the rate the trust funds received; TaxRateBasis), "
+ "applied to earnings capped at the contribution and benefit base "
+ "(SSAParameters.wage_base_for)"
+ ),
+ "self_employment": (
+ "not separated: every dollar of career earnings is taxed at the "
+ "employee-plus-employer rate"
+ ),
+ "interest_rate": (
+ "SSA's effective annual interest rate of the combined OASI and DI "
+ "trust funds (InterestSeries.EFFECTIVE_OASDI)"
+ ),
+ "timing": (
+ "a year's tax is credited at the end of that year and the present "
+ "value is taken at the end of the year of attaining 62: tax in "
+ "year t is multiplied by the product of (1 + i_s) for s = t+1..Y, "
+ "or divided by the product for s = Y+1..t when t > Y"
+ ),
+ "lifetime": (
+ "every career year in the frame, before and after 62 ('all the "
+ "payroll taxes that the individual paid over a lifetime')"
+ ),
+ "marital_state": (
+ "psid2010.marital_state_at at the end of each tax year; separated "
+ "counts as married (the repository's default)"
+ ),
+ "shared_years": (
+ "every year in which either spouse has a career row while the "
+ "person is married to that spouse; a spouse's missing year counts "
+ "as zero tax (counted)"
+ ),
+ "missing_spouse": (
+ "a married year whose spouse is not a joinable person in the "
+ "careers frame counts the person's own tax (OWN_ONLY, counted)"
+ ),
+ "unknown_marital_state": (
+ "a year whose marital state is 'unknown' counts own tax (counted)"
+ ),
+}
+
+
+# ---------------------------------------------------------------------------
+# Rate schedules
+# ---------------------------------------------------------------------------
+class _CombinedRateSource(Protocol):
+ def combined_for(self, year: int) -> float: ...
+
+
+class _InterestRateSource(Protocol):
+ def rate_for(self, year: int) -> float: ...
+
+
+def _sha256(raw: bytes) -> str:
+ return hashlib.sha256(raw).hexdigest()
+
+
+def _read_pinned_json(
+ path: Path, expected_sha256: str, schema: str
+) -> tuple[dict[str, Any], str]:
+ raw = Path(path).read_bytes()
+ digest = _sha256(raw)
+ if digest != expected_sha256:
+ raise ValueError(
+ f"{path} sha256 {digest} != pinned {expected_sha256}; refusing "
+ "a changed committed extraction."
+ )
+ document = json.loads(raw)
+ if document.get("schema_version") != schema:
+ raise ValueError(f"{path} schema_version is not {schema!r}.")
+ return document, digest
+
+
+def _year(value: object, label: str) -> int:
+ if isinstance(value, bool):
+ raise TypeError(f"{label} must be an integer year, not {value!r}.")
+ if isinstance(value, (int, np.integer)):
+ return int(value)
+ raise TypeError(f"{label} must be an integer year, not {value!r}.")
+
+
+@dataclass(frozen=True)
+class OASDITaxRates:
+ """Combined employee-plus-employer OASDI tax rates by calendar year.
+
+ ``combined_percent_by_year`` holds every year of the table's closed
+ periods; years from ``open_ended_from`` take the "and later" row. A
+ year before the first period (1937) raises.
+ """
+
+ combined_percent_by_year: Mapping[int, float]
+ open_ended_from: int
+ open_ended_combined_percent: float
+ basis: TaxRateBasis
+ provenance: Mapping[str, Any] = field(default_factory=dict)
+
+ def combined_percent_for(self, year: int) -> float:
+ """The combined rate in percent of taxable earnings."""
+ year = _year(year, "year")
+ if year >= self.open_ended_from:
+ return float(self.open_ended_combined_percent)
+ if year not in self.combined_percent_by_year:
+ raise KeyError(f"No OASDI tax rate for {year}.")
+ return float(self.combined_percent_by_year[year])
+
+ def combined_for(self, year: int) -> float:
+ """The combined rate as a fraction of taxable earnings.
+
+ The same interface as ``estimates.parameters.PayrollRateLegs``.
+ """
+ return self.combined_percent_for(year) / 100.0
+
+ @classmethod
+ def from_document(
+ cls,
+ document: Mapping[str, Any],
+ *,
+ basis: TaxRateBasis = TaxRateBasis.TRUST_FUND_RECEIVED,
+ provenance: Mapping[str, Any] | None = None,
+ ) -> OASDITaxRates:
+ """Build the schedule from an ``ssa_oasdi_tax_rates.v1`` document."""
+ basis = TaxRateBasis(basis)
+ by_year: dict[int, float] = {}
+ open_from: int | None = None
+ open_rate: float | None = None
+ for row in document["rows"]:
+ combined = 2.0 * float(row["employee_employer_each"]["total"])
+ first = int(row["first_year"])
+ if row["last_year"] is None:
+ open_from, open_rate = first, combined
+ continue
+ for year in range(first, int(row["last_year"]) + 1):
+ if year in by_year:
+ raise ValueError(f"OASDI tax year {year} listed twice.")
+ by_year[year] = combined
+ if open_from is None or open_rate is None:
+ raise ValueError("The OASDI schedule has no open-ended row.")
+ if basis is TaxRateBasis.EMPLOYEE_EMPLOYER_PAID:
+ for adjustment in document["paid_rate_adjustments"]:
+ year = int(adjustment["year"])
+ if year not in by_year:
+ raise ValueError(f"Paid-rate year {year} not tabled.")
+ by_year[year] = float(
+ adjustment["employee_effective_rate"]
+ ) + float(adjustment["employer_rate"])
+ return cls(
+ combined_percent_by_year=by_year,
+ open_ended_from=open_from,
+ open_ended_combined_percent=open_rate,
+ basis=basis,
+ provenance=dict(provenance or {}),
+ )
+
+
+def load_oasdi_tax_rates(
+ path: Path = OASDI_TAX_RATES_PATH,
+ *,
+ basis: TaxRateBasis = TaxRateBasis.TRUST_FUND_RECEIVED,
+ expected_sha256: str = OASDI_TAX_RATES_SHA256,
+) -> OASDITaxRates:
+ """Load SSA's OASDI tax-rate schedule (pinned committed extraction)."""
+ document, digest = _read_pinned_json(
+ path, expected_sha256, "ssa_oasdi_tax_rates.v1"
+ )
+ basis = TaxRateBasis(basis)
+ return OASDITaxRates.from_document(
+ document,
+ basis=basis,
+ provenance={
+ "file": _relative(path),
+ "sha256": digest,
+ "table": document["table"],
+ "rates_reflect": document["rates_reflect"],
+ "basis": basis.value,
+ },
+ )
+
+
+@dataclass(frozen=True)
+class TrustFundInterestRates:
+ """Annual trust fund interest rates (percent) by calendar year.
+
+ ``assumed_years`` are years added by :meth:`extended` under
+ ``assumption_source``; every other year is SSA's published rate.
+ """
+
+ percent_by_year: Mapping[int, float]
+ series: str
+ assumed_years: frozenset[int] = frozenset()
+ assumption_source: str | None = None
+ provenance: Mapping[str, Any] = field(default_factory=dict)
+
+ def covers(self, year: int) -> bool:
+ return _year(year, "year") in self.percent_by_year
+
+ def rate_for(self, year: int) -> float:
+ """The rate as a fraction; raises outside the series' coverage."""
+ year = _year(year, "year")
+ if year not in self.percent_by_year:
+ raise KeyError(
+ f"No {self.series} trust fund interest rate for {year}."
+ )
+ return float(self.percent_by_year[year]) / 100.0
+
+ def extended(
+ self, assumed_percent: Mapping[int, float], *, source: str
+ ) -> TrustFundInterestRates:
+ """A copy with assumed rates for uncovered years, named by source.
+
+ Refuses to overwrite a covered year, an empty source, or a rate
+ that is not finite and above -100 percent.
+ """
+ if not isinstance(source, str) or not source.strip():
+ raise ValueError("Assumed interest rates need a named source.")
+ overlap = sorted(
+ int(year) for year in assumed_percent if self.covers(int(year))
+ )
+ if overlap:
+ raise ValueError(f"Years {overlap} already have rates.")
+ added: dict[int, float] = {}
+ for year, value in assumed_percent.items():
+ rate = float(value)
+ if not math.isfinite(rate) or rate <= -100.0:
+ raise ValueError(f"Assumed rate {value!r} for {year}.")
+ added[_year(year, "year")] = rate
+ sources = [self.assumption_source, source.strip()]
+ return TrustFundInterestRates(
+ percent_by_year={**self.percent_by_year, **added},
+ series=self.series,
+ assumed_years=self.assumed_years | frozenset(added),
+ assumption_source="; ".join(s for s in sources if s),
+ provenance=dict(self.provenance),
+ )
+
+
+def load_trust_fund_interest_rates(
+ path: Path = TRUST_FUND_INTEREST_RATES_PATH,
+ *,
+ series: InterestSeries = InterestSeries.EFFECTIVE_OASDI,
+ expected_sha256: str = TRUST_FUND_INTEREST_RATES_SHA256,
+) -> TrustFundInterestRates:
+ """Load SSA's annual trust fund interest rates, 1940-2025."""
+ document, digest = _read_pinned_json(
+ path, expected_sha256, "ssa_trust_fund_interest_rates.v1"
+ )
+ series = InterestSeries(series)
+ by_year = {}
+ for year, row in document["data"].items():
+ value = row[series.value]
+ if value is None:
+ raise ValueError(f"{series.value} has no value for {year}.")
+ by_year[int(year)] = float(value)
+ return TrustFundInterestRates(
+ percent_by_year=by_year,
+ series=series.value,
+ provenance={
+ "file": _relative(path),
+ "sha256": digest,
+ "table": document["table"],
+ "series": series.value,
+ "definition": document["series"][series.value].get(
+ "definition", ""
+ ),
+ "latest_observation_year": document["latest_observation_year"],
+ },
+ )
+
+
+def _relative(path: Path) -> str:
+ try:
+ return str(Path(path).resolve().relative_to(_PROJECT_ROOT))
+ except ValueError:
+ return str(path)
+
+
+def accumulation_factor(
+ tax_year: int, reference_year: int, interest: _InterestRateSource
+) -> float:
+ """Value at the end of ``reference_year`` of 1 paid at the end of
+ ``tax_year``: the product of (1 + i_s) over s = tax_year+1 ..
+ reference_year, or its reciprocal over s = reference_year+1 .. tax_year
+ when the tax year is later."""
+ tax_year = _year(tax_year, "tax_year")
+ reference_year = _year(reference_year, "reference_year")
+ if tax_year == reference_year:
+ return 1.0
+ if tax_year < reference_year:
+ return math.prod(
+ 1.0 + interest.rate_for(year)
+ for year in range(tax_year + 1, reference_year + 1)
+ )
+ return 1.0 / math.prod(
+ 1.0 + interest.rate_for(year)
+ for year in range(reference_year + 1, tax_year + 1)
+ )
+
+
+# ---------------------------------------------------------------------------
+# Inputs
+# ---------------------------------------------------------------------------
+@dataclass(frozen=True)
+class MeasureResult:
+ """One measure for every person, with JSON-ready provenance."""
+
+ measure: str
+ frame: pd.DataFrame
+ provenance: dict[str, Any]
+
+
+def _integral(series: pd.Series, label: str) -> np.ndarray:
+ if series.isna().any():
+ raise ValueError(f"{label} has missing values.")
+ values = pd.to_numeric(series, errors="raise")
+ as_float = values.astype("float64").to_numpy()
+ if not np.all(np.isfinite(as_float)) or not np.all(
+ as_float == np.floor(as_float)
+ ):
+ raise ValueError(f"{label} must hold integers.")
+ return values.astype("int64").to_numpy()
+
+
+def _careers_frame(careers: pd.DataFrame) -> pd.DataFrame:
+ missing = {"person_id", "year", "earnings"} - set(careers.columns)
+ if missing:
+ raise ValueError(f"careers lacks columns {sorted(missing)}.")
+ frame = pd.DataFrame(
+ {
+ "person_id": _integral(careers["person_id"], "person_id"),
+ "year": _integral(careers["year"], "year"),
+ "earnings": pd.to_numeric(careers["earnings"], errors="raise")
+ .astype("float64")
+ .to_numpy(),
+ }
+ )
+ earnings = frame["earnings"].to_numpy()
+ if not np.all(np.isfinite(earnings)) or (earnings < 0).any():
+ raise ValueError("careers earnings must be finite and non-negative.")
+ if "provenance" in careers.columns:
+ frame["provenance"] = pd.array(
+ careers["provenance"].astype("string").to_numpy(),
+ dtype="string",
+ )
+ if frame.duplicated(["person_id", "year"]).any():
+ raise ValueError("careers has duplicate (person_id, year) rows.")
+ return frame
+
+
+def _persons_frame(persons: pd.DataFrame) -> pd.DataFrame:
+ missing = {"person_id", "birth_year"} - set(persons.columns)
+ if missing:
+ raise ValueError(f"persons lacks columns {sorted(missing)}.")
+ frame = pd.DataFrame(
+ {
+ "person_id": _integral(persons["person_id"], "person_id"),
+ "birth_year": _integral(persons["birth_year"], "birth_year"),
+ }
+ )
+ if frame["person_id"].duplicated().any():
+ raise ValueError("persons has duplicate person_id values.")
+ return frame
+
+
+def _histories(frame: pd.DataFrame) -> dict[int, dict[int, float]]:
+ histories: dict[int, dict[int, float]] = {}
+ for pid, year, earnings in zip(
+ frame["person_id"].to_numpy(),
+ frame["year"].to_numpy(),
+ frame["earnings"].to_numpy(),
+ strict=True,
+ ):
+ histories.setdefault(int(pid), {})[int(year)] = float(earnings)
+ return histories
+
+
+def _imputed_years(frame: pd.DataFrame) -> dict[int, set[int]] | None:
+ """Years whose provenance is not ``observed``, by person."""
+ if "provenance" not in frame.columns:
+ return None
+ flagged = frame[frame["provenance"].fillna("") != "observed"]
+ out: dict[int, set[int]] = {}
+ for pid, year in zip(flagged["person_id"], flagged["year"], strict=True):
+ out.setdefault(int(pid), set()).add(int(year))
+ return out
+
+
+def _frame_sha256(frame: pd.DataFrame) -> str:
+ return _sha256(
+ frame.to_csv(index=False, lineterminator="\n").encode("utf-8")
+ )
+
+
+def _input_digests(
+ careers: pd.DataFrame,
+ persons: pd.DataFrame,
+ episodes: pd.DataFrame | None = None,
+) -> dict[str, str]:
+ digests = {
+ "careers_sha256": _frame_sha256(
+ careers.sort_values(["person_id", "year"]).reset_index(drop=True)
+ ),
+ "persons_sha256": _frame_sha256(
+ persons.sort_values("person_id").reset_index(drop=True)
+ ),
+ }
+ if episodes is not None:
+ digests["marriage_episodes_sha256"] = _frame_sha256(
+ episodes.sort_values(["person_id", "marriage_order"])
+ .astype("string")
+ .reset_index(drop=True)
+ )
+ return digests
+
+
+def _counts(series: pd.Series) -> dict[str, int]:
+ values = series.astype("object").where(series.notna(), "none")
+ return {
+ str(key): int(value)
+ for key, value in sorted(values.value_counts().items())
+ }
+
+
+def _provenance(
+ measure: str,
+ frame: pd.DataFrame,
+ params: SSAParameters,
+ inputs: dict[str, str],
+ conventions: dict[str, Any],
+ extra: dict[str, Any] | None = None,
+) -> dict[str, Any]:
+ flag_columns = [
+ column
+ for column in frame.columns
+ if pd.api.types.is_bool_dtype(frame[column].dtype)
+ or column.startswith(("years_", "married_years_", "n_"))
+ ]
+ totals = {}
+ for column in flag_columns:
+ series = frame[column]
+ if pd.api.types.is_bool_dtype(series.dtype):
+ totals[column] = int(series.astype("boolean").fillna(False).sum())
+ else:
+ values = pd.to_numeric(series, errors="coerce")
+ totals[column] = int(values.fillna(0).sum())
+ record = {
+ "schema_version": SCHEMA_VERSION,
+ "measure": measure,
+ "n_persons": len(frame),
+ "status_counts": _counts(frame["status"]),
+ "reason_counts": _counts(frame["reason"]),
+ "flag_totals": totals,
+ "conventions": conventions,
+ "ssa_parameters": {"pe_us_revision": params.pe_us_revision},
+ "inputs": inputs,
+ "output_sha256": _frame_sha256(frame),
+ "pure_over_cohort_outputs": (
+ "computed from careers, persons and marriage-episode frames "
+ "only; no PSID file is opened"
+ ),
+ }
+ if extra:
+ record.update(extra)
+ return record
+
+
+def _nawi_missing(
+ years: Collection[int], indexing_year: int, params: SSAParameters
+) -> list[int]:
+ needed = {indexing_year} | {year for year in years if year < indexing_year}
+ return sorted(year for year in needed if year not in params.nawi)
+
+
+# ---------------------------------------------------------------------------
+# 1. Initial AIME at 62
+# ---------------------------------------------------------------------------
+_AIME_COLUMNS = (
+ "person_id",
+ "birth_year",
+ "aime",
+ "status",
+ "reason",
+ "basis",
+ "reaches_62_by_analysis_year",
+ "history_cutoff_year",
+ "first_earnings_year",
+ "last_earnings_year",
+ "history_starts_after_age_22",
+ "history_ends_before_cutoff",
+ "n_history_years",
+ "n_imputed_history_years",
+ "years_before_1951_dropped",
+ "computation_years",
+ "indexing_year",
+)
+
+
+def initial_aime_at_62(
+ careers: pd.DataFrame,
+ persons: pd.DataFrame,
+ params: SSAParameters,
+ *,
+ analysis_year: int,
+ convention: AimeConvention,
+) -> MeasureResult:
+ """Current-law AIME at age 62 for every person (MINT8's initial AIME).
+
+ The oracle's AIME (``ss.statutory_aime.oracle_aime``, unchanged) over
+ the person's career cut at ``birth_year + convention.last_earnings_age``
+ and never after ``analysis_year``. A person who attains 62 by the
+ analysis year has basis ``age_62``; a younger person's AIME uses the
+ history through the analysis year (the last observed year; closed-
+ cohort tests stop earnings in 2010) and has basis
+ ``provisional_through_last_observed``. The convention is required:
+ :data:`AIME_CONVENTIONS` names how each blind test computes an AIME.
+
+ Not computed, with a reason: no career rows; career rows only after
+ the cutoff (earnings before 62 unobserved, not zero); a statutory AIME
+ for a person attaining 62 before 1975 (the oracle refuses it); a NAWI
+ value the indexing needs. Years before 1951 cannot be computation
+ base years and are dropped, counted. Years inside the cutoff that the
+ careers frame lacks count as zero, as in the oracle; a first row after
+ age 22 (the PSID's first income year is 1967) is flagged
+ ``history_starts_after_age_22``.
+ """
+
+ analysis_year = _year(analysis_year, "analysis_year")
+ if not isinstance(convention, AimeConvention):
+ raise TypeError("convention must be an AimeConvention.")
+ convention_years = ComputationYears(convention.computation_years)
+ careers_frame = _careers_frame(careers)
+ persons_frame = _persons_frame(persons)
+ histories = _histories(careers_frame)
+ imputed = _imputed_years(careers_frame)
+ rows = []
+ for pid, birth in zip(
+ persons_frame["person_id"], persons_frame["birth_year"], strict=True
+ ):
+ pid, birth = int(pid), int(birth)
+ reaches = birth + _RETIREMENT_AGE <= analysis_year
+ cutoff = analysis_year
+ if convention.last_earnings_age is not None:
+ cutoff = min(cutoff, birth + int(convention.last_earnings_age))
+ history_all = histories.get(pid)
+ kept = {
+ year: value
+ for year, value in (history_all or {}).items()
+ if year <= cutoff
+ }
+ early = [year for year in kept if year < _FIRST_COMPUTATION_BASE_YEAR]
+ for year in early:
+ del kept[year]
+ first_year = min(kept) if kept else None
+ last_year = max(kept) if kept else None
+ # The year of attaining 60: benefits.indexed_history's year and
+ # statutory_aime.indexing_year's with no death or disability date.
+ indexing = birth + _AIME_INDEXING_AGE
+ n_years = (
+ statutory_aime.LEGACY_FIXED_COMPUTATION_YEARS
+ if convention_years is ComputationYears.LEGACY_FIXED_35
+ else None
+ )
+ row = {
+ "person_id": pid,
+ "birth_year": birth,
+ "aime": math.nan,
+ "status": NOT_COMPUTED,
+ "reason": None,
+ "basis": (
+ AimeBasis.AGE_62
+ if reaches
+ else AimeBasis.PROVISIONAL_THROUGH_LAST_OBSERVED
+ ).value,
+ "reaches_62_by_analysis_year": reaches,
+ "history_cutoff_year": cutoff,
+ "first_earnings_year": first_year,
+ "last_earnings_year": last_year,
+ "history_starts_after_age_22": (
+ first_year is not None and first_year > birth + _FIRST_AGE
+ ),
+ "history_ends_before_cutoff": (
+ last_year is not None and last_year < cutoff
+ ),
+ "n_history_years": len(kept),
+ "n_imputed_history_years": (
+ None
+ if imputed is None
+ else len(imputed.get(pid, set()) & set(kept))
+ ),
+ "years_before_1951_dropped": len(early),
+ "computation_years": n_years,
+ "indexing_year": indexing,
+ }
+ if not history_all:
+ row["reason"] = "no_career_rows"
+ elif not kept:
+ # Rows exist only after the cutoff (or before 1951): the
+ # earnings before 62 are unobserved, not zero.
+ row["reason"] = "no_career_rows_before_cutoff"
+ elif (
+ convention_years is ComputationYears.STATUTORY
+ and birth + _RETIREMENT_AGE
+ < statutory_aime.FIRST_AGE_62_YEAR_ENCODED
+ ):
+ row["reason"] = "statutory_oracle_refuses_attaining_62_before_1975"
+ elif missing := _nawi_missing(kept, indexing, params):
+ row["reason"] = f"nawi_unavailable:{missing[0]}"
+ else:
+ if convention_years is ComputationYears.STATUTORY:
+ row["computation_years"] = (
+ statutory_aime.benefit_computation_years(birth)
+ )
+ row["aime"] = float(
+ statutory_aime.oracle_aime(
+ kept, birth, params, computation_years=convention_years
+ )
+ )
+ row["status"] = COMPUTED
+ rows.append(row)
+ frame = pd.DataFrame(rows, columns=list(_AIME_COLUMNS))
+ frame = frame.astype(
+ {
+ "first_earnings_year": "Int64",
+ "last_earnings_year": "Int64",
+ "n_imputed_history_years": "Int64",
+ "computation_years": "Int64",
+ }
+ )
+ provenance = _provenance(
+ "initial_aime_at_62",
+ frame,
+ params,
+ _input_digests(careers_frame, persons_frame),
+ {
+ "analysis_year": analysis_year,
+ "aime_convention": {
+ "name": convention.name,
+ "computation_years": convention_years.value,
+ "last_earnings_age": convention.last_earnings_age,
+ "source": convention.source,
+ },
+ "oracle": "ss.statutory_aime.oracle_aime (unchanged)",
+ "definition": "mint8:initial_aime_quintile",
+ },
+ {"basis_counts": _counts(frame["basis"])},
+ )
+ return MeasureResult("initial_aime_at_62", frame, provenance)
+
+
+# ---------------------------------------------------------------------------
+# 2. Lifetime payroll tax at 62
+# ---------------------------------------------------------------------------
+def annual_payroll_taxes(
+ careers: pd.DataFrame,
+ params: SSAParameters,
+ rates: _CombinedRateSource,
+) -> pd.DataFrame:
+ """Each career year's OASDI payroll tax.
+
+ ``taxable_earnings = min(earnings, wage_base_for(year))`` and
+ ``tax = taxable_earnings * rates.combined_for(year)`` (the combined
+ employee and employer rate). Columns: ``person_id``, ``year``,
+ ``earnings``, ``taxable_earnings``, ``combined_rate``, ``tax``.
+ """
+
+ frame = _careers_frame(careers)
+ years = sorted(int(year) for year in frame["year"].unique())
+ base = {year: float(params.wage_base_for(year)) for year in years}
+ rate = {year: float(rates.combined_for(year)) for year in years}
+ taxable = np.minimum(
+ frame["earnings"].to_numpy(), frame["year"].map(base).to_numpy()
+ )
+ combined = frame["year"].map(rate).to_numpy()
+ out = frame[["person_id", "year", "earnings"]].copy()
+ out["taxable_earnings"] = taxable
+ out["combined_rate"] = combined
+ out["tax"] = taxable * combined
+ return out.reset_index(drop=True)
+
+
+def _episodes_by_person(
+ episodes: pd.DataFrame,
+) -> dict[int, pd.DataFrame]:
+ missing = set(_EPISODE_COLUMNS) - set(episodes.columns)
+ if missing:
+ raise ValueError(f"marriage_episodes lacks {sorted(missing)}.")
+ frame = episodes.copy()
+ frame["person_id"] = _integral(frame["person_id"], "episode person_id")
+ return {
+ int(pid): group.reset_index(drop=True)
+ for pid, group in frame.groupby("person_id", sort=True)
+ }
+
+
+def _spouse_id(value: Any) -> int | None:
+ return None if pd.isna(value) else int(value)
+
+
+class _MaritalYears:
+ """Year-end marital states through ``psid2010.marital_state_at``."""
+
+ def __init__(
+ self, episodes: pd.DataFrame, *, separated_is_married: bool
+ ) -> None:
+ from populace_dynamics.cohorts.psid2010 import marital_state_at
+
+ self._state_at = marital_state_at
+ self._by_person = _episodes_by_person(episodes)
+ self._empty = episodes.iloc[0:0]
+ self._separated_is_married = bool(separated_is_married)
+ self._cache: dict[tuple[int, int], tuple[str, int | None]] = {}
+
+ def has_history(self, pid: int) -> bool:
+ return pid in self._by_person
+
+ def spouses_ever(self, pid: int) -> set[int]:
+ group = self._by_person.get(pid)
+ if group is None:
+ return set()
+ return {
+ int(spouse)
+ for spouse in group["spouse_person_id"]
+ if not pd.isna(spouse)
+ }
+
+ def state(self, pid: int, year: int) -> tuple[str, int | None]:
+ key = (pid, year)
+ if key not in self._cache:
+ state = self._state_at(
+ self._by_person.get(pid, self._empty),
+ year,
+ separated_is_married=self._separated_is_married,
+ )
+ self._cache[key] = (
+ str(state["status"]),
+ _spouse_id(state["spouse_person_id"]),
+ )
+ return self._cache[key]
+
+
+@dataclass
+class _ShareCounts:
+ married_years_shared: int = 0
+ married_years_spouse_unavailable: int = 0
+ married_years_spouse_year_absent: int = 0
+ married_years_spouse_record_disagrees: int = 0
+ years_marital_unknown: int = 0
+
+
+def _shared_amounts(
+ pid: int,
+ own: Mapping[int, float],
+ spouse_amounts: Mapping[int, Mapping[int, float]],
+ marital: _MaritalYears,
+ candidate_years: Collection[int],
+) -> tuple[dict[int, float], _ShareCounts]:
+ """Per-year amounts with married years shared equally (MINT8 rule).
+
+ ``spouse_amounts`` maps a person to that person's amounts by year;
+ a person absent from it has no career (unavailable).
+ """
+
+ counts = _ShareCounts()
+ out: dict[int, float] = {}
+ for year in sorted(candidate_years):
+ status, spouse = marital.state(pid, year)
+ own_value = own.get(year)
+ if status == "married":
+ spouse_years = (
+ None if spouse is None else spouse_amounts.get(spouse)
+ )
+ if spouse_years is None:
+ counts.married_years_spouse_unavailable += 1
+ if own_value is not None:
+ out[year] = own_value
+ continue
+ if year not in spouse_years:
+ counts.married_years_spouse_year_absent += 1
+ if marital.has_history(spouse):
+ back_status, back_spouse = marital.state(spouse, year)
+ if back_status != "married" or back_spouse != pid:
+ counts.married_years_spouse_record_disagrees += 1
+ out[year] = (
+ (0.0 if own_value is None else own_value)
+ + spouse_years.get(year, 0.0)
+ ) / 2.0
+ counts.married_years_shared += 1
+ continue
+ if status == "unknown":
+ counts.years_marital_unknown += 1
+ if own_value is not None:
+ out[year] = own_value
+ return out, counts
+
+
+_PV_COLUMNS = (
+ "person_id",
+ "birth_year",
+ "pv_at_62",
+ "status",
+ "reason",
+ "reference_year",
+ "first_tax_year",
+ "last_tax_year",
+ "history_starts_after_age_22",
+ "n_tax_years",
+ "n_imputed_tax_years",
+)
+_SHARE_COLUMNS = (
+ "married_years_shared",
+ "married_years_spouse_unavailable",
+ "married_years_spouse_year_absent",
+ "married_years_spouse_record_disagrees",
+ "years_marital_unknown",
+ "marriage_history_absent",
+)
+
+
+def lifetime_payroll_tax_pv_at_62(
+ careers: pd.DataFrame,
+ persons: pd.DataFrame,
+ params: SSAParameters,
+ *,
+ shared: bool,
+ rates: _CombinedRateSource,
+ interest: _InterestRateSource,
+ marriage_episodes: pd.DataFrame | None = None,
+ separated_is_married: bool = True,
+ missing_spouse: MissingSpousePolicy = MissingSpousePolicy.OWN_ONLY,
+ missing_rate: MissingRatePolicy = MissingRatePolicy.REFUSE,
+ marriage_history_person_ids: Collection[int] | None = None,
+) -> MeasureResult:
+ """Present value at age 62 of current-law OASDI payroll taxes (MINT8).
+
+ Each career year's tax is the combined employee-plus-employer OASDI
+ rate (``rates.combined_for``; :func:`load_oasdi_tax_rates`) times
+ earnings capped at the contribution and benefit base
+ (``params.wage_base_for``). Taxes are accumulated (years before the
+ year of attaining 62, ``Y = birth_year + 62``) or discounted (later
+ years) to the end of year Y at the trust fund interest rate
+ (``interest.rate_for``; :func:`load_trust_fund_interest_rates`) by
+ :func:`accumulation_factor`, and summed with ``math.fsum``.
+
+ ``shared=True`` applies MINT8's rule: in each year the person is
+ married (year-end state from ``marriage_episodes`` through
+ ``psid2010.marital_state_at``), the year's tax is half the sum of the
+ two spouses' taxes, the spouse's from the spouse's own career rows; in
+ any other year it is the person's own. A married year whose spouse is
+ not in the careers frame follows ``missing_spouse`` and is counted.
+ The registered builder defaults are :data:`PAYROLL_TAX_BUILDER_DEFAULTS`.
+
+ Not computed, with a reason: no career rows; an interest rate outside
+ the series' coverage under ``missing_rate=NOT_COMPUTED`` (the default
+ refuses); a missing spouse under ``missing_spouse=NOT_COMPUTED``.
+ """
+
+ shared = bool(shared)
+ missing_spouse = MissingSpousePolicy(missing_spouse)
+ missing_rate = MissingRatePolicy(missing_rate)
+ if shared and marriage_episodes is None:
+ raise ValueError("shared=True needs marriage_episodes.")
+ careers_frame = _careers_frame(careers)
+ persons_frame = _persons_frame(persons)
+ taxes = annual_payroll_taxes(careers_frame, params, rates)
+ tax_by_person: dict[int, dict[int, float]] = {}
+ for pid, year, tax in zip(
+ taxes["person_id"], taxes["year"], taxes["tax"], strict=True
+ ):
+ tax_by_person.setdefault(int(pid), {})[int(year)] = float(tax)
+ imputed = _imputed_years(careers_frame)
+ marital = (
+ _MaritalYears(
+ marriage_episodes, separated_is_married=separated_is_married
+ )
+ if shared
+ else None
+ )
+ history_ids = (
+ None
+ if marriage_history_person_ids is None
+ else {int(pid) for pid in marriage_history_person_ids}
+ )
+
+ rows = []
+ uncovered: dict[int, list[int]] = {}
+ for pid, birth in zip(
+ persons_frame["person_id"], persons_frame["birth_year"], strict=True
+ ):
+ pid, birth = int(pid), int(birth)
+ reference = birth + _RETIREMENT_AGE
+ own = tax_by_person.get(pid, {})
+ row: dict[str, Any] = {
+ "person_id": pid,
+ "birth_year": birth,
+ "pv_at_62": math.nan,
+ "status": NOT_COMPUTED,
+ "reason": None,
+ "reference_year": reference,
+ "first_tax_year": min(own) if own else None,
+ "last_tax_year": max(own) if own else None,
+ "history_starts_after_age_22": bool(own)
+ and min(own) > birth + _FIRST_AGE,
+ "n_tax_years": len(own),
+ "n_imputed_tax_years": (
+ None if imputed is None else len(imputed.get(pid, set()))
+ ),
+ }
+ amounts: dict[int, float] = dict(own)
+ if shared:
+ assert marital is not None
+ spouses = marital.spouses_ever(pid)
+ candidates = set(own)
+ for spouse in spouses:
+ candidates |= set(tax_by_person.get(spouse, {}))
+ amounts, counts = _shared_amounts(
+ pid, own, tax_by_person, marital, candidates
+ )
+ row.update(vars(counts))
+ row["marriage_history_absent"] = (
+ None if history_ids is None else pid not in history_ids
+ )
+ if not own:
+ row["reason"] = "no_career_rows"
+ elif (
+ shared
+ and missing_spouse is MissingSpousePolicy.NOT_COMPUTED
+ and row["married_years_spouse_unavailable"] > 0
+ ):
+ row["reason"] = "spouse_career_unavailable"
+ else:
+ needed = set()
+ for year in amounts:
+ low, high = sorted((year, reference))
+ needed.update(range(low + 1, high + 1))
+ gaps = sorted(
+ year for year in needed if not _covers(interest, year)
+ )
+ if gaps:
+ uncovered[pid] = gaps
+ row["reason"] = f"interest_rate_unavailable:{gaps[0]}"
+ else:
+ row["pv_at_62"] = math.fsum(
+ amount * accumulation_factor(year, reference, interest)
+ for year, amount in sorted(amounts.items())
+ )
+ row["status"] = COMPUTED
+ rows.append(row)
+ if uncovered and missing_rate is MissingRatePolicy.REFUSE:
+ years = sorted({year for gaps in uncovered.values() for year in gaps})
+ raise ValueError(
+ f"{len(uncovered)} persons need trust fund interest rates for "
+ f"years {years[0]}-{years[-1]} that the series does not cover; "
+ "extend it with TrustFundInterestRates.extended(..., source=...) "
+ "or pass missing_rate=MissingRatePolicy.NOT_COMPUTED."
+ )
+ columns = list(_PV_COLUMNS) + (list(_SHARE_COLUMNS) if shared else [])
+ frame = pd.DataFrame(rows, columns=columns).astype(
+ {
+ "first_tax_year": "Int64",
+ "last_tax_year": "Int64",
+ "n_imputed_tax_years": "Int64",
+ }
+ )
+ if shared:
+ frame["marriage_history_absent"] = frame[
+ "marriage_history_absent"
+ ].astype("boolean")
+ measure = (
+ "lifetime_payroll_tax_pv_at_62_shared"
+ if shared
+ else "lifetime_payroll_tax_pv_at_62"
+ )
+ provenance = _provenance(
+ measure,
+ frame,
+ params,
+ _input_digests(
+ careers_frame, persons_frame, marriage_episodes if shared else None
+ ),
+ {
+ "shared": shared,
+ "builder_defaults": PAYROLL_TAX_BUILDER_DEFAULTS,
+ "separated_is_married": bool(separated_is_married),
+ "missing_spouse": missing_spouse.value,
+ "missing_rate": missing_rate.value,
+ "definition": (
+ "mint8:lifetime_payroll_tax_quintile_shared"
+ if shared
+ else "mint8:lifetime_payroll_tax_quintile"
+ ),
+ },
+ {
+ "tax_rates": dict(getattr(rates, "provenance", {}) or {}),
+ "interest_rates": {
+ **dict(getattr(interest, "provenance", {}) or {}),
+ "assumed_years": sorted(
+ getattr(interest, "assumed_years", frozenset())
+ ),
+ "assumption_source": getattr(
+ interest, "assumption_source", None
+ ),
+ },
+ },
+ )
+ return MeasureResult(measure, frame, provenance)
+
+
+def _covers(interest: _InterestRateSource, year: int) -> bool:
+ covers = getattr(interest, "covers", None)
+ if covers is not None:
+ return bool(covers(year))
+ try:
+ interest.rate_for(year)
+ except KeyError:
+ return False
+ return True
+
+
+# ---------------------------------------------------------------------------
+# 3. The Report's average indexed earnings at ages 22-62
+# ---------------------------------------------------------------------------
+_REPORT_COLUMNS = (
+ "person_id",
+ "birth_year",
+ "average_indexed_earnings",
+ "status",
+ "reason",
+ "index_year",
+ "n_ages_covered",
+ "n_ages_window",
+ "ages_complete",
+ "n_imputed_years",
+)
+
+
+def report_average_indexed_earnings_22_62(
+ careers: pd.DataFrame,
+ persons: pd.DataFrame,
+ params: SSAParameters,
+ *,
+ shared: bool,
+ marriage_episodes: pd.DataFrame | None = None,
+ conventions: ReportEarningsConventions | None = None,
+ marriage_history_person_ids: Collection[int] | None = None,
+) -> MeasureResult:
+ """Butrica and Uccello (2004)'s lifetime earnings, as recorded.
+
+ The average over ages ``first_age``-``last_age`` (22-62) of nominal
+ earnings indexed by NAWI to the year of attaining ``index_age``,
+ uncapped (the recorded own measure includes earnings above the taxable
+ maximum). ``shared=True`` averages, at each covered age the person is
+ married, the mean of the two spouses' indexed earnings. Every
+ unrecorded convention is a registered builder default
+ (:data:`REPORT_EARNINGS_BUILDER_DEFAULTS`), set in ``conventions``.
+
+ Not computed, with a reason: no career row at ages 22-62; a NAWI value
+ the indexing needs; a missing spouse under
+ ``missing_spouse=NOT_COMPUTED``.
+ """
+
+ conventions = conventions or ReportEarningsConventions()
+ shared = bool(shared)
+ if shared and marriage_episodes is None:
+ raise ValueError("shared=True needs marriage_episodes.")
+ if conventions.last_age < conventions.first_age:
+ raise ValueError("last_age precedes first_age.")
+ divisor = AverageDivisor(conventions.divisor)
+ missing_spouse = MissingSpousePolicy(conventions.missing_spouse)
+ careers_frame = _careers_frame(careers)
+ persons_frame = _persons_frame(persons)
+ histories = _histories(careers_frame)
+ imputed = _imputed_years(careers_frame)
+ marital = (
+ _MaritalYears(
+ marriage_episodes,
+ separated_is_married=conventions.separated_is_married,
+ )
+ if shared
+ else None
+ )
+ history_ids = (
+ None
+ if marriage_history_person_ids is None
+ else {int(pid) for pid in marriage_history_person_ids}
+ )
+ window_length = conventions.last_age - conventions.first_age + 1
+
+ def earnings_value(earnings: float, year: int) -> float:
+ if conventions.cap_at_taxable_maximum:
+ return min(earnings, float(params.wage_base_for(year)))
+ return earnings
+
+ rows = []
+ for pid, birth in zip(
+ persons_frame["person_id"], persons_frame["birth_year"], strict=True
+ ):
+ pid, birth = int(pid), int(birth)
+ first_year = birth + conventions.first_age
+ last_year = birth + conventions.last_age
+ index_year = birth + conventions.index_age
+ own = {
+ year: earnings_value(value, year)
+ for year, value in histories.get(pid, {}).items()
+ if first_year <= year <= last_year
+ }
+ row: dict[str, Any] = {
+ "person_id": pid,
+ "birth_year": birth,
+ "average_indexed_earnings": math.nan,
+ "status": NOT_COMPUTED,
+ "reason": None,
+ "index_year": index_year,
+ "n_ages_covered": len(own),
+ "n_ages_window": window_length,
+ "ages_complete": len(own) == window_length,
+ "n_imputed_years": (
+ None
+ if imputed is None
+ else len(imputed.get(pid, set()) & set(own))
+ ),
+ }
+ amounts: dict[int, float] = dict(own)
+ if shared:
+ assert marital is not None
+ spouse_amounts = {}
+ for spouse in marital.spouses_ever(pid):
+ if spouse in histories:
+ spouse_amounts[spouse] = {
+ year: earnings_value(value, year)
+ for year, value in histories[spouse].items()
+ if first_year <= year <= last_year
+ }
+ amounts, counts = _shared_amounts(
+ pid, own, spouse_amounts, marital, set(own)
+ )
+ row.update(vars(counts))
+ row["marriage_history_absent"] = (
+ None if history_ids is None else pid not in history_ids
+ )
+ if not own:
+ row["reason"] = "no_career_rows_at_ages_22_62"
+ elif (
+ shared
+ and missing_spouse is MissingSpousePolicy.NOT_COMPUTED
+ and row["married_years_spouse_unavailable"] > 0
+ ):
+ row["reason"] = "spouse_career_unavailable"
+ elif missing := _nawi_missing(amounts, index_year, params):
+ row["reason"] = f"nawi_unavailable:{missing[0]}"
+ else:
+ base = params.nawi[index_year]
+ indexed = [
+ (
+ value * base / params.nawi[year]
+ if year < index_year
+ else value
+ )
+ for year, value in sorted(amounts.items())
+ ]
+ count = (
+ len(own)
+ if divisor is AverageDivisor.COVERED_AGES
+ else (window_length)
+ )
+ row["average_indexed_earnings"] = math.fsum(indexed) / count
+ row["status"] = COMPUTED
+ rows.append(row)
+ columns = list(_REPORT_COLUMNS) + (list(_SHARE_COLUMNS) if shared else [])
+ frame = pd.DataFrame(rows, columns=columns).astype(
+ {"n_imputed_years": "Int64"}
+ )
+ if shared:
+ frame["marriage_history_absent"] = frame[
+ "marriage_history_absent"
+ ].astype("boolean")
+ measure = (
+ "report_average_indexed_earnings_22_62_shared"
+ if shared
+ else "report_average_indexed_earnings_22_62"
+ )
+ provenance = _provenance(
+ measure,
+ frame,
+ params,
+ _input_digests(
+ careers_frame, persons_frame, marriage_episodes if shared else None
+ ),
+ {
+ "shared": shared,
+ "conventions": {
+ "first_age": conventions.first_age,
+ "last_age": conventions.last_age,
+ "cap_at_taxable_maximum": conventions.cap_at_taxable_maximum,
+ "index_age": conventions.index_age,
+ "divisor": divisor.value,
+ "separated_is_married": conventions.separated_is_married,
+ "missing_spouse": missing_spouse.value,
+ },
+ "recorded": REPORT_EARNINGS_RECORDED,
+ "builder_defaults": REPORT_EARNINGS_BUILDER_DEFAULTS,
+ },
+ )
+ return MeasureResult(measure, frame, provenance)
+
+
+# ---------------------------------------------------------------------------
+# 4. Weighted quintiles
+# ---------------------------------------------------------------------------
+def _quintile_ranks(values: np.ndarray, weights: np.ndarray) -> np.ndarray:
+ """Rank 0 (lowest) to 4 (highest) for one cell, exactly.
+
+ Rows sharing a value form one group and always share a rank. Groups
+ are ordered by value; a group whose cumulative weight span is
+ [B, B + G] of the total W takes rank floor(5 * (B + G / 2) / W),
+ capped at 4: the quintile holding the midpoint of its span. Weights
+ are summed as exact fractions, so the ranks do not depend on row
+ order or on rounding.
+ """
+
+ unique, inverse = np.unique(values, return_inverse=True)
+ group_weights = [Fraction(0)] * len(unique)
+ for group, weight in zip(inverse, weights, strict=True):
+ group_weights[group] += Fraction(float(weight))
+ total = sum(group_weights, Fraction(0))
+ ranks = np.empty(len(unique), dtype=np.int64)
+ below = Fraction(0)
+ for index, weight in enumerate(group_weights):
+ twice_midpoint = 2 * below + weight
+ ranks[index] = min(
+ _N_QUINTILES - 1,
+ (_N_QUINTILES * twice_midpoint) // (2 * total),
+ )
+ below += weight
+ return ranks[inverse]
+
+
+def weighted_quintiles(
+ values: pd.Series,
+ weights: pd.Series,
+ *,
+ by: pd.Series | None = None,
+ labels: tuple[str, ...] = QUINTILE_LABELS,
+) -> pd.Series:
+ """Weighted quintile labels, in MINT order (``Highest`` .. ``Lowest``).
+
+ ``values`` (NaN = no value), ``weights`` (finite and positive for
+ every row with a value) and ``by`` (the cell key, e.g.
+ :func:`ten_year_birth_cohort`; ``None`` cuts the whole population)
+ must share one index. Quintiles are cut separately within each cell.
+
+ Ties: rows with the same value always share a label; the group takes
+ the quintile holding the midpoint of its cumulative weight span
+ (:func:`_quintile_ranks`). So with no ties each quintile holds 20
+ percent of the cell's weight to within the largest single weight's
+ share. MINT8 publishes no tie rule ("The dollar ranges are available
+ upon request"); this is the builder's registered rule.
+
+ Returns a categorical Series named ``quintile`` with categories
+ ``labels`` then :data:`NOT_COMPUTED`; a row without a value gets
+ :data:`NOT_COMPUTED`.
+ """
+
+ labels = tuple(labels)
+ if len(labels) != _N_QUINTILES or len(set(labels)) != _N_QUINTILES:
+ raise ValueError("labels must be five distinct strings.")
+ if NOT_COMPUTED in labels:
+ raise ValueError(f"{NOT_COMPUTED!r} cannot be a quintile label.")
+ values = pd.Series(values)
+ weights = pd.Series(weights)
+ if not values.index.equals(weights.index):
+ raise ValueError("values and weights must share one index.")
+ if by is not None:
+ by = pd.Series(by)
+ if not by.index.equals(values.index):
+ raise ValueError("by must share the index of values.")
+ numeric = pd.to_numeric(values, errors="raise").astype("float64")
+ present = numeric.notna().to_numpy()
+ array = numeric.to_numpy()
+ if np.isinf(array[present]).any():
+ raise ValueError("values must be finite or missing.")
+ weight_array = pd.to_numeric(weights, errors="raise").astype("float64")
+ weight_array = weight_array.to_numpy()
+ usable = weight_array[present]
+ if not np.all(np.isfinite(usable)) or (usable <= 0).any():
+ raise ValueError("weights must be finite and positive where valued.")
+ out = np.full(len(array), NOT_COMPUTED, dtype=object)
+ positions = np.flatnonzero(present)
+ if len(positions):
+ if by is None:
+ cells = {None: positions}
+ else:
+ keys = by.iloc[positions]
+ if keys.isna().any():
+ raise ValueError("by is missing for a row with a value.")
+ grouped = pd.Series(positions, index=keys.to_numpy()).groupby(
+ level=0, sort=True
+ )
+ cells = {key: group.to_numpy() for key, group in grouped}
+ label_array = np.array(labels, dtype=object)
+ for cell_positions in cells.values():
+ ranks = _quintile_ranks(
+ array[cell_positions], weight_array[cell_positions]
+ )
+ out[cell_positions] = label_array[_N_QUINTILES - 1 - ranks]
+ return pd.Series(
+ pd.Categorical(out, categories=[*labels, NOT_COMPUTED]),
+ index=values.index,
+ name="quintile",
+ )
+
+
+def ten_year_birth_cohort(birth_years: pd.Series) -> pd.Series:
+ """MINT8's 10-year birth cohort label, e.g. 1964 -> '1960–1969'."""
+ series = pd.Series(birth_years)
+ starts = (_integral(series, "birth_year") // 10) * 10
+ return pd.Series(
+ [f"{start}{_EN_DASH}{start + 9}" for start in starts],
+ index=series.index,
+ name="birth_cohort",
+ dtype="string",
+ )
+
+
+def quintile_cells(
+ birth_years: pd.Series, scope: QuintileScope
+) -> pd.Series | None:
+ """The ``by`` argument of :func:`weighted_quintiles` for a scope."""
+ scope = QuintileScope(scope)
+ if scope is QuintileScope.WHOLE_POPULATION:
+ return None
+ return ten_year_birth_cohort(birth_years)
+
+
+def quintile_summary(
+ labels: pd.Series,
+ weights: pd.Series,
+ *,
+ by: pd.Series | None = None,
+) -> pd.DataFrame:
+ """Unweighted count, weight and weight share by cell and label.
+
+ The unweighted counts are what MINT8's disclosure rule reads ("suppress
+ an entire characteristic subgroup if the sample size for any row in
+ that subgroup is less than 100 individuals").
+ """
+
+ labels = pd.Series(labels)
+ weights = pd.Series(weights)
+ if not labels.index.equals(weights.index):
+ raise ValueError("labels and weights must share one index.")
+ cell = (
+ pd.Series("all", index=labels.index, dtype="string")
+ if by is None
+ else pd.Series(by).astype("string")
+ )
+ frame = pd.DataFrame(
+ {
+ "cell": cell.to_numpy(),
+ "label": labels.astype("string").to_numpy(),
+ "weight": pd.to_numeric(weights).astype("float64").to_numpy(),
+ }
+ )
+ summary = (
+ frame.groupby(["cell", "label"], sort=True)
+ .agg(n=("weight", "size"), weight=("weight", "sum"))
+ .reset_index()
+ )
+ valued = summary[summary["label"] != NOT_COMPUTED]
+ totals = valued.groupby("cell")["weight"].sum()
+ summary["weight_share"] = [
+ (weight / totals[cell] if label != NOT_COMPUTED else math.nan)
+ for cell, label, weight in zip(
+ summary["cell"], summary["label"], summary["weight"], strict=True
+ )
+ ]
+ return summary
+
+
+# ---------------------------------------------------------------------------
+# Schemes
+# ---------------------------------------------------------------------------
+@dataclass(frozen=True)
+class QuintileDimension:
+ """One lifetime-earnings row group of a report's tables.
+
+ ``labels_high_to_low`` or ``scope`` ``None`` means the source does not
+ record it; :meth:`labels` and :meth:`cells` then refuse.
+ """
+
+ key: str
+ section: str
+ measure: str
+ shared: bool
+ scope: QuintileScope | None
+ labels_high_to_low: tuple[str, ...] | None
+ definition: str
+ source: str
+
+ def labels(self) -> tuple[str, ...]:
+ if self.labels_high_to_low is None:
+ raise ValueError(
+ f"{self.key}: the source does not record which label is "
+ "the highest quintile; register the order first."
+ )
+ return self.labels_high_to_low
+
+ def cells(self, birth_years: pd.Series) -> pd.Series | None:
+ if self.scope is None:
+ raise ValueError(
+ f"{self.key}: the source does not record the population "
+ "the quintiles are cut over; register it first."
+ )
+ return quintile_cells(birth_years, self.scope)
+
+
+@dataclass(frozen=True)
+class LifetimeEarningsScheme:
+ """A report's lifetime-earnings row groups (data, not code paths)."""
+
+ name: str
+ title: str
+ source: str
+ registered: bool
+ dimensions: tuple[QuintileDimension, ...]
+
+ def dimension(self, key: str) -> QuintileDimension:
+ for dimension in self.dimensions:
+ if dimension.key == key:
+ return dimension
+ raise KeyError(f"{self.name} has no dimension {key!r}.")
+
+
+#: Measure function and sharing of each MINT8 cohort-table row group.
+MINT8_DIMENSION_MEASURES: dict[str, tuple[str, bool]] = {
+ "initial_aime_quintile": ("initial_aime_at_62", False),
+ "lifetime_payroll_tax_quintile": ("lifetime_payroll_tax_pv_at_62", False),
+ "lifetime_payroll_tax_quintile_shared": (
+ "lifetime_payroll_tax_pv_at_62",
+ True,
+ ),
+}
+
+
+def load_mint8_definitions(
+ path: Path = MINT8_DEFINITIONS_PATH,
+ *,
+ expected_sha256: str = MINT8_DEFINITIONS_SHA256,
+) -> dict[str, Any]:
+ """The committed verbatim MINT8 definitions (pinned)."""
+ document, _ = _read_pinned_json(
+ path, expected_sha256, "mint8_lifetime_quintile_definitions.v1"
+ )
+ return document
+
+
+def load_mint8_scheme(
+ path: Path = MINT8_DEFINITIONS_PATH,
+ *,
+ expected_sha256: str = MINT8_DEFINITIONS_SHA256,
+) -> LifetimeEarningsScheme:
+ """MINT8's cohort-table lifetime-earnings scheme, from the data file."""
+ document = load_mint8_definitions(path, expected_sha256=expected_sha256)
+ labels = tuple(document["quintile_labels_high_to_low"])
+ if labels != QUINTILE_LABELS:
+ raise ValueError("MINT8 quintile labels differ from QUINTILE_LABELS.")
+ sections = document["cohort_table_row_groups"]
+ if set(sections) != set(MINT8_DIMENSION_MEASURES):
+ raise ValueError("MINT8 row groups differ from the measure map.")
+ dimensions = tuple(
+ QuintileDimension(
+ key=key,
+ section=sections[key],
+ measure=measure,
+ shared=shared,
+ scope=QuintileScope.TEN_YEAR_BIRTH_COHORT,
+ labels_high_to_low=labels,
+ definition=document["definitions"][key]["text"],
+ source=(
+ f"{document['document']} (dateCertified "
+ f"{document['date_certified']}), element id "
+ f"{document['definitions'][key]['element_id']}"
+ ),
+ )
+ for key, (measure, shared) in MINT8_DIMENSION_MEASURES.items()
+ )
+ return LifetimeEarningsScheme(
+ name="mint8",
+ title="MINT8 cohort tables (benefit/tax ratios, replacement rates)",
+ source=f"{_relative(path)} (sha256 {expected_sha256})",
+ registered=True,
+ dimensions=dimensions,
+ )
+
+
+def boomers2004_scheme() -> LifetimeEarningsScheme:
+ """Butrica and Uccello (2004)'s lifetime-earnings rows (unregistered).
+
+ Section and row labels come from Track U's named omissions
+ (``uniform_cut_tabulation.NOT_COMPUTED_REPORT_ROWS``). The order of
+ the quintile labels and the quintile population are not recorded, so
+ both dimensions refuse to label until a registration records them.
+ """
+
+ from populace_dynamics.estimates.uniform_cut_tabulation import (
+ NOT_COMPUTED_REPORT_ROWS,
+ )
+
+ dimensions = []
+ for kind in ("own", "shared"):
+ record = NOT_COMPUTED_REPORT_ROWS[f"lifetime_earnings_{kind}"]
+ dimensions.append(
+ QuintileDimension(
+ key=f"lifetime_earnings_{kind}",
+ section=record["section"],
+ measure="report_average_indexed_earnings_22_62",
+ shared=kind == "shared",
+ scope=None,
+ labels_high_to_low=None,
+ definition=REPORT_EARNINGS_RECORDED["measure"],
+ source=_REPORT_RECORD,
+ )
+ )
+ return LifetimeEarningsScheme(
+ name="boomers2004",
+ title="Butrica and Uccello (2004), Tables 19 and 21",
+ source=_REPORT_RECORD,
+ registered=False,
+ dimensions=tuple(dimensions),
+ )
+
+
+#: Scheme builders by name; add a report's scheme here as data.
+SCHEMES: dict[str, Callable[[], LifetimeEarningsScheme]] = {
+ "mint8": load_mint8_scheme,
+ "boomers2004": boomers2004_scheme,
+}
diff --git a/tests/estimates/test_lifetime_measure_sources.py b/tests/estimates/test_lifetime_measure_sources.py
new file mode 100644
index 00000000..6a75a943
--- /dev/null
+++ b/tests/estimates/test_lifetime_measure_sources.py
@@ -0,0 +1,266 @@
+"""The captured SSA sources behind the G2 lifetime-earnings measures.
+
+Reader-free: these tests read only committed source bodies and their
+extractions (SSA OACT tax and interest-rate pages, the MINT8 Table User
+Guide), never PSID data and never a comparator source.
+"""
+
+from __future__ import annotations
+
+import hashlib
+import json
+import math
+import sys
+from pathlib import Path
+
+import pandas as pd
+import pytest
+
+from populace_dynamics.engine.refit import validate_external_vintage
+from populace_dynamics.estimates import lifetime_measures as lm
+from populace_dynamics.ss.params import SSAParameters
+
+ROOT = Path(__file__).resolve().parents[2]
+SCRIPTS = ROOT / "scripts"
+EXTERNAL = ROOT / "data" / "external"
+VINTAGE_2014 = EXTERNAL / "ssa_effective_interest_rates_2014.json"
+
+if str(SCRIPTS) not in sys.path:
+ sys.path.insert(0, str(SCRIPTS))
+
+import extract_lifetime_measure_sources as extractor # noqa: E402
+
+
+@pytest.mark.parametrize("key", sorted(extractor.SOURCES))
+def test_sources_have_their_pinned_sha256_and_length(key):
+ spec = extractor.SOURCES[key]
+ raw = (EXTERNAL / spec["file"]).read_bytes()
+ assert hashlib.sha256(raw).hexdigest() == spec["sha256"]
+ assert len(raw) == spec["bytes"]
+
+
+def test_extractions_rebuild_byte_for_byte():
+ for path, document in extractor.build_all().items():
+ assert extractor.render(document) == path.read_text(encoding="utf-8")
+
+
+def test_module_pins_match_the_committed_extractions():
+ for path, pinned in (
+ (lm.OASDI_TAX_RATES_PATH, lm.OASDI_TAX_RATES_SHA256),
+ (
+ lm.TRUST_FUND_INTEREST_RATES_PATH,
+ lm.TRUST_FUND_INTEREST_RATES_SHA256,
+ ),
+ (lm.MINT8_DEFINITIONS_PATH, lm.MINT8_DEFINITIONS_SHA256),
+ ):
+ assert path.parent == EXTERNAL
+ assert hashlib.sha256(path.read_bytes()).hexdigest() == pinned
+
+
+def test_provenance_names_every_source_and_output():
+ record = json.loads(extractor.PROVENANCE_OUT.read_text())
+ assert record["schema_version"] == "external_source_provenance_set.v1"
+ assert set(record["sources"]) == set(extractor.SOURCES)
+ for key, entry in record["sources"].items():
+ assert entry["source_sha256"] == extractor.SOURCES[key]["sha256"]
+ assert entry["source_url"].startswith("https://www.ssa.gov/")
+ outputs = record["consumed_by"][
+ "scripts/extract_lifetime_measure_sources.py"
+ ]
+ assert sorted(outputs) == sorted(
+ f"data/external/{path.name}"
+ for path in (
+ extractor.INTEREST_OUT,
+ extractor.TAX_OUT,
+ extractor.MINT8_OUT,
+ )
+ )
+
+
+# ---------------------------------------------------------------------------
+# Interest rates
+# ---------------------------------------------------------------------------
+def test_interest_rates_cover_1940_2025_and_pin_rows():
+ data = json.loads(extractor.INTEREST_OUT.read_text())["data"]
+ assert [int(year) for year in data] == list(range(1940, 2026))
+ assert data["1940"] == {
+ "average_new_issue": 2.5,
+ "effective_oasdi": 2.4,
+ "effective_oasi": 2.4,
+ "effective_di": None,
+ }
+ assert data["1957"]["effective_di"] == 2.3
+ assert data["2025"] == {
+ "average_new_issue": 4.3,
+ "effective_oasdi": 2.6,
+ "effective_oasi": 2.5,
+ "effective_di": 3.6,
+ }
+
+
+def test_effective_rates_agree_with_the_committed_2014_vintage():
+ """Differential: two captures, 12 years apart, of the same SSA series."""
+ old = json.loads(VINTAGE_2014.read_text())["data"]
+ new = json.loads(extractor.INTEREST_OUT.read_text())["data"]
+ for year, row in old.items():
+ assert new[year]["effective_oasdi"] == row["oasdi"]
+ assert new[year]["effective_oasi"] == row["oasi"]
+ assert new[year]["effective_di"] == row["di"]
+
+
+def test_interest_file_is_a_post_boundary_vintage():
+ doc = json.loads(extractor.INTEREST_OUT.read_text())
+ assert doc["vintage_year"] == 2026
+ with pytest.raises(ValueError, match="post-T"):
+ validate_external_vintage(
+ doc["schema_version"], doc["vintage_year"], boundary_year=2014
+ )
+
+
+@pytest.mark.parametrize("series", list(lm.InterestSeries))
+def test_interest_loader_reads_each_series(series):
+ rates = lm.load_trust_fund_interest_rates(series=series)
+ assert rates.series == series.value
+ assert rates.covers(1940) and rates.covers(2025)
+ assert not rates.covers(2026)
+ expected = {
+ lm.InterestSeries.EFFECTIVE_OASDI: 0.026,
+ lm.InterestSeries.AVERAGE_NEW_ISSUE: 0.043,
+ }[series]
+ assert rates.rate_for(2025) == pytest.approx(expected)
+ assert rates.provenance["sha256"] == lm.TRUST_FUND_INTEREST_RATES_SHA256
+
+
+def test_loaders_refuse_a_changed_extraction(tmp_path):
+ changed = tmp_path / "rates.json"
+ changed.write_text(lm.OASDI_TAX_RATES_PATH.read_text() + " ")
+ with pytest.raises(ValueError, match="pinned"):
+ lm.load_oasdi_tax_rates(changed)
+ with pytest.raises(ValueError, match="pinned"):
+ lm.load_trust_fund_interest_rates(changed)
+
+
+# ---------------------------------------------------------------------------
+# Tax rates
+# ---------------------------------------------------------------------------
+def test_tax_schedule_rows_and_footnotes():
+ doc = json.loads(extractor.TAX_OUT.read_text())
+ assert doc["rates_reflect"] == (
+ "The rates shown reflect the amounts received by the trust funds."
+ )
+ assert doc["rows"][0]["period"] == "1937-49"
+ assert doc["rows"][-1]["period"] == "2019 and later b"
+ assert doc["rows"][-1]["last_year"] is None
+ assert sorted(doc["footnotes"]) == ["a", "b", "c", "d"]
+ assert doc["footnotes"]["a"].startswith("In 1984 only")
+ adjustments = {
+ row["year"]: row["employee_effective_rate"]
+ for row in doc["paid_rate_adjustments"]
+ }
+ assert adjustments == {1984: 5.4, 2011: 4.2, 2012: 4.2}
+
+
+@pytest.mark.parametrize(
+ ("year", "received", "paid"),
+ [
+ (1937, 2.0, 2.0),
+ (1949, 2.0, 2.0),
+ (1950, 3.0, 3.0),
+ (1968, 7.6, 7.6),
+ (1983, 10.8, 10.8),
+ (1984, 11.4, 11.1),
+ (1990, 12.4, 12.4),
+ (2010, 12.4, 12.4),
+ (2011, 12.4, 10.4),
+ (2012, 12.4, 10.4),
+ (2013, 12.4, 12.4),
+ (2017, 12.4, 12.4),
+ (2060, 12.4, 12.4),
+ ],
+)
+def test_combined_rates_by_basis(year, received, paid):
+ trust_fund = lm.load_oasdi_tax_rates()
+ employee_employer = lm.load_oasdi_tax_rates(
+ basis=lm.TaxRateBasis.EMPLOYEE_EMPLOYER_PAID
+ )
+ assert trust_fund.combined_percent_for(year) == pytest.approx(received)
+ assert employee_employer.combined_percent_for(year) == pytest.approx(paid)
+
+
+def test_tax_schedule_refuses_years_before_1937():
+ with pytest.raises(KeyError):
+ lm.load_oasdi_tax_rates().combined_for(1936)
+
+
+def test_pv_example_on_the_captured_schedules():
+ # INVENTED career: 10,000 in 2023 and 2024, born 1962 (Y = 2024), no
+ # binding base. Tax 12.4 percent = 1,240 each year; the 2023 tax
+ # earns the 2024 effective rate.
+ params = SSAParameters(
+ nawi={y: 10_000.0 for y in range(1951, 2071)},
+ wage_base={1937: 1.0e12},
+ pia_factors=(0.9, 0.32, 0.15),
+ fra_months_by_birth_year=[(1900, 804)],
+ early_monthly_rates=(5 / 900, 5 / 1200),
+ early_first_bracket_months=36,
+ pe_us_revision="INVENTED",
+ )
+ interest = lm.load_trust_fund_interest_rates()
+ careers = pd.DataFrame(
+ {"person_id": [1, 1], "year": [2023, 2024], "earnings": [1e4, 1e4]}
+ )
+ result = lm.lifetime_payroll_tax_pv_at_62(
+ careers,
+ pd.DataFrame({"person_id": [1], "birth_year": [1962]}),
+ params,
+ shared=False,
+ rates=lm.load_oasdi_tax_rates(),
+ interest=interest,
+ )
+ expected = 1_240.0 * (1.0 + interest.rate_for(2024)) + 1_240.0
+ assert result.frame.at[0, "pv_at_62"] == pytest.approx(expected)
+ record = result.provenance
+ assert record["tax_rates"]["sha256"] == lm.OASDI_TAX_RATES_SHA256
+ assert record["interest_rates"]["series"] == "effective_oasdi"
+ assert math.isclose(interest.rate_for(2024), 0.025)
+
+
+# ---------------------------------------------------------------------------
+# MINT8 definitions
+# ---------------------------------------------------------------------------
+def test_mint8_definitions_are_verbatim_and_certified():
+ doc = lm.load_mint8_definitions()
+ assert doc["date_certified"] == "2026-04-01"
+ definitions = doc["definitions"]
+ assert definitions["initial_aime_quintile"]["text"].startswith(
+ "Current-Law Initial AIME Quintile: Represents an individual's "
+ "average indexed monthly earnings (AIME) under current law at age 62"
+ )
+ assert (
+ "the payroll taxes paid while married are shared equally between "
+ "them" in definitions["lifetime_payroll_tax_quintile_shared"]["text"]
+ )
+ assert definitions["present_value_convention"]["text"].endswith(
+ "We use the Social Security Trust Fund interest rate to adjust "
+ "benefits and taxes to their present values at age 62."
+ )
+ assert "less than 100 individuals" in (
+ definitions["sample_size_restriction"]["text"]
+ )
+
+
+def test_mint8_scheme_is_registered_and_data_driven():
+ scheme = lm.load_mint8_scheme()
+ assert scheme.registered
+ assert [dimension.section for dimension in scheme.dimensions] == [
+ "Current-law initial AIME quintile",
+ "Lifetime payroll tax quintile",
+ "Lifetime payroll tax quintile (shared)",
+ ]
+ shared = scheme.dimension("lifetime_payroll_tax_quintile_shared")
+ assert shared.shared and shared.measure == "lifetime_payroll_tax_pv_at_62"
+ assert shared.labels() == lm.QUINTILE_LABELS
+ cells = shared.cells(pd.Series([1965, 1984]))
+ assert cells.tolist() == ["1960–1969", "1980–1989"]
+ provenance = lm.load_mint8_definitions()["cohort_table_label_provenance"]
+ assert provenance["label_file_sha256"].startswith("23fbfbc8")
diff --git a/tests/estimates/test_lifetime_measures.py b/tests/estimates/test_lifetime_measures.py
new file mode 100644
index 00000000..4b97d371
--- /dev/null
+++ b/tests/estimates/test_lifetime_measures.py
@@ -0,0 +1,1198 @@
+"""Lifetime-earnings measures and weighted quintiles (package G2).
+
+Every career, marriage record, wage index, contribution and benefit base,
+tax rate and interest rate in this module is INVENTED. No PSID file and
+no committed evidence file is read here; the captured SSA schedules are
+tested in ``test_lifetime_measure_sources.py``. Expected values come from
+the stated definitions (hand arithmetic in comments) or from independent
+reference implementations written in this module, never from the code
+under test.
+"""
+
+from __future__ import annotations
+
+import math
+import random
+from fractions import Fraction
+
+import numpy as np
+import pandas as pd
+import pytest
+from hypothesis import HealthCheck, assume, given, settings
+from hypothesis import strategies as st
+
+from populace_dynamics import scenario_benefits as sb
+from populace_dynamics.cola_track_a import benefits as track_a_benefits
+from populace_dynamics.estimates import lifetime_measures as lm
+from populace_dynamics.fra68_track import config as fra68_config
+from populace_dynamics.min_benefit_track_m import rules as track_m_rules
+from populace_dynamics.ss import benefits, statutory_aime
+from populace_dynamics.ss.params import SSAParameters
+from populace_dynamics.ss.statutory_aime import ComputationYears
+
+SETTINGS = settings(
+ max_examples=60,
+ deadline=None,
+ suppress_health_check=[HealthCheck.too_slow],
+)
+
+
+# ---------------------------------------------------------------------------
+# INVENTED parameters, schedules and frames
+# ---------------------------------------------------------------------------
+def _params(*, cap: float | None = None) -> SSAParameters:
+ """INVENTED: 4 percent wage growth from 1951; a rising base, or one
+ flat base ``cap`` when given."""
+ wage_base = (
+ {1937: float(cap)}
+ if cap is not None
+ else {1937: 3_000.0, 1975: 14_100.0, 1990: 51_300.0, 2010: 106_800}
+ )
+ return SSAParameters(
+ nawi={y: 2_800.0 * 1.04 ** (y - 1951) for y in range(1951, 2071)},
+ wage_base=wage_base,
+ pia_factors=(0.9, 0.32, 0.15),
+ fra_months_by_birth_year=[(1900, 792), (1955, 804)],
+ early_monthly_rates=(5 / 900, 5 / 1200),
+ early_first_bracket_months=36,
+ pe_us_revision="INVENTED",
+ delayed_credit_by_birth_year=[(1900, 0.08)],
+ )
+
+
+def _rates(percent: float = 12.0) -> lm.OASDITaxRates:
+ """INVENTED: one combined rate for every year from 1937."""
+ return lm.OASDITaxRates(
+ combined_percent_by_year={1937: percent},
+ open_ended_from=1938,
+ open_ended_combined_percent=percent,
+ basis=lm.TaxRateBasis.TRUST_FUND_RECEIVED,
+ provenance={"source": "INVENTED"},
+ )
+
+
+def _interest(
+ percent: float = 5.0, first: int = 1930, last: int = 2100
+) -> lm.TrustFundInterestRates:
+ """INVENTED: one interest rate for every year in [first, last]."""
+ return lm.TrustFundInterestRates(
+ percent_by_year={y: percent for y in range(first, last + 1)},
+ series="INVENTED",
+ provenance={"source": "INVENTED"},
+ )
+
+
+def _careers(histories: dict[int, dict[int, float]]) -> pd.DataFrame:
+ rows = [
+ (pid, year, value)
+ for pid, history in histories.items()
+ for year, value in sorted(history.items())
+ ]
+ return pd.DataFrame(rows, columns=["person_id", "year", "earnings"])
+
+
+def _persons(births: dict[int, int]) -> pd.DataFrame:
+ return pd.DataFrame(
+ {"person_id": list(births), "birth_year": list(births.values())}
+ )
+
+
+def _episodes(rows: list[tuple]) -> pd.DataFrame:
+ """INVENTED marriage episodes: (person, order, start, end, how,
+ spouse)."""
+ columns = [
+ "person_id",
+ "marriage_order",
+ "start_year",
+ "episode_end_year",
+ "how_ended",
+ "spouse_person_id",
+ ]
+ frame = pd.DataFrame(rows, columns=columns)
+ for column in ("start_year", "episode_end_year", "spouse_person_id"):
+ frame[column] = frame[column].astype("Int64")
+ return frame
+
+
+def _couple(start: int = 1950) -> pd.DataFrame:
+ """INVENTED: persons 1 and 2 married to each other from ``start``."""
+ return _episodes(
+ [
+ (1, 1, start, None, "intact", 2),
+ (2, 1, start, None, "intact", 1),
+ ]
+ )
+
+
+def _pv(result: lm.MeasureResult, pid: int) -> float:
+ frame = result.frame.set_index("person_id")
+ return float(frame.at[pid, "pv_at_62"])
+
+
+histories_strategy = st.dictionaries(
+ st.integers(1951, 2030),
+ st.floats(0, 300_000, allow_nan=False, allow_infinity=False),
+ min_size=1,
+ max_size=45,
+)
+
+
+# ---------------------------------------------------------------------------
+# 4. Weighted quintiles
+# ---------------------------------------------------------------------------
+def _reference_ranks(values: list[float], weights: list[float]) -> list[int]:
+ """Brute-force midpoint rule, exact (independent of the module)."""
+ total = sum(Fraction(w) for w in weights)
+ ranks = []
+ for value in values:
+ below = sum(
+ Fraction(w)
+ for v, w in zip(values, weights, strict=True)
+ if v < value
+ )
+ tied = sum(
+ Fraction(w)
+ for v, w in zip(values, weights, strict=True)
+ if v == value
+ )
+ midpoint = below + tied / 2
+ ranks.append(min(4, math.floor(5 * midpoint / total)))
+ return ranks
+
+
+def _rank(label: str) -> int:
+ return 4 - lm.QUINTILE_LABELS.index(label)
+
+
+weights_strategy = st.floats(0.01, 1_000.0, allow_nan=False)
+
+
+@SETTINGS
+@given(
+ st.lists(
+ st.tuples(st.integers(-50, 50).map(float), weights_strategy),
+ min_size=1,
+ max_size=60,
+ )
+)
+def test_quintiles_match_the_brute_force_reference(pairs):
+ values = [v for v, _ in pairs]
+ weights = [w for _, w in pairs]
+ labels = lm.weighted_quintiles(pd.Series(values), pd.Series(weights))
+ assert [_rank(label) for label in labels] == _reference_ranks(
+ values, weights
+ )
+
+
+@SETTINGS
+@given(
+ st.lists(
+ st.floats(-1e6, 1e6, allow_nan=False),
+ min_size=1,
+ max_size=80,
+ unique=True,
+ ),
+ st.data(),
+)
+def test_quintile_shares_are_a_fifth_within_the_largest_weight(values, data):
+ weights = data.draw(
+ st.lists(weights_strategy, min_size=len(values), max_size=len(values))
+ )
+ labels = lm.weighted_quintiles(pd.Series(values), pd.Series(weights))
+ total = math.fsum(weights)
+ bound = max(weights) / total + 1e-12
+ for label in lm.QUINTILE_LABELS:
+ share = (
+ math.fsum(
+ w
+ for w, got in zip(weights, labels, strict=True)
+ if got == label
+ )
+ / total
+ )
+ assert abs(share - 0.2) <= bound
+
+
+@pytest.mark.parametrize("m", [1, 2, 7, 40])
+def test_equal_weights_distinct_values_split_exactly(m):
+ values = pd.Series(
+ random.Random(m).sample(range(10_000), 5 * m), dtype=float
+ )
+ labels = lm.weighted_quintiles(values, pd.Series(np.ones(5 * m)))
+ assert (
+ labels.value_counts().reindex(lm.QUINTILE_LABELS).tolist() == [m] * 5
+ )
+ ordered = labels[values.sort_values().index].tolist()
+ assert ordered == [
+ label for label in reversed(lm.QUINTILE_LABELS) for _ in range(m)
+ ]
+
+
+@SETTINGS
+@given(
+ st.lists(
+ st.tuples(
+ st.integers(0, 30).map(float), st.integers(1, 10_000).map(float)
+ ),
+ min_size=1,
+ max_size=50,
+ ),
+ st.integers(-20, 20),
+ st.integers(1, 1_000),
+)
+def test_quintiles_are_invariant_to_exact_weight_scaling(pairs, power, factor):
+ values = pd.Series([v for v, _ in pairs])
+ weights = pd.Series([w for _, w in pairs])
+ base = lm.weighted_quintiles(values, weights)
+ # Powers of two and integer factors of integer weights are exact.
+ assert lm.weighted_quintiles(values, weights * 2.0**power).equals(base)
+ assert lm.weighted_quintiles(values, weights * float(factor)).equals(base)
+
+
+@SETTINGS
+@given(
+ st.lists(
+ st.tuples(st.integers(0, 20).map(float), weights_strategy),
+ min_size=1,
+ max_size=50,
+ ),
+ st.randoms(use_true_random=False),
+)
+def test_quintiles_are_invariant_to_row_order(pairs, rng):
+ frame = pd.DataFrame(pairs, columns=["value", "weight"])
+ frame["cell"] = [i % 3 for i in range(len(frame))]
+ base = lm.weighted_quintiles(
+ frame["value"], frame["weight"], by=frame["cell"]
+ )
+ order = list(frame.index)
+ rng.shuffle(order)
+ shuffled = frame.loc[order]
+ again = lm.weighted_quintiles(
+ shuffled["value"], shuffled["weight"], by=shuffled["cell"]
+ )
+ assert again.reindex(frame.index).equals(base)
+
+
+@SETTINGS
+@given(
+ st.lists(
+ st.tuples(
+ st.integers(0, 15).map(float),
+ weights_strategy,
+ st.sampled_from(["1950–1959", "1960–1969"]),
+ ),
+ min_size=1,
+ max_size=60,
+ )
+)
+def test_a_higher_value_never_gets_a_lower_quintile_in_its_cell(rows):
+ frame = pd.DataFrame(rows, columns=["value", "weight", "cell"])
+ labels = lm.weighted_quintiles(
+ frame["value"], frame["weight"], by=frame["cell"]
+ )
+ frame["rank"] = [_rank(label) for label in labels]
+ for _, cell in frame.groupby("cell"):
+ ordered = cell.sort_values("value", kind="stable")
+ assert ordered["rank"].is_monotonic_increasing
+ for _, tied in cell.groupby("value"):
+ assert tied["rank"].nunique() == 1
+
+
+@SETTINGS
+@given(
+ st.lists(
+ st.tuples(
+ st.one_of(st.none(), st.integers(0, 40).map(float)),
+ weights_strategy,
+ ),
+ min_size=1,
+ max_size=50,
+ )
+)
+def test_missing_values_and_only_they_are_not_computed(pairs):
+ values = pd.Series([np.nan if v is None else v for v, _ in pairs])
+ weights = pd.Series([w for _, w in pairs])
+ labels = lm.weighted_quintiles(values, weights)
+ assert ((labels == lm.NOT_COMPUTED) == values.isna()).all()
+ present = values.notna()
+ if present.any():
+ alone = lm.weighted_quintiles(values[present], weights[present])
+ assert labels[present].astype(str).equals(alone.astype(str))
+
+
+def test_cells_are_cut_independently():
+ values = pd.Series([1.0, 2.0, 3.0, 4.0, 5.0, 100.0, 200.0])
+ weights = pd.Series(np.ones(7))
+ by = pd.Series(["a"] * 5 + ["b"] * 2)
+ labels = lm.weighted_quintiles(values, weights, by=by)
+ a_only = lm.weighted_quintiles(values[:5], weights[:5])
+ assert labels[:5].astype(str).tolist() == a_only.astype(str).tolist()
+ # Cell b: two rows, midpoints at 1/4 and 3/4 of its weight.
+ assert labels[5:].tolist() == ["Second lowest", "Second highest"]
+
+
+def test_whole_population_and_cohort_scopes():
+ births = pd.Series([1958, 1961, 1969, 1970])
+ assert lm.quintile_cells(births, lm.QuintileScope.WHOLE_POPULATION) is None
+ cells = lm.quintile_cells(births, lm.QuintileScope.TEN_YEAR_BIRTH_COHORT)
+ assert cells.tolist() == [
+ "1950–1959",
+ "1960–1969",
+ "1960–1969",
+ "1970–1979",
+ ]
+
+
+@pytest.mark.parametrize(
+ ("values", "weights", "by", "message"),
+ [
+ ([1.0, 2.0], [1.0, 0.0], None, "positive"),
+ ([1.0, 2.0], [1.0, np.nan], None, "positive"),
+ ([1.0, np.inf], [1.0, 1.0], None, "finite"),
+ ([1.0, 2.0], [1.0, 1.0], [None, "a"], "by is missing"),
+ ],
+)
+def test_quintiles_refuse_bad_inputs(values, weights, by, message):
+ with pytest.raises(ValueError, match=message):
+ lm.weighted_quintiles(
+ pd.Series(values),
+ pd.Series(weights),
+ by=None if by is None else pd.Series(by),
+ )
+
+
+def test_quintiles_refuse_misaligned_indexes_and_bad_labels():
+ with pytest.raises(ValueError, match="share one index"):
+ lm.weighted_quintiles(
+ pd.Series([1.0], index=[0]), pd.Series([1.0], index=[1])
+ )
+ with pytest.raises(ValueError, match="five distinct"):
+ lm.weighted_quintiles(
+ pd.Series([1.0]), pd.Series([1.0]), labels=("a", "b")
+ )
+
+
+def test_quintile_summary_reports_counts_and_shares():
+ values = pd.Series([1.0, 2.0, 3.0, 4.0, 5.0, np.nan])
+ weights = pd.Series([1.0, 1.0, 1.0, 1.0, 1.0, 9.0])
+ labels = lm.weighted_quintiles(values, weights)
+ summary = lm.quintile_summary(labels, weights).set_index("label")
+ assert summary.loc["Highest", "n"] == 1
+ assert summary.loc["Highest", "weight_share"] == pytest.approx(0.2)
+ assert summary.loc[lm.NOT_COMPUTED, "weight"] == 9.0
+ assert math.isnan(summary.loc[lm.NOT_COMPUTED, "weight_share"])
+
+
+# ---------------------------------------------------------------------------
+# 1. Initial AIME at 62
+# ---------------------------------------------------------------------------
+@SETTINGS
+@given(
+ histories_strategy,
+ st.integers(1929, 1985),
+ st.integers(2000, 2060),
+ st.sampled_from(sorted(lm.AIME_CONVENTIONS)),
+)
+def test_aime_equals_the_oracle_on_the_same_cut_history(
+ history, birth, analysis_year, name
+):
+ convention = lm.AIME_CONVENTIONS[name]
+ params = _params()
+ result = lm.initial_aime_at_62(
+ _careers({7: history}),
+ _persons({7: birth}),
+ params,
+ analysis_year=analysis_year,
+ convention=convention,
+ )
+ row = result.frame.iloc[0]
+ cutoff = analysis_year
+ if convention.last_earnings_age is not None:
+ cutoff = min(cutoff, birth + convention.last_earnings_age)
+ kept = {y: e for y, e in history.items() if y <= cutoff}
+ assert row["history_cutoff_year"] == cutoff
+ assert row["n_history_years"] == len(kept)
+ assert bool(row["reaches_62_by_analysis_year"]) == (
+ birth + 62 <= analysis_year
+ )
+ if not kept:
+ # Rows only after the cutoff: unobserved before 62, not zero.
+ assert row["reason"] == "no_career_rows_before_cutoff"
+ return
+ assert row["status"] == lm.COMPUTED
+ assert row["first_earnings_year"] == min(kept)
+ assert bool(row["history_starts_after_age_22"]) == (min(kept) > birth + 22)
+ assert row["aime"] == statutory_aime.oracle_aime(
+ kept, birth, params, computation_years=convention.computation_years
+ )
+
+
+@SETTINGS
+@given(histories_strategy, st.integers(1929, 1985))
+def test_legacy_and_statutory_agree_for_births_from_1929(history, birth):
+ """Two implementations of one semantics: 415(b)(2) gives 35 years to
+ every worker born 1929 or later who is alive at 62."""
+ frames = [
+ lm.initial_aime_at_62(
+ _careers({1: history}),
+ _persons({1: birth}),
+ _params(),
+ analysis_year=2100,
+ convention=lm.AimeConvention(
+ name="t",
+ computation_years=years,
+ last_earnings_age=61,
+ source="",
+ ),
+ ).frame
+ for years in ComputationYears
+ ]
+ pd.testing.assert_series_equal(frames[0]["aime"], frames[1]["aime"])
+ assert frames[0]["status"].tolist() == frames[1]["status"].tolist()
+
+
+@SETTINGS
+@given(histories_strategy, st.integers(1913, 1985))
+def test_exercise_4_convention_equals_track_m_history_pia(history, birth):
+ """Differential against Track M's own old-age PIA record (rules.py),
+ for an entitlement in the year of attaining 62."""
+ assume(any(year <= birth + 61 for year in history))
+ params = _params()
+ record = track_m_rules.history_pia(
+ history,
+ birth_year=birth,
+ params=params,
+ basis=track_m_rules.BASIS_OLD_AGE,
+ window_year=birth + 62,
+ )
+ row = lm.initial_aime_at_62(
+ _careers({1: history}),
+ _persons({1: birth}),
+ params,
+ analysis_year=birth + 62,
+ convention=lm.AIME_CONVENTIONS["exercise_4_min_benefit"],
+ ).frame.iloc[0]
+ assert row["aime"] == record.aime
+
+
+@SETTINGS
+@given(histories_strategy, st.integers(1913, 1985), st.integers(2000, 2060))
+def test_exercise_1_convention_reproduces_the_track_a_age_62_pia(
+ history, birth, analysis_year
+):
+ """Differential against the Track A calculator's age-62 PIA call
+ (scenario_benefits.eligibility_pia_for_clock with Track A's
+ computation years over the history as supplied)."""
+ params = _params()
+ supplied = {y: e for y, e in history.items() if y <= analysis_year}
+ assume(supplied)
+ expected = sb.eligibility_pia_for_clock(
+ sb.WorkerClock.at_age_62(birth),
+ history=supplied,
+ birth_year=birth,
+ params=params,
+ computation_years=track_a_benefits.TRACK_A_COMPUTATION_YEARS,
+ )
+ row = lm.initial_aime_at_62(
+ _careers({1: history}),
+ _persons({1: birth}),
+ params,
+ analysis_year=analysis_year,
+ convention=lm.AIME_CONVENTIONS["exercise_1_cola"],
+ ).frame.iloc[0]
+ assert benefits.pia(row["aime"], birth + 62, params) == expected
+
+
+def test_named_conventions_match_the_blind_test_code():
+ conventions = lm.AIME_CONVENTIONS
+ assert (
+ conventions["exercise_1_cola"].computation_years
+ is track_a_benefits.TRACK_A_COMPUTATION_YEARS
+ )
+ assert (
+ conventions["exercise_3_fra68"].computation_years.value
+ == fra68_config.MAX_RULINGS["benefit_computation_years"]["ruling"]
+ )
+ assert conventions["exercise_4_min_benefit"].computation_years is (
+ ComputationYears.STATUTORY
+ )
+ assert conventions["mint8_initial_aime"].last_earnings_age == 61
+
+
+def test_aime_example_and_flags():
+ # INVENTED: flat wage index (indexing neutral), no binding base.
+ params = SSAParameters(
+ nawi={y: 10_000.0 for y in range(1951, 2071)},
+ wage_base={1937: 1.0e12},
+ pia_factors=(0.9, 0.32, 0.15),
+ fra_months_by_birth_year=[(1900, 792)],
+ early_monthly_rates=(5 / 900, 5 / 1200),
+ early_first_bracket_months=36,
+ pe_us_revision="INVENTED",
+ )
+ careers = _careers(
+ {
+ # 35 years of 42,000 before 62: AIME 42,000 / 12 = 3,500.
+ 1: {y: 42_000.0 for y in range(1975, 2010)} | {2012: 9e9},
+ # Younger than 62 in 2030: provisional, 20 years of 42,000
+ # over 35 years: floor(840,000 / 420) = 2,000.
+ 2: {y: 42_000.0 for y in range(1990, 2010)},
+ # Years before 1951 are dropped and counted.
+ 3: {1949: 5.0, 1950: 5.0, 1980: 4_200.0},
+ }
+ )
+ persons = _persons({1: 1950, 2: 1975, 3: 1925, 4: 1960, 5: 1910})
+ frame = lm.initial_aime_at_62(
+ careers,
+ persons,
+ params,
+ analysis_year=2030,
+ convention=lm.AIME_CONVENTIONS["mint8_initial_aime"],
+ ).frame.set_index("person_id")
+ assert frame.at[1, "aime"] == 3_500.0
+ assert frame.at[1, "basis"] == lm.AimeBasis.AGE_62.value
+ assert frame.at[1, "history_cutoff_year"] == 2011
+ assert frame.at[2, "aime"] == 2_000.0
+ assert frame.at[2, "basis"] == (
+ lm.AimeBasis.PROVISIONAL_THROUGH_LAST_OBSERVED.value
+ )
+ assert bool(frame.at[2, "history_ends_before_cutoff"])
+ # Born 1925: statutory 415(b)(2) gives 1951-1986 elapsed less 5 = 31.
+ assert frame.at[3, "years_before_1951_dropped"] == 2
+ assert frame.at[3, "computation_years"] == 31
+ assert frame.at[3, "aime"] == math.floor(4_200.0 / (31 * 12))
+ assert frame.at[4, "reason"] == "no_career_rows"
+ assert frame.at[5, "status"] == lm.NOT_COMPUTED
+
+
+def test_aime_refuses_pre_1975_and_missing_nawi_with_reasons():
+ params = _params()
+ careers = _careers({1: {1960: 1_000.0}, 2: {2000: 1_000.0}})
+ frame = lm.initial_aime_at_62(
+ careers,
+ _persons({1: 1910, 2: 2010}),
+ params,
+ analysis_year=2090,
+ convention=lm.AIME_CONVENTIONS["mint8_initial_aime"],
+ ).frame.set_index("person_id")
+ assert frame.at[1, "reason"] == (
+ "statutory_oracle_refuses_attaining_62_before_1975"
+ )
+ # Born 2010: indexing year 2070 is in the series, but 2072 is not.
+ assert frame.at[2, "status"] == lm.COMPUTED
+ frame = lm.initial_aime_at_62(
+ careers,
+ _persons({2: 2012}),
+ params,
+ analysis_year=2090,
+ convention=lm.AIME_CONVENTIONS["mint8_initial_aime"],
+ ).frame
+ assert frame.at[0, "reason"] == "nawi_unavailable:2072"
+
+
+@pytest.mark.parametrize(
+ ("careers", "message"),
+ [
+ (
+ pd.DataFrame(
+ {"person_id": [1, 1], "year": [2000, 2000], "earnings": [1, 2]}
+ ),
+ "duplicate",
+ ),
+ (
+ pd.DataFrame({"person_id": [1], "year": [2000], "earnings": [-1]}),
+ "non-negative",
+ ),
+ (
+ pd.DataFrame(
+ {"person_id": [1], "year": [2000.5], "earnings": [1]}
+ ),
+ "integers",
+ ),
+ (pd.DataFrame({"person_id": [1], "year": [2000]}), "lacks"),
+ ],
+)
+def test_measures_refuse_malformed_careers(careers, message):
+ with pytest.raises(ValueError, match=message):
+ lm.initial_aime_at_62(
+ careers,
+ _persons({1: 1950}),
+ _params(),
+ analysis_year=2030,
+ convention=lm.AIME_CONVENTIONS["mint8_initial_aime"],
+ )
+
+
+def test_measures_refuse_malformed_persons():
+ careers = _careers({1: {2000: 1.0}})
+ with pytest.raises(ValueError, match="duplicate person_id"):
+ lm.initial_aime_at_62(
+ careers,
+ pd.DataFrame({"person_id": [1, 1], "birth_year": [1950, 1950]}),
+ _params(),
+ analysis_year=2030,
+ convention=lm.AIME_CONVENTIONS["mint8_initial_aime"],
+ )
+ with pytest.raises(ValueError, match="missing"):
+ lm.initial_aime_at_62(
+ careers,
+ pd.DataFrame({"person_id": [1], "birth_year": [None]}),
+ _params(),
+ analysis_year=2030,
+ convention=lm.AIME_CONVENTIONS["mint8_initial_aime"],
+ )
+ with pytest.raises(TypeError, match="AimeConvention"):
+ lm.initial_aime_at_62(
+ careers,
+ _persons({1: 1950}),
+ _params(),
+ analysis_year=2030,
+ convention="mint8_initial_aime",
+ )
+
+
+# ---------------------------------------------------------------------------
+# 2. Lifetime payroll tax at 62
+# ---------------------------------------------------------------------------
+def _reference_pv(
+ history: dict[int, float],
+ birth: int,
+ params: SSAParameters,
+ rate: float,
+ interest: float,
+) -> float:
+ """Independent reference: constant rates, closed-form factors."""
+ reference = birth + 62
+ return math.fsum(
+ min(e, params.wage_base_for(y))
+ * rate
+ * (1.0 + interest) ** (reference - y)
+ for y, e in history.items()
+ )
+
+
+@SETTINGS
+@given(
+ histories_strategy,
+ st.integers(1920, 1990),
+ st.floats(1.0, 20.0),
+ st.floats(0.0, 10.0),
+)
+def test_pv_matches_the_closed_form_reference(history, birth, rate, interest):
+ params = _params()
+ result = lm.lifetime_payroll_tax_pv_at_62(
+ _careers({1: history}),
+ _persons({1: birth}),
+ params,
+ shared=False,
+ rates=_rates(rate),
+ interest=_interest(interest),
+ )
+ expected = _reference_pv(
+ history, birth, params, rate / 100.0, interest / 100.0
+ )
+ assert _pv(result, 1) == pytest.approx(expected, rel=1e-9, abs=1e-9)
+
+
+@SETTINGS
+@given(
+ st.dictionaries(
+ st.integers(1951, 2030),
+ st.floats(0, 1e6, allow_nan=False),
+ min_size=1,
+ max_size=40,
+ ),
+ st.integers(1920, 1990),
+)
+def test_pv_is_homogeneous_below_the_cap(history, birth):
+ params = _params(cap=1.0e12)
+ kwargs = {
+ "shared": False,
+ "rates": _rates(12.4),
+ "interest": _interest(3.0),
+ }
+ persons = _persons({1: birth})
+ once = lm.lifetime_payroll_tax_pv_at_62(
+ _careers({1: history}), persons, params, **kwargs
+ )
+ twice = lm.lifetime_payroll_tax_pv_at_62(
+ _careers({1: {y: 2.0 * e for y, e in history.items()}}),
+ persons,
+ params,
+ **kwargs,
+ )
+ zero = lm.lifetime_payroll_tax_pv_at_62(
+ _careers({1: dict.fromkeys(history, 0.0)}), persons, params, **kwargs
+ )
+ assert _pv(twice, 1) == 2.0 * _pv(once, 1)
+ assert _pv(zero, 1) == 0.0
+ double_rate = lm.lifetime_payroll_tax_pv_at_62(
+ _careers({1: history}),
+ persons,
+ params,
+ shared=False,
+ rates=_rates(24.8),
+ interest=_interest(3.0),
+ )
+ assert _pv(double_rate, 1) == 2.0 * _pv(once, 1)
+
+
+@SETTINGS
+@given(
+ st.dictionaries(
+ st.integers(1951, 2030),
+ st.floats(0, 1e6, allow_nan=False),
+ min_size=1,
+ max_size=40,
+ ),
+ st.integers(1920, 1990),
+)
+def test_earnings_above_the_cap_do_not_change_pv(extra, birth):
+ params = _params(cap=50_000.0)
+ persons = _persons({1: birth})
+ at_cap = {y: 50_000.0 for y in extra}
+ above = {y: 50_000.0 + e for y, e in extra.items()}
+ kwargs = {
+ "shared": False,
+ "rates": _rates(12.4),
+ "interest": _interest(4.0),
+ }
+ assert _pv(
+ lm.lifetime_payroll_tax_pv_at_62(
+ _careers({1: at_cap}), persons, params, **kwargs
+ ),
+ 1,
+ ) == _pv(
+ lm.lifetime_payroll_tax_pv_at_62(
+ _careers({1: above}), persons, params, **kwargs
+ ),
+ 1,
+ )
+
+
+@SETTINGS
+@given(histories_strategy, histories_strategy, st.integers(1920, 1985))
+def test_shared_pvs_of_a_couple_sum_to_their_own_pvs(a, b, birth):
+ """A couple with one birth year, married to each other in every year
+ either paid tax: the shared PVs sum to the own PVs."""
+ params = _params()
+ careers = _careers({1: a, 2: b})
+ persons = _persons({1: birth, 2: birth})
+ common = {"rates": _rates(12.4), "interest": _interest(4.0)}
+ own = lm.lifetime_payroll_tax_pv_at_62(
+ careers, persons, params, shared=False, **common
+ )
+ shared = lm.lifetime_payroll_tax_pv_at_62(
+ careers,
+ persons,
+ params,
+ shared=True,
+ marriage_episodes=_couple(1940),
+ **common,
+ )
+ assert _pv(shared, 1) + _pv(shared, 2) == pytest.approx(
+ _pv(own, 1) + _pv(own, 2), rel=1e-12, abs=1e-6
+ )
+ # Each spouse holds half the couple's taxes in every shared year.
+ assert _pv(shared, 1) == pytest.approx(_pv(shared, 2), rel=1e-12)
+ years = set(a) | set(b)
+ assert shared.frame["married_years_shared"].tolist() == [len(years)] * 2
+ # Person 1's spouse lacks a row in the years only person 1 worked.
+ assert shared.frame["married_years_spouse_year_absent"].tolist() == [
+ len(years - set(b)),
+ len(years - set(a)),
+ ]
+
+
+@SETTINGS
+@given(histories_strategy, st.integers(1920, 1985))
+def test_never_married_shared_equals_own(history, birth):
+ params = _params()
+ common = {"rates": _rates(12.4), "interest": _interest(4.0)}
+ careers = _careers({1: history})
+ persons = _persons({1: birth})
+ own = lm.lifetime_payroll_tax_pv_at_62(
+ careers, persons, params, shared=False, **common
+ )
+ shared = lm.lifetime_payroll_tax_pv_at_62(
+ careers,
+ persons,
+ params,
+ shared=True,
+ marriage_episodes=_episodes([]),
+ **common,
+ )
+ assert _pv(shared, 1) == _pv(own, 1)
+ assert shared.frame.at[0, "married_years_shared"] == 0
+
+
+def test_shared_example_by_hand():
+ # INVENTED: rate 10 percent, interest 0, no binding base. Person 1
+ # (born 1950) earns 100 in 1980-1989; person 2 earns 300 in 1985-1994.
+ # They marry in 1985 and divorce in 1990 (married at the end of
+ # 1985-1989). Own taxes: 1 -> 100, 2 -> 300. Person 1's shared
+ # taxes: 1980-84 own 10 each = 50; 1985-89 (10 + 30) / 2 = 20 each =
+ # 100: total 150. Person 2's: 1985-89 20 each = 100; 1990-94 own 30
+ # each = 150: total 250.
+ params = _params(cap=1.0e12)
+ careers = _careers(
+ {
+ 1: {y: 100.0 for y in range(1980, 1990)},
+ 2: {y: 300.0 for y in range(1985, 1995)},
+ }
+ )
+ episodes = _episodes(
+ [
+ (1, 1, 1985, 1990, "divorce", 2),
+ (2, 1, 1985, 1990, "divorce", 1),
+ ]
+ )
+ result = lm.lifetime_payroll_tax_pv_at_62(
+ careers,
+ _persons({1: 1950, 2: 1950}),
+ params,
+ shared=True,
+ rates=_rates(10.0),
+ interest=_interest(0.0),
+ marriage_episodes=episodes,
+ )
+ assert _pv(result, 1) == pytest.approx(150.0)
+ assert _pv(result, 2) == pytest.approx(250.0)
+ assert result.frame["married_years_shared"].tolist() == [5, 5]
+ assert result.frame["married_years_spouse_record_disagrees"].sum() == 0
+
+
+def test_missing_spouse_is_counted_or_refused():
+ params = _params(cap=1.0e12)
+ careers = _careers({1: {y: 100.0 for y in range(1980, 1990)}})
+ # Spouse 9 has no career; spouse NA is not a PSID person.
+ episodes = _episodes(
+ [
+ (1, 1, 1979, 1985, "divorce", 9),
+ (1, 2, 1986, None, "intact", None),
+ ]
+ )
+ common = {
+ "shared": True,
+ "rates": _rates(10.0),
+ "interest": _interest(0.0),
+ "marriage_episodes": episodes,
+ }
+ own_only = lm.lifetime_payroll_tax_pv_at_62(
+ careers, _persons({1: 1950}), params, **common
+ )
+ assert _pv(own_only, 1) == pytest.approx(100.0)
+ assert own_only.frame.at[0, "married_years_spouse_unavailable"] == 9
+ refused = lm.lifetime_payroll_tax_pv_at_62(
+ careers,
+ _persons({1: 1950}),
+ params,
+ missing_spouse=lm.MissingSpousePolicy.NOT_COMPUTED,
+ **common,
+ )
+ assert refused.frame.at[0, "reason"] == "spouse_career_unavailable"
+
+
+def test_disagreeing_spouse_records_are_counted():
+ params = _params(cap=1.0e12)
+ careers = _careers({1: {2000: 100.0}, 2: {2000: 300.0}})
+ episodes = _episodes([(1, 1, 1990, None, "intact", 2)])
+ episodes = pd.concat(
+ [episodes, _episodes([(2, 1, 1970, 1980, "divorce", 5)])]
+ )
+ result = lm.lifetime_payroll_tax_pv_at_62(
+ careers,
+ _persons({1: 1950, 2: 1950}),
+ params,
+ shared=True,
+ rates=_rates(10.0),
+ interest=_interest(0.0),
+ marriage_episodes=episodes,
+ marriage_history_person_ids=[1],
+ )
+ frame = result.frame.set_index("person_id")
+ assert frame.at[1, "married_years_spouse_record_disagrees"] == 1
+ assert _pv(result, 1) == pytest.approx(20.0)
+ assert _pv(result, 2) == pytest.approx(30.0)
+ assert not bool(frame.at[1, "marriage_history_absent"])
+ assert bool(frame.at[2, "marriage_history_absent"])
+
+
+def test_interest_gaps_refuse_or_mark_and_extension_is_named():
+ params = _params()
+ careers = _careers({1: {2000: 1_000.0}})
+ persons = _persons({1: 1970}) # reference year 2032
+ interest = _interest(3.0, first=1990, last=2025)
+ with pytest.raises(ValueError, match="2026-2032"):
+ lm.lifetime_payroll_tax_pv_at_62(
+ careers,
+ persons,
+ params,
+ shared=False,
+ rates=_rates(),
+ interest=interest,
+ )
+ marked = lm.lifetime_payroll_tax_pv_at_62(
+ careers,
+ persons,
+ params,
+ shared=False,
+ rates=_rates(),
+ interest=interest,
+ missing_rate=lm.MissingRatePolicy.NOT_COMPUTED,
+ )
+ assert marked.frame.at[0, "reason"] == "interest_rate_unavailable:2026"
+ extended = interest.extended(
+ dict.fromkeys(range(2026, 2033), 3.0), source="INVENTED test path"
+ )
+ computed = lm.lifetime_payroll_tax_pv_at_62(
+ careers,
+ persons,
+ params,
+ shared=False,
+ rates=_rates(),
+ interest=extended,
+ )
+ assert _pv(computed, 1) == pytest.approx(
+ 1_000.0 * 0.12 * 1.03**32, rel=1e-12
+ )
+ record = computed.provenance["interest_rates"]
+ assert record["assumed_years"] == list(range(2026, 2033))
+ assert record["assumption_source"] == "INVENTED test path"
+ with pytest.raises(ValueError, match="already have rates"):
+ interest.extended({2000: 1.0}, source="x")
+ with pytest.raises(ValueError, match="named source"):
+ interest.extended({2030: 1.0}, source=" ")
+
+
+@SETTINGS
+@given(st.integers(1940, 2100), st.integers(1940, 2100))
+def test_accumulation_factors_compose_and_invert(t, y):
+ interest = _interest(4.0, first=1940, last=2100)
+ forward = lm.accumulation_factor(t, y, interest)
+ backward = lm.accumulation_factor(y, t, interest)
+ assert forward * backward == pytest.approx(1.0, rel=1e-12)
+ assert forward == pytest.approx(1.04 ** (y - t), rel=1e-12)
+
+
+def test_tax_rate_bases_from_an_invented_document():
+ document = {
+ "rows": [
+ {
+ "first_year": 1937,
+ "last_year": 1983,
+ "employee_employer_each": {"total": 5.0},
+ },
+ {
+ "first_year": 1984,
+ "last_year": 1989,
+ "employee_employer_each": {"total": 6.0},
+ },
+ {
+ "first_year": 1990,
+ "last_year": None,
+ "employee_employer_each": {"total": 7.0},
+ },
+ ],
+ "paid_rate_adjustments": [
+ {
+ "year": 1984,
+ "employee_effective_rate": 5.5,
+ "employer_rate": 6.0,
+ }
+ ],
+ }
+ received = lm.OASDITaxRates.from_document(document)
+ paid = lm.OASDITaxRates.from_document(
+ document, basis=lm.TaxRateBasis.EMPLOYEE_EMPLOYER_PAID
+ )
+ assert received.combined_percent_for(1984) == 12.0
+ assert paid.combined_percent_for(1984) == 11.5
+ assert paid.combined_percent_for(1985) == 12.0
+ assert received.combined_percent_for(2050) == 14.0
+ with pytest.raises(KeyError):
+ received.combined_for(1936)
+
+
+def test_annual_taxes_cap_each_year_at_its_base():
+ params = _params()
+ taxes = lm.annual_payroll_taxes(
+ _careers({1: {1980: 50_000.0, 1995: 50_000.0}}), params, _rates(10.0)
+ )
+ assert taxes["taxable_earnings"].tolist() == [14_100.0, 50_000.0]
+ assert taxes["tax"].tolist() == pytest.approx([1_410.0, 5_000.0])
+
+
+# ---------------------------------------------------------------------------
+# 3. The Report's average indexed earnings at ages 22-62
+# ---------------------------------------------------------------------------
+@SETTINGS
+@given(histories_strategy, st.integers(1930, 1990))
+def test_report_average_matches_the_oracle_indexing(history, birth):
+ """Differential against ss.benefits.indexed_history (uncapped)."""
+ params = _params()
+ window = {
+ y: e for y, e in history.items() if birth + 22 <= y <= birth + 62
+ }
+ frame = lm.report_average_indexed_earnings_22_62(
+ _careers({1: history}), _persons({1: birth}), params, shared=False
+ ).frame
+ if not window:
+ assert frame.at[0, "reason"] == "no_career_rows_at_ages_22_62"
+ return
+ indexed = benefits.indexed_history(window, birth, params)
+ expected = math.fsum(indexed[y] for y in sorted(indexed)) / len(window)
+ assert frame.at[0, "average_indexed_earnings"] == expected
+ assert frame.at[0, "n_ages_covered"] == len(window)
+ assert bool(frame.at[0, "ages_complete"]) == (len(window) == 41)
+ all_ages = lm.report_average_indexed_earnings_22_62(
+ _careers({1: history}),
+ _persons({1: birth}),
+ params,
+ shared=False,
+ conventions=lm.ReportEarningsConventions(
+ divisor=lm.AverageDivisor.ALL_AGES
+ ),
+ ).frame
+ assert all_ages.at[0, "average_indexed_earnings"] == pytest.approx(
+ expected * len(window) / 41, rel=1e-12
+ )
+
+
+def test_report_average_example_and_cap_option():
+ # INVENTED: flat NAWI, so indexing is neutral. Born 1950: ages 22-62
+ # are 1972-2012. Rows in 1971 (age 21) and 2013 (age 63) fall out.
+ params = SSAParameters(
+ nawi={y: 10_000.0 for y in range(1951, 2071)},
+ wage_base={1937: 60_000.0},
+ pia_factors=(0.9, 0.32, 0.15),
+ fra_months_by_birth_year=[(1900, 792)],
+ early_monthly_rates=(5 / 900, 5 / 1200),
+ early_first_bracket_months=36,
+ pe_us_revision="INVENTED",
+ )
+ careers = _careers(
+ {1: {1971: 1e9, 1980: 100_000.0, 1990: 20_000.0, 2013: 1e9}}
+ )
+ own = lm.report_average_indexed_earnings_22_62(
+ careers, _persons({1: 1950}), params, shared=False
+ ).frame
+ assert own.at[0, "average_indexed_earnings"] == 60_000.0
+ capped = lm.report_average_indexed_earnings_22_62(
+ careers,
+ _persons({1: 1950}),
+ params,
+ shared=False,
+ conventions=lm.ReportEarningsConventions(cap_at_taxable_maximum=True),
+ ).frame
+ assert capped.at[0, "average_indexed_earnings"] == 40_000.0
+
+
+@SETTINGS
+@given(histories_strategy, histories_strategy, st.integers(1930, 1980))
+def test_report_shared_couple_conserves_the_sum_over_common_years(a, b, birth):
+ """With identical covered years, the couple's shared averages sum to
+ their own averages."""
+ years = sorted(set(a) | set(b))
+ a = {y: a.get(y, 0.0) for y in years}
+ b = {y: b.get(y, 0.0) for y in years}
+ params = _params()
+ careers = _careers({1: a, 2: b})
+ persons = _persons({1: birth, 2: birth})
+ own = lm.report_average_indexed_earnings_22_62(
+ careers, persons, params, shared=False
+ ).frame
+ shared = lm.report_average_indexed_earnings_22_62(
+ careers,
+ persons,
+ params,
+ shared=True,
+ marriage_episodes=_couple(1930),
+ ).frame
+ assert shared["status"].tolist() == own["status"].tolist()
+ if (own["status"] == lm.COMPUTED).all():
+ assert shared["average_indexed_earnings"].sum() == pytest.approx(
+ own["average_indexed_earnings"].sum(), rel=1e-12
+ )
+
+
+def test_report_conventions_are_recorded_or_registered():
+ provenance = lm.report_average_indexed_earnings_22_62(
+ _careers({1: {2000: 1.0}}),
+ _persons({1: 1960}),
+ _params(),
+ shared=False,
+ ).provenance
+ assert "ages 22-62" in provenance["conventions"]["recorded"]["measure"]
+ defaults = provenance["conventions"]["builder_defaults"]
+ for key in ("wage_index", "divisor", "shared_rule", "index_age"):
+ assert key in defaults
+ assert provenance["conventions"]["conventions"]["divisor"] == (
+ "covered_ages"
+ )
+
+
+# ---------------------------------------------------------------------------
+# Schemes and provenance
+# ---------------------------------------------------------------------------
+def test_boomers2004_scheme_is_unregistered_and_refuses():
+ from populace_dynamics.estimates.uniform_cut_tabulation import (
+ NOT_COMPUTED_REPORT_ROWS,
+ )
+
+ scheme = lm.boomers2004_scheme()
+ assert not scheme.registered
+ for dimension in scheme.dimensions:
+ record = NOT_COMPUTED_REPORT_ROWS[dimension.key]
+ assert dimension.section == record["section"]
+ with pytest.raises(ValueError, match="highest quintile"):
+ dimension.labels()
+ with pytest.raises(ValueError, match="population"):
+ dimension.cells(pd.Series([1940]))
+
+
+def test_results_carry_provenance_and_output_digest():
+ result = lm.lifetime_payroll_tax_pv_at_62(
+ _careers({1: {2000: 1_000.0}}),
+ _persons({1: 1950, 2: 1960}),
+ _params(),
+ shared=False,
+ rates=_rates(),
+ interest=_interest(),
+ )
+ record = result.provenance
+ assert record["schema_version"] == lm.SCHEMA_VERSION
+ assert record["status_counts"] == {"computed": 1, "not computed": 1}
+ assert record["reason_counts"] == {"no_career_rows": 1, "none": 1}
+ assert len(record["output_sha256"]) == 64
+ assert record["conventions"]["builder_defaults"] == (
+ lm.PAYROLL_TAX_BUILDER_DEFAULTS
+ )
+ again = lm.lifetime_payroll_tax_pv_at_62(
+ _careers({1: {2000: 1_000.0}}),
+ _persons({1: 1950, 2: 1960}),
+ _params(),
+ shared=False,
+ rates=_rates(),
+ interest=_interest(),
+ )
+ assert again.provenance == record
+
+
+def test_shared_measures_need_marriage_episodes():
+ with pytest.raises(ValueError, match="marriage_episodes"):
+ lm.lifetime_payroll_tax_pv_at_62(
+ _careers({1: {2000: 1.0}}),
+ _persons({1: 1950}),
+ _params(),
+ shared=True,
+ rates=_rates(),
+ interest=_interest(),
+ )
+ with pytest.raises(ValueError, match="marriage_episodes"):
+ lm.report_average_indexed_earnings_22_62(
+ _careers({1: {2000: 1.0}}),
+ _persons({1: 1950}),
+ _params(),
+ shared=True,
+ )
From 2b8b42ae28bf3a55933644c5e2b66be7c58dce5c Mon Sep 17 00:00:00 2001
From: Max Ghenis
Date: Fri, 2 Oct 2026 16:52:46 -0400
Subject: [PATCH 03/10] Add lifetime-earnings measures for the MINT breakdowns
estimates/lifetime_measures.py computes, from cohort outputs only:
current-law AIME at age 62 (MINT8's initial AIME), the present value at
62 of OASDI payroll taxes, own and shared between spouses for married
years (MINT8's lifetime payroll tax measures), the Butrica-Uccello
average of wage-indexed earnings at ages 22-62, and weighted quintiles
with MINT's labels, cut per birth cohort or over the whole population.
OASDI tax rates and trust fund effective interest rates are captured
from SSA with SHA-256 pins. Unrecorded conventions are exposed as named
builder defaults for the registration to fix.
Invariants tested: quintile weight bounds and monotonicity; invariance
to weight scaling and row order; zero earnings give zero tax; tax is
linear below the cap and flat above it; sharing conserves a couple's
total; AIME agrees with the ss oracle.
Co-Authored-By: Claude Opus 5.5
---
.../lifetime_measure_sources.provenance.json | 29 +-
.../mint8_lifetime_quintile_definitions.json | 4 +-
.../mint8_row_categories_2026.source.json | 1851 +++++++++++++++++
docs/design/lifetime_measures.md | 159 ++
scripts/extract_lifetime_measure_sources.py | 177 +-
.../estimates/lifetime_measures.py | 256 ++-
.../test_lifetime_measure_sources.py | 124 ++
tests/estimates/test_lifetime_measures.py | 9 +-
.../test_lifetime_measures_properties.py | 689 ++++++
9 files changed, 3244 insertions(+), 54 deletions(-)
create mode 100644 data/external/mint8_row_categories_2026.source.json
create mode 100644 docs/design/lifetime_measures.md
create mode 100644 tests/estimates/test_lifetime_measures_properties.py
diff --git a/data/external/lifetime_measure_sources.provenance.json b/data/external/lifetime_measure_sources.provenance.json
index d222d4a3..f5e21cf9 100644
--- a/data/external/lifetime_measure_sources.provenance.json
+++ b/data/external/lifetime_measure_sources.provenance.json
@@ -10,35 +10,54 @@
"source_url": "https://www.ssa.gov/oact/ProgData/annualinterestrates.html",
"document": "Average and Effective Interest Rates",
"source_sha256": "7c11df685faaf0602dd77c5e4e124d64709e8cffd58e463d938f35e7254a0d18",
- "source_length_bytes": 41498
+ "source_length_bytes": 41498,
+ "locator": "table caption: Average annual special-issue interest rates on new issues and effective annual interest rates (percent); Year row; Average and Effective columns; definition paragraphs above the table",
+ "acquisition": "Direct HTTPS GET of the live ssa.gov page with curl -A 'Wget/1.21.4' on 2026-10-01 (HTTP 200, content-type text/html; charset=UTF-8; response Date header Thu, 01 Oct 2026 22:31:33-34 GMT). The body is committed byte for byte. A per-request Akamai mPulse script in changes between fetches; an earlier fetch the same day parsed to the identical tables."
},
"effective_rates_1980_on": {
"committed_source_file": "data/external/ssa_effective_interest_rates_1980_2025.source.html",
"source_url": "https://www.ssa.gov/oact/ProgData/effectiveRates.html",
"document": "Effective Interest Rates",
"source_sha256": "eaa8870898da2907bbc39da55d6a4c14f349e38c518aefb0ff87be2ddae09322",
- "source_length_bytes": 39756
+ "source_length_bytes": 39756,
+ "locator": "table caption: Effective Interest Rates Earned By the Invested Assets of the OASI and DI Trust Funds [Percent]; Calendar year row; OASI, DI and OASDI columns",
+ "acquisition": "Direct HTTPS GET of the live ssa.gov page with curl -A 'Wget/1.21.4' on 2026-10-01 (HTTP 200, content-type text/html; charset=UTF-8; response Date header Thu, 01 Oct 2026 22:31:33-34 GMT). The body is committed byte for byte. A per-request Akamai mPulse script in changes between fetches; an earlier fetch the same day parsed to the identical tables."
},
"effective_rates_1940_1979": {
"committed_source_file": "data/external/ssa_effective_interest_rates_1940_1979.source.html",
"source_url": "https://www.ssa.gov/oact/ProgData/effectiveRts1940-79.html",
"document": "Effective Interest Rates, 1940-79 (historical document)",
"source_sha256": "fb16fdaff44b3427eec2216740734792aab3b94e2851acde57c26e4057094e3d",
- "source_length_bytes": 19037
+ "source_length_bytes": 19037,
+ "locator": "table caption: Estimated Effective Interest Rates Earned By the Assets of the OASI and DI Trust Funds, 1940-79[Percent]; Calendar year row in either panel; OASI, DI and OASDI columns",
+ "acquisition": "Direct HTTPS GET of the live ssa.gov page with curl -A 'Wget/1.21.4' on 2026-10-01 (HTTP 200, content-type text/html; charset=UTF-8; response Date header Thu, 01 Oct 2026 22:31:33-34 GMT). The body is committed byte for byte. A per-request Akamai mPulse script in changes between fetches; an earlier fetch the same day parsed to the identical tables."
},
"oasdi_tax_rates": {
"committed_source_file": "data/external/ssa_oasdi_tax_rates_2026.source.html",
"source_url": "https://www.ssa.gov/oact/ProgData/oasdiRates.html",
"document": "Social Security Tax Rates",
"source_sha256": "7aad3e899ab2ee4e2dfd91c2dc9208c79e40bf4cb7f3315867a11ba7efd8a9a4",
- "source_length_bytes": 41837
+ "source_length_bytes": 41837,
+ "locator": "table summary: Tax rate table for Social Security trust funds; Calendar years period row; OASI, DI and Total columns under employee/employer-each and self-employed headers; footnote anchors fna-fnd",
+ "acquisition": "Direct HTTPS GET of the live ssa.gov page with curl -A 'Wget/1.21.4' on 2026-10-01 (HTTP 200, content-type text/html; charset=UTF-8; response Date header Thu, 01 Oct 2026 22:31:33-34 GMT). The body is committed byte for byte. A per-request Akamai mPulse script in changes between fetches; an earlier fetch the same day parsed to the identical tables."
},
"mint8_user_guide": {
"committed_source_file": "data/external/ssa_mint8_user_guide_2026.source.html",
"source_url": "https://www.ssa.gov/policy/docs/projections/user-guide.html",
"document": "MINT8 Table User Guide",
"source_sha256": "8d5bc3f0de17831c2ed07383003257a63bb657df179d0c54868fda052f23f100",
- "source_length_bytes": 73967
+ "source_length_bytes": 73967,
+ "locator": "meta DCTERMS:dateCertified; paragraphs id=AIME, lifetime-tax, lifetime-tax-shared and taxes; other paragraphs identified by the exact sentences in MINT8_PARAGRAPHS",
+ "acquisition": "Direct HTTPS GET of the live ssa.gov page with curl -A 'Wget/1.21.4' on 2026-10-01 (HTTP 200, content-type text/html; charset=UTF-8; response Date header Thu, 01 Oct 2026 22:31:33-34 GMT). The body is committed byte for byte. A per-request Akamai mPulse script in changes between fetches; an earlier fetch the same day parsed to the identical tables."
+ },
+ "mint8_table_row_labels": {
+ "committed_source_file": "data/external/mint8_row_categories_2026.source.json",
+ "source_url": "https://www.ssa.gov/policy/docs/projections/policy-options/increase-payroll-tax-rate.html",
+ "document": "MINT8 payroll-tax option table labels (labels only)",
+ "source_sha256": "23fbfbc8dbc14b06144a83bf0583ea0a6c89505638f3f7dc5fd6b34a3a4ac650",
+ "source_length_bytes": 37524,
+ "locator": "label-only JSON: dateCertified; tables.13 through tables.20; groups whose group is Current-law initial AIME quintile, Lifetime payroll tax quintile or Lifetime payroll tax quintile (shared); their labels arrays",
+ "acquisition": "Copied byte for byte from the cleared MINT-categories lane's followup/inputs/mint/mint8_row_categories.json; label-only extraction of the archived source named in MINT8_TABLE_LABEL_PROVENANCE, never a table body."
}
},
"consumed_by": {
diff --git a/data/external/mint8_lifetime_quintile_definitions.json b/data/external/mint8_lifetime_quintile_definitions.json
index bab32025..2e2df976 100644
--- a/data/external/mint8_lifetime_quintile_definitions.json
+++ b/data/external/mint8_lifetime_quintile_definitions.json
@@ -55,13 +55,15 @@
"date_certified": "2026-04-01",
"label_file": "mint8_row_categories.json (MINT-categories lane)",
"label_file_sha256": "23fbfbc8dbc14b06144a83bf0583ea0a6c89505638f3f7dc5fd6b34a3a4ac650",
+ "committed_label_file": "data/external/mint8_row_categories_2026.source.json",
"tables": "13-20 (benefit/tax ratios and initial replacement rates)",
"note": "Labels only. The lane's parser emitted th, caption and heading text and no data cell; the option table itself is not committed and supplies no input to this package."
},
"build": {
"built_by": "scripts/extract_lifetime_measure_sources.py",
"sources": [
- "mint8_user_guide"
+ "mint8_user_guide",
+ "mint8_table_row_labels"
],
"provenance_file": "data/external/lifetime_measure_sources.provenance.json"
}
diff --git a/data/external/mint8_row_categories_2026.source.json b/data/external/mint8_row_categories_2026.source.json
new file mode 100644
index 00000000..887adf55
--- /dev/null
+++ b/data/external/mint8_row_categories_2026.source.json
@@ -0,0 +1,1851 @@
+{
+ "source_url": "https://www.ssa.gov/policy/docs/projections/policy-options/increase-payroll-tax-rate.html",
+ "retrieved_via": "http://web.archive.org/web/20260519190449id_/https://www.ssa.gov/policy/docs/projections/policy-options/increase-payroll-tax-rate.html",
+ "raw_sha256": "3cfb6eb636cad6d062eb22734c533fb349a0bc4761014eee8f2ccb0eb6e28d8c",
+ "dateCertified": "2026-04-01",
+ "note": "LABELS ONLY. No data cells were extracted. Table numbers are document order on the page.",
+ "tables": {
+ "1": {
+ "caption": "Projected Effects of Proposal on Social Security Benefits in 2030 POPULATION: Current-law beneficiaries aged 60 or older (characteristics)",
+ "columns": [
+ "Characteristic",
+ "Percent of population with a—",
+ "Percent change in Social Security benefits at the—",
+ "Benefit decrease",
+ "Benefit increase",
+ "10th %ile",
+ "Median",
+ "90th %ile"
+ ],
+ "groups": [
+ {
+ "group": "Total",
+ "labels": []
+ },
+ {
+ "group": "Sex",
+ "labels": [
+ "Female",
+ "Male"
+ ]
+ },
+ {
+ "group": "Race and ethnicity",
+ "labels": [
+ "Hispanic or Latino, any race",
+ "White, non-Hispanic",
+ "Black or African American, non-Hispanic",
+ "All other races, non-Hispanic"
+ ]
+ },
+ {
+ "group": "Country of birth",
+ "labels": [
+ "United States",
+ "Other countries"
+ ]
+ },
+ {
+ "group": "Age",
+ "labels": [
+ "60–69",
+ "70–79",
+ "80–89",
+ "90 or older"
+ ]
+ },
+ {
+ "group": "Marital status",
+ "labels": [
+ "Married",
+ "Divorced",
+ "Widowed",
+ "Never married"
+ ]
+ },
+ {
+ "group": "Highest education level",
+ "labels": [
+ "Graduate",
+ "Bachelor",
+ "Associate",
+ "High school",
+ "Less than high school"
+ ]
+ },
+ {
+ "group": "Current-law poverty status",
+ "labels": [
+ "Above poverty",
+ "In poverty"
+ ]
+ },
+ {
+ "group": "Current-law household income quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Current-law benefit type",
+ "labels": [
+ "Retired worker only",
+ "Widow(er) (includes dually entitled)",
+ "Spousal (includes dually entitled)",
+ "Disabled worker only"
+ ]
+ }
+ ]
+ },
+ "2": {
+ "caption": "Projected Effects of Proposal on Social Security Benefits in 2050 POPULATION: Current-law beneficiaries aged 60 or older (characteristics)",
+ "columns": [
+ "Characteristic",
+ "Percent of population with a—",
+ "Percent change in Social Security benefits at the—",
+ "Benefit decrease",
+ "Benefit increase",
+ "10th %ile",
+ "Median",
+ "90th %ile"
+ ],
+ "groups": [
+ {
+ "group": "Total",
+ "labels": []
+ },
+ {
+ "group": "Sex",
+ "labels": [
+ "Female",
+ "Male"
+ ]
+ },
+ {
+ "group": "Race and ethnicity",
+ "labels": [
+ "Hispanic or Latino, any race",
+ "White, non-Hispanic",
+ "Black or African American, non-Hispanic",
+ "All other races, non-Hispanic"
+ ]
+ },
+ {
+ "group": "Country of birth",
+ "labels": [
+ "United States",
+ "Other countries"
+ ]
+ },
+ {
+ "group": "Age",
+ "labels": [
+ "60–69",
+ "70–79",
+ "80–89",
+ "90 or older"
+ ]
+ },
+ {
+ "group": "Marital status",
+ "labels": [
+ "Married",
+ "Divorced",
+ "Widowed",
+ "Never married"
+ ]
+ },
+ {
+ "group": "Highest education level",
+ "labels": [
+ "Graduate",
+ "Bachelor",
+ "Associate",
+ "High school",
+ "Less than high school"
+ ]
+ },
+ {
+ "group": "Current-law poverty status",
+ "labels": [
+ "Above poverty",
+ "In poverty"
+ ]
+ },
+ {
+ "group": "Current-law household income quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Current-law benefit type",
+ "labels": [
+ "Retired worker only",
+ "Widow(er) (includes dually entitled)",
+ "Spousal (includes dually entitled)",
+ "Disabled worker only"
+ ]
+ }
+ ]
+ },
+ "3": {
+ "caption": "Projected Effects of Proposal on Social Security Benefits in 2070 POPULATION: Current-law beneficiaries aged 60 or older (characteristics)",
+ "columns": [
+ "Characteristic",
+ "Percent of population with a—",
+ "Percent change in Social Security benefits at the—",
+ "Benefit decrease",
+ "Benefit increase",
+ "10th %ile",
+ "Median",
+ "90th %ile"
+ ],
+ "groups": [
+ {
+ "group": "Total",
+ "labels": []
+ },
+ {
+ "group": "Sex",
+ "labels": [
+ "Female",
+ "Male"
+ ]
+ },
+ {
+ "group": "Race and ethnicity",
+ "labels": [
+ "Hispanic or Latino, any race",
+ "White, non-Hispanic",
+ "Black or African American, non-Hispanic",
+ "All other races, non-Hispanic"
+ ]
+ },
+ {
+ "group": "Country of birth",
+ "labels": [
+ "United States",
+ "Other countries"
+ ]
+ },
+ {
+ "group": "Age",
+ "labels": [
+ "60–69",
+ "70–79",
+ "80–89",
+ "90 or older"
+ ]
+ },
+ {
+ "group": "Marital status",
+ "labels": [
+ "Married",
+ "Divorced",
+ "Widowed",
+ "Never married"
+ ]
+ },
+ {
+ "group": "Highest education level",
+ "labels": [
+ "Graduate",
+ "Bachelor",
+ "Associate",
+ "High school",
+ "Less than high school"
+ ]
+ },
+ {
+ "group": "Current-law poverty status",
+ "labels": [
+ "Above poverty",
+ "In poverty"
+ ]
+ },
+ {
+ "group": "Current-law household income quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Current-law benefit type",
+ "labels": [
+ "Retired worker only",
+ "Widow(er) (includes dually entitled)",
+ "Spousal (includes dually entitled)",
+ "Disabled worker only"
+ ]
+ }
+ ]
+ },
+ "4": {
+ "caption": "Projected Effects of Proposal on Social Security Taxes Paid in 2030 POPULATION: Current-law payroll taxpayers aged 31 or older (characteristics)",
+ "columns": [
+ "Characteristic",
+ "Percent of population with a—",
+ "Percent change in Social Security taxes paid at the—",
+ "Change in taxes paid (in 2024$) at the—",
+ "Tax decrease",
+ "Tax increase",
+ "10th %ile",
+ "Median",
+ "90th %ile",
+ "10th %ile",
+ "Median",
+ "90th %ile"
+ ],
+ "groups": [
+ {
+ "group": "Total",
+ "labels": []
+ },
+ {
+ "group": "Sex",
+ "labels": [
+ "Female",
+ "Male"
+ ]
+ },
+ {
+ "group": "Race and ethnicity",
+ "labels": [
+ "Hispanic or Latino, any race",
+ "White, non-Hispanic",
+ "Black or African American, non-Hispanic",
+ "All other races, non-Hispanic"
+ ]
+ },
+ {
+ "group": "Country of birth",
+ "labels": [
+ "United States",
+ "Other countries"
+ ]
+ },
+ {
+ "group": "Age",
+ "labels": [
+ "31–39",
+ "40–49",
+ "50–59",
+ "60–69",
+ "70 or older"
+ ]
+ },
+ {
+ "group": "Marital status",
+ "labels": [
+ "Married",
+ "Divorced",
+ "Widowed",
+ "Never married"
+ ]
+ },
+ {
+ "group": "Highest education level",
+ "labels": [
+ "Graduate",
+ "Bachelor",
+ "Associate",
+ "High school",
+ "Less than high school"
+ ]
+ },
+ {
+ "group": "Current-law household income quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Current-law payroll taxes quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ }
+ ]
+ },
+ "5": {
+ "caption": "Projected Effects of Proposal on Social Security Taxes Paid in 2050 POPULATION: Current-law payroll taxpayers aged 31 or older (characteristics)",
+ "columns": [
+ "Characteristic",
+ "Percent of population with a—",
+ "Percent change in Social Security taxes paid at the—",
+ "Change in taxes paid (in 2024$) at the—",
+ "Tax decrease",
+ "Tax increase",
+ "10th %ile",
+ "Median",
+ "90th %ile",
+ "10th %ile",
+ "Median",
+ "90th %ile"
+ ],
+ "groups": [
+ {
+ "group": "Total",
+ "labels": []
+ },
+ {
+ "group": "Sex",
+ "labels": [
+ "Female",
+ "Male"
+ ]
+ },
+ {
+ "group": "Race and ethnicity",
+ "labels": [
+ "Hispanic or Latino, any race",
+ "White, non-Hispanic",
+ "Black or African American, non-Hispanic",
+ "All other races, non-Hispanic"
+ ]
+ },
+ {
+ "group": "Country of birth",
+ "labels": [
+ "United States",
+ "Other countries"
+ ]
+ },
+ {
+ "group": "Age",
+ "labels": [
+ "31–39",
+ "40–49",
+ "50–59",
+ "60–69",
+ "70 or older"
+ ]
+ },
+ {
+ "group": "Marital status",
+ "labels": [
+ "Married",
+ "Divorced",
+ "Widowed",
+ "Never married"
+ ]
+ },
+ {
+ "group": "Highest education level",
+ "labels": [
+ "Graduate",
+ "Bachelor",
+ "Associate",
+ "High school",
+ "Less than high school"
+ ]
+ },
+ {
+ "group": "Current-law household income quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Current-law payroll taxes quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ }
+ ]
+ },
+ "6": {
+ "caption": "Projected Effects of Proposal on Social Security Taxes Paid in 2070 POPULATION: Current-law payroll taxpayers aged 31 or older (characteristics)",
+ "columns": [
+ "Characteristic",
+ "Percent of population with a—",
+ "Percent change in Social Security taxes paid at the—",
+ "Change in taxes paid (in 2024$) at the—",
+ "Tax decrease",
+ "Tax increase",
+ "10th %ile",
+ "Median",
+ "90th %ile",
+ "10th %ile",
+ "Median",
+ "90th %ile"
+ ],
+ "groups": [
+ {
+ "group": "Total",
+ "labels": []
+ },
+ {
+ "group": "Sex",
+ "labels": [
+ "Female",
+ "Male"
+ ]
+ },
+ {
+ "group": "Race and ethnicity",
+ "labels": [
+ "Hispanic or Latino, any race",
+ "White, non-Hispanic",
+ "Black or African American, non-Hispanic",
+ "All other races, non-Hispanic"
+ ]
+ },
+ {
+ "group": "Country of birth",
+ "labels": [
+ "United States",
+ "Other countries"
+ ]
+ },
+ {
+ "group": "Age",
+ "labels": [
+ "31–39",
+ "40–49",
+ "50–59",
+ "60–69",
+ "70 or older"
+ ]
+ },
+ {
+ "group": "Marital status",
+ "labels": [
+ "Married",
+ "Divorced",
+ "Widowed",
+ "Never married"
+ ]
+ },
+ {
+ "group": "Highest education level",
+ "labels": [
+ "Graduate",
+ "Bachelor",
+ "Associate",
+ "High school",
+ "Less than high school"
+ ]
+ },
+ {
+ "group": "Current-law household income quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Current-law payroll taxes quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ }
+ ]
+ },
+ "7": {
+ "caption": "Projected Effects of Proposal on Household Income in 2030 POPULATION: Current-law beneficiaries aged 60 or older (characteristics)",
+ "columns": [
+ "Characteristic",
+ "Percent of population with an—",
+ "Percent change in household income at the—",
+ "Income decrease",
+ "Income increase",
+ "10th %ile",
+ "Median",
+ "90th %ile"
+ ],
+ "groups": [
+ {
+ "group": "Total",
+ "labels": []
+ },
+ {
+ "group": "Sex",
+ "labels": [
+ "Female",
+ "Male"
+ ]
+ },
+ {
+ "group": "Race and ethnicity",
+ "labels": [
+ "Hispanic or Latino, any race",
+ "White, non-Hispanic",
+ "Black or African American, non-Hispanic",
+ "All other races, non-Hispanic"
+ ]
+ },
+ {
+ "group": "Country of birth",
+ "labels": [
+ "United States",
+ "Other countries"
+ ]
+ },
+ {
+ "group": "Age",
+ "labels": [
+ "60–69",
+ "70–79",
+ "80–89",
+ "90 or older"
+ ]
+ },
+ {
+ "group": "Marital status",
+ "labels": [
+ "Married",
+ "Divorced",
+ "Widowed",
+ "Never married"
+ ]
+ },
+ {
+ "group": "Highest education level",
+ "labels": [
+ "Graduate",
+ "Bachelor",
+ "Associate",
+ "High school",
+ "Less than high school"
+ ]
+ },
+ {
+ "group": "Current-law poverty status",
+ "labels": [
+ "Above poverty",
+ "In poverty"
+ ]
+ },
+ {
+ "group": "Current-law household income quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Current-law benefit type",
+ "labels": [
+ "Retired worker only",
+ "Widow(er) (includes dually entitled)",
+ "Spousal (includes dually entitled)",
+ "Disabled worker only"
+ ]
+ }
+ ]
+ },
+ "8": {
+ "caption": "Projected Effects of Proposal on Household Income in 2050 POPULATION: Current-law beneficiaries aged 60 or older (characteristics)",
+ "columns": [
+ "Characteristic",
+ "Percent of population with an—",
+ "Percent change in household income at the—",
+ "Income decrease",
+ "Income increase",
+ "10th %ile",
+ "Median",
+ "90th %ile"
+ ],
+ "groups": [
+ {
+ "group": "Total",
+ "labels": []
+ },
+ {
+ "group": "Sex",
+ "labels": [
+ "Female",
+ "Male"
+ ]
+ },
+ {
+ "group": "Race and ethnicity",
+ "labels": [
+ "Hispanic or Latino, any race",
+ "White, non-Hispanic",
+ "Black or African American, non-Hispanic",
+ "All other races, non-Hispanic"
+ ]
+ },
+ {
+ "group": "Country of birth",
+ "labels": [
+ "United States",
+ "Other countries"
+ ]
+ },
+ {
+ "group": "Age",
+ "labels": [
+ "60–69",
+ "70–79",
+ "80–89",
+ "90 or older"
+ ]
+ },
+ {
+ "group": "Marital status",
+ "labels": [
+ "Married",
+ "Divorced",
+ "Widowed",
+ "Never married"
+ ]
+ },
+ {
+ "group": "Highest education level",
+ "labels": [
+ "Graduate",
+ "Bachelor",
+ "Associate",
+ "High school",
+ "Less than high school"
+ ]
+ },
+ {
+ "group": "Current-law poverty status",
+ "labels": [
+ "Above poverty",
+ "In poverty"
+ ]
+ },
+ {
+ "group": "Current-law household income quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Current-law benefit type",
+ "labels": [
+ "Retired worker only",
+ "Widow(er) (includes dually entitled)",
+ "Spousal (includes dually entitled)",
+ "Disabled worker only"
+ ]
+ }
+ ]
+ },
+ "9": {
+ "caption": "Projected Effects of Proposal on Household Income in 2070 POPULATION: Current-law beneficiaries aged 60 or older (characteristics)",
+ "columns": [
+ "Characteristic",
+ "Percent of population with an—",
+ "Percent change in household income at the—",
+ "Income decrease",
+ "Income increase",
+ "10th %ile",
+ "Median",
+ "90th %ile"
+ ],
+ "groups": [
+ {
+ "group": "Total",
+ "labels": []
+ },
+ {
+ "group": "Sex",
+ "labels": [
+ "Female",
+ "Male"
+ ]
+ },
+ {
+ "group": "Race and ethnicity",
+ "labels": [
+ "Hispanic or Latino, any race",
+ "White, non-Hispanic",
+ "Black or African American, non-Hispanic",
+ "All other races, non-Hispanic"
+ ]
+ },
+ {
+ "group": "Country of birth",
+ "labels": [
+ "United States",
+ "Other countries"
+ ]
+ },
+ {
+ "group": "Age",
+ "labels": [
+ "60–69",
+ "70–79",
+ "80–89",
+ "90 or older"
+ ]
+ },
+ {
+ "group": "Marital status",
+ "labels": [
+ "Married",
+ "Divorced",
+ "Widowed",
+ "Never married"
+ ]
+ },
+ {
+ "group": "Highest education level",
+ "labels": [
+ "Graduate",
+ "Bachelor",
+ "Associate",
+ "High school",
+ "Less than high school"
+ ]
+ },
+ {
+ "group": "Current-law poverty status",
+ "labels": [
+ "Above poverty",
+ "In poverty"
+ ]
+ },
+ {
+ "group": "Current-law household income quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Current-law benefit type",
+ "labels": [
+ "Retired worker only",
+ "Widow(er) (includes dually entitled)",
+ "Spousal (includes dually entitled)",
+ "Disabled worker only"
+ ]
+ }
+ ]
+ },
+ "10": {
+ "caption": "Projected Effects of Proposal on Official Poverty Measure in 2030 POPULATION: Current-law beneficiaries aged 60 or older (characteristics)",
+ "columns": [
+ "Characteristic",
+ "Official poverty rate",
+ "Number of population in poverty (in thousands)",
+ "Percent change in the number in poverty",
+ "Under current law",
+ "With proposal",
+ "Under current law",
+ "With proposal",
+ "Change"
+ ],
+ "groups": [
+ {
+ "group": "Total",
+ "labels": []
+ },
+ {
+ "group": "Sex",
+ "labels": [
+ "Female",
+ "Male"
+ ]
+ },
+ {
+ "group": "Race and ethnicity",
+ "labels": [
+ "Hispanic or Latino, any race",
+ "White, non-Hispanic",
+ "Black or African American, non-Hispanic",
+ "All other races, non-Hispanic"
+ ]
+ },
+ {
+ "group": "Country of birth",
+ "labels": [
+ "United States",
+ "Other countries"
+ ]
+ },
+ {
+ "group": "Age",
+ "labels": [
+ "60–69",
+ "70–79",
+ "80–89",
+ "90 or older"
+ ]
+ },
+ {
+ "group": "Marital status",
+ "labels": [
+ "Married",
+ "Divorced",
+ "Widowed",
+ "Never married"
+ ]
+ },
+ {
+ "group": "Highest education level",
+ "labels": [
+ "Graduate",
+ "Bachelor",
+ "Associate",
+ "High school",
+ "Less than high school"
+ ]
+ },
+ {
+ "group": "Current-law poverty status",
+ "labels": [
+ "Above poverty",
+ "In poverty"
+ ]
+ },
+ {
+ "group": "Current-law benefit type",
+ "labels": [
+ "Retired worker only",
+ "Widow(er) (includes dually entitled)",
+ "Spousal (includes dually entitled)",
+ "Disabled worker only"
+ ]
+ }
+ ]
+ },
+ "11": {
+ "caption": "Projected Effects of Proposal on Official Poverty Measure in 2050 POPULATION: Current-law beneficiaries aged 60 or older (characteristics)",
+ "columns": [
+ "Characteristic",
+ "Official poverty rate",
+ "Number of population in poverty (in thousands)",
+ "Percent change in the number in poverty",
+ "Under current law",
+ "With proposal",
+ "Under current law",
+ "With proposal",
+ "Change"
+ ],
+ "groups": [
+ {
+ "group": "Total",
+ "labels": []
+ },
+ {
+ "group": "Sex",
+ "labels": [
+ "Female",
+ "Male"
+ ]
+ },
+ {
+ "group": "Race and ethnicity",
+ "labels": [
+ "Hispanic or Latino, any race",
+ "White, non-Hispanic",
+ "Black or African American, non-Hispanic",
+ "All other races, non-Hispanic"
+ ]
+ },
+ {
+ "group": "Country of birth",
+ "labels": [
+ "United States",
+ "Other countries"
+ ]
+ },
+ {
+ "group": "Age",
+ "labels": [
+ "60–69",
+ "70–79",
+ "80–89",
+ "90 or older"
+ ]
+ },
+ {
+ "group": "Marital status",
+ "labels": [
+ "Married",
+ "Divorced",
+ "Widowed",
+ "Never married"
+ ]
+ },
+ {
+ "group": "Highest education level",
+ "labels": [
+ "Graduate",
+ "Bachelor",
+ "Associate",
+ "High school",
+ "Less than high school"
+ ]
+ },
+ {
+ "group": "Current-law poverty status",
+ "labels": [
+ "Above poverty",
+ "In poverty"
+ ]
+ },
+ {
+ "group": "Current-law benefit type",
+ "labels": [
+ "Retired worker only",
+ "Widow(er) (includes dually entitled)",
+ "Spousal (includes dually entitled)",
+ "Disabled worker only"
+ ]
+ }
+ ]
+ },
+ "12": {
+ "caption": "Projected Effects of Proposal on Official Poverty Measure in 2070 POPULATION: Current-law beneficiaries aged 60 or older (characteristics)",
+ "columns": [
+ "Characteristic",
+ "Official poverty rate",
+ "Number of population in poverty (in thousands)",
+ "Percent change in the number in poverty",
+ "Under current law",
+ "With proposal",
+ "Under current law",
+ "With proposal",
+ "Change"
+ ],
+ "groups": [
+ {
+ "group": "Total",
+ "labels": []
+ },
+ {
+ "group": "Sex",
+ "labels": [
+ "Female",
+ "Male"
+ ]
+ },
+ {
+ "group": "Race and ethnicity",
+ "labels": [
+ "Hispanic or Latino, any race",
+ "White, non-Hispanic",
+ "Black or African American, non-Hispanic",
+ "All other races, non-Hispanic"
+ ]
+ },
+ {
+ "group": "Country of birth",
+ "labels": [
+ "United States",
+ "Other countries"
+ ]
+ },
+ {
+ "group": "Age",
+ "labels": [
+ "60–69",
+ "70–79",
+ "80–89",
+ "90 or older"
+ ]
+ },
+ {
+ "group": "Marital status",
+ "labels": [
+ "Married",
+ "Divorced",
+ "Widowed",
+ "Never married"
+ ]
+ },
+ {
+ "group": "Highest education level",
+ "labels": [
+ "Graduate",
+ "Bachelor",
+ "Associate",
+ "High school",
+ "Less than high school"
+ ]
+ },
+ {
+ "group": "Current-law poverty status",
+ "labels": [
+ "Above poverty",
+ "In poverty"
+ ]
+ },
+ {
+ "group": "Current-law benefit type",
+ "labels": [
+ "Retired worker only",
+ "Widow(er) (includes dually entitled)",
+ "Spousal (includes dually entitled)",
+ "Disabled worker only"
+ ]
+ }
+ ]
+ },
+ "13": {
+ "caption": "Projected Effects of Proposal on Benefit/Tax Ratios POPULATION: Workers born 1960–1969 with a benefit/tax ratio (characteristics)",
+ "columns": [
+ "Characteristic",
+ "Percent of population with a—",
+ "Percent change in benefit/tax ratio at the—",
+ "Benefit/tax ratio under current law at the—",
+ "Benefit/tax ratio with proposal at the—",
+ "Ratio decrease",
+ "Ratio increase",
+ "10th %ile",
+ "Median",
+ "90th %ile",
+ "10th %ile",
+ "Median",
+ "90th %ile",
+ "10th %ile",
+ "Median",
+ "90th %ile"
+ ],
+ "groups": [
+ {
+ "group": "Total",
+ "labels": []
+ },
+ {
+ "group": "Sex",
+ "labels": [
+ "Female",
+ "Male"
+ ]
+ },
+ {
+ "group": "Race and ethnicity",
+ "labels": [
+ "Hispanic or Latino, any race",
+ "White, non-Hispanic",
+ "Black or African American, non-Hispanic",
+ "All other races, non-Hispanic"
+ ]
+ },
+ {
+ "group": "Country of birth",
+ "labels": [
+ "United States",
+ "Other countries"
+ ]
+ },
+ {
+ "group": "Highest education level",
+ "labels": [
+ "Graduate",
+ "Bachelor",
+ "Associate",
+ "High school",
+ "Less than high school"
+ ]
+ },
+ {
+ "group": "Current-law initial AIME quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Lifetime payroll tax quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Lifetime payroll tax quintile (shared)",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ }
+ ]
+ },
+ "14": {
+ "caption": "Projected Effects of Proposal on Benefit/Tax Ratios POPULATION: Workers born 1980–1989 with a benefit/tax ratio (characteristics)",
+ "columns": [
+ "Characteristic",
+ "Percent of population with a—",
+ "Percent change in benefit/tax ratio at the—",
+ "Benefit/tax ratio under current law at the—",
+ "Benefit/tax ratio with proposal at the—",
+ "Ratio decrease",
+ "Ratio increase",
+ "10th %ile",
+ "Median",
+ "90th %ile",
+ "10th %ile",
+ "Median",
+ "90th %ile",
+ "10th %ile",
+ "Median",
+ "90th %ile"
+ ],
+ "groups": [
+ {
+ "group": "Total",
+ "labels": []
+ },
+ {
+ "group": "Sex",
+ "labels": [
+ "Female",
+ "Male"
+ ]
+ },
+ {
+ "group": "Race and ethnicity",
+ "labels": [
+ "Hispanic or Latino, any race",
+ "White, non-Hispanic",
+ "Black or African American, non-Hispanic",
+ "All other races, non-Hispanic"
+ ]
+ },
+ {
+ "group": "Country of birth",
+ "labels": [
+ "United States",
+ "Other countries"
+ ]
+ },
+ {
+ "group": "Highest education level",
+ "labels": [
+ "Graduate",
+ "Bachelor",
+ "Associate",
+ "High school",
+ "Less than high school"
+ ]
+ },
+ {
+ "group": "Current-law initial AIME quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Lifetime payroll tax quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Lifetime payroll tax quintile (shared)",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ }
+ ]
+ },
+ "15": {
+ "caption": "Projected Effects of Proposal on Benefit/Tax Ratios POPULATION: Workers born 2000–2009 with a benefit/tax ratio (characteristics)",
+ "columns": [
+ "Characteristic",
+ "Percent of population with a—",
+ "Percent change in benefit/tax ratio at the—",
+ "Benefit/tax ratio under current law at the—",
+ "Benefit/tax ratio with proposal at the—",
+ "Ratio decrease",
+ "Ratio increase",
+ "10th %ile",
+ "Median",
+ "90th %ile",
+ "10th %ile",
+ "Median",
+ "90th %ile",
+ "10th %ile",
+ "Median",
+ "90th %ile"
+ ],
+ "groups": [
+ {
+ "group": "Total",
+ "labels": []
+ },
+ {
+ "group": "Sex",
+ "labels": [
+ "Female",
+ "Male"
+ ]
+ },
+ {
+ "group": "Race and ethnicity",
+ "labels": [
+ "Hispanic or Latino, any race",
+ "White, non-Hispanic",
+ "Black or African American, non-Hispanic",
+ "All other races, non-Hispanic"
+ ]
+ },
+ {
+ "group": "Country of birth",
+ "labels": [
+ "United States",
+ "Other countries"
+ ]
+ },
+ {
+ "group": "Highest education level",
+ "labels": [
+ "Graduate",
+ "Bachelor",
+ "Associate",
+ "High school",
+ "Less than high school"
+ ]
+ },
+ {
+ "group": "Current-law initial AIME quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Lifetime payroll tax quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Lifetime payroll tax quintile (shared)",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ }
+ ]
+ },
+ "16": {
+ "caption": "Projected Effects of Proposal on Benefit/Tax Ratios POPULATION: Workers born 2020–2029 with a benefit/tax ratio (characteristics)",
+ "columns": [
+ "Characteristic",
+ "Percent of population with a—",
+ "Percent change in benefit/tax ratio at the—",
+ "Benefit/tax ratio under current law at the—",
+ "Benefit/tax ratio with proposal at the—",
+ "Ratio decrease",
+ "Ratio increase",
+ "10th %ile",
+ "Median",
+ "90th %ile",
+ "10th %ile",
+ "Median",
+ "90th %ile",
+ "10th %ile",
+ "Median",
+ "90th %ile"
+ ],
+ "groups": [
+ {
+ "group": "Total",
+ "labels": []
+ },
+ {
+ "group": "Sex",
+ "labels": [
+ "Female",
+ "Male"
+ ]
+ },
+ {
+ "group": "Race and ethnicity",
+ "labels": [
+ "Hispanic or Latino, any race",
+ "White, non-Hispanic",
+ "Black or African American, non-Hispanic",
+ "All other races, non-Hispanic"
+ ]
+ },
+ {
+ "group": "Country of birth",
+ "labels": [
+ "United States",
+ "Other countries"
+ ]
+ },
+ {
+ "group": "Highest education level",
+ "labels": [
+ "Graduate",
+ "Bachelor",
+ "Associate",
+ "High school",
+ "Less than high school"
+ ]
+ },
+ {
+ "group": "Current-law initial AIME quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Lifetime payroll tax quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Lifetime payroll tax quintile (shared)",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ }
+ ]
+ },
+ "17": {
+ "caption": "Projected Effects of Proposal on Initial Replacement Rates POPULATION: Current-law beneficiaries born 1960–1969 with a replacement rate (characteristics)",
+ "columns": [
+ "Characteristic",
+ "Percent of population with a—",
+ "Percent change in initial replacement rate at the—",
+ "Initial replacement rate under current law at the—",
+ "Initial replacement rate with proposal at the—",
+ "Rate decrease",
+ "Rate increase",
+ "10th %ile",
+ "Median",
+ "90th %ile",
+ "10th %ile",
+ "Median",
+ "90th %ile",
+ "10th %ile",
+ "Median",
+ "90th %ile"
+ ],
+ "groups": [
+ {
+ "group": "Total",
+ "labels": []
+ },
+ {
+ "group": "Sex",
+ "labels": [
+ "Female",
+ "Male"
+ ]
+ },
+ {
+ "group": "Race and ethnicity",
+ "labels": [
+ "Hispanic or Latino, any race",
+ "White, non-Hispanic",
+ "Black or African American, non-Hispanic",
+ "All other races, non-Hispanic"
+ ]
+ },
+ {
+ "group": "Country of birth",
+ "labels": [
+ "United States",
+ "Other countries"
+ ]
+ },
+ {
+ "group": "Highest education level",
+ "labels": [
+ "Graduate",
+ "Bachelor",
+ "Associate",
+ "High school",
+ "Less than high school"
+ ]
+ },
+ {
+ "group": "Current-law initial AIME quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Lifetime payroll tax quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Lifetime payroll tax quintile (shared)",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ }
+ ]
+ },
+ "18": {
+ "caption": "Projected Effects of Proposal on Initial Replacement Rates POPULATION: Current-law beneficiaries born 1980–1989 with a replacement rate (characteristics)",
+ "columns": [
+ "Characteristic",
+ "Percent of population with a—",
+ "Percent change in initial replacement rate at the—",
+ "Initial replacement rate under current law at the—",
+ "Initial replacement rate with proposal at the—",
+ "Rate decrease",
+ "Rate increase",
+ "10th %ile",
+ "Median",
+ "90th %ile",
+ "10th %ile",
+ "Median",
+ "90th %ile",
+ "10th %ile",
+ "Median",
+ "90th %ile"
+ ],
+ "groups": [
+ {
+ "group": "Total",
+ "labels": []
+ },
+ {
+ "group": "Sex",
+ "labels": [
+ "Female",
+ "Male"
+ ]
+ },
+ {
+ "group": "Race and ethnicity",
+ "labels": [
+ "Hispanic or Latino, any race",
+ "White, non-Hispanic",
+ "Black or African American, non-Hispanic",
+ "All other races, non-Hispanic"
+ ]
+ },
+ {
+ "group": "Country of birth",
+ "labels": [
+ "United States",
+ "Other countries"
+ ]
+ },
+ {
+ "group": "Highest education level",
+ "labels": [
+ "Graduate",
+ "Bachelor",
+ "Associate",
+ "High school",
+ "Less than high school"
+ ]
+ },
+ {
+ "group": "Current-law initial AIME quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Lifetime payroll tax quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Lifetime payroll tax quintile (shared)",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ }
+ ]
+ },
+ "19": {
+ "caption": "Projected Effects of Proposal on Initial Replacement Rates POPULATION: Current-law beneficiaries born 2000–2009 with a replacement rate (characteristics)",
+ "columns": [
+ "Characteristic",
+ "Percent of population with a—",
+ "Percent change in initial replacement rate at the—",
+ "Initial replacement rate under current law at the—",
+ "Initial replacement rate with proposal at the—",
+ "Rate decrease",
+ "Rate increase",
+ "10th %ile",
+ "Median",
+ "90th %ile",
+ "10th %ile",
+ "Median",
+ "90th %ile",
+ "10th %ile",
+ "Median",
+ "90th %ile"
+ ],
+ "groups": [
+ {
+ "group": "Total",
+ "labels": []
+ },
+ {
+ "group": "Sex",
+ "labels": [
+ "Female",
+ "Male"
+ ]
+ },
+ {
+ "group": "Race and ethnicity",
+ "labels": [
+ "Hispanic or Latino, any race",
+ "White, non-Hispanic",
+ "Black or African American, non-Hispanic",
+ "All other races, non-Hispanic"
+ ]
+ },
+ {
+ "group": "Country of birth",
+ "labels": [
+ "United States",
+ "Other countries"
+ ]
+ },
+ {
+ "group": "Highest education level",
+ "labels": [
+ "Graduate",
+ "Bachelor",
+ "Associate",
+ "High school",
+ "Less than high school"
+ ]
+ },
+ {
+ "group": "Current-law initial AIME quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Lifetime payroll tax quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Lifetime payroll tax quintile (shared)",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ }
+ ]
+ },
+ "20": {
+ "caption": "Projected Effects of Proposal on Initial Replacement Rates POPULATION: Current-law beneficiaries born 2020–2029 with a replacement rate (characteristics)",
+ "columns": [
+ "Characteristic",
+ "Percent of population with a—",
+ "Percent change in initial replacement rate at the—",
+ "Initial replacement rate under current law at the—",
+ "Initial replacement rate with proposal at the—",
+ "Rate decrease",
+ "Rate increase",
+ "10th %ile",
+ "Median",
+ "90th %ile",
+ "10th %ile",
+ "Median",
+ "90th %ile",
+ "10th %ile",
+ "Median",
+ "90th %ile"
+ ],
+ "groups": [
+ {
+ "group": "Total",
+ "labels": []
+ },
+ {
+ "group": "Sex",
+ "labels": [
+ "Female",
+ "Male"
+ ]
+ },
+ {
+ "group": "Race and ethnicity",
+ "labels": [
+ "Hispanic or Latino, any race",
+ "White, non-Hispanic",
+ "Black or African American, non-Hispanic",
+ "All other races, non-Hispanic"
+ ]
+ },
+ {
+ "group": "Country of birth",
+ "labels": [
+ "United States",
+ "Other countries"
+ ]
+ },
+ {
+ "group": "Highest education level",
+ "labels": [
+ "Graduate",
+ "Bachelor",
+ "Associate",
+ "High school",
+ "Less than high school"
+ ]
+ },
+ {
+ "group": "Current-law initial AIME quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Lifetime payroll tax quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Lifetime payroll tax quintile (shared)",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ }
+ ]
+ }
+ }
+}
\ No newline at end of file
diff --git a/docs/design/lifetime_measures.md b/docs/design/lifetime_measures.md
new file mode 100644
index 00000000..40459bb3
--- /dev/null
+++ b/docs/design/lifetime_measures.md
@@ -0,0 +1,159 @@
+# G2 lifetime measures for NASI group breakdowns
+
+`populace_dynamics.estimates.lifetime_measures` computes pure measures
+over already-built cohort frames. It opens no PSID file and performs no
+real-data tabulation. Development and validation use INVENTED careers.
+Any real-data group outcomes require the separate issue #42 registration.
+
+## Inputs and results
+
+Careers have `person_id`, `year`, `earnings`, and optionally `provenance`.
+There must be one row per person-year, with finite nonnegative nominal
+earnings. Persons have unique `person_id` and `birth_year`; other columns,
+including sex, are accepted but not used. Marriage episodes use the
+already-loaded cohort marriage-episode schema, optionally including
+`separation_year`. No reader is added.
+
+Each measure returns a frozen `MeasureResult(measure, frame, provenance)`.
+The frame contains one row per supplied person, in supplied order, with
+the value, `status`, `reason`, and coverage flags. Missing histories are
+`not computed` with missing values; an observed zero is a computed zero.
+SHA-256 provenance pins normalized input frames, supplied NAWI and wage
+base schedules, and result frames. Payroll results also pin applied tax
+rates and available required interest rates, including custom schedules.
+
+## AIME conventions
+
+`initial_aime_at_62(careers, persons, params, *, analysis_year, convention)`
+requires an explicit `AIME_CONVENTIONS` entry or `AimeConvention`:
+
+| Name | Computation years | Earnings cutoff |
+| --- | --- | --- |
+| `mint8_initial_aime` | Statutory 415(b) | Year of attaining 61 |
+| `exercise_1_cola` | Legacy fixed 35 | Supplied career through analysis year |
+| `exercise_3_fra68` | Legacy fixed 35 | Supplied career through analysis year |
+| `exercise_4_min_benefit` | Statutory 415(b), old-age at 62 | Year of attaining 61 |
+
+The statutory choice composes `ss.statutory_aime.oracle_aime`; the
+legacy choice dispatches to the unchanged `ss.benefits.aime`. Track U
+uses reported benefits and has no AIME oracle convention. The Track M
+choice matches an old-age entitlement at 62, not a worker's observed DI
+or death-basis AIME. Statutory pre-1975 age-62 cases remain unsupported
+and receive a reason rather than an invented formula.
+
+People younger than 62 at the analysis year use earnings through that
+year and are flagged `provisional_through_last_observed`. Earnings are
+never projected. NAWI still indexes to age 60, so the caller must supply
+that wage-index path. Left and right history censoring and imputed rows
+are counted. `n_history_years_after_age_61` identifies legacy full-career
+AIMEs containing supplied earnings after the initial-entitlement cutoff.
+
+## Payroll taxes and age-62 present value
+
+`lifetime_payroll_tax_pv_at_62(careers, persons, params, *, shared, rates,
+interest, marriage_episodes=None, ...)` applies the combined employee
+and employer OASDI rate to earnings capped at the contribution and
+benefit base. Every supplied career year enters, including years after
+62. A year's payment is valued at year end; reference year is birth year
+plus 62. Accumulate with rates for payment year + 1 through reference
+year, or discount with reference year + 1 through payment year.
+
+The captured schedules supply two explicit tax bases. Default
+`TRUST_FUND_RECEIVED` doubles SSA's employee/employer "each" rate.
+`EMPLOYEE_EMPLOYER_PAID` accounts for the employee credits in 1984 and
+2011–2012. Default interest is combined OASDI effective annual interest;
+the annual average new-issue rate is an explicit alternative. These
+choices, timing, self-employment treatment, and sharing conventions are
+recorded as registered builder defaults for the downstream registration.
+MINT's wording is taxes "paid": `EMPLOYEE_EMPLOYER_PAID` is the literal
+interpretation in credit years. The default receipt basis includes general
+revenue replacements in those years, so the downstream registration must
+state its selected basis rather than claim those two bases are identical.
+
+Shared taxes use the year-end marriage state from the unchanged cohort
+helper. Over married years, each person receives half the sum of both
+spouses' annual taxes, using the union of their career years. Missing
+spouse careers use flagged `OWN_ONLY`, or `NOT_COMPUTED` if requested.
+Absent individual years within an available career contribute zero and
+are counted separately. Separated people count as married by default.
+Unknown states, absent history-roster membership, overlapping marriages,
+and disagreeing reciprocal spouse links are exposed in flags/defaults.
+The latest-start marriage wins in overlaps; each person's recorded
+history governs when reciprocal links disagree.
+
+With reciprocal links, sharing conserves the couple's annual taxes.
+Their two PVs conserve the total at the same reference date: same-birth
+spouses can be added directly; different-age spouses must first be
+valued at a common date. Disagreeing histories do not guarantee this.
+
+Historical interest covers 1940–2025. Uncovered years refuse by default
+or produce `not computed` if explicitly requested. Extend coverage only
+with `TrustFundInterestRates.extended(..., source="named assumption")`;
+there is no silently repeated future rate. Invalid rates fail explicitly.
+
+## Report average and schemes
+
+`report_average_indexed_earnings_22_62(careers, persons, params, *, shared,
+marriage_episodes=None, conventions=None, ...)` follows the recorded
+Butrica–Uccello measure: average wage-indexed earnings at ages 22–62,
+uncapped for the own measure, including uncovered labor earnings.
+
+Unrecorded conventions are named in `REPORT_EARNINGS_BUILDER_DEFAULTS`:
+inclusive ages, NAWI indexing to 60, nominal earnings thereafter, average
+over ages with own career rows, shared uncapped earnings, year-end
+marriage, and missing-spouse treatment. Unlike payroll union-year sharing,
+the default report average uses own covered-age support. The optional
+`ALL_AGES` divisor treats absent ages as zero over 41 years. Coverage and
+imputation flags accompany either choice.
+
+`load_mint8_scheme()` reads the pinned category definitions for initial
+AIME, own lifetime payroll tax, and shared lifetime payroll tax.
+`quintile_cells(birth_years, scope)` exposes whole-population or ten-year
+birth-cohort cells. `boomers2004_scheme()` retains the report's named
+dimensions but refuses label order and quintile population until those
+unrecorded conventions are specified. Scheme metadata does not authorize
+a real-data run.
+
+## Quintile invariants
+
+`weighted_quintiles(values, weights, *, by=None, labels=QUINTILE_LABELS)`
+returns categorical labels in MINT order: Highest, Second highest,
+Middle, Second lowest, Lowest, followed by `not computed` for missing
+values. Valued rows need finite positive weights and nonmissing cell keys.
+Indexes must match; cells are cut independently.
+
+Tied values stay together. Each tie group's weighted cumulative midpoint
+determines its quintile. Exact fractional weight sums make the result
+independent of row order. A midpoint within 1e-12 rank units of a boundary
+takes the upper quintile, absorbing floating-point multiplication error
+under weight scaling. Without ties, each quintile's deviation from 20%
+is bounded by the largest row's weight share. Equal weights and a row
+count divisible by five give exactly 20%. Higher values cannot receive
+a lower quintile within a cell.
+
+`quintile_summary` retains empty categorical quintiles as zero-case rows.
+The `not computed` row is a coverage diagnostic; exclude it when applying
+SSA's suppression rule to the five official quintile rows. If valued
+weight totals zero, shares remain undefined.
+
+## Captures and integration handoff
+
+`scripts/extract_lifetime_measure_sources.py --check` reproduces pinned
+JSON from captured SSA OACT tax/interest pages and the MINT user guide,
+and checks effective interest against both per-fund historical pages.
+The cleared row-category capture contains labels only, never data cells.
+Source hashes and element/table locators accompany every extraction.
+The captured MINT source is certified 2026-04-01; the dispatch brief's
+2025-10-01 description is an older vintage.
+
+Tests: `test_lifetime_measures.py` and
+`test_lifetime_measures_properties.py` are `unit`;
+`test_lifetime_measure_sources.py` is `artifact`. They include Hypothesis
+invariants, independent PV/quintile formulas, and differential checks
+against both SS oracles and the Track A/Track M benefit conventions.
+
+The integrator must add the new source module to
+`POST_REVIEW_SOURCE_EXCLUSIONS` and its pinned tuple test, update
+`tests/tier_counts.json`, and add Hypothesis to dev extras if absent.
+Those shared files are intentionally outside this package's write scope.
+Hypothesis is already present in this checkout's dev extras.
diff --git a/scripts/extract_lifetime_measure_sources.py b/scripts/extract_lifetime_measure_sources.py
index 7bf41ed1..e16ee583 100644
--- a/scripts/extract_lifetime_measure_sources.py
+++ b/scripts/extract_lifetime_measure_sources.py
@@ -14,7 +14,7 @@
* the verbatim MINT8 definitions of the lifetime-earnings quintile rows
(SSA, "MINT8 Table User Guide", ``user-guide.html``).
-Every source is the exact HTTP 200 response body fetched from ssa.gov on
+Every HTML source is the exact HTTP 200 response body fetched from ssa.gov on
2026-10-01 with ``curl -A 'Wget/1.21.4'`` (ssa.gov refuses some default
clients). The pages carry a per-request Akamai mPulse script in
````, so a second fetch has different bytes; two fetches taken
@@ -27,18 +27,23 @@
(``engine.refit.validate_external_vintage`` rejects vintage 2026).
The MINT8 table row-group labels ("Current-law initial AIME quintile",
-...) are not on the user guide page. They are transcribed from the
-label-only extraction of SSA's payroll-tax option table made by the
-MINT-categories lane (no data cell was extracted); its provenance is
-recorded in :data:`MINT8_TABLE_LABEL_PROVENANCE`.
+...) are not on the user guide page. They are read from the committed,
+SHA-256-pinned label-only extraction of SSA's payroll-tax option table
+made by the MINT-categories lane (no data cell was extracted); its
+provenance is recorded in :data:`MINT8_TABLE_LABEL_PROVENANCE`. The source
+guide and label extraction are certified 2026-04-01, a later version
+than the 2025-10-01 guide identified in the G2 brief.
Run from the repository root::
.venv/bin/python scripts/extract_lifetime_measure_sources.py
+
+Add ``--check`` to verify all artifacts without rewriting them.
"""
from __future__ import annotations
+import argparse
import hashlib
import json
import re
@@ -110,6 +115,56 @@
),
"bytes": 73967,
},
+ "mint8_table_row_labels": {
+ "file": "mint8_row_categories_2026.source.json",
+ "url": (
+ "https://www.ssa.gov/policy/docs/projections/policy-options/"
+ "increase-payroll-tax-rate.html"
+ ),
+ "title": "MINT8 payroll-tax option table labels (labels only)",
+ "sha256": (
+ "23fbfbc8dbc14b06144a83bf0583ea0a6c89505638f3f7dc5fd6b34a3a4ac650"
+ ),
+ "bytes": 37524,
+ },
+}
+
+#: Every rate is located by its named table, year/period row and column;
+#: definitions are located by paragraph id or an exact identifying sentence.
+SOURCE_LOCATORS = {
+ "trust_fund_interest_rates": (
+ "table caption: Average annual special-issue interest rates on "
+ "new issues and effective annual interest rates (percent); Year "
+ "row; Average and Effective columns; definition paragraphs above "
+ "the table"
+ ),
+ "effective_rates_1980_on": (
+ "table caption: Effective Interest Rates Earned By the Invested "
+ "Assets of the OASI and DI Trust Funds [Percent]; Calendar year "
+ "row; OASI, DI and OASDI columns"
+ ),
+ "effective_rates_1940_1979": (
+ "table caption: Estimated Effective Interest Rates Earned By the "
+ "Assets of the OASI and DI Trust Funds, 1940-79[Percent]; Calendar "
+ "year row in either panel; OASI, DI and OASDI columns"
+ ),
+ "oasdi_tax_rates": (
+ "table summary: Tax rate table for Social Security trust funds; "
+ "Calendar years period row; OASI, DI and Total columns under "
+ "employee/employer-each and self-employed headers; footnote "
+ "anchors fna-fnd"
+ ),
+ "mint8_user_guide": (
+ "meta DCTERMS:dateCertified; paragraphs id=AIME, lifetime-tax, "
+ "lifetime-tax-shared and taxes; other paragraphs identified by "
+ "the exact sentences in MINT8_PARAGRAPHS"
+ ),
+ "mint8_table_row_labels": (
+ "label-only JSON: dateCertified; tables.13 through tables.20; "
+ "groups whose group is Current-law initial AIME quintile, "
+ "Lifetime payroll tax quintile or Lifetime payroll tax quintile "
+ "(shared); their labels arrays"
+ ),
}
INTEREST_OUT = EXTERNAL / "ssa_trust_fund_interest_rates_2026.json"
@@ -257,6 +312,9 @@
"label_file_sha256": (
"23fbfbc8dbc14b06144a83bf0583ea0a6c89505638f3f7dc5fd6b34a3a4ac650"
),
+ "committed_label_file": (
+ "data/external/mint8_row_categories_2026.source.json"
+ ),
"tables": "13-20 (benefit/tax ratios and initial replacement rates)",
"note": (
"Labels only. The lane's parser emitted th, caption and heading "
@@ -325,6 +383,8 @@ def handle_starttag(
match = re.fullmatch(r"fn([a-z])", attributes.get("name") or "")
if match:
self._footnote = match.group(1)
+ if self._footnote in self.footnotes:
+ raise ValueError(f"duplicate footnote {self._footnote}")
self.footnotes[self._footnote] = []
elif tag == "br" and self._footnote is not None:
self.footnotes[self._footnote].append(" ")
@@ -409,6 +469,8 @@ def parse_interest_rates() -> dict[int, dict[str, float | None]]:
"average_new_issue": _percent(cells[1]),
"effective_oasdi": _percent(cells[2]),
}
+ if any(value is None for value in rates[year].values()):
+ raise ValueError(f"{year}: missing combined interest rate")
expected = list(range(FIRST_INTEREST_YEAR, LATEST_INTEREST_YEAR + 1))
if sorted(rates) != expected:
raise ValueError(f"years {sorted(rates)} != {expected}")
@@ -421,7 +483,10 @@ def parse_interest_rates() -> dict[int, dict[str, float | None]]:
raise ValueError("expected two per-fund header rows (1980 on)")
for cells in recent.rows:
if len(cells) == 4 and re.fullmatch(r"\d{4}", cells[0]):
- per_fund[int(cells[0])] = [_percent(cell) for cell in cells[1:]]
+ year = int(cells[0])
+ if year in per_fund:
+ raise ValueError(f"duplicate per-fund year {year}")
+ per_fund[year] = [_percent(cell) for cell in cells[1:]]
early = _parse("effective_rates_1940_1979")
if EFFECTIVE_1940_CAPTION not in early.captions:
raise ValueError(f"caption {EFFECTIVE_1940_CAPTION!r} not found")
@@ -444,6 +509,10 @@ def parse_interest_rates() -> dict[int, dict[str, float | None]]:
raise ValueError("per-fund pages do not cover 1940-2025 exactly")
for year in expected:
oasi, di, oasdi = per_fund[year]
+ if oasi is None or oasdi is None:
+ raise ValueError(f"{year}: missing OASI/OASDI interest rate")
+ if (di is None) != (year < 1957):
+ raise ValueError(f"{year}: DI must be absent only before 1957")
if oasdi != rates[year]["effective_oasdi"]:
raise ValueError(
f"{year}: combined effective "
@@ -538,6 +607,8 @@ def _period(label: str) -> tuple[int, int | None, list[str]]:
else:
last = first
notes = re.findall(r"[a-d]", match.group(4) or "")
+ if len(notes) != len(set(notes)):
+ raise ValueError(f"duplicate footnotes in period {label!r}")
return first, last, notes
@@ -562,11 +633,23 @@ def parse_tax_rates() -> tuple[list[dict[str, Any]], dict[str, str]]:
self_employed = dict(
zip(("oasi", "di", "total"), values[3:], strict=True)
)
- if each["total"] is None:
+ if each["total"] is None or each["oasi"] is None:
raise ValueError(f"{cells[0]}: no employee/employer total")
parts = [part for part in (each["oasi"], each["di"]) if part]
if abs(sum(parts) - each["total"]) > 1e-9:
raise ValueError(f"{cells[0]}: OASI + DI != total")
+ if any(value is not None for value in self_employed.values()):
+ if self_employed["total"] is None or self_employed["oasi"] is None:
+ raise ValueError(f"{cells[0]}: incomplete self-employed rate")
+ parts = [
+ part
+ for part in (self_employed["oasi"], self_employed["di"])
+ if part is not None
+ ]
+ if abs(sum(parts) - self_employed["total"]) > 1e-9:
+ raise ValueError(
+ f"{cells[0]}: self-employed OASI + DI != total"
+ )
rows.append(
{
"period": cells[0],
@@ -577,6 +660,8 @@ def parse_tax_rates() -> tuple[list[dict[str, Any]], dict[str, str]]:
"self_employed": self_employed,
}
)
+ if not rows:
+ raise ValueError("the tax table contains no rate rows")
expected_first = FIRST_TAX_YEAR
for index, row in enumerate(rows):
if row["first_year"] != expected_first:
@@ -674,6 +759,46 @@ def build_tax() -> dict[str, Any]:
# ---------------------------------------------------------------------------
# MINT8 definitions
# ---------------------------------------------------------------------------
+def parse_mint8_labels() -> tuple[dict[str, str], tuple[str, ...]]:
+ """Read only the cleared label JSON and verify every cohort table.
+
+ The cleared extraction contains headings, captions and labels, never
+ outcome cells. All eight cohort tables must have the same lifetime
+ dimension labels and quintile ordering. The archived option-table
+ body is neither read nor needed to reproduce this artifact.
+ """
+ document = json.loads(read_source("mint8_table_row_labels"))
+ source_record = MINT8_TABLE_LABEL_PROVENANCE
+ for field, provenance_field in (
+ ("source_url", "source_url"),
+ ("retrieved_via", "retrieved_via"),
+ ("raw_sha256", "raw_sha256"),
+ ("dateCertified", "date_certified"),
+ ):
+ if document.get(field) != source_record[provenance_field]:
+ raise ValueError(f"label source {field} disagrees with provenance")
+ if not document.get("note", "").startswith("LABELS ONLY."):
+ raise ValueError("the label extraction must be marked LABELS ONLY")
+ names: dict[str, str] = {}
+ labels: tuple[str, ...] | None = None
+ for number in range(13, 21):
+ groups = document["tables"][str(number)]["groups"]
+ for key, section in MINT8_COHORT_ROW_GROUPS.items():
+ matches = [group for group in groups if group["group"] == section]
+ if len(matches) != 1:
+ raise ValueError(
+ f"table {number}: {section!r} has {len(matches)} matches"
+ )
+ names[key] = matches[0]["group"]
+ actual = tuple(matches[0]["labels"])
+ if actual != MINT8_QUINTILE_LABELS:
+ raise ValueError(f"table {number}: quintile labels changed")
+ labels = actual
+ if labels is None:
+ raise ValueError("no cohort quintile labels were extracted")
+ return names, labels
+
+
def build_mint8() -> dict[str, Any]:
"""Verbatim MINT8 definitions used by the lifetime-earnings rows."""
page = _parse("mint8_user_guide")
@@ -683,6 +808,7 @@ def build_mint8() -> dict[str, Any]:
)
if not certified:
raise ValueError("dateCertified meta tag not found")
+ row_groups, quintile_labels = parse_mint8_labels()
definitions: dict[str, dict[str, str]] = {}
for key, locator in MINT8_PARAGRAPHS.items():
if "element_id" in locator:
@@ -705,12 +831,12 @@ def build_mint8() -> dict[str, Any]:
"document": SOURCES["mint8_user_guide"]["title"],
"date_certified": certified.group(1),
"definitions": definitions,
- "cohort_table_row_groups": MINT8_COHORT_ROW_GROUPS,
- "quintile_labels_high_to_low": list(MINT8_QUINTILE_LABELS),
+ "cohort_table_row_groups": row_groups,
+ "quintile_labels_high_to_low": list(quintile_labels),
"cohort_table_label_provenance": MINT8_TABLE_LABEL_PROVENANCE,
"build": {
"built_by": "scripts/extract_lifetime_measure_sources.py",
- "sources": ["mint8_user_guide"],
+ "sources": ["mint8_user_guide", "mint8_table_row_labels"],
"provenance_file": (
"data/external/lifetime_measure_sources.provenance.json"
),
@@ -719,7 +845,7 @@ def build_mint8() -> dict[str, Any]:
def build_provenance() -> dict[str, Any]:
- """One provenance record for the five committed source bodies."""
+ """Provenance for the five HTML captures and cleared label extraction."""
for key in SOURCES:
read_source(key)
return {
@@ -735,6 +861,15 @@ def build_provenance() -> dict[str, Any]:
"document": spec["title"],
"source_sha256": spec["sha256"],
"source_length_bytes": spec["bytes"],
+ "locator": SOURCE_LOCATORS[key],
+ "acquisition": (
+ "Copied byte for byte from the cleared MINT-categories "
+ "lane's followup/inputs/mint/mint8_row_categories.json; "
+ "label-only extraction of the archived source named "
+ "in MINT8_TABLE_LABEL_PROVENANCE, never a table body."
+ if key == "mint8_table_row_labels"
+ else FETCH_METHOD
+ ),
}
for key, spec in SOURCES.items()
},
@@ -768,10 +903,26 @@ def build_all() -> dict[Path, dict[str, Any]]:
def main() -> None:
+ parser = argparse.ArgumentParser(description=__doc__)
+ parser.add_argument(
+ "--check",
+ action="store_true",
+ help="Verify byte identity without writing any artifact.",
+ )
+ args = parser.parse_args()
for path, document in build_all().items():
- path.write_text(render(document), encoding="utf-8")
+ rendered = render(document)
+ if args.check:
+ if (
+ not path.is_file()
+ or path.read_text(encoding="utf-8") != rendered
+ ):
+ parser.error(f"{path.relative_to(ROOT)} needs rebuilding")
+ else:
+ path.write_text(rendered, encoding="utf-8")
digest = hashlib.sha256(path.read_bytes()).hexdigest()
- print(f"wrote {path.relative_to(ROOT)} sha256 {digest}")
+ action = "verified" if args.check else "wrote"
+ print(f"{action} {path.relative_to(ROOT)} sha256 {digest}")
if __name__ == "__main__":
diff --git a/src/populace_dynamics/estimates/lifetime_measures.py b/src/populace_dynamics/estimates/lifetime_measures.py
index 5739e3fa..0f3b8040 100644
--- a/src/populace_dynamics/estimates/lifetime_measures.py
+++ b/src/populace_dynamics/estimates/lifetime_measures.py
@@ -61,7 +61,8 @@
1. Quintiles: within a cell with no tied values every quintile holds
20 percent of the weight to within the largest row's weight share;
- labels are invariant to exact positive weight scaling and to row
+ labels are invariant to positive weight scaling within the documented
+ floating-point boundary tolerance and to row
order; a higher value never gets a lower quintile in the same cell;
missing values get :data:`NOT_COMPUTED` and nothing else does.
2. Payroll tax present value: zero earnings give zero; doubling earnings
@@ -82,6 +83,7 @@
import json
import math
from collections.abc import Callable, Collection, Mapping
+from copy import deepcopy
from dataclasses import dataclass, field
from enum import Enum
from fractions import Fraction
@@ -163,7 +165,7 @@
#: Verbatim MINT8 Table User Guide definitions and the cohort-table labels.
MINT8_DEFINITIONS_PATH = _EXTERNAL / "mint8_lifetime_quintile_definitions.json"
MINT8_DEFINITIONS_SHA256 = (
- "0df7298e0655eecdd4aca5b04b41ec9c520228baa7329080e7208950e7769608"
+ "857d82d78371483d69ffb1402df8f2229928629cf5e657dcb5ee497f853ffa74"
)
#: MINT8's quintile labels in its table order (always these five).
@@ -180,6 +182,10 @@
COMPUTED = "computed"
_N_QUINTILES = 5
+# Snap a midpoint within this distance of an integer quintile boundary
+# to that boundary. This prevents float multiplication from moving an
+# exact boundary across quintiles when all weights are rescaled.
+_QUINTILE_BOUNDARY_TOLERANCE = 1e-12
_RETIREMENT_AGE = 62
#: The age at which the career frames' coverage starts at the earliest
#: (estimates/career.py: max(1968, birth year + 22)); a later first row
@@ -420,6 +426,22 @@ class ReportEarningsConventions:
"counts as married (the repository's default)"
),
"missing_spouse": "own earnings for that year, counted (OWN_ONLY)",
+ "missing_spouse_year": (
+ "a spouse with a career but no row in this year contributes zero; "
+ "the year is counted as married_years_spouse_year_absent"
+ ),
+ "missing_own_year": (
+ "COVERED_AGES uses only the person's own observed career years; "
+ "spouse-only years are outside this report average"
+ ),
+ "reciprocal_history_disagreement": (
+ "use each person's recorded spouse; disagreement with an available "
+ "spouse history is counted and couple conservation is not promised"
+ ),
+ "multiple_marriages_in_force": (
+ "use the latest-start marriage as marital_state_at does; count "
+ "years_multiple_marriages_in_force"
+ ),
"quintile_order_and_population": (
"not registered: whether '1st Quintile' is the lowest and over "
"which population the quintiles are cut are unrecorded, so the "
@@ -469,6 +491,22 @@ class ReportEarningsConventions:
"unknown_marital_state": (
"a year whose marital state is 'unknown' counts own tax (counted)"
),
+ "missing_spouse_year": (
+ "a spouse with a career but no row in this year contributes zero; "
+ "the year is counted as married_years_spouse_year_absent"
+ ),
+ "missing_own_year": (
+ "share over the union of both spouses' career years while married; "
+ "an absent own year contributes zero and is counted"
+ ),
+ "reciprocal_history_disagreement": (
+ "use each person's recorded spouse; disagreement with an available "
+ "spouse history is counted and couple conservation is not promised"
+ ),
+ "multiple_marriages_in_force": (
+ "use the latest-start marriage as marital_state_at does; count "
+ "years_multiple_marriages_in_force"
+ ),
}
@@ -716,15 +754,19 @@ def accumulation_factor(
reference_year = _year(reference_year, "reference_year")
if tax_year == reference_year:
return 1.0
- if tax_year < reference_year:
- return math.prod(
- 1.0 + interest.rate_for(year)
- for year in range(tax_year + 1, reference_year + 1)
- )
- return 1.0 / math.prod(
- 1.0 + interest.rate_for(year)
- for year in range(reference_year + 1, tax_year + 1)
- )
+ low, high = sorted((tax_year, reference_year))
+ growth = []
+ for year in range(low + 1, high + 1):
+ rate = float(interest.rate_for(year))
+ if not math.isfinite(rate) or rate <= -1.0:
+ raise ValueError(
+ f"Interest rate for {year} must be finite and above -1."
+ )
+ growth.append(1.0 + rate)
+ factor = math.prod(growth)
+ if not math.isfinite(factor) or factor <= 0:
+ raise ValueError("Interest accumulation is outside finite range.")
+ return factor if tax_year < reference_year else 1.0 / factor
# ---------------------------------------------------------------------------
@@ -825,6 +867,7 @@ def _input_digests(
careers: pd.DataFrame,
persons: pd.DataFrame,
episodes: pd.DataFrame | None = None,
+ marriage_history_person_ids: Collection[int] | None = None,
) -> dict[str, str]:
digests = {
"careers_sha256": _frame_sha256(
@@ -840,6 +883,12 @@ def _input_digests(
.astype("string")
.reset_index(drop=True)
)
+ if marriage_history_person_ids is not None:
+ digests["marriage_history_person_ids_sha256"] = _sha256(
+ json.dumps(sorted(set(marriage_history_person_ids))).encode(
+ "utf-8"
+ )
+ )
return digests
@@ -881,7 +930,11 @@ def _provenance(
"reason_counts": _counts(frame["reason"]),
"flag_totals": totals,
"conventions": conventions,
- "ssa_parameters": {"pe_us_revision": params.pe_us_revision},
+ "ssa_parameters": {
+ "pe_us_revision": params.pe_us_revision,
+ "nawi_sha256": _schedule_sha256(params.nawi),
+ "wage_base_sha256": _schedule_sha256(params.wage_base),
+ },
"inputs": inputs,
"output_sha256": _frame_sha256(frame),
"pure_over_cohort_outputs": (
@@ -891,16 +944,46 @@ def _provenance(
}
if extra:
record.update(extra)
- return record
+ # A caller can edit its audit record without changing module defaults
+ # or another call's provenance.
+ return deepcopy(record)
+
+
+def _schedule_sha256(schedule: Mapping[int, float]) -> str:
+ """Pin supplied parameter values, including replaced projection paths."""
+ return _sha256(
+ json.dumps(
+ {
+ str(year): float(value)
+ for year, value in sorted(schedule.items())
+ },
+ sort_keys=True,
+ allow_nan=False,
+ ).encode("utf-8")
+ )
def _nawi_missing(
years: Collection[int], indexing_year: int, params: SSAParameters
) -> list[int]:
needed = {indexing_year} | {year for year in years if year < indexing_year}
+ for year in needed & params.nawi.keys():
+ value = float(params.nawi[year])
+ if not math.isfinite(value) or value <= 0:
+ raise ValueError(f"NAWI for {year} must be finite and positive.")
return sorted(year for year in needed if year not in params.nawi)
+def _wage_base(year: int, params: SSAParameters) -> float:
+ """Validate the supplied contribution and benefit base before use."""
+ base = float(params.wage_base_for(year))
+ if not math.isfinite(base) or base < 0:
+ raise ValueError(
+ f"Wage base for {year} must be finite and non-negative."
+ )
+ return base
+
+
# ---------------------------------------------------------------------------
# 1. Initial AIME at 62
# ---------------------------------------------------------------------------
@@ -918,6 +1001,7 @@ def _nawi_missing(
"history_starts_after_age_22",
"history_ends_before_cutoff",
"n_history_years",
+ "n_history_years_after_age_61",
"n_imputed_history_years",
"years_before_1951_dropped",
"computation_years",
@@ -957,6 +1041,8 @@ def initial_aime_at_62(
analysis_year = _year(analysis_year, "analysis_year")
if not isinstance(convention, AimeConvention):
raise TypeError("convention must be an AimeConvention.")
+ if convention.last_earnings_age is not None:
+ _year(convention.last_earnings_age, "last_earnings_age")
convention_years = ComputationYears(convention.computation_years)
careers_frame = _careers_frame(careers)
persons_frame = _persons_frame(persons)
@@ -1012,6 +1098,9 @@ def initial_aime_at_62(
last_year is not None and last_year < cutoff
),
"n_history_years": len(kept),
+ "n_history_years_after_age_61": sum(
+ year > birth + 61 for year in kept
+ ),
"n_imputed_history_years": (
None
if imputed is None
@@ -1036,6 +1125,8 @@ def initial_aime_at_62(
elif missing := _nawi_missing(kept, indexing, params):
row["reason"] = f"nawi_unavailable:{missing[0]}"
else:
+ for year in kept:
+ _wage_base(year, params)
if convention_years is ComputationYears.STATUTORY:
row["computation_years"] = (
statutory_aime.benefit_computation_years(birth)
@@ -1070,6 +1161,12 @@ def initial_aime_at_62(
"source": convention.source,
},
"oracle": "ss.statutory_aime.oracle_aime (unchanged)",
+ "legacy_history_window": (
+ "last_earnings_age=None reproduces the blind-test oracle "
+ "on supplied careers through analysis_year; post-61 years "
+ "are counted explicitly, so this is a legacy career AIME "
+ "rather than an initial-at-62 observation where present"
+ ),
"definition": "mint8:initial_aime_quintile",
},
{"basis_counts": _counts(frame["basis"])},
@@ -1095,8 +1192,13 @@ def annual_payroll_taxes(
frame = _careers_frame(careers)
years = sorted(int(year) for year in frame["year"].unique())
- base = {year: float(params.wage_base_for(year)) for year in years}
+ base = {year: _wage_base(year, params) for year in years}
rate = {year: float(rates.combined_for(year)) for year in years}
+ for year in years:
+ if not math.isfinite(rate[year]) or rate[year] < 0:
+ raise ValueError(
+ f"Tax rate for {year} must be finite and non-negative."
+ )
taxable = np.minimum(
frame["earnings"].to_numpy(), frame["year"].map(base).to_numpy()
)
@@ -1116,6 +1218,19 @@ def _episodes_by_person(
raise ValueError(f"marriage_episodes lacks {sorted(missing)}.")
frame = episodes.copy()
frame["person_id"] = _integral(frame["person_id"], "episode person_id")
+ for column in (
+ "marriage_order",
+ "start_year",
+ "episode_end_year",
+ "spouse_person_id",
+ "separation_year",
+ ):
+ if column in frame:
+ present = frame[column].notna()
+ _integral(frame.loc[present, column], f"episode {column}")
+ ordered = frame[frame["marriage_order"].notna()]
+ if ordered.duplicated(["person_id", "marriage_order"]).any():
+ raise ValueError("marriage_episodes has duplicate marriage orders.")
return {
int(pid): group.reset_index(drop=True)
for pid, group in frame.groupby("person_id", sort=True)
@@ -1139,6 +1254,7 @@ def __init__(
self._empty = episodes.iloc[0:0]
self._separated_is_married = bool(separated_is_married)
self._cache: dict[tuple[int, int], tuple[str, int | None]] = {}
+ self._multiple_in_force: dict[tuple[int, int], bool] = {}
def has_history(self, pid: int) -> bool:
return pid in self._by_person
@@ -1165,8 +1281,13 @@ def state(self, pid: int, year: int) -> tuple[str, int | None]:
str(state["status"]),
_spouse_id(state["spouse_person_id"]),
)
+ self._multiple_in_force[key] = bool(state["multiple_in_force"])
return self._cache[key]
+ def multiple_in_force(self, pid: int, year: int) -> bool:
+ self.state(pid, year)
+ return self._multiple_in_force[(pid, year)]
+
@dataclass
class _ShareCounts:
@@ -1174,6 +1295,8 @@ class _ShareCounts:
married_years_spouse_unavailable: int = 0
married_years_spouse_year_absent: int = 0
married_years_spouse_record_disagrees: int = 0
+ married_years_own_year_absent: int = 0
+ years_multiple_marriages_in_force: int = 0
years_marital_unknown: int = 0
@@ -1194,6 +1317,9 @@ def _shared_amounts(
out: dict[int, float] = {}
for year in sorted(candidate_years):
status, spouse = marital.state(pid, year)
+ counts.years_multiple_marriages_in_force += int(
+ marital.multiple_in_force(pid, year)
+ )
own_value = own.get(year)
if status == "married":
spouse_years = (
@@ -1206,6 +1332,8 @@ def _shared_amounts(
continue
if year not in spouse_years:
counts.married_years_spouse_year_absent += 1
+ if own_value is None:
+ counts.married_years_own_year_absent += 1
if marital.has_history(spouse):
back_status, back_spouse = marital.state(spouse, year)
if back_status != "married" or back_spouse != pid:
@@ -1241,6 +1369,8 @@ def _shared_amounts(
"married_years_spouse_unavailable",
"married_years_spouse_year_absent",
"married_years_spouse_record_disagrees",
+ "married_years_own_year_absent",
+ "years_multiple_marriages_in_force",
"years_marital_unknown",
"marriage_history_absent",
)
@@ -1284,7 +1414,8 @@ def lifetime_payroll_tax_pv_at_62(
refuses); a missing spouse under ``missing_spouse=NOT_COMPUTED``.
"""
- shared = bool(shared)
+ if not isinstance(shared, bool):
+ raise TypeError("shared must be a bool.")
missing_spouse = MissingSpousePolicy(missing_spouse)
missing_rate = MissingRatePolicy(missing_rate)
if shared and marriage_episodes is None:
@@ -1313,6 +1444,7 @@ def lifetime_payroll_tax_pv_at_62(
rows = []
uncovered: dict[int, list[int]] = {}
+ required_interest_years: set[int] = set()
for pid, birth in zip(
persons_frame["person_id"], persons_frame["birth_year"], strict=True
):
@@ -1362,6 +1494,7 @@ def lifetime_payroll_tax_pv_at_62(
for year in amounts:
low, high = sorted((year, reference))
needed.update(range(low + 1, high + 1))
+ required_interest_years.update(needed)
gaps = sorted(
year for year in needed if not _covers(interest, year)
)
@@ -1405,7 +1538,10 @@ def lifetime_payroll_tax_pv_at_62(
frame,
params,
_input_digests(
- careers_frame, persons_frame, marriage_episodes if shared else None
+ careers_frame,
+ persons_frame,
+ marriage_episodes if shared else None,
+ history_ids if shared else None,
),
{
"shared": shared,
@@ -1420,7 +1556,15 @@ def lifetime_payroll_tax_pv_at_62(
),
},
{
- "tax_rates": dict(getattr(rates, "provenance", {}) or {}),
+ "tax_rates": {
+ **dict(getattr(rates, "provenance", {}) or {}),
+ "basis": getattr(rates, "basis", None),
+ "applied_rates_sha256": _schedule_sha256(
+ dict(
+ zip(taxes["year"], taxes["combined_rate"], strict=True)
+ )
+ ),
+ },
"interest_rates": {
**dict(getattr(interest, "provenance", {}) or {}),
"assumed_years": sorted(
@@ -1429,6 +1573,15 @@ def lifetime_payroll_tax_pv_at_62(
"assumption_source": getattr(
interest, "assumption_source", None
),
+ "series": getattr(interest, "series", None),
+ "required_years": sorted(required_interest_years),
+ "available_rates_sha256": _schedule_sha256(
+ {
+ year: interest.rate_for(year)
+ for year in required_interest_years
+ if _covers(interest, year)
+ }
+ ),
},
},
)
@@ -1489,7 +1642,10 @@ def report_average_indexed_earnings_22_62(
"""
conventions = conventions or ReportEarningsConventions()
- shared = bool(shared)
+ if not isinstance(shared, bool):
+ raise TypeError("shared must be a bool.")
+ for name in ("first_age", "last_age", "index_age"):
+ _year(getattr(conventions, name), name)
if shared and marriage_episodes is None:
raise ValueError("shared=True needs marriage_episodes.")
if conventions.last_age < conventions.first_age:
@@ -1517,7 +1673,7 @@ def report_average_indexed_earnings_22_62(
def earnings_value(earnings: float, year: int) -> float:
if conventions.cap_at_taxable_maximum:
- return min(earnings, float(params.wage_base_for(year)))
+ return min(earnings, _wage_base(year, params))
return earnings
rows = []
@@ -1613,7 +1769,10 @@ def earnings_value(earnings: float, year: int) -> float:
frame,
params,
_input_digests(
- careers_frame, persons_frame, marriage_episodes if shared else None
+ careers_frame,
+ persons_frame,
+ marriage_episodes if shared else None,
+ history_ids if shared else None,
),
{
"shared": shared,
@@ -1642,9 +1801,11 @@ def _quintile_ranks(values: np.ndarray, weights: np.ndarray) -> np.ndarray:
Rows sharing a value form one group and always share a rank. Groups
are ordered by value; a group whose cumulative weight span is
[B, B + G] of the total W takes rank floor(5 * (B + G / 2) / W),
- capped at 4: the quintile holding the midpoint of its span. Weights
- are summed as exact fractions, so the ranks do not depend on row
- order or on rounding.
+ capped at 4: the quintile holding the midpoint of its span. Weights
+ are summed as exact fractions. A normalized midpoint within 1e-12
+ of an integer boundary (in rank units, 0-5) takes the upper quintile.
+ This tolerance absorbs float multiplication error during rescaling;
+ exact sums keep the result independent of row order.
"""
unique, inverse = np.unique(values, return_inverse=True)
@@ -1656,10 +1817,13 @@ def _quintile_ranks(values: np.ndarray, weights: np.ndarray) -> np.ndarray:
below = Fraction(0)
for index, weight in enumerate(group_weights):
twice_midpoint = 2 * below + weight
- ranks[index] = min(
- _N_QUINTILES - 1,
- (_N_QUINTILES * twice_midpoint) // (2 * total),
- )
+ midpoint_rank = Fraction(_N_QUINTILES * twice_midpoint, 2 * total)
+ nearest = round(midpoint_rank)
+ if abs(midpoint_rank - nearest) <= _QUINTILE_BOUNDARY_TOLERANCE:
+ rank = nearest
+ else:
+ rank = math.floor(midpoint_rank)
+ ranks[index] = min(_N_QUINTILES - 1, rank)
below += weight
return ranks[inverse]
@@ -1684,6 +1848,8 @@ def weighted_quintiles(
percent of the cell's weight to within the largest single weight's
share. MINT8 publishes no tie rule ("The dollar ranges are available
upon request"); this is the builder's registered rule.
+ A midpoint within 1e-12 rank units of a boundary takes the upper
+ quintile, so floating-point weight rescaling preserves boundary labels.
Returns a categorical Series named ``quintile`` with categories
``labels`` then :data:`NOT_COMPUTED`; a row without a value gets
@@ -1691,7 +1857,11 @@ def weighted_quintiles(
"""
labels = tuple(labels)
- if len(labels) != _N_QUINTILES or len(set(labels)) != _N_QUINTILES:
+ if (
+ len(labels) != _N_QUINTILES
+ or len(set(labels)) != _N_QUINTILES
+ or not all(isinstance(label, str) for label in labels)
+ ):
raise ValueError("labels must be five distinct strings.")
if NOT_COMPUTED in labels:
raise ValueError(f"{NOT_COMPUTED!r} cannot be a quintile label.")
@@ -1772,12 +1942,20 @@ def quintile_summary(
The unweighted counts are what MINT8's disclosure rule reads ("suppress
an entire characteristic subgroup if the sample size for any row in
that subgroup is less than 100 individuals").
+ Categorical inputs retain empty label rows with zero cases, so an
+ empty quintile remains visible to that rule. Zero total valued weight
+ gives missing shares.
"""
labels = pd.Series(labels)
weights = pd.Series(weights)
if not labels.index.equals(weights.index):
raise ValueError("labels and weights must share one index.")
+ if by is not None and not pd.Series(by).index.equals(labels.index):
+ raise ValueError("by must share the index of labels.")
+ numeric_weights = pd.to_numeric(weights).astype("float64")
+ if not np.all(np.isfinite(numeric_weights)) or (numeric_weights < 0).any():
+ raise ValueError("summary weights must be finite and non-negative.")
cell = (
pd.Series("all", index=labels.index, dtype="string")
if by is None
@@ -1787,18 +1965,28 @@ def quintile_summary(
{
"cell": cell.to_numpy(),
"label": labels.astype("string").to_numpy(),
- "weight": pd.to_numeric(weights).astype("float64").to_numpy(),
+ "weight": numeric_weights.to_numpy(),
}
)
- summary = (
- frame.groupby(["cell", "label"], sort=True)
- .agg(n=("weight", "size"), weight=("weight", "sum"))
- .reset_index()
+ summary = frame.groupby(["cell", "label"], sort=True).agg(
+ n=("weight", "size"), weight=("weight", "sum")
)
+ if isinstance(labels.dtype, pd.CategoricalDtype):
+ cells = sorted(frame["cell"].dropna().unique())
+ label_order = [str(label) for label in labels.cat.categories]
+ complete = pd.MultiIndex.from_product(
+ [cells, label_order], names=["cell", "label"]
+ )
+ summary = summary.reindex(complete, fill_value=0)
+ summary = summary.reset_index()
valued = summary[summary["label"] != NOT_COMPUTED]
totals = valued.groupby("cell")["weight"].sum()
summary["weight_share"] = [
- (weight / totals[cell] if label != NOT_COMPUTED else math.nan)
+ (
+ weight / totals[cell]
+ if label != NOT_COMPUTED and totals.get(cell, 0) > 0
+ else math.nan
+ )
for cell, label, weight in zip(
summary["cell"], summary["label"], summary["weight"], strict=True
)
diff --git a/tests/estimates/test_lifetime_measure_sources.py b/tests/estimates/test_lifetime_measure_sources.py
index 6a75a943..65f809d7 100644
--- a/tests/estimates/test_lifetime_measure_sources.py
+++ b/tests/estimates/test_lifetime_measure_sources.py
@@ -64,6 +64,8 @@ def test_provenance_names_every_source_and_output():
for key, entry in record["sources"].items():
assert entry["source_sha256"] == extractor.SOURCES[key]["sha256"]
assert entry["source_url"].startswith("https://www.ssa.gov/")
+ assert entry["locator"] == extractor.SOURCE_LOCATORS[key]
+ assert entry["acquisition"]
outputs = record["consumed_by"][
"scripts/extract_lifetime_measure_sources.py"
]
@@ -77,6 +79,28 @@ def test_provenance_names_every_source_and_output():
)
+def test_extractor_refuses_changed_source_bytes(tmp_path, monkeypatch):
+ spec = extractor.SOURCES["oasdi_tax_rates"]
+ changed = tmp_path / spec["file"]
+ changed.write_bytes((EXTERNAL / spec["file"]).read_bytes() + b" ")
+ monkeypatch.setattr(extractor, "EXTERNAL", tmp_path)
+ with pytest.raises(ValueError, match="re-verify the source"):
+ extractor.read_source("oasdi_tax_rates")
+
+
+def _tamper_page(monkeypatch, key, change):
+ """INVENTED parser perturbation; the committed bytes stay unchanged."""
+ original = extractor._parse
+
+ def parse(source_key):
+ page = original(source_key)
+ if source_key == key:
+ change(page)
+ return page
+
+ monkeypatch.setattr(extractor, "_parse", parse)
+
+
# ---------------------------------------------------------------------------
# Interest rates
# ---------------------------------------------------------------------------
@@ -117,6 +141,44 @@ def test_interest_file_is_a_post_boundary_vintage():
)
+@pytest.mark.parametrize(
+ ("source", "match"),
+ [
+ ("trust_fund_interest_rates", "duplicate year"),
+ ("effective_rates_1980_on", "duplicate per-fund year"),
+ ("effective_rates_1940_1979", "duplicate per-fund year"),
+ ],
+)
+def test_interest_parser_refuses_duplicate_year(monkeypatch, source, match):
+ def duplicate(page):
+ row = next(row for row in page.rows if row and row[0].startswith("19"))
+ page.rows.append(row[:])
+
+ _tamper_page(monkeypatch, source, duplicate)
+ with pytest.raises(ValueError, match=match):
+ extractor.parse_interest_rates()
+
+
+def test_interest_parser_refuses_disagreeing_tables(monkeypatch):
+ def disagree(page):
+ row = next(row for row in page.rows if row[0] == "2025")
+ row[-1] = "99.9" # INVENTED disagreement, not a source rate.
+
+ _tamper_page(monkeypatch, "effective_rates_1980_on", disagree)
+ with pytest.raises(ValueError, match="combined effective"):
+ extractor.parse_interest_rates()
+
+
+def test_interest_parser_refuses_missing_required_rate(monkeypatch):
+ def missing(page):
+ row = next(row for row in page.rows if row[0] == "2025")
+ row[1] = "--"
+
+ _tamper_page(monkeypatch, "trust_fund_interest_rates", missing)
+ with pytest.raises(ValueError, match="missing combined interest rate"):
+ extractor.parse_interest_rates()
+
+
@pytest.mark.parametrize("series", list(lm.InterestSeries))
def test_interest_loader_reads_each_series(series):
rates = lm.load_trust_fund_interest_rates(series=series)
@@ -192,6 +254,36 @@ def test_tax_schedule_refuses_years_before_1937():
lm.load_oasdi_tax_rates().combined_for(1936)
+@pytest.mark.parametrize("problem", ["no_rows", "gap", "self_total"])
+def test_tax_parser_refuses_invalid_table(monkeypatch, problem):
+ def change(page):
+ rows = [row for row in page.rows if len(row) == 7]
+ if problem == "no_rows":
+ page.rows = [row for row in page.rows if row not in rows]
+ elif problem == "gap":
+ page.rows.remove(rows[0])
+ else:
+ row = next(row for row in rows if row[-1] != "--")
+ row[-1] = "99.9" # INVENTED arithmetic mismatch.
+
+ _tamper_page(monkeypatch, "oasdi_tax_rates", change)
+ match = {
+ "no_rows": "no rate rows",
+ "gap": "not contiguous",
+ "self_total": "self-employed OASI",
+ }[problem]
+ with pytest.raises(ValueError, match=match):
+ extractor.parse_tax_rates()
+
+
+def test_credit_rate_requires_its_exact_source_quote():
+ rows, footnotes = extractor.parse_tax_rates()
+ # INVENTED quote corruption; 5.4 is the captured source value.
+ footnotes["a"] = footnotes["a"].replace("5.4 percent", "5.6 percent")
+ with pytest.raises(ValueError, match="footnote a quote not found"):
+ extractor._paid_adjustments(rows, footnotes)
+
+
def test_pv_example_on_the_captured_schedules():
# INVENTED career: 10,000 in 2023 and 2024, born 1962 (Y = 2024), no
# binding base. Tax 12.4 percent = 1,240 each year; the 2023 tax
@@ -264,3 +356,35 @@ def test_mint8_scheme_is_registered_and_data_driven():
assert cells.tolist() == ["1960–1969", "1980–1989"]
provenance = lm.load_mint8_definitions()["cohort_table_label_provenance"]
assert provenance["label_file_sha256"].startswith("23fbfbc8")
+ path = ROOT / provenance["committed_label_file"]
+ assert hashlib.sha256(path.read_bytes()).hexdigest() == (
+ provenance["label_file_sha256"]
+ )
+
+
+@pytest.mark.parametrize("problem", ["date", "duplicate", "order"])
+def test_mint_label_parser_refuses_changed_labels(monkeypatch, problem):
+ original = extractor.read_source
+ document = json.loads(original("mint8_table_row_labels"))
+ if problem == "date":
+ document["dateCertified"] = "INVENTED"
+ else:
+ groups = document["tables"]["13"]["groups"]
+ group = next(
+ group
+ for group in groups
+ if group["group"] == "Lifetime payroll tax quintile"
+ )
+ if problem == "duplicate":
+ groups.append(group)
+ else:
+ group["labels"].reverse()
+
+ def changed_source(key):
+ if key == "mint8_table_row_labels":
+ return json.dumps(document)
+ return original(key)
+
+ monkeypatch.setattr(extractor, "read_source", changed_source)
+ with pytest.raises(ValueError):
+ extractor.parse_mint8_labels()
diff --git a/tests/estimates/test_lifetime_measures.py b/tests/estimates/test_lifetime_measures.py
index 4b97d371..42e7542d 100644
--- a/tests/estimates/test_lifetime_measures.py
+++ b/tests/estimates/test_lifetime_measures.py
@@ -156,7 +156,14 @@ def _reference_ranks(values: list[float], weights: list[float]) -> list[int]:
if v == value
)
midpoint = below + tied / 2
- ranks.append(min(4, math.floor(5 * midpoint / total)))
+ normalized = 5 * midpoint / total
+ boundary = round(normalized)
+ rank = (
+ boundary
+ if abs(normalized - boundary) <= 1e-12
+ else math.floor(normalized)
+ )
+ ranks.append(min(4, rank))
return ranks
diff --git a/tests/estimates/test_lifetime_measures_properties.py b/tests/estimates/test_lifetime_measures_properties.py
new file mode 100644
index 00000000..1dbc4c7f
--- /dev/null
+++ b/tests/estimates/test_lifetime_measures_properties.py
@@ -0,0 +1,689 @@
+"""Adversarial lifetime-measure invariants on INVENTED inputs (unit tier).
+
+No source data is loaded. Parameters, careers, marriage episodes, and rate
+schedules are all INVENTED. Different-age spouses conserve shared taxes
+after their age-62 present values are expressed at a common date. Missing
+observations remain distinct from observed zero earnings; post-62 AIME
+earnings follow the explicitly selected oracle convention.
+"""
+
+from __future__ import annotations
+
+from dataclasses import dataclass, replace
+
+import pandas as pd
+import pytest
+from hypothesis import given, settings
+from hypothesis import strategies as st
+
+from populace_dynamics.estimates import lifetime_measures as lm
+from populace_dynamics.ss import benefits, statutory_aime
+from populace_dynamics.ss.params import SSAParameters
+
+PROPERTY_SETTINGS = settings(max_examples=25, deadline=None)
+
+
+@dataclass(frozen=True)
+class _Rates:
+ """INVENTED combined employee/employer payroll rate."""
+
+ value: float = 0.1
+
+ def combined_for(self, year: int) -> float:
+ return self.value
+
+
+@dataclass(frozen=True)
+class _Interest:
+ """INVENTED constant annual interest rate."""
+
+ value: float = 0.03
+
+ def rate_for(self, year: int) -> float:
+ return self.value
+
+
+def _params() -> SSAParameters:
+ """INVENTED flat wage index and contribution/benefit base."""
+ return SSAParameters(
+ nawi={year: 100.0 for year in range(1951, 2101)},
+ wage_base={1937: 10_000.0},
+ pia_factors=(0.9, 0.32, 0.15),
+ fra_months_by_birth_year=[(1900, 792)],
+ early_monthly_rates=(5 / 900, 5 / 1200),
+ early_first_bracket_months=36,
+ pe_us_revision="INVENTED",
+ )
+
+
+def _persons(birth1: int = 1950, birth2: int = 1950) -> pd.DataFrame:
+ return pd.DataFrame({"person_id": [1, 2], "birth_year": [birth1, birth2]})
+
+
+def _careers(rows: list[tuple[int, int, float]]) -> pd.DataFrame:
+ frame = pd.DataFrame(rows, columns=["person_id", "year", "earnings"])
+ frame["provenance"] = "INVENTED"
+ return frame
+
+
+def _episodes() -> pd.DataFrame:
+ """INVENTED reciprocal marriage, intact from 1970 onward."""
+ return pd.DataFrame(
+ {
+ "person_id": [1, 2],
+ "marriage_order": [1, 1],
+ "start_year": pd.array([1970, 1970], dtype="Int64"),
+ "episode_end_year": pd.array([None, None], dtype="Int64"),
+ "how_ended": ["intact", "intact"],
+ "spouse_person_id": pd.array([2, 1], dtype="Int64"),
+ }
+ )
+
+
+@PROPERTY_SETTINGS
+@given(
+ st.integers(1940, 1955),
+ st.integers(1, 10),
+ st.lists(st.integers(0, 9_000), min_size=8, max_size=8),
+)
+def test_different_age_couple_conserves_pv_at_a_common_date(
+ birth, gap, earnings
+):
+ """Sharing conserves the couple's sum at one valuation date.
+
+ Comparing a raw sum of values at different age-62 years would compare
+ amounts measured at different dates, and is not a conservation law.
+ """
+ careers = _careers(
+ [
+ (1, 1985 + offset, value)
+ for offset, value in enumerate(earnings[:4])
+ ]
+ + [
+ (2, 1987 + offset, value)
+ for offset, value in enumerate(earnings[4:])
+ ]
+ )
+ persons = _persons(birth, birth + gap)
+ args = {
+ "rates": _Rates(),
+ "interest": _Interest(),
+ }
+ own = lm.lifetime_payroll_tax_pv_at_62(
+ careers, persons, _params(), shared=False, **args
+ ).frame.set_index("person_id")
+ shared = lm.lifetime_payroll_tax_pv_at_62(
+ careers,
+ persons,
+ _params(),
+ shared=True,
+ marriage_episodes=_episodes(),
+ **args,
+ ).frame.set_index("person_id")
+ common_year = birth + gap + 62
+
+ def common_value(frame: pd.DataFrame) -> float:
+ return sum(
+ frame.at[pid, "pv_at_62"]
+ * 1.03 ** (common_year - frame.at[pid, "reference_year"])
+ for pid in [1, 2]
+ )
+
+ assert shared["status"].tolist() == [lm.COMPUTED, lm.COMPUTED]
+ assert common_value(shared) == pytest.approx(
+ common_value(own), rel=1e-12, abs=1e-12
+ )
+
+
+def test_different_age_raw_pv_sums_are_not_a_conservation_law():
+ careers = _careers([(1, 2000, 1_000.0), (2, 2000, 3_000.0)])
+ persons = _persons(1940, 1950)
+ own = lm.lifetime_payroll_tax_pv_at_62(
+ careers,
+ persons,
+ _params(),
+ shared=False,
+ rates=_Rates(),
+ interest=_Interest(),
+ ).frame
+ shared = lm.lifetime_payroll_tax_pv_at_62(
+ careers,
+ persons,
+ _params(),
+ shared=True,
+ marriage_episodes=_episodes(),
+ rates=_Rates(),
+ interest=_Interest(),
+ ).frame
+ assert shared["pv_at_62"].sum() != pytest.approx(own["pv_at_62"].sum())
+
+
+@pytest.mark.parametrize("measure", ["aime", "payroll", "report"])
+def test_measures_leave_input_frames_and_parameters_unchanged(measure):
+ careers = _careers([(2, 2001, 3_000.0), (1, 2000, 1_000.0)])
+ persons = _persons()
+ episodes = _episodes().iloc[::-1].copy()
+ params = _params()
+ snapshots = [
+ frame.copy(deep=True) for frame in [careers, persons, episodes]
+ ]
+ nawi, wage_base = dict(params.nawi), dict(params.wage_base)
+ if measure == "aime":
+ lm.initial_aime_at_62(
+ careers,
+ persons,
+ params,
+ analysis_year=2030,
+ convention=lm.AIME_CONVENTIONS["mint8_initial_aime"],
+ )
+ elif measure == "payroll":
+ lm.lifetime_payroll_tax_pv_at_62(
+ careers,
+ persons,
+ params,
+ shared=True,
+ marriage_episodes=episodes,
+ rates=_Rates(),
+ interest=_Interest(),
+ )
+ else:
+ lm.report_average_indexed_earnings_22_62(
+ careers,
+ persons,
+ params,
+ shared=True,
+ marriage_episodes=episodes,
+ )
+ for original, snapshot in zip(
+ [careers, persons, episodes], snapshots, strict=True
+ ):
+ pd.testing.assert_frame_equal(original, snapshot)
+ assert params.nawi == nawi
+ assert params.wage_base == wage_base
+
+
+def test_payroll_shares_union_years_and_counts_absent_spouse_years():
+ """An absent spouse-year uses the registered zero-year convention.
+
+ Both spouses have careers; neither is counted as unavailable. Each
+ receives half of each observed tax year, including spouse-only years.
+ """
+ result = lm.lifetime_payroll_tax_pv_at_62(
+ _careers([(1, 2000, 1_000.0), (2, 2001, 3_000.0)]),
+ _persons(),
+ _params(),
+ shared=True,
+ marriage_episodes=_episodes(),
+ rates=_Rates(),
+ interest=_Interest(0.0),
+ )
+ assert result.frame["pv_at_62"].tolist() == [200.0, 200.0]
+ assert result.frame["married_years_shared"].tolist() == [2, 2]
+ assert result.frame["married_years_spouse_year_absent"].tolist() == [1, 1]
+ assert result.frame["married_years_own_year_absent"].tolist() == [1, 1]
+ assert result.frame["married_years_spouse_unavailable"].tolist() == [0, 0]
+
+
+def test_report_sharing_uses_own_covered_age_support_and_counts_gaps():
+ """The report's covered-age divisor does not fill a missing own row."""
+ result = lm.report_average_indexed_earnings_22_62(
+ _careers([(1, 2000, 1_000.0), (2, 2001, 3_000.0)]),
+ _persons(),
+ _params(),
+ shared=True,
+ marriage_episodes=_episodes(),
+ )
+ assert result.frame["average_indexed_earnings"].tolist() == [
+ 500.0,
+ 1_500.0,
+ ]
+ assert result.frame["n_ages_covered"].tolist() == [1, 1]
+ assert result.frame["married_years_spouse_year_absent"].tolist() == [1, 1]
+
+
+def test_zero_earnings_are_computed_but_an_absent_history_is_not():
+ careers = _careers([(1, 2000, 0.0)])
+ persons = _persons()
+ results_and_values = [
+ (
+ lm.initial_aime_at_62(
+ careers,
+ persons,
+ _params(),
+ analysis_year=2030,
+ convention=lm.AIME_CONVENTIONS["mint8_initial_aime"],
+ ),
+ "aime",
+ ),
+ (
+ lm.lifetime_payroll_tax_pv_at_62(
+ careers,
+ persons,
+ _params(),
+ shared=False,
+ rates=_Rates(),
+ interest=_Interest(),
+ ),
+ "pv_at_62",
+ ),
+ (
+ lm.report_average_indexed_earnings_22_62(
+ careers, persons, _params(), shared=False
+ ),
+ "average_indexed_earnings",
+ ),
+ ]
+ for result, column in results_and_values:
+ frame = result.frame.set_index("person_id")
+ assert frame.at[1, "status"] == lm.COMPUTED
+ assert frame.at[1, column] == 0.0
+ assert frame.at[2, "status"] == lm.NOT_COMPUTED
+ assert pd.isna(frame.at[2, column])
+
+
+@pytest.mark.parametrize("name", list(lm.AIME_CONVENTIONS))
+def test_post62_earnings_follow_named_oracle_and_are_counted(name):
+ """Differential: initial statutory cutoff and legacy full history."""
+ params = _params()
+ history = {2000: 1_000.0, 2012: 5_000.0, 2015: 9_000.0}
+ convention = lm.AIME_CONVENTIONS[name]
+ row = lm.initial_aime_at_62(
+ _careers([(1, year, value) for year, value in history.items()]),
+ _persons(),
+ params,
+ analysis_year=2020,
+ convention=convention,
+ ).frame.iloc[0]
+ if convention.last_earnings_age is None:
+ expected = benefits.aime(history, 1950, params)
+ assert row["n_history_years_after_age_61"] == 2
+ else:
+ expected = statutory_aime.aime({2000: 1_000.0}, 1950, params)
+ assert row["n_history_years_after_age_61"] == 0
+ assert row["aime"] == expected
+
+
+@PROPERTY_SETTINGS
+@given(
+ st.lists(st.integers(1, 10_000), min_size=2, max_size=25),
+ st.floats(0.01, 10_000.0, allow_nan=False, allow_infinity=False),
+)
+def test_quintiles_are_invariant_to_inexact_positive_weight_scaling(
+ weights, factor
+):
+ """Scale changes preserve labels even when floats round differently."""
+ values = pd.Series(range(len(weights)), dtype="float64")
+ base = pd.Series(weights, dtype="float64")
+ pd.testing.assert_series_equal(
+ lm.weighted_quintiles(values, base),
+ lm.weighted_quintiles(values, base * factor),
+ )
+
+
+def test_quintile_midpoint_boundary_survives_decimal_weight_scaling():
+ """The first group's midpoint is exactly the twenty-percent cut.
+
+ Binary representations of 1.2 and 1.8 must not move the row below
+ that cut after multiplying the weights by the positive factor 0.3.
+ """
+ values = pd.Series([1.0, 2.0])
+ weights = pd.Series([4.0, 6.0])
+ original = lm.weighted_quintiles(values, weights)
+ scaled = lm.weighted_quintiles(values, weights * 0.3)
+ assert original.tolist() == ["Second lowest", "Second highest"]
+ pd.testing.assert_series_equal(original, scaled)
+
+
+def test_overlapping_marriages_use_latest_start_and_flag_ambiguity():
+ episodes = pd.concat(
+ [
+ _episodes(),
+ pd.DataFrame(
+ {
+ "person_id": [1],
+ "marriage_order": [2],
+ "start_year": pd.array([1990], dtype="Int64"),
+ "episode_end_year": pd.array([None], dtype="Int64"),
+ "how_ended": ["intact"],
+ "spouse_person_id": pd.array([3], dtype="Int64"),
+ }
+ ),
+ ],
+ ignore_index=True,
+ )
+ result = lm.lifetime_payroll_tax_pv_at_62(
+ _careers([(1, 2000, 1_000.0), (2, 2000, 3_000.0), (3, 2000, 5_000.0)]),
+ _persons(),
+ _params(),
+ shared=True,
+ marriage_episodes=episodes,
+ rates=_Rates(),
+ interest=_Interest(0.0),
+ )
+ frame = result.frame.set_index("person_id")
+ assert frame.at[1, "pv_at_62"] == 300.0
+ assert frame.at[1, "years_multiple_marriages_in_force"] == 1
+
+
+@pytest.mark.parametrize("measure", ["payroll", "report"])
+def test_unrecorded_sharing_conventions_are_explicit(measure):
+ if measure == "payroll":
+ defaults = lm.PAYROLL_TAX_BUILDER_DEFAULTS
+ else:
+ defaults = lm.REPORT_EARNINGS_BUILDER_DEFAULTS
+ for key in (
+ "reciprocal_history_disagreement",
+ "missing_spouse_year",
+ "missing_own_year",
+ ):
+ assert defaults[key].strip()
+
+
+@pytest.mark.parametrize("measure", ["aime", "payroll", "report"])
+@pytest.mark.parametrize("schedule", ["nawi", "wage_base"])
+def test_parameter_provenance_pins_changed_paths_with_same_revision(
+ measure, schedule
+):
+ """Replacing a parameter path cannot retain its old source fingerprint.
+
+ An alternate wage projection or cap path may be supplied without
+ changing the parameter repository revision. Every result distinguishes
+ those paths, including parameters that do not affect that measure.
+ """
+ careers = _careers([(1, 2000, 15_000.0)])
+ persons = _persons()
+ original = _params()
+ if schedule == "nawi":
+ changed = replace(original, nawi={**original.nawi, 2010: 200.0})
+ else:
+ changed = replace(
+ original, wage_base={**original.wage_base, 1970: 5_000.0}
+ )
+
+ def run(params: SSAParameters) -> lm.MeasureResult:
+ if measure == "aime":
+ return lm.initial_aime_at_62(
+ careers,
+ persons,
+ params,
+ analysis_year=2030,
+ convention=lm.AIME_CONVENTIONS["mint8_initial_aime"],
+ )
+ if measure == "payroll":
+ return lm.lifetime_payroll_tax_pv_at_62(
+ careers,
+ persons,
+ params,
+ shared=False,
+ rates=_Rates(),
+ interest=_Interest(),
+ )
+ return lm.report_average_indexed_earnings_22_62(
+ careers, persons, params, shared=False
+ )
+
+ before, after = run(original), run(changed)
+ before_source = before.provenance["ssa_parameters"]
+ after_source = after.provenance["ssa_parameters"]
+ assert before_source["pe_us_revision"] == after_source["pe_us_revision"]
+ assert (
+ before_source[f"{schedule}_sha256"]
+ != after_source[f"{schedule}_sha256"]
+ )
+ other = "wage_base" if schedule == "nawi" else "nawi"
+ assert before_source[f"{other}_sha256"] == after_source[f"{other}_sha256"]
+ if (measure == "report" and schedule == "wage_base") or (
+ measure == "payroll" and schedule == "nawi"
+ ):
+ pd.testing.assert_frame_equal(before.frame, after.frame)
+ else:
+ assert (
+ before.provenance["output_sha256"]
+ != after.provenance["output_sha256"]
+ )
+
+
+@pytest.mark.parametrize("schedule", ["tax", "interest"])
+def test_payroll_provenance_pins_supplied_rates_without_metadata(schedule):
+ """Callable rate sources with identical empty metadata remain distinct."""
+ careers = _careers([(1, 2000, 1_000.0)])
+ persons = _persons()
+ original_rates, changed_rates = _Rates(), _Rates(0.2)
+ original_interest, changed_interest = _Interest(), _Interest(0.06)
+ before = lm.lifetime_payroll_tax_pv_at_62(
+ careers,
+ persons,
+ _params(),
+ shared=False,
+ rates=original_rates,
+ interest=original_interest,
+ )
+ after = lm.lifetime_payroll_tax_pv_at_62(
+ careers,
+ persons,
+ _params(),
+ shared=False,
+ rates=changed_rates if schedule == "tax" else original_rates,
+ interest=(
+ changed_interest if schedule == "interest" else original_interest
+ ),
+ )
+ hashes = {
+ "tax": ("tax_rates", "applied_rates_sha256"),
+ "interest": ("interest_rates", "available_rates_sha256"),
+ }
+ source, key = hashes[schedule]
+ assert before.provenance[source][key] != after.provenance[source][key]
+ other = "interest" if schedule == "tax" else "tax"
+ source, key = hashes[other]
+ assert before.provenance[source][key] == after.provenance[source][key]
+ assert before.frame.at[0, "pv_at_62"] != after.frame.at[0, "pv_at_62"]
+
+
+def test_quintile_summary_refuses_misaligned_cell_index():
+ values = pd.Series([100.0, 200.0], index=[1, 2])
+ weights = pd.Series([1.0, 1.0], index=[1, 2])
+ labels = lm.weighted_quintiles(values, weights)
+ with pytest.raises(ValueError, match="by must share"):
+ lm.quintile_summary(
+ labels, weights, by=pd.Series(["a", "b"], index=[2, 1])
+ )
+
+
+@pytest.mark.parametrize("measure", ["aime", "report"])
+@pytest.mark.parametrize("year", [2000, 2010])
+@pytest.mark.parametrize("value", [0.0, -1.0, float("nan"), float("inf")])
+def test_indexing_refuses_nonpositive_or_nonfinite_nawi(measure, year, value):
+ """Both source-year and indexing-year NAWI must be finite and positive."""
+ original = _params()
+ params = replace(original, nawi={**original.nawi, year: value})
+ careers = _careers([(1, 2000, 1_000.0)])
+ with pytest.raises(ValueError, match="NAWI.*finite and positive"):
+ if measure == "aime":
+ lm.initial_aime_at_62(
+ careers,
+ _persons(),
+ params,
+ analysis_year=2030,
+ convention=lm.AIME_CONVENTIONS["mint8_initial_aime"],
+ )
+ else:
+ lm.report_average_indexed_earnings_22_62(
+ careers, _persons(), params, shared=False
+ )
+
+
+@pytest.mark.parametrize(
+ "measure", ["aime_statutory", "aime_legacy", "report"]
+)
+def test_negative_wage_base_is_refused_by_every_capped_measure(measure):
+ params = replace(_params(), wage_base={1937: -1.0})
+ careers = _careers([(1, 2000, 1_000.0)])
+ with pytest.raises(ValueError, match="Wage base.*finite and non-negative"):
+ if measure.startswith("aime"):
+ name = (
+ "mint8_initial_aime"
+ if measure == "aime_statutory"
+ else "exercise_1_cola"
+ )
+ lm.initial_aime_at_62(
+ careers,
+ _persons(),
+ params,
+ analysis_year=2030,
+ convention=lm.AIME_CONVENTIONS[name],
+ )
+ else:
+ lm.report_average_indexed_earnings_22_62(
+ careers,
+ _persons(),
+ params,
+ shared=False,
+ conventions=lm.ReportEarningsConventions(
+ cap_at_taxable_maximum=True
+ ),
+ )
+
+
+@pytest.mark.parametrize(
+ "column",
+ [
+ "person_id",
+ "marriage_order",
+ "start_year",
+ "episode_end_year",
+ "spouse_person_id",
+ "separation_year",
+ ],
+)
+def test_fractional_episode_coordinates_are_refused(column):
+ """A fractional year or identifier must never be silently truncated."""
+ episodes = _episodes()
+ if column == "separation_year":
+ episodes[column] = pd.Series([None, None], dtype="float64")
+ else:
+ episodes[column] = episodes[column].astype("float64")
+ episodes.at[0, column] = 2000.5
+ with pytest.raises(ValueError, match="must hold integers"):
+ lm.lifetime_payroll_tax_pv_at_62(
+ _careers([(1, 2000, 1_000.0), (2, 2000, 3_000.0)]),
+ _persons(),
+ _params(),
+ shared=True,
+ marriage_episodes=episodes,
+ rates=_Rates(),
+ interest=_Interest(),
+ )
+
+
+def test_duplicate_episode_orders_are_refused():
+ episodes = _episodes()
+ episodes = pd.concat([episodes, episodes.iloc[[0]]], ignore_index=True)
+ with pytest.raises(ValueError, match="duplicate marriage orders"):
+ lm.lifetime_payroll_tax_pv_at_62(
+ _careers([(1, 2000, 1_000.0), (2, 2000, 3_000.0)]),
+ _persons(),
+ _params(),
+ shared=True,
+ marriage_episodes=episodes,
+ rates=_Rates(),
+ interest=_Interest(),
+ )
+
+
+@pytest.mark.parametrize("rate", [float("nan"), float("inf"), -0.1])
+def test_annual_taxes_refuse_invalid_schedule_rates(rate):
+ with pytest.raises(ValueError):
+ lm.annual_payroll_taxes(
+ _careers([(1, 2000, 1_000.0)]), _params(), _Rates(rate)
+ )
+
+
+@pytest.mark.parametrize("rate", [float("nan"), float("inf"), -1.0, -1.01])
+@pytest.mark.parametrize("years", [(2000, 2001), (2001, 2000)])
+def test_accumulation_refuses_invalid_interest_in_both_directions(rate, years):
+ with pytest.raises(ValueError):
+ lm.accumulation_factor(*years, _Interest(rate))
+
+
+def test_returned_provenance_cannot_change_later_results():
+ def run():
+ return lm.report_average_indexed_earnings_22_62(
+ _careers([(1, 2000, 1_000.0)]),
+ _persons(),
+ _params(),
+ shared=False,
+ )
+
+ first = run()
+ expected = first.provenance["conventions"]["builder_defaults"]["divisor"]
+ first.provenance["conventions"]["builder_defaults"]["divisor"] = "edited"
+ assert run().provenance["conventions"]["builder_defaults"]["divisor"] == (
+ expected
+ )
+
+
+def test_summary_keeps_empty_quintiles_for_disclosure():
+ weights = pd.Series([1.0])
+ labels = lm.weighted_quintiles(pd.Series([100.0]), weights)
+ summary = lm.quintile_summary(labels, weights).set_index("label")
+ for label in lm.QUINTILE_LABELS:
+ assert summary.at[label, "n"] == int(label == "Middle")
+ assert summary.at["Lowest", "weight"] == 0
+ assert summary.at["Lowest", "weight_share"] == 0
+
+
+def test_summary_all_missing_values_has_undefined_weight_shares():
+ weights = pd.Series([1.0])
+ labels = lm.weighted_quintiles(pd.Series([float("nan")]), weights)
+ summary = lm.quintile_summary(labels, weights)
+ assert summary["weight_share"].isna().all()
+ assert summary.loc[summary["label"] == lm.NOT_COMPUTED, "n"].item() == 1
+
+
+@pytest.mark.parametrize("measure", ["payroll", "report"])
+def test_marriage_history_roster_is_pinned(measure):
+ def run(roster):
+ args = {
+ "shared": True,
+ "marriage_episodes": _episodes(),
+ "marriage_history_person_ids": roster,
+ }
+ if measure == "payroll":
+ return lm.lifetime_payroll_tax_pv_at_62(
+ _careers([(1, 2000, 1_000.0)]),
+ _persons(),
+ _params(),
+ rates=_Rates(),
+ interest=_Interest(),
+ **args,
+ )
+ return lm.report_average_indexed_earnings_22_62(
+ _careers([(1, 2000, 1_000.0)]), _persons(), _params(), **args
+ )
+
+ before, reordered, changed = run([1, 2]), run([2, 1]), run([1])
+ assert before.provenance["inputs"] == reordered.provenance["inputs"]
+ key = "marriage_history_person_ids_sha256"
+ assert (
+ before.provenance["inputs"][key] != changed.provenance["inputs"][key]
+ )
+
+
+@pytest.mark.parametrize("measure", ["payroll", "report"])
+def test_shared_switch_refuses_truthy_text(measure):
+ args = {"shared": "False"}
+ with pytest.raises(TypeError, match="shared must be a bool"):
+ if measure == "payroll":
+ lm.lifetime_payroll_tax_pv_at_62(
+ _careers([(1, 2000, 1_000.0)]),
+ _persons(),
+ _params(),
+ rates=_Rates(),
+ interest=_Interest(),
+ **args,
+ )
+ else:
+ lm.report_average_indexed_earnings_22_62(
+ _careers([(1, 2000, 1_000.0)]), _persons(), _params(), **args
+ )
From d49785d7bcf1bca116fe8eb3cad491ceb6de6577 Mon Sep 17 00:00:00 2001
From: Max Ghenis
Date: Fri, 2 Oct 2026 16:54:41 -0400
Subject: [PATCH 04/10] Add PSID race, education and country-of-birth
attributes for breakdowns
data/group_attributes_psid.py reads, with exact label checks and
documented code domains, the family-file race and Hispanic-origin
reports of reference persons and spouses, the individual file's
completed education by wave, and country of birth where the PSID asks
it. cohorts/group_attributes.py turns them into a person-keyed side
frame for the 2009 and 2011 projection cohorts, the age-67
observations and the Track M 2023 universe, mapped to MINT8's
categories and to the Butrica-Uccello report rows. Conflicts across
waves resolve deterministically; people never asked (other family-unit
members) get an explicit, counted unknown. The side frame has its own
file audit and SHA-256 seals, so the existing cohort frames and their
pinned digests do not change.
Invariants tested: every requested person keeps exactly one row;
labels, domains and source pins are enforced; resolution is
deterministic; mutation is detected.
Co-Authored-By: Claude Opus 5.5
---
data/external/group_category_schemes_v1.json | 674 ++
...id_group_attribute_codebook_values_v1.json | 9920 +++++++++++++++++
.../ssa_mint8_payroll_option_row_labels.json | 1851 +++
...l_option_row_labels.source.provenance.json | 11 +
.../ssa_mint8_table_user_guide.source.html | 633 ++
...t8_table_user_guide.source.provenance.json | 14 +
docs/design/nasi_g1_group_attributes.md | 522 +
.../build_group_attribute_codebook_values.py | 214 +
.../cohorts/group_attributes.py | 1102 ++
.../data/group_attributes_psid.py | 2000 ++++
tests/cohorts/test_group_attributes.py | 792 ++
.../test_group_attributes_integration.py | 182 +
tests/cohorts/test_group_category_schemes.py | 296 +
tests/data/test_group_attributes_psid.py | 507 +
14 files changed, 18718 insertions(+)
create mode 100644 data/external/group_category_schemes_v1.json
create mode 100644 data/external/psid_group_attribute_codebook_values_v1.json
create mode 100644 data/external/ssa_mint8_payroll_option_row_labels.json
create mode 100644 data/external/ssa_mint8_payroll_option_row_labels.source.provenance.json
create mode 100644 data/external/ssa_mint8_table_user_guide.source.html
create mode 100644 data/external/ssa_mint8_table_user_guide.source.provenance.json
create mode 100644 docs/design/nasi_g1_group_attributes.md
create mode 100644 scripts/build_group_attribute_codebook_values.py
create mode 100644 src/populace_dynamics/cohorts/group_attributes.py
create mode 100644 src/populace_dynamics/data/group_attributes_psid.py
create mode 100644 tests/cohorts/test_group_attributes.py
create mode 100644 tests/cohorts/test_group_attributes_integration.py
create mode 100644 tests/cohorts/test_group_category_schemes.py
create mode 100644 tests/data/test_group_attributes_psid.py
diff --git a/data/external/group_category_schemes_v1.json b/data/external/group_category_schemes_v1.json
new file mode 100644
index 00000000..ed725078
--- /dev/null
+++ b/data/external/group_category_schemes_v1.json
@@ -0,0 +1,674 @@
+{
+ "schema_version": "group_category_schemes.v1",
+ "note": "Report category schemes for the person attributes built by populace_dynamics.cohorts.group_attributes. Each scheme maps a harmonized attribute to a report's row labels. A value a source does not place is left unassigned with a named reason, never guessed. Labels are verbatim from the cited source; every rule that is this file's reading rather than the source's own definition is listed under 'assumptions'.",
+ "sources": {
+ "mint8_user_guide": {
+ "title": "Table User Guide—Modeling Income in the Near Term (MINT) 8",
+ "url": "https://www.ssa.gov/policy/docs/projections/user-guide.html",
+ "date_certified": "2025-10-01",
+ "capture": "https://web.archive.org/web/20260419231110id_/https://www.ssa.gov/policy/docs/projections/user-guide.html",
+ "committed_file": "data/external/ssa_mint8_table_user_guide.source.html",
+ "sha256": "278d5d19c1b50f1d354db1ada515288af563c035a67eb16fb971d25700fb94e9",
+ "locator": "Definitions—Table Rows and Columns > Characteristic Subgroups—Table Rows"
+ },
+ "mint8_row_labels": {
+ "title": "MINT 8 policy-option tables, 'Increase payroll tax rate': row labels only",
+ "url": "https://www.ssa.gov/policy/docs/projections/policy-options/increase-payroll-tax-rate.html",
+ "capture": "http://web.archive.org/web/20260519190449id_/https://www.ssa.gov/policy/docs/projections/policy-options/increase-payroll-tax-rate.html",
+ "committed_file": "data/external/ssa_mint8_payroll_option_row_labels.json",
+ "sha256": "23fbfbc8dbc14b06144a83bf0583ea0a6c89505638f3f7dc5fd6b34a3a4ac650",
+ "locator": "Tables 1–3 (benefits), 7–9 (household income), 10–12 (official poverty), 13–16 (benefit/tax ratio), and 17–20 (initial replacement rate): captions and row labels only; no data cells."
+ },
+ "boomers2004_report_rows": {
+ "title": "Butrica and Uccello (2004), How Will Boomers Fare at Retirement?, Tables 19 and 21 row labels as recorded in this repository",
+ "repository_locator": "src/populace_dynamics/estimates/uniform_cut_tabulation.py NOT_COMPUTED_REPORT_ROWS['race_ethnicity'] and ['education']; docs/design/boomers2004_uniform_cut_comparison.md, 'Report rows not computed in v1 (named omissions)'",
+ "note": "The repository records the Report's row labels as printed (cleared extract) and, for education, that the Report does not say whether 'High school graduate' includes some college. It records no other definition of these rows. The Report itself was not opened for this file."
+ }
+ },
+ "schemes": {
+ "mint8": {
+ "title": "SSA MINT 8 characteristic subgroups",
+ "dimensions": {
+ "race_ethnicity": {
+ "column": "race_ethnicity_mint8",
+ "group_label": "Race and ethnicity",
+ "source": "mint8_user_guide",
+ "quote": "Race/Ethnicity: We list “Hispanic or Latino, any race” first; the rest of the groups (White, Black or African American, and All other races) are non-Hispanic.",
+ "labels_source": "mint8_row_labels",
+ "categories": {
+ "hispanic": "Hispanic or Latino, any race",
+ "white_non_hispanic": "White, non-Hispanic",
+ "black_non_hispanic": "Black or African American, non-Hispanic",
+ "other_non_hispanic": "All other races, non-Hispanic"
+ },
+ "multiple_races": "other_non_hispanic",
+ "assumptions": [
+ "SSA's guide does not say where a non-Hispanic person who reports two or more races goes. This scheme places such a person in 'All other races, non-Hispanic', reading 'White' and 'Black or African American' as single-race groups (the 'alone' tabulation of the OMB 1997 standard). This is an explicit builder convention pending registration; it is not an SSA definition."
+ ]
+ },
+ "education": {
+ "column": "education_mint8",
+ "group_label": "Highest education level",
+ "source": "mint8_user_guide",
+ "quote": "Highest Education Level: Reported number of years of education. Graduate means more than 16 years of education, Bachelor means 16 years of education, Associate means 14–15 years of education, High school means 12–13 years of education, and Less than high school means less than 12 years of education.",
+ "labels_source": "mint8_row_labels",
+ "bands": [
+ {
+ "label": "Graduate",
+ "min": 17,
+ "max": null
+ },
+ {
+ "label": "Bachelor",
+ "min": 16,
+ "max": 16
+ },
+ {
+ "label": "Associate",
+ "min": 14,
+ "max": 15
+ },
+ {
+ "label": "High school",
+ "min": 12,
+ "max": 13
+ },
+ {
+ "label": "Less than high school",
+ "min": 0,
+ "max": 11
+ }
+ ],
+ "unresolved_bands": [],
+ "assumptions": [
+ "The PSID's top code 17 is 'at least some post-graduate work', its only value above 16, so it is the guide's 'more than 16 years' (Graduate)."
+ ]
+ },
+ "country_of_birth": {
+ "column": "country_of_birth_mint8",
+ "group_label": "Country of birth",
+ "source": "mint8_user_guide",
+ "quote": "Country of Birth: We differentiate between the United States and other countries.",
+ "labels_source": "mint8_row_labels",
+ "categories": {
+ "united_states": "United States",
+ "foreign_country": "Other countries"
+ },
+ "unresolved": {
+ "us_territory": "SSA's guide does not say whether a person born in a U.S. territory counts as born in the United States or in other countries."
+ },
+ "assumptions": []
+ }
+ },
+ "table_profiles": {
+ "annual_benefits": {
+ "source": "mint8_row_labels",
+ "locator": "tables 1, 2, 3",
+ "row_groups": [
+ {
+ "group": "Total",
+ "labels": []
+ },
+ {
+ "group": "Sex",
+ "labels": [
+ "Female",
+ "Male"
+ ]
+ },
+ {
+ "group": "Race and ethnicity",
+ "labels": [
+ "Hispanic or Latino, any race",
+ "White, non-Hispanic",
+ "Black or African American, non-Hispanic",
+ "All other races, non-Hispanic"
+ ]
+ },
+ {
+ "group": "Country of birth",
+ "labels": [
+ "United States",
+ "Other countries"
+ ]
+ },
+ {
+ "group": "Age",
+ "labels": [
+ "60–69",
+ "70–79",
+ "80–89",
+ "90 or older"
+ ]
+ },
+ {
+ "group": "Marital status",
+ "labels": [
+ "Married",
+ "Divorced",
+ "Widowed",
+ "Never married"
+ ]
+ },
+ {
+ "group": "Highest education level",
+ "labels": [
+ "Graduate",
+ "Bachelor",
+ "Associate",
+ "High school",
+ "Less than high school"
+ ]
+ },
+ {
+ "group": "Current-law poverty status",
+ "labels": [
+ "Above poverty",
+ "In poverty"
+ ]
+ },
+ {
+ "group": "Current-law household income quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Current-law benefit type",
+ "labels": [
+ "Retired worker only",
+ "Widow(er) (includes dually entitled)",
+ "Spousal (includes dually entitled)",
+ "Disabled worker only"
+ ]
+ }
+ ],
+ "analysis_years": [
+ 2030,
+ 2050,
+ 2070
+ ],
+ "population": "Current-law beneficiaries aged 60 or older"
+ },
+ "annual_household_income": {
+ "source": "mint8_row_labels",
+ "locator": "tables 7, 8, 9",
+ "row_groups": [
+ {
+ "group": "Total",
+ "labels": []
+ },
+ {
+ "group": "Sex",
+ "labels": [
+ "Female",
+ "Male"
+ ]
+ },
+ {
+ "group": "Race and ethnicity",
+ "labels": [
+ "Hispanic or Latino, any race",
+ "White, non-Hispanic",
+ "Black or African American, non-Hispanic",
+ "All other races, non-Hispanic"
+ ]
+ },
+ {
+ "group": "Country of birth",
+ "labels": [
+ "United States",
+ "Other countries"
+ ]
+ },
+ {
+ "group": "Age",
+ "labels": [
+ "60–69",
+ "70–79",
+ "80–89",
+ "90 or older"
+ ]
+ },
+ {
+ "group": "Marital status",
+ "labels": [
+ "Married",
+ "Divorced",
+ "Widowed",
+ "Never married"
+ ]
+ },
+ {
+ "group": "Highest education level",
+ "labels": [
+ "Graduate",
+ "Bachelor",
+ "Associate",
+ "High school",
+ "Less than high school"
+ ]
+ },
+ {
+ "group": "Current-law poverty status",
+ "labels": [
+ "Above poverty",
+ "In poverty"
+ ]
+ },
+ {
+ "group": "Current-law household income quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Current-law benefit type",
+ "labels": [
+ "Retired worker only",
+ "Widow(er) (includes dually entitled)",
+ "Spousal (includes dually entitled)",
+ "Disabled worker only"
+ ]
+ }
+ ],
+ "analysis_years": [
+ 2030,
+ 2050,
+ 2070
+ ],
+ "population": "Current-law beneficiaries aged 60 or older"
+ },
+ "annual_official_poverty": {
+ "source": "mint8_row_labels",
+ "locator": "tables 10, 11, 12",
+ "row_groups": [
+ {
+ "group": "Total",
+ "labels": []
+ },
+ {
+ "group": "Sex",
+ "labels": [
+ "Female",
+ "Male"
+ ]
+ },
+ {
+ "group": "Race and ethnicity",
+ "labels": [
+ "Hispanic or Latino, any race",
+ "White, non-Hispanic",
+ "Black or African American, non-Hispanic",
+ "All other races, non-Hispanic"
+ ]
+ },
+ {
+ "group": "Country of birth",
+ "labels": [
+ "United States",
+ "Other countries"
+ ]
+ },
+ {
+ "group": "Age",
+ "labels": [
+ "60–69",
+ "70–79",
+ "80–89",
+ "90 or older"
+ ]
+ },
+ {
+ "group": "Marital status",
+ "labels": [
+ "Married",
+ "Divorced",
+ "Widowed",
+ "Never married"
+ ]
+ },
+ {
+ "group": "Highest education level",
+ "labels": [
+ "Graduate",
+ "Bachelor",
+ "Associate",
+ "High school",
+ "Less than high school"
+ ]
+ },
+ {
+ "group": "Current-law poverty status",
+ "labels": [
+ "Above poverty",
+ "In poverty"
+ ]
+ },
+ {
+ "group": "Current-law benefit type",
+ "labels": [
+ "Retired worker only",
+ "Widow(er) (includes dually entitled)",
+ "Spousal (includes dually entitled)",
+ "Disabled worker only"
+ ]
+ }
+ ],
+ "analysis_years": [
+ 2030,
+ 2050,
+ 2070
+ ],
+ "population": "Current-law beneficiaries aged 60 or older"
+ },
+ "cohort_benefit_tax_ratio": {
+ "source": "mint8_row_labels",
+ "locator": "tables 13, 14, 15, 16",
+ "row_groups": [
+ {
+ "group": "Total",
+ "labels": []
+ },
+ {
+ "group": "Sex",
+ "labels": [
+ "Female",
+ "Male"
+ ]
+ },
+ {
+ "group": "Race and ethnicity",
+ "labels": [
+ "Hispanic or Latino, any race",
+ "White, non-Hispanic",
+ "Black or African American, non-Hispanic",
+ "All other races, non-Hispanic"
+ ]
+ },
+ {
+ "group": "Country of birth",
+ "labels": [
+ "United States",
+ "Other countries"
+ ]
+ },
+ {
+ "group": "Highest education level",
+ "labels": [
+ "Graduate",
+ "Bachelor",
+ "Associate",
+ "High school",
+ "Less than high school"
+ ]
+ },
+ {
+ "group": "Current-law initial AIME quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Lifetime payroll tax quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Lifetime payroll tax quintile (shared)",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ }
+ ],
+ "birth_cohorts": [
+ [
+ 1960,
+ 1969
+ ],
+ [
+ 1980,
+ 1989
+ ],
+ [
+ 2000,
+ 2009
+ ],
+ [
+ 2020,
+ 2029
+ ]
+ ],
+ "population": "Workers with a benefit/tax ratio"
+ },
+ "cohort_initial_replacement_rate": {
+ "source": "mint8_row_labels",
+ "locator": "tables 17, 18, 19, 20",
+ "row_groups": [
+ {
+ "group": "Total",
+ "labels": []
+ },
+ {
+ "group": "Sex",
+ "labels": [
+ "Female",
+ "Male"
+ ]
+ },
+ {
+ "group": "Race and ethnicity",
+ "labels": [
+ "Hispanic or Latino, any race",
+ "White, non-Hispanic",
+ "Black or African American, non-Hispanic",
+ "All other races, non-Hispanic"
+ ]
+ },
+ {
+ "group": "Country of birth",
+ "labels": [
+ "United States",
+ "Other countries"
+ ]
+ },
+ {
+ "group": "Highest education level",
+ "labels": [
+ "Graduate",
+ "Bachelor",
+ "Associate",
+ "High school",
+ "Less than high school"
+ ]
+ },
+ {
+ "group": "Current-law initial AIME quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Lifetime payroll tax quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Lifetime payroll tax quintile (shared)",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ }
+ ],
+ "birth_cohorts": [
+ [
+ 1960,
+ 1969
+ ],
+ [
+ 1980,
+ 1989
+ ],
+ [
+ 2000,
+ 2009
+ ],
+ [
+ 2020,
+ 2029
+ ]
+ ],
+ "population": "Current-law beneficiaries with an initial replacement rate"
+ }
+ },
+ "lifetime_earnings_dimensions": {
+ "initial_aime": {
+ "group_label": "Current-law initial AIME quintile",
+ "column": "initial_aime_quintile_mint8",
+ "source": "mint8_user_guide",
+ "locator": "Definitions—Table Rows and Columns > Characteristic Subgroups—Table Rows > Current-Law Initial AIME Quintile (#AIME)",
+ "definition": "Average indexed monthly earnings under current law at age 62, the earliest eligibility age for retired-worker benefits.",
+ "reference_age": 62,
+ "quintile_population": "each birth cohort",
+ "labels_source": "mint8_row_labels",
+ "labels_locator": "tables 13–20, group Current-law initial AIME quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ],
+ "availability": "Supplied current-law measures only. These side-frame readers do not invent missing lifetime earnings, tax histories, trust-fund interest rates, or quintile thresholds."
+ },
+ "payroll_tax_own": {
+ "group_label": "Lifetime payroll tax quintile",
+ "column": "lifetime_payroll_tax_quintile_mint8",
+ "source": "mint8_user_guide",
+ "locator": "Definitions—Table Rows and Columns > Characteristic Subgroups—Table Rows > Lifetime Payroll Tax Quintile (#lifetime-tax); Explanation of Table Types > Benefit/Tax Ratios (present-value paragraph)",
+ "definition": "Present value at age 62 of the individual's current-law payroll taxes over a lifetime, adjusted using the Social Security Trust Fund interest rate.",
+ "reference_age": 62,
+ "discount_rate": "Social Security Trust Fund interest rate",
+ "quintile_population": "each birth cohort",
+ "labels_source": "mint8_row_labels",
+ "labels_locator": "tables 13–20, group Lifetime payroll tax quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ],
+ "availability": "Supplied current-law measures only. These side-frame readers do not invent missing lifetime earnings, tax histories, trust-fund interest rates, or quintile thresholds."
+ },
+ "payroll_tax_shared": {
+ "group_label": "Lifetime payroll tax quintile (shared)",
+ "column": "lifetime_payroll_tax_shared_quintile_mint8",
+ "source": "mint8_user_guide",
+ "locator": "Definitions—Table Rows and Columns > Characteristic Subgroups—Table Rows > Lifetime Payroll Tax Quintile (Shared) (#lifetime-tax-shared); Explanation of Table Types > Benefit/Tax Ratios (present-value paragraph)",
+ "definition": "Present value at age 62 of current-law payroll taxes, adjusted using the Social Security Trust Fund interest rate. Taxes paid while married are shared equally between spouses; unmarried years use only the individual's taxes. Never-married individuals have the same own and shared values.",
+ "reference_age": 62,
+ "discount_rate": "Social Security Trust Fund interest rate",
+ "quintile_population": "each birth cohort",
+ "married_year_rule": "(own payroll taxes + spouse payroll taxes) / 2",
+ "unmarried_year_rule": "own payroll taxes",
+ "labels_source": "mint8_row_labels",
+ "labels_locator": "tables 13–20, group Lifetime payroll tax quintile (shared)",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ],
+ "availability": "Supplied current-law measures only. These side-frame readers do not invent missing lifetime earnings, tax histories, trust-fund interest rates, or quintile thresholds."
+ }
+ },
+ "unresolved_definitions": {
+ "separated_marital_status": "The guide lists Married, Divorced, Widowed and Never married but does not assign separated people.",
+ "country_of_birth_us_territories": "SSA's guide does not say whether a person born in a U.S. territory counts as born in the United States or in other countries.",
+ "multiracial_non_hispanic": "The guide does not specify a multiple-race rule; the attribute scheme states a builder convention pending registration.",
+ "lifetime_quintile_ties": "The guide does not specify how equal measure values at a weighted quintile boundary are assigned. No tie rule or dollar threshold is inferred."
+ }
+ },
+ "boomers2004": {
+ "title": "Butrica and Uccello (2004) Tables 19 and 21 rows (Track U's named omissions)",
+ "dimensions": {
+ "race_ethnicity": {
+ "column": "race_ethnicity_report4",
+ "group_label": "Race/Ethnicity",
+ "source": "boomers2004_report_rows",
+ "quote": null,
+ "labels_source": "boomers2004_report_rows",
+ "categories": {
+ "white_non_hispanic": "White, non-hispanic",
+ "black_non_hispanic": "Black, non-hispanic",
+ "hispanic": "Hispanic",
+ "other_non_hispanic": "Other"
+ },
+ "multiple_races": "other_non_hispanic",
+ "assumptions": [
+ "The repository records only the row labels. Because the White and Black rows are marked non-Hispanic, this scheme reads 'Hispanic' as Hispanic of any race and 'Other' as every other non-Hispanic person; both readings are provisional, not Report definitions.",
+ "Persons reporting two or more races are placed in 'Other', as in scheme mint8; the Report's rule is not recorded."
+ ]
+ },
+ "education": {
+ "column": "education_report3",
+ "group_label": "Education",
+ "source": "boomers2004_report_rows",
+ "quote": null,
+ "labels_source": "boomers2004_report_rows",
+ "bands": [],
+ "unresolved_bands": [
+ {
+ "min": 0,
+ "max": 17,
+ "status": "definition_not_recorded",
+ "reason": "The cleared repository records only these three row labels and the uncertainty about some college. It records no numeric years-of-schooling definitions for any row, so no cutoff is inferred."
+ }
+ ],
+ "assumptions": [],
+ "row_labels": [
+ "High school dropout",
+ "High school graduate",
+ "College graduate"
+ ]
+ }
+ }
+ }
+ }
+}
diff --git a/data/external/psid_group_attribute_codebook_values_v1.json b/data/external/psid_group_attribute_codebook_values_v1.json
new file mode 100644
index 00000000..4ad28808
--- /dev/null
+++ b/data/external/psid_group_attribute_codebook_values_v1.json
@@ -0,0 +1,9920 @@
+{
+ "extraction": "pdftotext version 26.09.0 -layout; value rows are the codebook's 'Count / % / Value/Range Code / Value/Range Text' table with continuation lines joined (counts and percents dropped); page is the 1-based PDF page of the variable header",
+ "family": {
+ "1985": {
+ "codebook": {
+ "path": "family/1985/FAM1985_codebook.pdf",
+ "sha256": "ef9cdf2eafccf167b5ab0495cab085877dce4787ccb9a2c01fdf4a868d13909a"
+ },
+ "formats_sas": null,
+ "variables": {
+ "V11937": {
+ "label": "G31 SPANISH DESCENT-HEAD",
+ "page": 277,
+ "values": [
+ {
+ "code": 0,
+ "text": "Inap.: is not Spanish/Hispanic"
+ },
+ {
+ "code": 1,
+ "text": "Mexican"
+ },
+ {
+ "code": 2,
+ "text": "Mexican American"
+ },
+ {
+ "code": 3,
+ "text": "Chicano"
+ },
+ {
+ "code": 4,
+ "text": "Puerto Rican"
+ },
+ {
+ "code": 5,
+ "text": "Cuban"
+ },
+ {
+ "code": 6,
+ "text": "Combination; more than 1 mention"
+ },
+ {
+ "code": 7,
+ "text": "Other Spanish"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "V11938": {
+ "label": "G32 RACE OF HEAD (1 MEN)",
+ "page": 277,
+ "values": [
+ {
+ "code": 0,
+ "text": "Inap.: no second mention"
+ },
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 8,
+ "text": "More than 2 mentions"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "V11939": {
+ "label": "G32 RACE OF HEAD (2 MEN)",
+ "page": 278,
+ "values": [
+ {
+ "code": 0,
+ "text": "Inap.: no second mention"
+ },
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 8,
+ "text": "More than 2 mentions"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "V12292": {
+ "label": "N31 SPANISH DESCENT-WIFE",
+ "page": 413,
+ "values": [
+ {
+ "code": 0,
+ "text": "Inap.: is not Spanish/Hispanic; Head is male, no Wife/\"Wife\" in FU now (V11999=2); Head is female (V12261=2)"
+ },
+ {
+ "code": 1,
+ "text": "Mexican"
+ },
+ {
+ "code": 2,
+ "text": "Mexican American"
+ },
+ {
+ "code": 3,
+ "text": "Chicano"
+ },
+ {
+ "code": 4,
+ "text": "Puerto Rican"
+ },
+ {
+ "code": 5,
+ "text": "Cuban"
+ },
+ {
+ "code": 6,
+ "text": "Combination; more than 1 mention"
+ },
+ {
+ "code": 7,
+ "text": "Other Spanish"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "V12293": {
+ "label": "N32 RACE OF WIFE (1 MEN)",
+ "page": 413,
+ "values": [
+ {
+ "code": 0,
+ "text": "Inap.: Head is male, no Wife/\"Wife\" in FU now (V11999=2); Head is female (V12261=2)"
+ },
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 8,
+ "text": "More than 2 mentions"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "V12294": {
+ "label": "N32 RACE OF WIFE (2 MEN)",
+ "page": 414,
+ "values": [
+ {
+ "code": 0,
+ "text": "Inap.: no second mention; Head is male, no Wife/ \"Wife\" in FU now (V11999=2); Head is female (V12261=2)"
+ },
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 8,
+ "text": "More than 2 mentions"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ }
+ }
+ },
+ "1986": {
+ "codebook": {
+ "path": "family/1986/FAM1986_codebook.pdf",
+ "sha256": "6483a31953e3f969b8757c97779ff6f2c45734b54a2806e8fd506db72c220112"
+ },
+ "formats_sas": null,
+ "variables": {
+ "V13499": {
+ "label": "K18 SPANISH DESCENT WF",
+ "page": 344,
+ "values": [
+ {
+ "code": 0,
+ "text": "Inap.: is not Spanish/Hispanic; no Wife/\"Wife\" in FU (V13484=5)"
+ },
+ {
+ "code": 1,
+ "text": "Mexican"
+ },
+ {
+ "code": 2,
+ "text": "Mexican American"
+ },
+ {
+ "code": 3,
+ "text": "Chicano"
+ },
+ {
+ "code": 4,
+ "text": "Puerto Rican"
+ },
+ {
+ "code": 5,
+ "text": "Cuban"
+ },
+ {
+ "code": 6,
+ "text": "Combination; more than one mention"
+ },
+ {
+ "code": 7,
+ "text": "Other Spanish"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "V13500": {
+ "label": "K19 RACE OF WIFE 1",
+ "page": 345,
+ "values": [
+ {
+ "code": 0,
+ "text": "Inap.: no Wife/\"Wife\" in FU (V13484=5)"
+ },
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "V13501": {
+ "label": "K19 RACE OF WIFE 2",
+ "page": 345,
+ "values": [
+ {
+ "code": 0,
+ "text": "Inap.: no second mention; no Wife/\"Wife\" in FU (V13484=5)"
+ },
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 8,
+ "text": "More than two mentions"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "V13564": {
+ "label": "L31 SPANISH DESCENT HD",
+ "page": 376,
+ "values": [
+ {
+ "code": 0,
+ "text": "Inap.: is not Spanish/Hispanic"
+ },
+ {
+ "code": 1,
+ "text": "Mexican"
+ },
+ {
+ "code": 2,
+ "text": "Mexican American"
+ },
+ {
+ "code": 3,
+ "text": "Chicano"
+ },
+ {
+ "code": 4,
+ "text": "Puerto Rican"
+ },
+ {
+ "code": 5,
+ "text": "Cuban"
+ },
+ {
+ "code": 6,
+ "text": "Combination; more than 1 mention"
+ },
+ {
+ "code": 7,
+ "text": "Other Spanish"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "V13565": {
+ "label": "L32 RACE OF HEAD 1",
+ "page": 376,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "V13566": {
+ "label": "L32 RACE OF HEAD 2",
+ "page": 377,
+ "values": [
+ {
+ "code": 0,
+ "text": "Inap.: no second mention"
+ },
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 8,
+ "text": "More than two mentions"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ }
+ }
+ },
+ "1987": {
+ "codebook": {
+ "path": "family/1987/FAM1987_codebook.pdf",
+ "sha256": "86065d5cb127766a27037735ee428d1cec5afeeac554bfc9cc2cabfebd300056"
+ },
+ "formats_sas": null,
+ "variables": {
+ "V14546": {
+ "label": "K18 SPANISH DESCENT WF",
+ "page": 296,
+ "values": [
+ {
+ "code": 0,
+ "text": "Inap.: is not Spanish/Hispanic; no Wife/\"Wife\" in FU (V14531=5)"
+ },
+ {
+ "code": 1,
+ "text": "Mexican"
+ },
+ {
+ "code": 2,
+ "text": "Mexican American"
+ },
+ {
+ "code": 3,
+ "text": "Chicano"
+ },
+ {
+ "code": 4,
+ "text": "Puerto Rican"
+ },
+ {
+ "code": 5,
+ "text": "Cuban"
+ },
+ {
+ "code": 6,
+ "text": "Combination; more than one mention"
+ },
+ {
+ "code": 7,
+ "text": "Other Spanish"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "V14547": {
+ "label": "K19 RACE OF WIFE 1",
+ "page": 296,
+ "values": [
+ {
+ "code": 0,
+ "text": "Inap.: no Wife/\"Wife\" in FU (V14531=5)"
+ },
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "V14548": {
+ "label": "K19 RACE OF WIFE 2",
+ "page": 297,
+ "values": [
+ {
+ "code": 0,
+ "text": "Inap.: no second mention; no Wife/\"Wife\" in FU (V14531=5)"
+ },
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 8,
+ "text": "More than two mentions"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "V14611": {
+ "label": "L31 SPANISH DESCENT HD",
+ "page": 326,
+ "values": [
+ {
+ "code": 0,
+ "text": "Inap.: is not Spanish/Hispanic"
+ },
+ {
+ "code": 1,
+ "text": "Mexican"
+ },
+ {
+ "code": 2,
+ "text": "Mexican American"
+ },
+ {
+ "code": 3,
+ "text": "Chicano"
+ },
+ {
+ "code": 4,
+ "text": "Puerto Rican"
+ },
+ {
+ "code": 5,
+ "text": "Cuban"
+ },
+ {
+ "code": 6,
+ "text": "Combination; more than 1 mention"
+ },
+ {
+ "code": 7,
+ "text": "Other Spanish"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "V14612": {
+ "label": "L32 RACE OF HEAD 1",
+ "page": 326,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "V14613": {
+ "label": "L32 RACE OF HEAD 2",
+ "page": 327,
+ "values": [
+ {
+ "code": 0,
+ "text": "Inap.: no second mention"
+ },
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 8,
+ "text": "More than two mentions"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ }
+ }
+ },
+ "1988": {
+ "codebook": {
+ "path": "family/1988/FAM1988_codebook.pdf",
+ "sha256": "4e2a75c05ee4b2ca9d44a31063096567f45c97285127a124d2d98a913f772c8f"
+ },
+ "formats_sas": null,
+ "variables": {
+ "V16020": {
+ "label": "K18 SPANISH DESCENT WF",
+ "page": 406,
+ "values": [
+ {
+ "code": 0,
+ "text": "Inap.: is not Spanish/Hispanic; no Wife/\"Wife\" in FU (V16005=5)"
+ },
+ {
+ "code": 1,
+ "text": "Mexican"
+ },
+ {
+ "code": 2,
+ "text": "Mexican American"
+ },
+ {
+ "code": 3,
+ "text": "Chicano"
+ },
+ {
+ "code": 4,
+ "text": "Puerto Rican"
+ },
+ {
+ "code": 5,
+ "text": "Cuban"
+ },
+ {
+ "code": 6,
+ "text": "Combination; more than one mention"
+ },
+ {
+ "code": 7,
+ "text": "Other Spanish"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "V16021": {
+ "label": "K19 RACE OF WIFE 1",
+ "page": 406,
+ "values": [
+ {
+ "code": 0,
+ "text": "Inap.: no Wife/\"Wife\" in FU (V16005=5)"
+ },
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "V16022": {
+ "label": "K19 RACE OF WIFE 2",
+ "page": 407,
+ "values": [
+ {
+ "code": 0,
+ "text": "Inap.: no second mention; no Wife/\"Wife\" in FU (V16005=5)"
+ },
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 8,
+ "text": "More than two mentions"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "V16085": {
+ "label": "L31 SPANISH DESCENT HD",
+ "page": 437,
+ "values": [
+ {
+ "code": 0,
+ "text": "Inap.: is not Spanish/Hispanic"
+ },
+ {
+ "code": 1,
+ "text": "Mexican"
+ },
+ {
+ "code": 2,
+ "text": "Mexican American"
+ },
+ {
+ "code": 3,
+ "text": "Chicano"
+ },
+ {
+ "code": 4,
+ "text": "Puerto Rican"
+ },
+ {
+ "code": 5,
+ "text": "Cuban"
+ },
+ {
+ "code": 6,
+ "text": "Combination; more than 1 mention"
+ },
+ {
+ "code": 7,
+ "text": "Other Spanish"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "V16086": {
+ "label": "L32 RACE OF HEAD 1",
+ "page": 438,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "V16087": {
+ "label": "L32 RACE OF HEAD 2",
+ "page": 438,
+ "values": [
+ {
+ "code": 0,
+ "text": "Inap.: no second mention"
+ },
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 8,
+ "text": "More than two mentions"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ }
+ }
+ },
+ "1989": {
+ "codebook": {
+ "path": "family/1989/FAM1989_codebook.pdf",
+ "sha256": "1dff0473a8df0b7564f1716fd01ef67d276920105d17a986f6f167de5e1a049a"
+ },
+ "formats_sas": null,
+ "variables": {
+ "V17417": {
+ "label": "K18 SPANISH DESCENT WF",
+ "page": 358,
+ "values": [
+ {
+ "code": 0,
+ "text": "Inap.: is not Spanish/Hispanic; no Wife/\"Wife\" in FU (V17402=5)"
+ },
+ {
+ "code": 1,
+ "text": "Mexican"
+ },
+ {
+ "code": 2,
+ "text": "Mexican American"
+ },
+ {
+ "code": 3,
+ "text": "Chicano"
+ },
+ {
+ "code": 4,
+ "text": "Puerto Rican"
+ },
+ {
+ "code": 5,
+ "text": "Cuban"
+ },
+ {
+ "code": 6,
+ "text": "Combination; more than one mention"
+ },
+ {
+ "code": 7,
+ "text": "Other Spanish"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "V17418": {
+ "label": "K19 RACE OF WIFE 1",
+ "page": 359,
+ "values": [
+ {
+ "code": 0,
+ "text": "Inap.: no Wife/\"Wife\" in FU (V17402=5)"
+ },
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "V17419": {
+ "label": "K19 RACE OF WIFE 2",
+ "page": 359,
+ "values": [
+ {
+ "code": 0,
+ "text": "Inap.: no second mention; no Wife/\"Wife\" in FU (V17402=5)"
+ },
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 8,
+ "text": "More than two mentions"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "V17482": {
+ "label": "L31 SPANISH DESCENT HD",
+ "page": 389,
+ "values": [
+ {
+ "code": 0,
+ "text": "Inap.: is not Spanish/Hispanic"
+ },
+ {
+ "code": 1,
+ "text": "Mexican"
+ },
+ {
+ "code": 2,
+ "text": "Mexican American"
+ },
+ {
+ "code": 3,
+ "text": "Chicano"
+ },
+ {
+ "code": 4,
+ "text": "Puerto Rican"
+ },
+ {
+ "code": 5,
+ "text": "Cuban"
+ },
+ {
+ "code": 6,
+ "text": "Combination; more than 1 mention"
+ },
+ {
+ "code": 7,
+ "text": "Other Spanish"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "V17483": {
+ "label": "L32 RACE OF HEAD 1",
+ "page": 389,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "V17484": {
+ "label": "L32 RACE OF HEAD 2",
+ "page": 390,
+ "values": [
+ {
+ "code": 0,
+ "text": "Inap.: no second mention"
+ },
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 8,
+ "text": "More than two mentions"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ }
+ }
+ },
+ "1990": {
+ "codebook": {
+ "path": "family/1990/FAM1990_codebook.pdf",
+ "sha256": "4e6e0d3c22eedc50137b3d8a8b33e949064681b5bf42275d51da4b349734b32f"
+ },
+ "formats_sas": null,
+ "variables": {
+ "V18748": {
+ "label": "L18 SPANISH DESCENT WF",
+ "page": 326,
+ "values": [
+ {
+ "code": 0,
+ "text": "Inap.: is not Spanish/Hispanic; no Wife/\"Wife\" in FU (V18733=5)"
+ },
+ {
+ "code": 1,
+ "text": "Mexican"
+ },
+ {
+ "code": 2,
+ "text": "Mexican American"
+ },
+ {
+ "code": 3,
+ "text": "Chicano"
+ },
+ {
+ "code": 4,
+ "text": "Puerto Rican"
+ },
+ {
+ "code": 5,
+ "text": "Cuban"
+ },
+ {
+ "code": 6,
+ "text": "Combination; more than one mention"
+ },
+ {
+ "code": 7,
+ "text": "Other Spanish; Hispanic; Latino"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "V18749": {
+ "label": "L19 RACE OF WIFE 1",
+ "page": 327,
+ "values": [
+ {
+ "code": 0,
+ "text": "Inap.: no Wife/\"Wife\" in FU (V18733=5)"
+ },
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "V18750": {
+ "label": "L19 RACE OF WIFE 2",
+ "page": 327,
+ "values": [
+ {
+ "code": 0,
+ "text": "Inap.: no second mention; no Wife/\"Wife\" in FU (V18733=5)"
+ },
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 8,
+ "text": "More than two mentions"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "V18813": {
+ "label": "M31 SPANISH DESCENT HD",
+ "page": 355,
+ "values": [
+ {
+ "code": 0,
+ "text": "Inap.: is not Spanish/Hispanic"
+ },
+ {
+ "code": 1,
+ "text": "Mexican"
+ },
+ {
+ "code": 2,
+ "text": "Mexican American"
+ },
+ {
+ "code": 3,
+ "text": "Chicano"
+ },
+ {
+ "code": 4,
+ "text": "Puerto Rican"
+ },
+ {
+ "code": 5,
+ "text": "Cuban"
+ },
+ {
+ "code": 6,
+ "text": "Combination; more than one mention"
+ },
+ {
+ "code": 7,
+ "text": "Other Spanish; Hispanic; Latino"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "V18814": {
+ "label": "M32 RACE OF HEAD 1",
+ "page": 355,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "V18815": {
+ "label": "M32 RACE OF HEAD 2",
+ "page": 356,
+ "values": [
+ {
+ "code": 0,
+ "text": "Inap.: no second mention"
+ },
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 8,
+ "text": "More than two mentions"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ }
+ }
+ },
+ "1991": {
+ "codebook": {
+ "path": "family/1991/FAM1991_codebook.pdf",
+ "sha256": "cb25d86256acb143cbc7b6f400702dc6a20787af9476dfbe9fc277ccecca424c"
+ },
+ "formats_sas": null,
+ "variables": {
+ "V20048": {
+ "label": "K18 SPANISH DESCENT WF",
+ "page": 326,
+ "values": [
+ {
+ "code": 0,
+ "text": "Inap.: is not Spanish/Hispanic; no Wife/\"Wife\" in FU (V20033=5)"
+ },
+ {
+ "code": 1,
+ "text": "Mexican"
+ },
+ {
+ "code": 2,
+ "text": "Mexican American"
+ },
+ {
+ "code": 3,
+ "text": "Chicano"
+ },
+ {
+ "code": 4,
+ "text": "Puerto Rican"
+ },
+ {
+ "code": 5,
+ "text": "Cuban"
+ },
+ {
+ "code": 6,
+ "text": "Combination; more than one mention"
+ },
+ {
+ "code": 7,
+ "text": "Other Spanish; Hispanic; Latino"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "V20049": {
+ "label": "K19 RACE OF WIFE 1",
+ "page": 327,
+ "values": [
+ {
+ "code": 0,
+ "text": "Inap.: no Wife/\"Wife\" in FU (V20033=5)"
+ },
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "V20050": {
+ "label": "K19 RACE OF WIFE 2",
+ "page": 327,
+ "values": [
+ {
+ "code": 0,
+ "text": "Inap.: no second mention; no Wife/\"Wife\" in FU (V20033=5)"
+ },
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 8,
+ "text": "More than two mentions"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "V20113": {
+ "label": "L31 SPANISH DESCENT HD",
+ "page": 351,
+ "values": [
+ {
+ "code": 0,
+ "text": "Inap.: is not Spanish/Hispanic"
+ },
+ {
+ "code": 1,
+ "text": "Mexican"
+ },
+ {
+ "code": 2,
+ "text": "Mexican American"
+ },
+ {
+ "code": 3,
+ "text": "Chicano"
+ },
+ {
+ "code": 4,
+ "text": "Puerto Rican"
+ },
+ {
+ "code": 5,
+ "text": "Cuban"
+ },
+ {
+ "code": 6,
+ "text": "Combination; more than one mention"
+ },
+ {
+ "code": 7,
+ "text": "Other Spanish; Hispanic; Latino"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "V20114": {
+ "label": "L32 RACE OF HEAD 1",
+ "page": 351,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "V20115": {
+ "label": "L32 RACE OF HEAD 2",
+ "page": 352,
+ "values": [
+ {
+ "code": 0,
+ "text": "Inap.: no second mention"
+ },
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 8,
+ "text": "More than two mentions"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ }
+ }
+ },
+ "1992": {
+ "codebook": {
+ "path": "family/1992/FAM1992_codebook.pdf",
+ "sha256": "249b4c1a9a333792df66511c4fc63ae2ac51a28bb764f38aff4bc04a6d389d52"
+ },
+ "formats_sas": null,
+ "variables": {
+ "V21354": {
+ "label": "L18 SPANISH DESCENT WF",
+ "page": 329,
+ "values": [
+ {
+ "code": 0,
+ "text": "Inap.: is not Spanish/Hispanic; no Wife/\"Wife\" in FU (V21339=5)"
+ },
+ {
+ "code": 1,
+ "text": "Mexican"
+ },
+ {
+ "code": 2,
+ "text": "Mexican American"
+ },
+ {
+ "code": 3,
+ "text": "Chicano"
+ },
+ {
+ "code": 4,
+ "text": "Puerto Rican"
+ },
+ {
+ "code": 5,
+ "text": "Cuban"
+ },
+ {
+ "code": 6,
+ "text": "Combination; more than one mention"
+ },
+ {
+ "code": 7,
+ "text": "Other Spanish; Hispanic; Latino"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "V21355": {
+ "label": "L19 RACE OF WIFE 1",
+ "page": 329,
+ "values": [
+ {
+ "code": 0,
+ "text": "Inap.: no Wife/\"Wife\" in FU (V21339=5)"
+ },
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "V21356": {
+ "label": "L19 RACE OF WIFE 2",
+ "page": 329,
+ "values": [
+ {
+ "code": 0,
+ "text": "Inap.: no second mention; no Wife/\"Wife\" in FU (V21339=5)"
+ },
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 8,
+ "text": "More than two mentions"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "V21419": {
+ "label": "M31 SPANISH DESCENT HD",
+ "page": 353,
+ "values": [
+ {
+ "code": 0,
+ "text": "Inap.: is not Spanish/Hispanic"
+ },
+ {
+ "code": 1,
+ "text": "Mexican"
+ },
+ {
+ "code": 2,
+ "text": "Mexican American"
+ },
+ {
+ "code": 3,
+ "text": "Chicano"
+ },
+ {
+ "code": 4,
+ "text": "Puerto Rican"
+ },
+ {
+ "code": 5,
+ "text": "Cuban"
+ },
+ {
+ "code": 6,
+ "text": "Combination; more than one mention"
+ },
+ {
+ "code": 7,
+ "text": "Other Spanish; Hispanic; Latino"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "V21420": {
+ "label": "M32 RACE OF HEAD 1",
+ "page": 354,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "V21421": {
+ "label": "M32 RACE OF HEAD 2",
+ "page": 354,
+ "values": [
+ {
+ "code": 0,
+ "text": "Inap.: no second mention"
+ },
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 8,
+ "text": "More than two mentions"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ }
+ }
+ },
+ "1993": {
+ "codebook": {
+ "path": "family/1993/fam1993_codebook.pdf",
+ "sha256": "a4659bad3b0d2a4ec0040f9ad63994060d1d98c5924327853dfc7899ac172411"
+ },
+ "formats_sas": null,
+ "variables": {
+ "V23211": {
+ "label": "K18 WTR WF OF SPANISH DESCENT",
+ "page": 409,
+ "values": [
+ {
+ "code": 0,
+ "text": "Inap.: is not Spanish/Hispanic; no wife/\"wife\" in FU"
+ },
+ {
+ "code": 1,
+ "text": "Mexican"
+ },
+ {
+ "code": 2,
+ "text": "Mexican American"
+ },
+ {
+ "code": 3,
+ "text": "Chicano"
+ },
+ {
+ "code": 4,
+ "text": "Puerto Rican"
+ },
+ {
+ "code": 5,
+ "text": "Cuban"
+ },
+ {
+ "code": 6,
+ "text": "Combination; more than one mention"
+ },
+ {
+ "code": 7,
+ "text": "Other Spanish; Hispanic; Latino"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "V23212": {
+ "label": "K19 RACE OF WF-1ST MENTION",
+ "page": 409,
+ "values": [
+ {
+ "code": 0,
+ "text": "Inap.: no wife/\"wife\" in FU"
+ },
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "V23213": {
+ "label": "K19 RACE OF WF-2ND MENTION",
+ "page": 410,
+ "values": [
+ {
+ "code": 0,
+ "text": "Inap.: no second mention; no wife/\"wife\" in FU"
+ },
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 8,
+ "text": "More than two mentions"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "V23275": {
+ "label": "L31 WTR HD OF SPANISH DESCENT",
+ "page": 432,
+ "values": [
+ {
+ "code": 0,
+ "text": "Inap.: is not Spanish/Hispanic"
+ },
+ {
+ "code": 1,
+ "text": "Mexican"
+ },
+ {
+ "code": 2,
+ "text": "Mexican American"
+ },
+ {
+ "code": 3,
+ "text": "Chicano"
+ },
+ {
+ "code": 4,
+ "text": "Puerto Rican"
+ },
+ {
+ "code": 5,
+ "text": "Cuban"
+ },
+ {
+ "code": 6,
+ "text": "Combination; more than one mention"
+ },
+ {
+ "code": 7,
+ "text": "Other Spanish; Hispanic; Latino"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "V23276": {
+ "label": "L32 RACE OF HD-1ST MENTION",
+ "page": 432,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "V23277": {
+ "label": "L32 RACE OF HD-2ND MENTION",
+ "page": 433,
+ "values": [
+ {
+ "code": 0,
+ "text": "Inap.: no second mention"
+ },
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 8,
+ "text": "More than two mentions"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "V23333": {
+ "label": "COMPLETED ED-HD 1993",
+ "page": 459,
+ "values": [
+ {
+ "code": 0,
+ "text": "Inap: completed no grades of school"
+ },
+ {
+ "range": [
+ 1,
+ 16
+ ],
+ "text": "Actual number"
+ },
+ {
+ "code": 17,
+ "text": "At least some post-graduate work"
+ },
+ {
+ "code": 99,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "V23334": {
+ "label": "COMPLETED ED-WF 1993",
+ "page": 460,
+ "values": [
+ {
+ "code": 0,
+ "text": "Inap: completed no grades of school; no wife/\"wife\" in FU"
+ },
+ {
+ "range": [
+ 1,
+ 16
+ ],
+ "text": "Actual number"
+ },
+ {
+ "code": 17,
+ "text": "At least some post-graduate work"
+ },
+ {
+ "code": 99,
+ "text": "NA; DK"
+ }
+ ]
+ }
+ }
+ },
+ "1994": {
+ "codebook": {
+ "path": "family/1994/FAM1994ER_codebook_public.pdf",
+ "sha256": "ecf8667f3b6d2a5ac7e312ecd9761a00ae7c1d2db6bee2f8af049e495ff1cf56"
+ },
+ "formats_sas": null,
+ "variables": {
+ "ER3880": {
+ "label": "K18 SPANISH DESCENT 1 WF",
+ "page": 545,
+ "values": [
+ {
+ "code": 1,
+ "text": "Mexican"
+ },
+ {
+ "code": 2,
+ "text": "Mexican American"
+ },
+ {
+ "code": 3,
+ "text": "Chicano"
+ },
+ {
+ "code": 4,
+ "text": "Puerto Rican"
+ },
+ {
+ "code": 5,
+ "text": "Cuban"
+ },
+ {
+ "code": 6,
+ "text": "Combination; more than one mention"
+ },
+ {
+ "code": 7,
+ "text": "Other Spanish; Hispanic; Latino"
+ },
+ {
+ "code": 8,
+ "text": "DK"
+ },
+ {
+ "code": 9,
+ "text": "NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no wife/\"wife\" in FU"
+ }
+ ]
+ },
+ "ER3883": {
+ "label": "K19 RACE OF WIFE 1",
+ "page": 546,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 8,
+ "text": "DK"
+ },
+ {
+ "code": 9,
+ "text": "NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no wife/\"wife\" in FU"
+ }
+ ]
+ },
+ "ER3884": {
+ "label": "K19 RACE OF WIFE 2",
+ "page": 546,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 8,
+ "text": "DK"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no second mention; no wife/\"wife\" in FU; DK/NA to first mention"
+ }
+ ]
+ },
+ "ER3885": {
+ "label": "K19 RACE OF WIFE 3",
+ "page": 547,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no third mention; no wife/\"wife\" in FU; DK/NA to first mention"
+ }
+ ]
+ },
+ "ER3941": {
+ "label": "L31 SPANISH DESCENT 1 HD",
+ "page": 566,
+ "values": [
+ {
+ "code": 1,
+ "text": "Mexican"
+ },
+ {
+ "code": 2,
+ "text": "Mexican American"
+ },
+ {
+ "code": 3,
+ "text": "Chicano"
+ },
+ {
+ "code": 4,
+ "text": "Puerto Rican"
+ },
+ {
+ "code": 5,
+ "text": "Cuban"
+ },
+ {
+ "code": 6,
+ "text": "Combination; more than one mention"
+ },
+ {
+ "code": 7,
+ "text": "Other Spanish; Hispanic; Latino"
+ },
+ {
+ "code": 8,
+ "text": "DK"
+ },
+ {
+ "code": 9,
+ "text": "NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: not Spanish, Hispanic or Latino"
+ }
+ ]
+ },
+ "ER3944": {
+ "label": "L32 RACE OF HEAD 1",
+ "page": 567,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 8,
+ "text": "DK"
+ },
+ {
+ "code": 9,
+ "text": "NA; refused"
+ }
+ ]
+ },
+ "ER3945": {
+ "label": "L32 RACE OF HEAD 2",
+ "page": 568,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 8,
+ "text": "DK"
+ },
+ {
+ "code": 9,
+ "text": "NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.; no second mention; DK/NA to first mention"
+ }
+ ]
+ },
+ "ER3946": {
+ "label": "L32 RACE OF HEAD 3",
+ "page": 568,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 0,
+ "text": "Inap.; no third mention; DK/NA to first mention"
+ }
+ ]
+ },
+ "ER4158": {
+ "label": "COMPLETED ED-HD",
+ "page": 669,
+ "values": [
+ {
+ "code": 0,
+ "text": "Completed no grades of school"
+ },
+ {
+ "range": [
+ 1,
+ 16
+ ],
+ "text": "Actual number"
+ },
+ {
+ "code": 17,
+ "text": "At least some post-graduate work"
+ },
+ {
+ "code": 99,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "ER4159": {
+ "label": "COMPLETED ED-WF",
+ "page": 669,
+ "values": [
+ {
+ "range": [
+ 1,
+ 16
+ ],
+ "text": "Actual number"
+ },
+ {
+ "code": 17,
+ "text": "At least some post-graduate work"
+ },
+ {
+ "code": 99,
+ "text": "NA; DK"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no wife/\"wife\" in FU; completed no grades of school"
+ }
+ ]
+ }
+ }
+ },
+ "1995": {
+ "codebook": {
+ "path": "family/1995/FAM1995ER_codebook_public.pdf",
+ "sha256": "c4a161e935244487d26d7a581ff68b1c7f844c8032fe756eda82c035f51cb985"
+ },
+ "formats_sas": null,
+ "variables": {
+ "ER6750": {
+ "label": "K18 SPANISH DESCENT 1 WF",
+ "page": 496,
+ "values": [
+ {
+ "code": 1,
+ "text": "Mexican"
+ },
+ {
+ "code": 2,
+ "text": "Mexican American"
+ },
+ {
+ "code": 3,
+ "text": "Chicano"
+ },
+ {
+ "code": 4,
+ "text": "Puerto Rican"
+ },
+ {
+ "code": 5,
+ "text": "Cuban"
+ },
+ {
+ "code": 6,
+ "text": "Combination; more than one mention"
+ },
+ {
+ "code": 7,
+ "text": "Other Spanish; Hispanic; Latino"
+ },
+ {
+ "code": 8,
+ "text": "DK"
+ },
+ {
+ "code": 9,
+ "text": "NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no wife/\"wife\" in FU"
+ }
+ ]
+ },
+ "ER6753": {
+ "label": "K19 RACE OF WIFE 1",
+ "page": 497,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 8,
+ "text": "DK"
+ },
+ {
+ "code": 9,
+ "text": "NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no wife/\"wife\" in FU"
+ }
+ ]
+ },
+ "ER6754": {
+ "label": "K19 RACE OF WIFE 2",
+ "page": 498,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 8,
+ "text": "DK"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no second mention; no wife/\"wife\" in FU; NA, DK to first mention"
+ }
+ ]
+ },
+ "ER6755": {
+ "label": "K19 RACE OF WIFE 3",
+ "page": 498,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no third mention; no wife/\"wife\" in FU; NA, DK to first mention"
+ }
+ ]
+ },
+ "ER6811": {
+ "label": "L31 SPANISH DESCENT 1 HD",
+ "page": 517,
+ "values": [
+ {
+ "code": 1,
+ "text": "Mexican"
+ },
+ {
+ "code": 2,
+ "text": "Mexican American"
+ },
+ {
+ "code": 3,
+ "text": "Chicano"
+ },
+ {
+ "code": 4,
+ "text": "Puerto Rican"
+ },
+ {
+ "code": 5,
+ "text": "Cuban"
+ },
+ {
+ "code": 6,
+ "text": "Combination; more than one mention"
+ },
+ {
+ "code": 7,
+ "text": "Other Spanish; Hispanic; Latino"
+ },
+ {
+ "code": 8,
+ "text": "DK"
+ },
+ {
+ "code": 9,
+ "text": "NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: not Spanish, Hispanic or Latino"
+ }
+ ]
+ },
+ "ER6814": {
+ "label": "L32 RACE OF HEAD 1",
+ "page": 518,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 8,
+ "text": "DK"
+ },
+ {
+ "code": 9,
+ "text": "NA; refused"
+ }
+ ]
+ },
+ "ER6815": {
+ "label": "L32 RACE OF HEAD 2",
+ "page": 519,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 8,
+ "text": "DK"
+ },
+ {
+ "code": 9,
+ "text": "NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.; no second mention; NA, DK to first mention"
+ }
+ ]
+ },
+ "ER6816": {
+ "label": "L32 RACE OF HEAD 3",
+ "page": 519,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ },
+ {
+ "code": 0,
+ "text": "Inap.; no third mention; NA, DK to first mention"
+ }
+ ]
+ },
+ "ER6998": {
+ "label": "COMPLETED ED-HD",
+ "page": 612,
+ "values": [
+ {
+ "code": 0,
+ "text": "Completed no grades of school"
+ },
+ {
+ "range": [
+ 1,
+ 16
+ ],
+ "text": "Actual number"
+ },
+ {
+ "code": 17,
+ "text": "At least some post-graduate work"
+ },
+ {
+ "code": 99,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "ER6999": {
+ "label": "COMPLETED ED-WF",
+ "page": 612,
+ "values": [
+ {
+ "range": [
+ 1,
+ 16
+ ],
+ "text": "Actual number"
+ },
+ {
+ "code": 17,
+ "text": "At least some post-graduate work"
+ },
+ {
+ "code": 99,
+ "text": "NA; DK"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no wife/\"wife\" in FU; completed no grades of school"
+ }
+ ]
+ }
+ }
+ },
+ "1996": {
+ "codebook": {
+ "path": "family/1996/FAM1996ER_codebook_public.pdf",
+ "sha256": "aa52222228e4ecc3d7309c39c61929a7603951615c6a80b3ba3423e7f3ba8300"
+ },
+ "formats_sas": null,
+ "variables": {
+ "ER8996": {
+ "label": "K18 SPANISH DESCENT 1 WF",
+ "page": 604,
+ "values": [
+ {
+ "code": 1,
+ "text": "Mexican"
+ },
+ {
+ "code": 2,
+ "text": "Mexican American"
+ },
+ {
+ "code": 3,
+ "text": "Chicano"
+ },
+ {
+ "code": 4,
+ "text": "Puerto Rican"
+ },
+ {
+ "code": 5,
+ "text": "Cuban"
+ },
+ {
+ "code": 6,
+ "text": "Combination; more than one mention"
+ },
+ {
+ "code": 7,
+ "text": "Other Spanish; Hispanic; Latino"
+ },
+ {
+ "code": 8,
+ "text": "DK"
+ },
+ {
+ "code": 9,
+ "text": "NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no wife/\"wife\" in FU"
+ }
+ ]
+ },
+ "ER8999": {
+ "label": "K19 RACE OF WIFE 1",
+ "page": 605,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 8,
+ "text": "DK"
+ },
+ {
+ "code": 9,
+ "text": "NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no wife/\"wife\" in FU"
+ }
+ ]
+ },
+ "ER9000": {
+ "label": "K19 RACE OF WIFE 2",
+ "page": 605,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 8,
+ "text": "DK"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no wife/\"wife\" in FU; no second mention; NA, DK to first mention"
+ }
+ ]
+ },
+ "ER9001": {
+ "label": "K19 RACE OF WIFE 3",
+ "page": 606,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no wife/\"wife\" in FU; no third mention; NA, DK to first mention"
+ }
+ ]
+ },
+ "ER9057": {
+ "label": "L31 SPANISH DESCENT 1 HD",
+ "page": 625,
+ "values": [
+ {
+ "code": 1,
+ "text": "Mexican"
+ },
+ {
+ "code": 2,
+ "text": "Mexican American"
+ },
+ {
+ "code": 3,
+ "text": "Chicano"
+ },
+ {
+ "code": 4,
+ "text": "Puerto Rican"
+ },
+ {
+ "code": 5,
+ "text": "Cuban"
+ },
+ {
+ "code": 6,
+ "text": "Combinations; more than one mention"
+ },
+ {
+ "code": 7,
+ "text": "Other Spanish; Hispanic; Latino"
+ },
+ {
+ "code": 8,
+ "text": "DK"
+ },
+ {
+ "code": 9,
+ "text": "NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: not Spanish, Hispanic or Latino"
+ }
+ ]
+ },
+ "ER9060": {
+ "label": "L32 RACE OF HEAD 1",
+ "page": 626,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 8,
+ "text": "DK"
+ },
+ {
+ "code": 9,
+ "text": "NA; refused"
+ }
+ ]
+ },
+ "ER9061": {
+ "label": "L32 RACE OF HEAD 2",
+ "page": 627,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 8,
+ "text": "DK"
+ },
+ {
+ "code": 9,
+ "text": "NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.; no second mention; NA, DK to first mention"
+ }
+ ]
+ },
+ "ER9062": {
+ "label": "L32 RACE OF HEAD 3",
+ "page": 627,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ },
+ {
+ "code": 0,
+ "text": "Inap.; no third mention; NA, DK to first mention"
+ }
+ ]
+ },
+ "ER9249": {
+ "label": "COMPLETED ED-HD",
+ "page": 719,
+ "values": [
+ {
+ "code": 0,
+ "text": "Completed no grades of school"
+ },
+ {
+ "range": [
+ 1,
+ 16
+ ],
+ "text": "Actual number"
+ },
+ {
+ "code": 17,
+ "text": "At least some post-graduate work"
+ },
+ {
+ "code": 99,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "ER9250": {
+ "label": "COMPLETED ED-WF",
+ "page": 720,
+ "values": [
+ {
+ "range": [
+ 1,
+ 16
+ ],
+ "text": "Actual number"
+ },
+ {
+ "code": 17,
+ "text": "At least some post-graduate work"
+ },
+ {
+ "code": 99,
+ "text": "NA; DK"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no wife/\"wife\" in FU; completed no grades of school"
+ }
+ ]
+ }
+ }
+ },
+ "1997": {
+ "codebook": {
+ "path": "family/1997/FAM1997ER_codebook_public.pdf",
+ "sha256": "9716779d9b1102fb5be611e620525759a6cb8676714ba9722a34da8c7287bb3d"
+ },
+ "formats_sas": null,
+ "variables": {
+ "ER11760": {
+ "label": "K34/87 RACE OF WIFE 1",
+ "page": 470,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 8,
+ "text": "DK"
+ },
+ {
+ "code": 9,
+ "text": "NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no wife/\"wife\" in FU"
+ }
+ ]
+ },
+ "ER11761": {
+ "label": "K34/87 RACE OF WIFE 2",
+ "page": 471,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no wife/\"wife\" in FU; no second mention; NA, DK to first mention"
+ }
+ ]
+ },
+ "ER11762": {
+ "label": "K34/87 RACE OF WIFE 3",
+ "page": 471,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no wife/\"wife\" in FU; no third mention; NA, DK to first mention"
+ }
+ ]
+ },
+ "ER11763": {
+ "label": "K34/87 RACE OF WIFE 4",
+ "page": 471,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no wife/\"wife\" in FU; no fourth mention; NA, DK to first mention"
+ }
+ ]
+ },
+ "ER11848": {
+ "label": "L40/95 RACE OF HEAD 1",
+ "page": 504,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 8,
+ "text": "DK"
+ },
+ {
+ "code": 9,
+ "text": "NA; refused"
+ }
+ ]
+ },
+ "ER11849": {
+ "label": "L40/95 RACE OF HEAD 2",
+ "page": 504,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 0,
+ "text": "Inap.; no second mention; NA, DK to first mention"
+ }
+ ]
+ },
+ "ER11850": {
+ "label": "L40/95 RACE OF HEAD 3",
+ "page": 504,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 0,
+ "text": "Inap.; no third mention; NA, DK to first mention"
+ }
+ ]
+ },
+ "ER11851": {
+ "label": "L40/95 RACE OF HEAD 4",
+ "page": 505,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "Native American"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Color besides black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no fourth mention; NA, DK to first mention"
+ }
+ ]
+ },
+ "ER12222": {
+ "label": "COMPLETED ED-HD",
+ "page": 660,
+ "values": [
+ {
+ "code": 0,
+ "text": "Completed no grades of school"
+ },
+ {
+ "range": [
+ 1,
+ 16
+ ],
+ "text": "Actual number"
+ },
+ {
+ "code": 17,
+ "text": "At least some post-graduate work"
+ },
+ {
+ "code": 99,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "ER12223": {
+ "label": "COMPLETED ED-WF",
+ "page": 660,
+ "values": [
+ {
+ "range": [
+ 1,
+ 16
+ ],
+ "text": "Actual number"
+ },
+ {
+ "code": 17,
+ "text": "At least some post-graduate work"
+ },
+ {
+ "code": 99,
+ "text": "NA; DK"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no wife/\"wife\" in FU; completed no grades of school"
+ }
+ ]
+ }
+ }
+ },
+ "1999": {
+ "codebook": {
+ "path": "family/1999/FAM1999ER_codebook.pdf",
+ "sha256": "5accefc4c1b50b3447b1a674ac2750de94568c538873f9bab34c2049c80c6107"
+ },
+ "formats_sas": null,
+ "variables": {
+ "ER15836": {
+ "label": "K34/87 RACE OF WIFE 1",
+ "page": 818,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 8,
+ "text": "DK"
+ },
+ {
+ "code": 9,
+ "text": "NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no wife/\"wife\" in FU"
+ }
+ ]
+ },
+ "ER15837": {
+ "label": "K34/87 RACE OF WIFE 2",
+ "page": 819,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no wife/\"wife\" in FU; no second mention; NA, DK to first mention"
+ }
+ ]
+ },
+ "ER15838": {
+ "label": "K34/87 RACE OF WIFE 3",
+ "page": 819,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no wife/\"wife\" in FU; no third mention; NA, DK to first mention"
+ }
+ ]
+ },
+ "ER15839": {
+ "label": "K34/87 RACE OF WIFE 4",
+ "page": 819,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no wife/\"wife\" in FU; no fourth mention; NA, DK to first mention"
+ }
+ ]
+ },
+ "ER15928": {
+ "label": "L40/95 RACE OF HEAD 1",
+ "page": 858,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 8,
+ "text": "DK"
+ },
+ {
+ "code": 9,
+ "text": "NA; refused"
+ }
+ ]
+ },
+ "ER15929": {
+ "label": "L40/95 RACE OF HEAD 2",
+ "page": 858,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no second mention; NA, DK to first mention"
+ }
+ ]
+ },
+ "ER15930": {
+ "label": "L40/95 RACE OF HEAD 3",
+ "page": 858,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no third mention; NA, DK to first mention"
+ }
+ ]
+ },
+ "ER15931": {
+ "label": "L40/95 RACE OF HEAD 4",
+ "page": 859,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "Native American"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Color besides black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no fourth mention; NA, DK to first mention"
+ }
+ ]
+ },
+ "ER16516": {
+ "label": "COMPLETED ED-HD",
+ "page": 1079,
+ "values": [
+ {
+ "code": 0,
+ "text": "Completed no grades of school"
+ },
+ {
+ "range": [
+ 1,
+ 16
+ ],
+ "text": "Actual number"
+ },
+ {
+ "code": 17,
+ "text": "At least some post-graduate work"
+ },
+ {
+ "code": 99,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "ER16517": {
+ "label": "COMPLETED ED-WF",
+ "page": 1079,
+ "values": [
+ {
+ "range": [
+ 1,
+ 16
+ ],
+ "text": "Actual number"
+ },
+ {
+ "code": 17,
+ "text": "At least some post-graduate work"
+ },
+ {
+ "code": 99,
+ "text": "NA; DK"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no wife/\"wife\" in FU; completed no grades of school"
+ }
+ ]
+ }
+ }
+ },
+ "2001": {
+ "codebook": {
+ "path": "family/2001/FAM2001ER_codebook.pdf",
+ "sha256": "13ae45ce24125c2ac30c15e62f1efeddc7842dc924b3415fe5aaa058666029a0"
+ },
+ "formats_sas": null,
+ "variables": {
+ "ER19897": {
+ "label": "K34/87 RACE OF WIFE 1",
+ "page": 821,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 8,
+ "text": "DK"
+ },
+ {
+ "code": 9,
+ "text": "NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no wife/\"wife\" in FU"
+ }
+ ]
+ },
+ "ER19898": {
+ "label": "K34/87 RACE OF WIFE 2",
+ "page": 822,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no wife/\"wife\" in FU; no second mention; NA, DK to first mention"
+ }
+ ]
+ },
+ "ER19899": {
+ "label": "K34/87 RACE OF WIFE 3",
+ "page": 822,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no wife/\"wife\" in FU; no third mention; NA, DK to first mention"
+ }
+ ]
+ },
+ "ER19900": {
+ "label": "K34/87 RACE OF WIFE 4",
+ "page": 822,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no wife/\"wife\" in FU; no fourth mention; NA, DK to first mention"
+ }
+ ]
+ },
+ "ER19989": {
+ "label": "L40/95 RACE OF HEAD 1",
+ "page": 860,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 8,
+ "text": "DK"
+ },
+ {
+ "code": 9,
+ "text": "NA; refused"
+ }
+ ]
+ },
+ "ER19990": {
+ "label": "L40/95 RACE OF HEAD 2",
+ "page": 861,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no second mention; NA, DK to first mention"
+ }
+ ]
+ },
+ "ER19991": {
+ "label": "L40/95 RACE OF HEAD 3",
+ "page": 861,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "American Indian, Aleut, Eskimo"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Mentions Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Mentions color other than black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no third mention; NA, DK to first mention"
+ }
+ ]
+ },
+ "ER19992": {
+ "label": "L40/95 RACE OF HEAD 4",
+ "page": 862,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "Native American"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Color besides black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no fourth mention; NA, DK to first mention"
+ }
+ ]
+ },
+ "ER20457": {
+ "label": "COMPLETED ED-HD",
+ "page": 1037,
+ "values": [
+ {
+ "code": 0,
+ "text": "Completed no grades of school"
+ },
+ {
+ "range": [
+ 1,
+ 16
+ ],
+ "text": "Actual number"
+ },
+ {
+ "code": 17,
+ "text": "At least some post-graduate work"
+ },
+ {
+ "code": 99,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "ER20458": {
+ "label": "COMPLETED ED-WF",
+ "page": 1038,
+ "values": [
+ {
+ "range": [
+ 1,
+ 16
+ ],
+ "text": "Actual number"
+ },
+ {
+ "code": 17,
+ "text": "At least some post-graduate work"
+ },
+ {
+ "code": 99,
+ "text": "NA; DK"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no wife/\"wife\" in FU; completed no grades of school"
+ }
+ ]
+ }
+ }
+ },
+ "2003": {
+ "codebook": {
+ "path": "family/2003/FAM2003ER_codebook.pdf",
+ "sha256": "5a69ee0e605aec1305a23931330da97d75dd21c32d183b8d67c3b6597f23a201"
+ },
+ "formats_sas": null,
+ "variables": {
+ "ER23334": {
+ "label": "K34/87 RACE OF WIFE 1",
+ "page": 673,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "Native American"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Color besides black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no wife/\"wife\" in FU"
+ }
+ ]
+ },
+ "ER23335": {
+ "label": "K34/87 RACE OF WIFE 2",
+ "page": 674,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "Native American"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Color besides black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no wife/\"wife\" in FU; no second mention; NA, DK to first mention"
+ }
+ ]
+ },
+ "ER23336": {
+ "label": "K34/87 RACE OF WIFE 3",
+ "page": 674,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "Native American"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Color besides black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no wife/\"wife\" in FU; fewer than three mentions; NA, DK to first mention"
+ }
+ ]
+ },
+ "ER23337": {
+ "label": "K34/87 RACE OF WIFE 4",
+ "page": 674,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "Native American"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Color besides black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no wife/\"wife\" in FU; fewer than four mentions; NA, DK to first mention"
+ }
+ ]
+ },
+ "ER23426": {
+ "label": "L40/95 RACE OF HEAD 1",
+ "page": 714,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "Native American"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Color besides black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "ER23427": {
+ "label": "L40/95 RACE OF HEAD 2",
+ "page": 714,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "Native American"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Color besides black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no second mention; NA, DK to first mention"
+ }
+ ]
+ },
+ "ER23428": {
+ "label": "L40/95 RACE OF HEAD 3",
+ "page": 714,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "Native American"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Color besides black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: fewer than three mentions; NA, DK to first mention"
+ }
+ ]
+ },
+ "ER23429": {
+ "label": "L40/95 RACE OF HEAD 4",
+ "page": 715,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black"
+ },
+ {
+ "code": 3,
+ "text": "Native American"
+ },
+ {
+ "code": 4,
+ "text": "Asian, Pacific Islander"
+ },
+ {
+ "code": 5,
+ "text": "Latino origin or descent"
+ },
+ {
+ "code": 6,
+ "text": "Color besides black or white"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: fewer than four mentions; NA, DK to first mention"
+ }
+ ]
+ },
+ "ER24148": {
+ "label": "COMPLETED ED-HD",
+ "page": 1004,
+ "values": [
+ {
+ "code": 0,
+ "text": "Completed no grades of school"
+ },
+ {
+ "range": [
+ 1,
+ 16
+ ],
+ "text": "Actual number"
+ },
+ {
+ "code": 17,
+ "text": "At least some post-graduate work"
+ },
+ {
+ "code": 99,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "ER24149": {
+ "label": "COMPLETED ED-WF",
+ "page": 1005,
+ "values": [
+ {
+ "range": [
+ 1,
+ 16
+ ],
+ "text": "Actual number"
+ },
+ {
+ "code": 17,
+ "text": "At least some post-graduate work"
+ },
+ {
+ "code": 99,
+ "text": "NA; DK"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no wife/\"wife\" in FU; completed no grades of school"
+ }
+ ]
+ }
+ }
+ },
+ "2005": {
+ "codebook": {
+ "path": "family/2005/FAM2005ER_codebook.pdf",
+ "sha256": "eec8de3c04ac88ffce0495622f0cbe0ac621780b201ce6bd0b4b54344a3a6243"
+ },
+ "formats_sas": null,
+ "variables": {
+ "ER27296": {
+ "label": "K33A SPANISH DESCENT-WIFE",
+ "page": 665,
+ "values": [
+ {
+ "code": 1,
+ "text": "Mexican"
+ },
+ {
+ "code": 2,
+ "text": "Mexican American"
+ },
+ {
+ "code": 3,
+ "text": "Chicano"
+ },
+ {
+ "code": 4,
+ "text": "Puerto Rican"
+ },
+ {
+ "code": 5,
+ "text": "Cuban"
+ },
+ {
+ "code": 7,
+ "text": "Other Spanish; Hispanic; Latino"
+ },
+ {
+ "code": 8,
+ "text": "DK"
+ },
+ {
+ "code": 9,
+ "text": "NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no wife/\"wife\" in FU; not Spanish, Hispanic or Latino"
+ }
+ ]
+ },
+ "ER27297": {
+ "label": "K34 RACE OF WIFE-MENTION 1",
+ "page": 666,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no wife/\"wife\" in FU"
+ }
+ ]
+ },
+ "ER27298": {
+ "label": "K34 RACE OF WIFE-MENTION 2",
+ "page": 666,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no wife/\"wife\" in FU; no second mention; NA, DK to first mention"
+ }
+ ]
+ },
+ "ER27299": {
+ "label": "K34 RACE OF WIFE-MENTION 3",
+ "page": 666,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no wife/\"wife\" in FU; fewer than three mentions; NA, DK to first mention"
+ }
+ ]
+ },
+ "ER27300": {
+ "label": "K34 RACE OF WIFE-MENTION 4",
+ "page": 667,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no wife/\"wife\" in FU; fewer than four mentions; NA, DK to first mention"
+ }
+ ]
+ },
+ "ER27392": {
+ "label": "L39A SPANISH DESCENT-HEAD",
+ "page": 708,
+ "values": [
+ {
+ "code": 1,
+ "text": "Mexican"
+ },
+ {
+ "code": 2,
+ "text": "Mexican American"
+ },
+ {
+ "code": 3,
+ "text": "Chicano"
+ },
+ {
+ "code": 4,
+ "text": "Puerto Rican"
+ },
+ {
+ "code": 5,
+ "text": "Cuban"
+ },
+ {
+ "code": 7,
+ "text": "Other Spanish; Hispanic; Latino"
+ },
+ {
+ "code": 8,
+ "text": "DK"
+ },
+ {
+ "code": 9,
+ "text": "NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: not Spanish, Hispanic or Latino"
+ }
+ ]
+ },
+ "ER27393": {
+ "label": "L40 RACE OF HEAD-MENTION 1",
+ "page": 708,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 0,
+ "text": "Wild code"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "ER27394": {
+ "label": "L40 RACE OF HEAD-MENTION 2",
+ "page": 708,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no second mention; NA, DK to first mention"
+ }
+ ]
+ },
+ "ER27395": {
+ "label": "L40 RACE OF HEAD-MENTION 3",
+ "page": 709,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: fewer than three mentions; NA, DK to first mention"
+ }
+ ]
+ },
+ "ER27396": {
+ "label": "L40 RACE OF HEAD-MENTION 4",
+ "page": 709,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: fewer than four mentions; NA, DK to first mention"
+ }
+ ]
+ },
+ "ER28047": {
+ "label": "COMPLETED ED-HD",
+ "page": 961,
+ "values": [
+ {
+ "code": 0,
+ "text": "Completed no grades of school"
+ },
+ {
+ "range": [
+ 1,
+ 16
+ ],
+ "text": "Actual number"
+ },
+ {
+ "code": 17,
+ "text": "At least some post-graduate work"
+ },
+ {
+ "code": 99,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "ER28048": {
+ "label": "COMPLETED ED-WF",
+ "page": 962,
+ "values": [
+ {
+ "range": [
+ 1,
+ 16
+ ],
+ "text": "Actual number"
+ },
+ {
+ "code": 17,
+ "text": "At least some post-graduate work"
+ },
+ {
+ "code": 99,
+ "text": "NA; DK"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no wife/\"wife\" in FU; completed no grades of school"
+ }
+ ]
+ }
+ }
+ },
+ "2007": {
+ "codebook": {
+ "path": "family/2007/FAM2007ER_codebook.pdf",
+ "sha256": "d92bc05151cf1b77513e4520578402305d356a669c7998f10f91ad48ac66c89e"
+ },
+ "formats_sas": null,
+ "variables": {
+ "ER40471": {
+ "label": "K39 SPANISH DESCENT-WIFE",
+ "page": 1357,
+ "values": [
+ {
+ "code": 1,
+ "text": "Mexican"
+ },
+ {
+ "code": 2,
+ "text": "Mexican American"
+ },
+ {
+ "code": 3,
+ "text": "Chicano"
+ },
+ {
+ "code": 4,
+ "text": "Puerto Rican"
+ },
+ {
+ "code": 5,
+ "text": "Cuban"
+ },
+ {
+ "code": 7,
+ "text": "Other Spanish; Hispanic; Latino"
+ },
+ {
+ "code": 8,
+ "text": "DK"
+ },
+ {
+ "code": 9,
+ "text": "NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no wife/\"wife\" in FU; not Spanish, Hispanic or Latino"
+ }
+ ]
+ },
+ "ER40472": {
+ "label": "K40 RACE OF WIFE-MENTION 1",
+ "page": 1358,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no wife/\"wife\" in FU;"
+ }
+ ]
+ },
+ "ER40473": {
+ "label": "K40 RACE OF WIFE-MENTION 2",
+ "page": 1358,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no wife/\"wife\" in FU; no second mention; NA, DK to first mention"
+ }
+ ]
+ },
+ "ER40474": {
+ "label": "K40 RACE OF WIFE-MENTION 3",
+ "page": 1358,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no wife/\"wife\" in FU; fewer than three mentions; NA, DK to first mention"
+ }
+ ]
+ },
+ "ER40475": {
+ "label": "K40 RACE OF WIFE-MENTION 4",
+ "page": 1359,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no wife/\"wife\" in FU; fewer than four mentions; NA, DK to first mention"
+ }
+ ]
+ },
+ "ER40564": {
+ "label": "L39 SPANISH DESCENT-HEAD",
+ "page": 1399,
+ "values": [
+ {
+ "code": 1,
+ "text": "Mexican"
+ },
+ {
+ "code": 2,
+ "text": "Mexican American"
+ },
+ {
+ "code": 3,
+ "text": "Chicano"
+ },
+ {
+ "code": 4,
+ "text": "Puerto Rican"
+ },
+ {
+ "code": 5,
+ "text": "Cuban"
+ },
+ {
+ "code": 7,
+ "text": "Other Spanish; Hispanic; Latino"
+ },
+ {
+ "code": 8,
+ "text": "DK"
+ },
+ {
+ "code": 9,
+ "text": "NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: not Spanish, Hispanic or Latino"
+ }
+ ]
+ },
+ "ER40565": {
+ "label": "L40 RACE OF HEAD-MENTION 1",
+ "page": 1399,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 0,
+ "text": "Wild code"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "ER40566": {
+ "label": "L40 RACE OF HEAD-MENTION 2",
+ "page": 1399,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no second mention; NA, DK to first mention"
+ }
+ ]
+ },
+ "ER40567": {
+ "label": "L40 RACE OF HEAD-MENTION 3",
+ "page": 1400,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: fewer than three mentions; NA, DK to first mention"
+ }
+ ]
+ },
+ "ER40568": {
+ "label": "L40 RACE OF HEAD-MENTION 4",
+ "page": 1400,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "NA; DK"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: fewer than four mentions; NA, DK to first mention"
+ }
+ ]
+ },
+ "ER41037": {
+ "label": "COMPLETED ED-HD",
+ "page": 1590,
+ "values": [
+ {
+ "code": 0,
+ "text": "Completed no grades of school"
+ },
+ {
+ "range": [
+ 1,
+ 16
+ ],
+ "text": "Actual number"
+ },
+ {
+ "code": 17,
+ "text": "At least some post-graduate work"
+ },
+ {
+ "code": 99,
+ "text": "NA; DK"
+ }
+ ]
+ },
+ "ER41038": {
+ "label": "COMPLETED ED-WF",
+ "page": 1591,
+ "values": [
+ {
+ "range": [
+ 1,
+ 16
+ ],
+ "text": "Actual number"
+ },
+ {
+ "code": 17,
+ "text": "At least some post-graduate work"
+ },
+ {
+ "code": 99,
+ "text": "NA; DK"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no wife/\"wife\" in FU; completed no grades of school"
+ }
+ ]
+ }
+ }
+ },
+ "2009": {
+ "codebook": {
+ "path": "family/2009/FAM2009ER_codebook.pdf",
+ "sha256": "2639c0cad9d5b6e1a08eae57c578d9698a249ed331475b79df525a0caa7c4c6b"
+ },
+ "formats_sas": null,
+ "variables": {
+ "ER46448": {
+ "label": "K39 SPANISH DESCENT-WIFE",
+ "page": 1458,
+ "values": [
+ {
+ "code": 1,
+ "text": "Mexican"
+ },
+ {
+ "code": 2,
+ "text": "Mexican American"
+ },
+ {
+ "code": 3,
+ "text": "Chicano"
+ },
+ {
+ "code": 4,
+ "text": "Puerto Rican"
+ },
+ {
+ "code": 5,
+ "text": "Cuban"
+ },
+ {
+ "code": 7,
+ "text": "Other Spanish; Hispanic; Latino"
+ },
+ {
+ "code": 8,
+ "text": "DK"
+ },
+ {
+ "code": 9,
+ "text": "NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no wife/\"wife\" in FU (ER42019=0); not Spanish, Hispanic or Latino"
+ }
+ ]
+ },
+ "ER46449": {
+ "label": "K40 RACE OF WIFE-MENTION 1",
+ "page": 1459,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no wife/\"wife\" in FU (ER42019=0);"
+ }
+ ]
+ },
+ "ER46450": {
+ "label": "K40 RACE OF WIFE-MENTION 2",
+ "page": 1459,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no wife/\"wife\" in FU (ER42019=0); no second mention; NA, DK to first mention"
+ }
+ ]
+ },
+ "ER46451": {
+ "label": "K40 RACE OF WIFE-MENTION 3",
+ "page": 1459,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no wife/\"wife\" in FU (ER42019=0); fewer than three mentions; NA, DK to first mention"
+ }
+ ]
+ },
+ "ER46452": {
+ "label": "K40 RACE OF WIFE-MENTION 4",
+ "page": 1460,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no wife/\"wife\" in FU (ER42019=0); fewer than four mentions; NA, DK to first mention"
+ }
+ ]
+ },
+ "ER46542": {
+ "label": "L39 SPANISH DESCENT-HEAD",
+ "page": 1500,
+ "values": [
+ {
+ "code": 1,
+ "text": "Mexican"
+ },
+ {
+ "code": 2,
+ "text": "Mexican American"
+ },
+ {
+ "code": 3,
+ "text": "Chicano"
+ },
+ {
+ "code": 4,
+ "text": "Puerto Rican"
+ },
+ {
+ "code": 5,
+ "text": "Cuban"
+ },
+ {
+ "code": 7,
+ "text": "Other Spanish; Hispanic; Latino"
+ },
+ {
+ "code": 8,
+ "text": "DK"
+ },
+ {
+ "code": 9,
+ "text": "NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: not Spanish, Hispanic or Latino"
+ }
+ ]
+ },
+ "ER46543": {
+ "label": "L40 RACE OF HEAD-MENTION 1",
+ "page": 1501,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA"
+ }
+ ]
+ },
+ "ER46544": {
+ "label": "L40 RACE OF HEAD-MENTION 2",
+ "page": 1501,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no second mention; NA, DK to first mention"
+ }
+ ]
+ },
+ "ER46545": {
+ "label": "L40 RACE OF HEAD-MENTION 3",
+ "page": 1501,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: fewer than three mentions; NA, DK to first mention"
+ }
+ ]
+ },
+ "ER46546": {
+ "label": "L40 RACE OF HEAD-MENTION 4",
+ "page": 1502,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: fewer than four mentions; NA, DK to first mention"
+ }
+ ]
+ },
+ "ER46981": {
+ "label": "COMPLETED ED-HD",
+ "page": 1671,
+ "values": [
+ {
+ "code": 0,
+ "text": "Completed no grades of school"
+ },
+ {
+ "range": [
+ 1,
+ 16
+ ],
+ "text": "Actual number"
+ },
+ {
+ "code": 17,
+ "text": "At least some post-graduate work"
+ },
+ {
+ "code": 99,
+ "text": "DK; NA"
+ }
+ ]
+ },
+ "ER46982": {
+ "label": "COMPLETED ED-WF",
+ "page": 1672,
+ "values": [
+ {
+ "range": [
+ 1,
+ 16
+ ],
+ "text": "Actual number"
+ },
+ {
+ "code": 17,
+ "text": "At least some post-graduate work"
+ },
+ {
+ "code": 99,
+ "text": "DK; NA"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no wife/\"wife\" in FU (ER42019=0); completed no grades of school"
+ }
+ ]
+ }
+ }
+ },
+ "2011": {
+ "codebook": {
+ "path": "family/2011/FAM2011ER_codebook.pdf",
+ "sha256": "00c0517569b94efcc7fc6594529a263feb9e3bb9a88efff3becf5b8aaeebdd3c"
+ },
+ "formats_sas": null,
+ "variables": {
+ "ER51809": {
+ "label": "K39 SPANISH DESCENT-WIFE",
+ "page": 1610,
+ "values": [
+ {
+ "code": 1,
+ "text": "Mexican"
+ },
+ {
+ "code": 2,
+ "text": "Mexican American"
+ },
+ {
+ "code": 3,
+ "text": "Chicano"
+ },
+ {
+ "code": 4,
+ "text": "Puerto Rican"
+ },
+ {
+ "code": 5,
+ "text": "Cuban"
+ },
+ {
+ "code": 7,
+ "text": "Other Spanish; Hispanic; Latino"
+ },
+ {
+ "code": 8,
+ "text": "DK"
+ },
+ {
+ "code": 9,
+ "text": "NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no Wife/\"Wife\" in FU (ER47319=0); not Spanish, Hispanic or Latino"
+ }
+ ]
+ },
+ "ER51810": {
+ "label": "K40 RACE OF WIFE-MENTION 1",
+ "page": 1610,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 8,
+ "text": "DK"
+ },
+ {
+ "code": 9,
+ "text": "NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no Wife/\"Wife\" in FU (ER47319=0)"
+ }
+ ]
+ },
+ "ER51811": {
+ "label": "K40 RACE OF WIFE-MENTION 2",
+ "page": 1611,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 8,
+ "text": "DK"
+ },
+ {
+ "code": 9,
+ "text": "NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no Wife/\"Wife\" in FU (ER47319=0); NA, DK, RF to first mention (ER51810=8 or 9); no second mention"
+ }
+ ]
+ },
+ "ER51812": {
+ "label": "K40 RACE OF WIFE-MENTION 3",
+ "page": 1611,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 8,
+ "text": "DK"
+ },
+ {
+ "code": 9,
+ "text": "NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: fewer than three mentions; no Wife/\"Wife\" in FU (ER47319=0); NA, DK, RF to first mention (ER51810=8 or 9)"
+ }
+ ]
+ },
+ "ER51813": {
+ "label": "K40 RACE OF WIFE-MENTION 4",
+ "page": 1611,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 8,
+ "text": "DK"
+ },
+ {
+ "code": 9,
+ "text": "NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: fewer than four mentions; no Wife/\"Wife\" in FU (ER47319=0); NA, DK, RF to first mention (ER51810=8 or 9)"
+ }
+ ]
+ },
+ "ER51903": {
+ "label": "L39 SPANISH DESCENT-HEAD",
+ "page": 1655,
+ "values": [
+ {
+ "code": 1,
+ "text": "Mexican"
+ },
+ {
+ "code": 2,
+ "text": "Mexican American"
+ },
+ {
+ "code": 3,
+ "text": "Chicano"
+ },
+ {
+ "code": 4,
+ "text": "Puerto Rican"
+ },
+ {
+ "code": 5,
+ "text": "Cuban"
+ },
+ {
+ "code": 7,
+ "text": "Other Spanish; Hispanic; Latino"
+ },
+ {
+ "code": 8,
+ "text": "DK"
+ },
+ {
+ "code": 9,
+ "text": "NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: not Spanish, Hispanic or Latino"
+ }
+ ]
+ },
+ "ER51904": {
+ "label": "L40 RACE OF HEAD-MENTION 1",
+ "page": 1655,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 8,
+ "text": "DK"
+ },
+ {
+ "code": 9,
+ "text": "NA; refused"
+ }
+ ]
+ },
+ "ER51905": {
+ "label": "L40 RACE OF HEAD-MENTION 2",
+ "page": 1656,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 8,
+ "text": "DK"
+ },
+ {
+ "code": 9,
+ "text": "NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: NA, DK, RF to first mention (ER51904=8 or 9); no second mention"
+ }
+ ]
+ },
+ "ER51906": {
+ "label": "L40 RACE OF HEAD-MENTION 3",
+ "page": 1656,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 8,
+ "text": "DK"
+ },
+ {
+ "code": 9,
+ "text": "NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: fewer than three mentions; NA, DK, RF to first mention (ER51904=8 or 9)"
+ }
+ ]
+ },
+ "ER51907": {
+ "label": "L40 RACE OF HEAD-MENTION 4",
+ "page": 1656,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 8,
+ "text": "DK"
+ },
+ {
+ "code": 9,
+ "text": "NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: fewer than four mentions; NA, DK, RF to first mention (ER51904=8 or 9)"
+ }
+ ]
+ },
+ "ER52405": {
+ "label": "COMPLETED ED-HD",
+ "page": 1861,
+ "values": [
+ {
+ "code": 0,
+ "text": "Completed no grades of school"
+ },
+ {
+ "range": [
+ 1,
+ 16
+ ],
+ "text": "Actual number"
+ },
+ {
+ "code": 17,
+ "text": "At least some post-graduate work"
+ },
+ {
+ "code": 99,
+ "text": "DK; NA"
+ }
+ ]
+ },
+ "ER52406": {
+ "label": "COMPLETED ED-WF",
+ "page": 1862,
+ "values": [
+ {
+ "range": [
+ 1,
+ 16
+ ],
+ "text": "Actual number"
+ },
+ {
+ "code": 17,
+ "text": "At least some post-graduate work"
+ },
+ {
+ "code": 99,
+ "text": "DK; NA"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no Wife/\"Wife\" in FU (ER47319=0); completed no grades of school"
+ }
+ ]
+ }
+ }
+ },
+ "2013": {
+ "codebook": {
+ "path": "family/2013/FAM2013ER_codebook.pdf",
+ "sha256": "40cc2de6af801b94b93518f7252e24c7c95098e713fc97f9ead1af74c4ce25d6"
+ },
+ "formats_sas": null,
+ "variables": {
+ "ER57541": {
+ "label": "K33 STATE WIFE WAS BORN",
+ "page": 1618,
+ "values": [
+ {
+ "range": [
+ 1,
+ 56
+ ],
+ "text": "Actual state (FIPS code)"
+ },
+ {
+ "code": 99,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: U.S. territory or foreign country; no Wife/\"Wife\" in FU (ER54305=5)"
+ }
+ ]
+ },
+ "ER57542": {
+ "label": "K33YR YEAR CAME TO UNITED STATES-WF",
+ "page": 1619,
+ "values": [
+ {
+ "range": [
+ 1901,
+ 2013
+ ],
+ "text": "Actual year"
+ },
+ {
+ "code": 9997,
+ "text": "Not living in the United States"
+ },
+ {
+ "code": 9998,
+ "text": "DK"
+ },
+ {
+ "code": 9999,
+ "text": "NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no Wife/\"Wife\" in FU (ER54305=5); Wife/\"Wife\" was born in the United States or U.S. territory"
+ }
+ ]
+ },
+ "ER57548": {
+ "label": "K39 SPANISH DESCENT-WIFE",
+ "page": 1620,
+ "values": [
+ {
+ "code": 1,
+ "text": "Mexican"
+ },
+ {
+ "code": 2,
+ "text": "Mexican American"
+ },
+ {
+ "code": 3,
+ "text": "Chicano"
+ },
+ {
+ "code": 4,
+ "text": "Puerto Rican"
+ },
+ {
+ "code": 5,
+ "text": "Cuban"
+ },
+ {
+ "code": 7,
+ "text": "Other Spanish; Hispanic; Latino"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no Wife/\"Wife\" in FU (ER54305=5); not Spanish, Hispanic or Latino"
+ }
+ ]
+ },
+ "ER57549": {
+ "label": "K40 RACE OF WIFE-MENTION 1",
+ "page": 1621,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no Wife/\"Wife\" in FU (ER54305=5)"
+ }
+ ]
+ },
+ "ER57550": {
+ "label": "K40 RACE OF WIFE-MENTION 2",
+ "page": 1621,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no Wife/\"Wife\" in FU (ER54305=5); DK, NA, or RF to first mention (ER57549=8 or 9); no second mention"
+ }
+ ]
+ },
+ "ER57551": {
+ "label": "K40 RACE OF WIFE-MENTION 3",
+ "page": 1621,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: fewer than three mentions; no Wife/\"Wife\" in FU (ER54305=5); DK, NA, or RF to first mention (ER57549=8 or 9)"
+ }
+ ]
+ },
+ "ER57552": {
+ "label": "K40 RACE OF WIFE-MENTION 4",
+ "page": 1622,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: fewer than four mentions; no Wife/\"Wife\" in FU (ER54305=5); DK, NA, or RF to first mention (ER57549=8 or 9)"
+ }
+ ]
+ },
+ "ER57651": {
+ "label": "L33 STATE HEAD WAS BORN",
+ "page": 1675,
+ "values": [
+ {
+ "range": [
+ 1,
+ 56
+ ],
+ "text": "Actual state (FIPS code)"
+ },
+ {
+ "code": 99,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: U.S. territory or foreign country"
+ }
+ ]
+ },
+ "ER57652": {
+ "label": "L33YR YEAR CAME TO UNITED STATES-HD",
+ "page": 1675,
+ "values": [
+ {
+ "range": [
+ 1901,
+ 2013
+ ],
+ "text": "Actual year"
+ },
+ {
+ "code": 9997,
+ "text": "Not living in the United States"
+ },
+ {
+ "code": 9998,
+ "text": "DK"
+ },
+ {
+ "code": 9999,
+ "text": "NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: Head was born in the United States or U.S. territory"
+ }
+ ]
+ },
+ "ER57658": {
+ "label": "L39 SPANISH DESCENT-HEAD",
+ "page": 1677,
+ "values": [
+ {
+ "code": 1,
+ "text": "Mexican"
+ },
+ {
+ "code": 2,
+ "text": "Mexican American"
+ },
+ {
+ "code": 3,
+ "text": "Chicano"
+ },
+ {
+ "code": 4,
+ "text": "Puerto Rican"
+ },
+ {
+ "code": 5,
+ "text": "Cuban"
+ },
+ {
+ "code": 7,
+ "text": "Other Spanish; Hispanic; Latino"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: not Spanish, Hispanic or Latino"
+ }
+ ]
+ },
+ "ER57659": {
+ "label": "L40 RACE OF HEAD-MENTION 1",
+ "page": 1677,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ }
+ ]
+ },
+ "ER57660": {
+ "label": "L40 RACE OF HEAD-MENTION 2",
+ "page": 1678,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: DK, NA, or RF to first mention (ER57659=8 or 9); no second mention"
+ }
+ ]
+ },
+ "ER57661": {
+ "label": "L40 RACE OF HEAD-MENTION 3",
+ "page": 1678,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: fewer than three mentions; DK, NA, or RF to first mention (ER57659=8 or 9)"
+ }
+ ]
+ },
+ "ER57662": {
+ "label": "L40 RACE OF HEAD-MENTION 4",
+ "page": 1678,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: fewer than four mentions; DK, NA, or RF to first mention (ER57659=8 or 9)"
+ }
+ ]
+ },
+ "ER58223": {
+ "label": "COMPLETED ED-HD",
+ "page": 1876,
+ "values": [
+ {
+ "code": 0,
+ "text": "Completed no grades of school"
+ },
+ {
+ "range": [
+ 1,
+ 17
+ ],
+ "text": "Actual number"
+ },
+ {
+ "code": 99,
+ "text": "DK; NA"
+ }
+ ]
+ },
+ "ER58224": {
+ "label": "COMPLETED ED-WF",
+ "page": 1877,
+ "values": [
+ {
+ "range": [
+ 1,
+ 17
+ ],
+ "text": "Actual number"
+ },
+ {
+ "code": 99,
+ "text": "DK; NA"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no Wife/\"Wife\" in FU (ER54305=5); completed no grades of school"
+ }
+ ]
+ }
+ }
+ },
+ "2015": {
+ "codebook": {
+ "path": "family/2015/FAM2015ER_codebook.pdf",
+ "sha256": "7dcb71f823e309a286b904531410243d58cde48b8f50079401856de9ac762238"
+ },
+ "formats_sas": null,
+ "variables": {
+ "ER64663": {
+ "label": "K33 STATE SPOUSE WAS BORN",
+ "page": 1677,
+ "values": [
+ {
+ "range": [
+ 1,
+ 56
+ ],
+ "text": "Actual state (FIPS code)"
+ },
+ {
+ "code": 99,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no Spouse/Partner in FU (ER61347=5); U.S. territory or foreign country"
+ }
+ ]
+ },
+ "ER64664": {
+ "label": "K33YR YEAR CAME TO UNITED STATES-SP",
+ "page": 1677,
+ "values": [
+ {
+ "range": [
+ 1920,
+ 2015
+ ],
+ "text": "Actual year"
+ },
+ {
+ "code": 9997,
+ "text": "Not living in the United States"
+ },
+ {
+ "code": 9999,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no Spouse/Partner in FU (ER61347=5); Spouse/Partner was born in the United States or U.S. territory"
+ }
+ ]
+ },
+ "ER64670": {
+ "label": "K39 SPANISH DESCENT-SPOUSE",
+ "page": 1679,
+ "values": [
+ {
+ "code": 1,
+ "text": "Mexican"
+ },
+ {
+ "code": 2,
+ "text": "Mexican American"
+ },
+ {
+ "code": 3,
+ "text": "Chicano"
+ },
+ {
+ "code": 4,
+ "text": "Puerto Rican"
+ },
+ {
+ "code": 5,
+ "text": "Cuban"
+ },
+ {
+ "code": 7,
+ "text": "Other Spanish; Hispanic; Latino"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no Spouse/Partner in FU (ER61347=5); not Spanish, Hispanic or Latino"
+ }
+ ]
+ },
+ "ER64671": {
+ "label": "K40 RACE OF SPOUSE-MENTION 1",
+ "page": 1679,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no Spouse/Partner in FU (ER61347=5)"
+ }
+ ]
+ },
+ "ER64672": {
+ "label": "K40 RACE OF SPOUSE-MENTION 2",
+ "page": 1680,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no Spouse/Partner in FU (ER61347=5); DK, NA, or RF to first mention (ER64671=9); no second mention"
+ }
+ ]
+ },
+ "ER64673": {
+ "label": "K40 RACE OF SPOUSE-MENTION 3",
+ "page": 1680,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no Spouse/Partner in FU (ER61347=5); DK, NA, or RF to first mention (ER64671=9); fewer than three mentions"
+ }
+ ]
+ },
+ "ER64674": {
+ "label": "K40 RACE OF SPOUSE-MENTION 4",
+ "page": 1680,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no Spouse/Partner in FU (ER61347=5); DK, NA, or RF to first mention (ER64671=9); fewer than four mentions"
+ }
+ ]
+ },
+ "ER64802": {
+ "label": "L33 STATE HEAD WAS BORN",
+ "page": 1769,
+ "values": [
+ {
+ "range": [
+ 1,
+ 56
+ ],
+ "text": "Actual state (FIPS code)"
+ },
+ {
+ "code": 99,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: U.S. territory or foreign country"
+ }
+ ]
+ },
+ "ER64803": {
+ "label": "L33YR YEAR CAME TO UNITED STATES-HD",
+ "page": 1769,
+ "values": [
+ {
+ "range": [
+ 1901,
+ 2015
+ ],
+ "text": "Actual year"
+ },
+ {
+ "code": 9997,
+ "text": "Not living in the United States"
+ },
+ {
+ "code": 9999,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: Head was born in the United States or U.S. territory"
+ }
+ ]
+ },
+ "ER64809": {
+ "label": "L39 SPANISH DESCENT-HEAD",
+ "page": 1771,
+ "values": [
+ {
+ "code": 1,
+ "text": "Mexican"
+ },
+ {
+ "code": 2,
+ "text": "Mexican American"
+ },
+ {
+ "code": 3,
+ "text": "Chicano"
+ },
+ {
+ "code": 4,
+ "text": "Puerto Rican"
+ },
+ {
+ "code": 5,
+ "text": "Cuban"
+ },
+ {
+ "code": 7,
+ "text": "Other Spanish; Hispanic; Latino"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: not Spanish, Hispanic or Latino"
+ }
+ ]
+ },
+ "ER64810": {
+ "label": "L40 RACE OF HEAD-MENTION 1",
+ "page": 1771,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ }
+ ]
+ },
+ "ER64811": {
+ "label": "L40 RACE OF HEAD-MENTION 2",
+ "page": 1772,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: DK, NA, or RF to first mention (ER64810=9); no second mention"
+ }
+ ]
+ },
+ "ER64812": {
+ "label": "L40 RACE OF HEAD-MENTION 3",
+ "page": 1772,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: DK, NA, or RF to first mention (ER64810=9); fewer than three mentions"
+ }
+ ]
+ },
+ "ER64813": {
+ "label": "L40 RACE OF HEAD-MENTION 4",
+ "page": 1772,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: DK, NA, or RF to first mention (ER64810=9); fewer than four mentions"
+ }
+ ]
+ },
+ "ER65459": {
+ "label": "COMPLETED ED-HD",
+ "page": 2017,
+ "values": [
+ {
+ "code": 0,
+ "text": "Completed no grades of school"
+ },
+ {
+ "range": [
+ 1,
+ 17
+ ],
+ "text": "Actual number"
+ },
+ {
+ "code": 99,
+ "text": "DK; NA"
+ }
+ ]
+ },
+ "ER65460": {
+ "label": "COMPLETED ED-SP",
+ "page": 2018,
+ "values": [
+ {
+ "range": [
+ 1,
+ 17
+ ],
+ "text": "Actual number"
+ },
+ {
+ "code": 99,
+ "text": "DK; NA"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: completed no grades of school; no Spouse/Partner in FU (ER61347=5)"
+ }
+ ]
+ }
+ }
+ },
+ "2017": {
+ "codebook": {
+ "path": "family/2017/FAM2017ER_codebook.pdf",
+ "sha256": "19bfba9f7a3fecc82a18711466e2b2a5f00b49c9d3db0f9f8587795fad2bf8c1"
+ },
+ "formats_sas": null,
+ "variables": {
+ "ER70736": {
+ "label": "K33 STATE SPOUSE WAS BORN",
+ "page": 1708,
+ "values": [
+ {
+ "range": [
+ 1,
+ 56
+ ],
+ "text": "Actual state (FIPS code)"
+ },
+ {
+ "code": 99,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no Spouse/Partner in FU (ER67399=5); U.S. territory or foreign country"
+ }
+ ]
+ },
+ "ER70737": {
+ "label": "K33YR YEAR CAME TO UNITED STATES-SP",
+ "page": 1708,
+ "values": [
+ {
+ "range": [
+ 1920,
+ 2017
+ ],
+ "text": "Actual year"
+ },
+ {
+ "code": 9997,
+ "text": "Not living in the United States"
+ },
+ {
+ "code": 9999,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no Spouse/Partner in FU (ER67399=5); Spouse/Partner was born in the United States or U.S. territory"
+ }
+ ]
+ },
+ "ER70743": {
+ "label": "K39 SPANISH DESCENT-SPOUSE",
+ "page": 1709,
+ "values": [
+ {
+ "code": 1,
+ "text": "Mexican"
+ },
+ {
+ "code": 2,
+ "text": "Mexican American"
+ },
+ {
+ "code": 3,
+ "text": "Chicano"
+ },
+ {
+ "code": 4,
+ "text": "Puerto Rican"
+ },
+ {
+ "code": 5,
+ "text": "Cuban"
+ },
+ {
+ "code": 7,
+ "text": "Other Spanish; Hispanic; Latino"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no Spouse/Partner in FU (ER67399=5); not Spanish, Hispanic or Latino"
+ }
+ ]
+ },
+ "ER70744": {
+ "label": "K40 RACE OF SPOUSE-MENTION 1",
+ "page": 1710,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no Spouse/Partner in FU (ER67399=5)"
+ }
+ ]
+ },
+ "ER70745": {
+ "label": "K40 RACE OF SPOUSE-MENTION 2",
+ "page": 1710,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no Spouse/Partner in FU (ER67399=5); DK, NA, or RF to first mention (ER70744=9); no second mention"
+ }
+ ]
+ },
+ "ER70746": {
+ "label": "K40 RACE OF SPOUSE-MENTION 3",
+ "page": 1710,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no Spouse/Partner in FU (ER67399=5); DK, NA, or RF to first mention (ER70744=9); fewer than three mentions"
+ }
+ ]
+ },
+ "ER70747": {
+ "label": "K40 RACE OF SPOUSE-MENTION 4",
+ "page": 1711,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no Spouse/Partner in FU (ER67399=5); DK, NA, or RF to first mention (ER70744=9); fewer than four mentions"
+ }
+ ]
+ },
+ "ER70874": {
+ "label": "L33 STATE REFERENCE PERSON WAS BORN",
+ "page": 1799,
+ "values": [
+ {
+ "range": [
+ 1,
+ 56
+ ],
+ "text": "Actual state (FIPS code)"
+ },
+ {
+ "code": 99,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: U.S. territory or foreign country"
+ }
+ ]
+ },
+ "ER70875": {
+ "label": "L33YR YEAR CAME TO UNITED STATES-RP",
+ "page": 1799,
+ "values": [
+ {
+ "range": [
+ 1901,
+ 2017
+ ],
+ "text": "Actual year"
+ },
+ {
+ "code": 9997,
+ "text": "Not living in the United States"
+ },
+ {
+ "code": 9999,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: Reference Person was born in the United States or U.S. territory"
+ }
+ ]
+ },
+ "ER70881": {
+ "label": "L39 SPANISH DESCENT-RP",
+ "page": 1801,
+ "values": [
+ {
+ "code": 1,
+ "text": "Mexican"
+ },
+ {
+ "code": 2,
+ "text": "Mexican American"
+ },
+ {
+ "code": 3,
+ "text": "Chicano"
+ },
+ {
+ "code": 4,
+ "text": "Puerto Rican"
+ },
+ {
+ "code": 5,
+ "text": "Cuban"
+ },
+ {
+ "code": 7,
+ "text": "Other Spanish; Hispanic; Latino"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: not Spanish, Hispanic or Latino"
+ }
+ ]
+ },
+ "ER70882": {
+ "label": "L40 RACE OF REFERENCE PERSON-MENTION 1",
+ "page": 1802,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ }
+ ]
+ },
+ "ER70883": {
+ "label": "L40 RACE OF REFERENCE PERSON-MENTION 2",
+ "page": 1802,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: DK, NA, or RF to first mention (ER70882=9); no second mention"
+ }
+ ]
+ },
+ "ER70884": {
+ "label": "L40 RACE OF REFERENCE PERSON-MENTION 3",
+ "page": 1802,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: DK, NA, or RF to first mention (ER70882=9); fewer than three mentions"
+ }
+ ]
+ },
+ "ER70885": {
+ "label": "L40 RACE OF REFERENCE PERSON-MENTION 4",
+ "page": 1803,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: DK, NA, or RF to first mention (ER70882=9); fewer than four mentions"
+ }
+ ]
+ },
+ "ER71538": {
+ "label": "COMPLETED ED-RP",
+ "page": 2071,
+ "values": [
+ {
+ "code": 0,
+ "text": "Completed no grades of school"
+ },
+ {
+ "range": [
+ 1,
+ 17
+ ],
+ "text": "Actual number"
+ },
+ {
+ "code": 99,
+ "text": "DK; NA"
+ }
+ ]
+ },
+ "ER71539": {
+ "label": "COMPLETED ED-SP",
+ "page": 2072,
+ "values": [
+ {
+ "range": [
+ 1,
+ 17
+ ],
+ "text": "Actual number"
+ },
+ {
+ "code": 99,
+ "text": "DK; NA"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: completed no grades of school; no Spouse/Partner in FU (ER67399=5)"
+ }
+ ]
+ }
+ }
+ },
+ "2019": {
+ "codebook": {
+ "path": "family/2019/fam2019er_codebook.pdf",
+ "sha256": "ce637e37e0f5e306dd7f10455e85069b56c9b13d5c138d3eb4b9c8d6a8452406"
+ },
+ "formats_sas": null,
+ "variables": {
+ "ER76744": {
+ "label": "K33 STATE SPOUSE WAS BORN",
+ "page": 1706,
+ "values": [
+ {
+ "range": [
+ 1,
+ 56
+ ],
+ "text": "Actual state (FIPS code)"
+ },
+ {
+ "code": 99,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no Spouse/Partner in FU (ER73422=5); U.S. territory or foreign country"
+ }
+ ]
+ },
+ "ER76745": {
+ "label": "K33YR YEAR CAME TO UNITED STATES-SP",
+ "page": 1706,
+ "values": [
+ {
+ "range": [
+ 1920,
+ 2019
+ ],
+ "text": "Actual year"
+ },
+ {
+ "code": 9997,
+ "text": "Not living in the United States"
+ },
+ {
+ "code": 9999,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no Spouse/Partner in FU (ER73422=5); Spouse/Partner was born in the United States or U.S. territory"
+ }
+ ]
+ },
+ "ER76751": {
+ "label": "K39 SPANISH DESCENT-SPOUSE",
+ "page": 1707,
+ "values": [
+ {
+ "code": 1,
+ "text": "Mexican"
+ },
+ {
+ "code": 2,
+ "text": "Mexican American"
+ },
+ {
+ "code": 3,
+ "text": "Chicano"
+ },
+ {
+ "code": 4,
+ "text": "Puerto Rican"
+ },
+ {
+ "code": 5,
+ "text": "Cuban"
+ },
+ {
+ "code": 7,
+ "text": "Other Spanish; Hispanic; Latino"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no Spouse/Partner in FU (ER73422=5); not Spanish, Hispanic or Latino"
+ }
+ ]
+ },
+ "ER76752": {
+ "label": "K40 RACE OF SPOUSE-MENTION 1",
+ "page": 1708,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no Spouse/Partner in FU (ER73422=5)"
+ }
+ ]
+ },
+ "ER76753": {
+ "label": "K40 RACE OF SPOUSE-MENTION 2",
+ "page": 1708,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no Spouse/Partner in FU (ER73422=5); DK, NA, or RF to first mention (ER76752=9); no second mention"
+ }
+ ]
+ },
+ "ER76754": {
+ "label": "K40 RACE OF SPOUSE-MENTION 3",
+ "page": 1708,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no Spouse/Partner in FU (ER73422=5); DK, NA, or RF to first mention (ER76752=9); fewer than three mentions"
+ }
+ ]
+ },
+ "ER76755": {
+ "label": "K40 RACE OF SPOUSE-MENTION 4",
+ "page": 1709,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no Spouse/Partner in FU (ER73422=5); DK, NA, or RF to first mention (ER76752=9); fewer than four mentions"
+ }
+ ]
+ },
+ "ER76889": {
+ "label": "L33 STATE REFERENCE PERSON WAS BORN",
+ "page": 1791,
+ "values": [
+ {
+ "range": [
+ 1,
+ 56
+ ],
+ "text": "Actual state (FIPS code)"
+ },
+ {
+ "code": 99,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: U.S. territory or foreign country"
+ }
+ ]
+ },
+ "ER76890": {
+ "label": "L33YR YEAR CAME TO UNITED STATES-RP",
+ "page": 1791,
+ "values": [
+ {
+ "range": [
+ 1901,
+ 2019
+ ],
+ "text": "Actual year"
+ },
+ {
+ "code": 9997,
+ "text": "Not living in the United States"
+ },
+ {
+ "code": 9999,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: Reference Person was born in the United States or U.S. territory"
+ }
+ ]
+ },
+ "ER76896": {
+ "label": "L39 SPANISH DESCENT-RP",
+ "page": 1792,
+ "values": [
+ {
+ "code": 1,
+ "text": "Mexican"
+ },
+ {
+ "code": 2,
+ "text": "Mexican American"
+ },
+ {
+ "code": 3,
+ "text": "Chicano"
+ },
+ {
+ "code": 4,
+ "text": "Puerto Rican"
+ },
+ {
+ "code": 5,
+ "text": "Cuban"
+ },
+ {
+ "code": 7,
+ "text": "Other Spanish; Hispanic; Latino"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: not Spanish, Hispanic or Latino"
+ }
+ ]
+ },
+ "ER76897": {
+ "label": "L40 RACE OF REFERENCE PERSON-MENTION 1",
+ "page": 1793,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ }
+ ]
+ },
+ "ER76898": {
+ "label": "L40 RACE OF REFERENCE PERSON-MENTION 2",
+ "page": 1793,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: DK, NA, or RF to first mention (ER76897=9); no second mention"
+ }
+ ]
+ },
+ "ER76899": {
+ "label": "L40 RACE OF REFERENCE PERSON-MENTION 3",
+ "page": 1793,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: DK, NA, or RF to first mention (ER76897=9); fewer than three mentions"
+ }
+ ]
+ },
+ "ER76900": {
+ "label": "L40 RACE OF REFERENCE PERSON-MENTION 4",
+ "page": 1794,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: DK, NA, or RF to first mention (ER76897=9); fewer than four mentions"
+ }
+ ]
+ },
+ "ER77599": {
+ "label": "COMPLETED ED-RP",
+ "page": 2056,
+ "values": [
+ {
+ "code": 0,
+ "text": "Completed no grades of school"
+ },
+ {
+ "range": [
+ 1,
+ 17
+ ],
+ "text": "Actual number"
+ },
+ {
+ "code": 99,
+ "text": "DK; NA"
+ }
+ ]
+ },
+ "ER77600": {
+ "label": "COMPLETED ED-SP",
+ "page": 2057,
+ "values": [
+ {
+ "range": [
+ 1,
+ 17
+ ],
+ "text": "Actual number"
+ },
+ {
+ "code": 99,
+ "text": "DK; NA"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: completed no grades of school; no Spouse/Partner in FU (ER73422=5)"
+ }
+ ]
+ }
+ }
+ },
+ "2021": {
+ "codebook": {
+ "path": "family/2021/FAM2021ER_codebook.pdf",
+ "sha256": "e3b88ef394952fa05b9b98eb7f05bf9a99896840515942562aa37399e1bcd462"
+ },
+ "formats_sas": {
+ "path": "family/2021/FAM2021ER_formats.sas",
+ "sha256": "0f38201437ece15f8c5ac1b8c773e36e8f2db25e362aa0434b75a34cea3e0e99"
+ },
+ "variables": {
+ "ER81009": {
+ "label": "K33 STATE SPOUSE WAS BORN",
+ "page": 1048,
+ "values": [
+ {
+ "range": [
+ 1,
+ 56
+ ],
+ "text": "Actual state (FIPS code)"
+ },
+ {
+ "code": 99,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no Spouse/Partner in FU (ER79524=5); U.S. territory or foreign country"
+ }
+ ]
+ },
+ "ER81010": {
+ "label": "K33YR YEAR CAME TO UNITED STATES-SP",
+ "page": 1048,
+ "values": [
+ {
+ "range": [
+ 1920,
+ 2021
+ ],
+ "text": "Actual year"
+ },
+ {
+ "code": 9997,
+ "text": "Not living in the United States"
+ },
+ {
+ "code": 9999,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no Spouse/Partner in FU (ER79524=5); Spouse/Partner was born in the United States or U.S. territory"
+ }
+ ]
+ },
+ "ER81016": {
+ "label": "K39 SPANISH DESCENT-SPOUSE",
+ "page": 1050,
+ "values": [
+ {
+ "code": 1,
+ "text": "Mexican"
+ },
+ {
+ "code": 2,
+ "text": "Mexican American"
+ },
+ {
+ "code": 3,
+ "text": "Chicano"
+ },
+ {
+ "code": 4,
+ "text": "Puerto Rican"
+ },
+ {
+ "code": 5,
+ "text": "Cuban"
+ },
+ {
+ "code": 7,
+ "text": "Other Spanish; Hispanic; Latino"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no Spouse/Partner in FU (ER79524=5); not Spanish, Hispanic or Latino"
+ }
+ ]
+ },
+ "ER81017": {
+ "label": "K40 RACE OF SPOUSE-MENTION 1",
+ "page": 1050,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no Spouse/Partner in FU (ER79524=5)"
+ }
+ ]
+ },
+ "ER81018": {
+ "label": "K40 RACE OF SPOUSE-MENTION 2",
+ "page": 1050,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no Spouse/Partner in FU (ER79524=5); DK, NA, or RF to first mention (ER81017=9); no second mention"
+ }
+ ]
+ },
+ "ER81019": {
+ "label": "K40 RACE OF SPOUSE-MENTION 3",
+ "page": 1051,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no Spouse/Partner in FU (ER79524=5); DK, NA, or RF to first mention (ER81017=9); fewer than three mentions"
+ }
+ ]
+ },
+ "ER81020": {
+ "label": "K40 RACE OF SPOUSE-MENTION 4",
+ "page": 1051,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no Spouse/Partner in FU (ER79524=5); DK, NA, or RF to first mention (ER81017=9); fewer than four mentions"
+ }
+ ]
+ },
+ "ER81136": {
+ "label": "L33 STATE REFERENCE PERSON WAS BORN",
+ "page": 1122,
+ "values": [
+ {
+ "range": [
+ 1,
+ 56
+ ],
+ "text": "Actual state (FIPS code)"
+ },
+ {
+ "code": 99,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: U.S. territory or foreign country"
+ }
+ ]
+ },
+ "ER81137": {
+ "label": "L33YR YEAR CAME TO UNITED STATES-RP",
+ "page": 1122,
+ "values": [
+ {
+ "range": [
+ 1901,
+ 2021
+ ],
+ "text": "Actual year"
+ },
+ {
+ "code": 9997,
+ "text": "Not living in the United States"
+ },
+ {
+ "code": 9999,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: Reference Person was born in the United States or U.S. territory"
+ }
+ ]
+ },
+ "ER81143": {
+ "label": "L39 SPANISH DESCENT-RP",
+ "page": 1123,
+ "values": [
+ {
+ "code": 1,
+ "text": "Mexican"
+ },
+ {
+ "code": 2,
+ "text": "Mexican American"
+ },
+ {
+ "code": 3,
+ "text": "Chicano"
+ },
+ {
+ "code": 4,
+ "text": "Puerto Rican"
+ },
+ {
+ "code": 5,
+ "text": "Cuban"
+ },
+ {
+ "code": 7,
+ "text": "Other Spanish; Hispanic; Latino"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: not Spanish, Hispanic or Latino"
+ }
+ ]
+ },
+ "ER81144": {
+ "label": "L40 RACE OF REFERENCE PERSON-MENTION 1",
+ "page": 1124,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ }
+ ]
+ },
+ "ER81145": {
+ "label": "L40 RACE OF REFERENCE PERSON-MENTION 2",
+ "page": 1124,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: DK, NA, or RF to first mention (ER81144=9); no second mention"
+ }
+ ]
+ },
+ "ER81146": {
+ "label": "L40 RACE OF REFERENCE PERSON-MENTION 3",
+ "page": 1124,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: DK, NA, or RF to first mention (ER81144=9); fewer than three mentions"
+ }
+ ]
+ },
+ "ER81147": {
+ "label": "L40 RACE OF REFERENCE PERSON-MENTION 4",
+ "page": 1125,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: DK, NA, or RF to first mention (ER81144=9); fewer than four mentions"
+ }
+ ]
+ },
+ "ER81926": {
+ "label": "COMPLETED ED-RP",
+ "page": 1420,
+ "values": [
+ {
+ "code": 0,
+ "text": "Completed no grades of school"
+ },
+ {
+ "range": [
+ 1,
+ 17
+ ],
+ "text": "Actual number"
+ },
+ {
+ "code": 99,
+ "text": "DK; NA"
+ }
+ ]
+ },
+ "ER81927": {
+ "label": "COMPLETED ED-SP",
+ "page": 1421,
+ "values": [
+ {
+ "range": [
+ 1,
+ 17
+ ],
+ "text": "Actual number"
+ },
+ {
+ "code": 99,
+ "text": "DK; NA"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: completed no grades of school; no Spouse/Partner in FU (ER79524=5)"
+ }
+ ]
+ }
+ }
+ },
+ "2023": {
+ "codebook": {
+ "path": "family/2023/FAM2023ER_codebook.pdf",
+ "sha256": "85a4078ebb79146d023df3d9d4e2d2ef6017eced849656ca39f4b740cb5a5688"
+ },
+ "formats_sas": {
+ "path": "family/2023/FAM2023ER_formats.sas",
+ "sha256": "a1c2c4531369f2857de5502169ac8d70ab12f79c859a650eaeceeb18c36ba17f"
+ },
+ "variables": {
+ "ER84986": {
+ "label": "K33 STATE SPOUSE WAS BORN",
+ "page": 1039,
+ "values": [
+ {
+ "range": [
+ 1,
+ 56
+ ],
+ "text": "Actual state (FIPS code)"
+ },
+ {
+ "code": 99,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no Spouse/Partner in FU (ER83493=5); U.S. territory or foreign country"
+ }
+ ]
+ },
+ "ER84987": {
+ "label": "K33YR YEAR CAME TO UNITED STATES-SP",
+ "page": 1039,
+ "values": [
+ {
+ "range": [
+ 1920,
+ 2023
+ ],
+ "text": "Actual year"
+ },
+ {
+ "code": 9997,
+ "text": "Not living in the United States"
+ },
+ {
+ "code": 9999,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no Spouse/Partner in FU (ER83493=5); Spouse/Partner was born in the United States or U.S. territory"
+ }
+ ]
+ },
+ "ER84993": {
+ "label": "K39 SPANISH DESCENT-SPOUSE",
+ "page": 1041,
+ "values": [
+ {
+ "code": 1,
+ "text": "Mexican"
+ },
+ {
+ "code": 2,
+ "text": "Mexican American"
+ },
+ {
+ "code": 3,
+ "text": "Chicano"
+ },
+ {
+ "code": 4,
+ "text": "Puerto Rican"
+ },
+ {
+ "code": 5,
+ "text": "Cuban"
+ },
+ {
+ "code": 7,
+ "text": "Other Spanish; Hispanic; Latino"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no Spouse/Partner in FU (ER83493=5); not Spanish, Hispanic or Latino"
+ }
+ ]
+ },
+ "ER84994": {
+ "label": "K40 RACE OF SPOUSE-MENTION 1",
+ "page": 1041,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no Spouse/Partner in FU (ER83493=5)"
+ }
+ ]
+ },
+ "ER84995": {
+ "label": "K40 RACE OF SPOUSE-MENTION 2",
+ "page": 1041,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no Spouse/Partner in FU (ER83493=5); DK, NA, or RF to first mention (ER84994=9); no second mention"
+ }
+ ]
+ },
+ "ER84996": {
+ "label": "K40 RACE OF SPOUSE-MENTION 3",
+ "page": 1042,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no Spouse/Partner in FU (ER83493=5); DK, NA, or RF to first mention (ER84994=9); fewer than three mentions"
+ }
+ ]
+ },
+ "ER84997": {
+ "label": "K40 RACE OF SPOUSE-MENTION 4",
+ "page": 1042,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: no Spouse/Partner in FU (ER83493=5); DK, NA, or RF to first mention (ER84994=9); fewer than four mentions"
+ }
+ ]
+ },
+ "ER85113": {
+ "label": "L33 STATE REFERENCE PERSON WAS BORN",
+ "page": 1113,
+ "values": [
+ {
+ "range": [
+ 1,
+ 56
+ ],
+ "text": "Actual state (FIPS code)"
+ },
+ {
+ "code": 99,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: U.S. territory or foreign country"
+ }
+ ]
+ },
+ "ER85114": {
+ "label": "L33YR YEAR CAME TO UNITED STATES-RP",
+ "page": 1113,
+ "values": [
+ {
+ "range": [
+ 1901,
+ 2023
+ ],
+ "text": "Actual year"
+ },
+ {
+ "code": 9997,
+ "text": "Not living in the United States"
+ },
+ {
+ "code": 9999,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: Reference Person was born in the United States or U.S. territory"
+ }
+ ]
+ },
+ "ER85120": {
+ "label": "L39 SPANISH DESCENT-RP",
+ "page": 1114,
+ "values": [
+ {
+ "code": 1,
+ "text": "Mexican"
+ },
+ {
+ "code": 2,
+ "text": "Mexican American"
+ },
+ {
+ "code": 3,
+ "text": "Chicano"
+ },
+ {
+ "code": 4,
+ "text": "Puerto Rican"
+ },
+ {
+ "code": 5,
+ "text": "Cuban"
+ },
+ {
+ "code": 7,
+ "text": "Other Spanish; Hispanic; Latino"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: not Spanish, Hispanic or Latino"
+ }
+ ]
+ },
+ "ER85121": {
+ "label": "L40 RACE OF REFERENCE PERSON-MENTION 1",
+ "page": 1115,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 9,
+ "text": "DK; NA; refused"
+ }
+ ]
+ },
+ "ER85122": {
+ "label": "L40 RACE OF REFERENCE PERSON-MENTION 2",
+ "page": 1115,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: DK, NA, or RF to first mention (ER85121=9); no second mention"
+ }
+ ]
+ },
+ "ER85123": {
+ "label": "L40 RACE OF REFERENCE PERSON-MENTION 3",
+ "page": 1115,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: DK, NA, or RF to first mention (ER85121=9); fewer than three mentions"
+ }
+ ]
+ },
+ "ER85124": {
+ "label": "L40 RACE OF REFERENCE PERSON-MENTION 4",
+ "page": 1116,
+ "values": [
+ {
+ "code": 1,
+ "text": "White"
+ },
+ {
+ "code": 2,
+ "text": "Black, African-American, or Negro"
+ },
+ {
+ "code": 3,
+ "text": "American Indian or Alaska Native"
+ },
+ {
+ "code": 4,
+ "text": "Asian"
+ },
+ {
+ "code": 5,
+ "text": "Native Hawaiian or Pacific Islander"
+ },
+ {
+ "code": 7,
+ "text": "Other"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: DK, NA, or RF to first mention (ER85121=9); fewer than four mentions"
+ }
+ ]
+ },
+ "ER85780": {
+ "label": "COMPLETED ED-RP",
+ "page": 1338,
+ "values": [
+ {
+ "code": 0,
+ "text": "Completed no grades of school"
+ },
+ {
+ "range": [
+ 1,
+ 17
+ ],
+ "text": "Actual number"
+ },
+ {
+ "code": 99,
+ "text": "DK; NA"
+ }
+ ]
+ },
+ "ER85781": {
+ "label": "COMPLETED ED-SP",
+ "page": 1339,
+ "values": [
+ {
+ "range": [
+ 1,
+ 17
+ ],
+ "text": "Actual number"
+ },
+ {
+ "code": 99,
+ "text": "DK; NA"
+ },
+ {
+ "code": 0,
+ "text": "Inap.: completed no grades of school; no Spouse/Partner in FU (ER83493=5)"
+ }
+ ]
+ }
+ }
+ }
+ },
+ "generated_by": "scripts/build_group_attribute_codebook_values.py",
+ "note": "Documentation only: labels, value codes and their text from the staged PSID family codebook PDFs. No data value or count is recorded.",
+ "schema_version": "psid_group_attribute_codebook.v1"
+}
diff --git a/data/external/ssa_mint8_payroll_option_row_labels.json b/data/external/ssa_mint8_payroll_option_row_labels.json
new file mode 100644
index 00000000..887adf55
--- /dev/null
+++ b/data/external/ssa_mint8_payroll_option_row_labels.json
@@ -0,0 +1,1851 @@
+{
+ "source_url": "https://www.ssa.gov/policy/docs/projections/policy-options/increase-payroll-tax-rate.html",
+ "retrieved_via": "http://web.archive.org/web/20260519190449id_/https://www.ssa.gov/policy/docs/projections/policy-options/increase-payroll-tax-rate.html",
+ "raw_sha256": "3cfb6eb636cad6d062eb22734c533fb349a0bc4761014eee8f2ccb0eb6e28d8c",
+ "dateCertified": "2026-04-01",
+ "note": "LABELS ONLY. No data cells were extracted. Table numbers are document order on the page.",
+ "tables": {
+ "1": {
+ "caption": "Projected Effects of Proposal on Social Security Benefits in 2030 POPULATION: Current-law beneficiaries aged 60 or older (characteristics)",
+ "columns": [
+ "Characteristic",
+ "Percent of population with a—",
+ "Percent change in Social Security benefits at the—",
+ "Benefit decrease",
+ "Benefit increase",
+ "10th %ile",
+ "Median",
+ "90th %ile"
+ ],
+ "groups": [
+ {
+ "group": "Total",
+ "labels": []
+ },
+ {
+ "group": "Sex",
+ "labels": [
+ "Female",
+ "Male"
+ ]
+ },
+ {
+ "group": "Race and ethnicity",
+ "labels": [
+ "Hispanic or Latino, any race",
+ "White, non-Hispanic",
+ "Black or African American, non-Hispanic",
+ "All other races, non-Hispanic"
+ ]
+ },
+ {
+ "group": "Country of birth",
+ "labels": [
+ "United States",
+ "Other countries"
+ ]
+ },
+ {
+ "group": "Age",
+ "labels": [
+ "60–69",
+ "70–79",
+ "80–89",
+ "90 or older"
+ ]
+ },
+ {
+ "group": "Marital status",
+ "labels": [
+ "Married",
+ "Divorced",
+ "Widowed",
+ "Never married"
+ ]
+ },
+ {
+ "group": "Highest education level",
+ "labels": [
+ "Graduate",
+ "Bachelor",
+ "Associate",
+ "High school",
+ "Less than high school"
+ ]
+ },
+ {
+ "group": "Current-law poverty status",
+ "labels": [
+ "Above poverty",
+ "In poverty"
+ ]
+ },
+ {
+ "group": "Current-law household income quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Current-law benefit type",
+ "labels": [
+ "Retired worker only",
+ "Widow(er) (includes dually entitled)",
+ "Spousal (includes dually entitled)",
+ "Disabled worker only"
+ ]
+ }
+ ]
+ },
+ "2": {
+ "caption": "Projected Effects of Proposal on Social Security Benefits in 2050 POPULATION: Current-law beneficiaries aged 60 or older (characteristics)",
+ "columns": [
+ "Characteristic",
+ "Percent of population with a—",
+ "Percent change in Social Security benefits at the—",
+ "Benefit decrease",
+ "Benefit increase",
+ "10th %ile",
+ "Median",
+ "90th %ile"
+ ],
+ "groups": [
+ {
+ "group": "Total",
+ "labels": []
+ },
+ {
+ "group": "Sex",
+ "labels": [
+ "Female",
+ "Male"
+ ]
+ },
+ {
+ "group": "Race and ethnicity",
+ "labels": [
+ "Hispanic or Latino, any race",
+ "White, non-Hispanic",
+ "Black or African American, non-Hispanic",
+ "All other races, non-Hispanic"
+ ]
+ },
+ {
+ "group": "Country of birth",
+ "labels": [
+ "United States",
+ "Other countries"
+ ]
+ },
+ {
+ "group": "Age",
+ "labels": [
+ "60–69",
+ "70–79",
+ "80–89",
+ "90 or older"
+ ]
+ },
+ {
+ "group": "Marital status",
+ "labels": [
+ "Married",
+ "Divorced",
+ "Widowed",
+ "Never married"
+ ]
+ },
+ {
+ "group": "Highest education level",
+ "labels": [
+ "Graduate",
+ "Bachelor",
+ "Associate",
+ "High school",
+ "Less than high school"
+ ]
+ },
+ {
+ "group": "Current-law poverty status",
+ "labels": [
+ "Above poverty",
+ "In poverty"
+ ]
+ },
+ {
+ "group": "Current-law household income quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Current-law benefit type",
+ "labels": [
+ "Retired worker only",
+ "Widow(er) (includes dually entitled)",
+ "Spousal (includes dually entitled)",
+ "Disabled worker only"
+ ]
+ }
+ ]
+ },
+ "3": {
+ "caption": "Projected Effects of Proposal on Social Security Benefits in 2070 POPULATION: Current-law beneficiaries aged 60 or older (characteristics)",
+ "columns": [
+ "Characteristic",
+ "Percent of population with a—",
+ "Percent change in Social Security benefits at the—",
+ "Benefit decrease",
+ "Benefit increase",
+ "10th %ile",
+ "Median",
+ "90th %ile"
+ ],
+ "groups": [
+ {
+ "group": "Total",
+ "labels": []
+ },
+ {
+ "group": "Sex",
+ "labels": [
+ "Female",
+ "Male"
+ ]
+ },
+ {
+ "group": "Race and ethnicity",
+ "labels": [
+ "Hispanic or Latino, any race",
+ "White, non-Hispanic",
+ "Black or African American, non-Hispanic",
+ "All other races, non-Hispanic"
+ ]
+ },
+ {
+ "group": "Country of birth",
+ "labels": [
+ "United States",
+ "Other countries"
+ ]
+ },
+ {
+ "group": "Age",
+ "labels": [
+ "60–69",
+ "70–79",
+ "80–89",
+ "90 or older"
+ ]
+ },
+ {
+ "group": "Marital status",
+ "labels": [
+ "Married",
+ "Divorced",
+ "Widowed",
+ "Never married"
+ ]
+ },
+ {
+ "group": "Highest education level",
+ "labels": [
+ "Graduate",
+ "Bachelor",
+ "Associate",
+ "High school",
+ "Less than high school"
+ ]
+ },
+ {
+ "group": "Current-law poverty status",
+ "labels": [
+ "Above poverty",
+ "In poverty"
+ ]
+ },
+ {
+ "group": "Current-law household income quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Current-law benefit type",
+ "labels": [
+ "Retired worker only",
+ "Widow(er) (includes dually entitled)",
+ "Spousal (includes dually entitled)",
+ "Disabled worker only"
+ ]
+ }
+ ]
+ },
+ "4": {
+ "caption": "Projected Effects of Proposal on Social Security Taxes Paid in 2030 POPULATION: Current-law payroll taxpayers aged 31 or older (characteristics)",
+ "columns": [
+ "Characteristic",
+ "Percent of population with a—",
+ "Percent change in Social Security taxes paid at the—",
+ "Change in taxes paid (in 2024$) at the—",
+ "Tax decrease",
+ "Tax increase",
+ "10th %ile",
+ "Median",
+ "90th %ile",
+ "10th %ile",
+ "Median",
+ "90th %ile"
+ ],
+ "groups": [
+ {
+ "group": "Total",
+ "labels": []
+ },
+ {
+ "group": "Sex",
+ "labels": [
+ "Female",
+ "Male"
+ ]
+ },
+ {
+ "group": "Race and ethnicity",
+ "labels": [
+ "Hispanic or Latino, any race",
+ "White, non-Hispanic",
+ "Black or African American, non-Hispanic",
+ "All other races, non-Hispanic"
+ ]
+ },
+ {
+ "group": "Country of birth",
+ "labels": [
+ "United States",
+ "Other countries"
+ ]
+ },
+ {
+ "group": "Age",
+ "labels": [
+ "31–39",
+ "40–49",
+ "50–59",
+ "60–69",
+ "70 or older"
+ ]
+ },
+ {
+ "group": "Marital status",
+ "labels": [
+ "Married",
+ "Divorced",
+ "Widowed",
+ "Never married"
+ ]
+ },
+ {
+ "group": "Highest education level",
+ "labels": [
+ "Graduate",
+ "Bachelor",
+ "Associate",
+ "High school",
+ "Less than high school"
+ ]
+ },
+ {
+ "group": "Current-law household income quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Current-law payroll taxes quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ }
+ ]
+ },
+ "5": {
+ "caption": "Projected Effects of Proposal on Social Security Taxes Paid in 2050 POPULATION: Current-law payroll taxpayers aged 31 or older (characteristics)",
+ "columns": [
+ "Characteristic",
+ "Percent of population with a—",
+ "Percent change in Social Security taxes paid at the—",
+ "Change in taxes paid (in 2024$) at the—",
+ "Tax decrease",
+ "Tax increase",
+ "10th %ile",
+ "Median",
+ "90th %ile",
+ "10th %ile",
+ "Median",
+ "90th %ile"
+ ],
+ "groups": [
+ {
+ "group": "Total",
+ "labels": []
+ },
+ {
+ "group": "Sex",
+ "labels": [
+ "Female",
+ "Male"
+ ]
+ },
+ {
+ "group": "Race and ethnicity",
+ "labels": [
+ "Hispanic or Latino, any race",
+ "White, non-Hispanic",
+ "Black or African American, non-Hispanic",
+ "All other races, non-Hispanic"
+ ]
+ },
+ {
+ "group": "Country of birth",
+ "labels": [
+ "United States",
+ "Other countries"
+ ]
+ },
+ {
+ "group": "Age",
+ "labels": [
+ "31–39",
+ "40–49",
+ "50–59",
+ "60–69",
+ "70 or older"
+ ]
+ },
+ {
+ "group": "Marital status",
+ "labels": [
+ "Married",
+ "Divorced",
+ "Widowed",
+ "Never married"
+ ]
+ },
+ {
+ "group": "Highest education level",
+ "labels": [
+ "Graduate",
+ "Bachelor",
+ "Associate",
+ "High school",
+ "Less than high school"
+ ]
+ },
+ {
+ "group": "Current-law household income quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Current-law payroll taxes quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ }
+ ]
+ },
+ "6": {
+ "caption": "Projected Effects of Proposal on Social Security Taxes Paid in 2070 POPULATION: Current-law payroll taxpayers aged 31 or older (characteristics)",
+ "columns": [
+ "Characteristic",
+ "Percent of population with a—",
+ "Percent change in Social Security taxes paid at the—",
+ "Change in taxes paid (in 2024$) at the—",
+ "Tax decrease",
+ "Tax increase",
+ "10th %ile",
+ "Median",
+ "90th %ile",
+ "10th %ile",
+ "Median",
+ "90th %ile"
+ ],
+ "groups": [
+ {
+ "group": "Total",
+ "labels": []
+ },
+ {
+ "group": "Sex",
+ "labels": [
+ "Female",
+ "Male"
+ ]
+ },
+ {
+ "group": "Race and ethnicity",
+ "labels": [
+ "Hispanic or Latino, any race",
+ "White, non-Hispanic",
+ "Black or African American, non-Hispanic",
+ "All other races, non-Hispanic"
+ ]
+ },
+ {
+ "group": "Country of birth",
+ "labels": [
+ "United States",
+ "Other countries"
+ ]
+ },
+ {
+ "group": "Age",
+ "labels": [
+ "31–39",
+ "40–49",
+ "50–59",
+ "60–69",
+ "70 or older"
+ ]
+ },
+ {
+ "group": "Marital status",
+ "labels": [
+ "Married",
+ "Divorced",
+ "Widowed",
+ "Never married"
+ ]
+ },
+ {
+ "group": "Highest education level",
+ "labels": [
+ "Graduate",
+ "Bachelor",
+ "Associate",
+ "High school",
+ "Less than high school"
+ ]
+ },
+ {
+ "group": "Current-law household income quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Current-law payroll taxes quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ }
+ ]
+ },
+ "7": {
+ "caption": "Projected Effects of Proposal on Household Income in 2030 POPULATION: Current-law beneficiaries aged 60 or older (characteristics)",
+ "columns": [
+ "Characteristic",
+ "Percent of population with an—",
+ "Percent change in household income at the—",
+ "Income decrease",
+ "Income increase",
+ "10th %ile",
+ "Median",
+ "90th %ile"
+ ],
+ "groups": [
+ {
+ "group": "Total",
+ "labels": []
+ },
+ {
+ "group": "Sex",
+ "labels": [
+ "Female",
+ "Male"
+ ]
+ },
+ {
+ "group": "Race and ethnicity",
+ "labels": [
+ "Hispanic or Latino, any race",
+ "White, non-Hispanic",
+ "Black or African American, non-Hispanic",
+ "All other races, non-Hispanic"
+ ]
+ },
+ {
+ "group": "Country of birth",
+ "labels": [
+ "United States",
+ "Other countries"
+ ]
+ },
+ {
+ "group": "Age",
+ "labels": [
+ "60–69",
+ "70–79",
+ "80–89",
+ "90 or older"
+ ]
+ },
+ {
+ "group": "Marital status",
+ "labels": [
+ "Married",
+ "Divorced",
+ "Widowed",
+ "Never married"
+ ]
+ },
+ {
+ "group": "Highest education level",
+ "labels": [
+ "Graduate",
+ "Bachelor",
+ "Associate",
+ "High school",
+ "Less than high school"
+ ]
+ },
+ {
+ "group": "Current-law poverty status",
+ "labels": [
+ "Above poverty",
+ "In poverty"
+ ]
+ },
+ {
+ "group": "Current-law household income quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Current-law benefit type",
+ "labels": [
+ "Retired worker only",
+ "Widow(er) (includes dually entitled)",
+ "Spousal (includes dually entitled)",
+ "Disabled worker only"
+ ]
+ }
+ ]
+ },
+ "8": {
+ "caption": "Projected Effects of Proposal on Household Income in 2050 POPULATION: Current-law beneficiaries aged 60 or older (characteristics)",
+ "columns": [
+ "Characteristic",
+ "Percent of population with an—",
+ "Percent change in household income at the—",
+ "Income decrease",
+ "Income increase",
+ "10th %ile",
+ "Median",
+ "90th %ile"
+ ],
+ "groups": [
+ {
+ "group": "Total",
+ "labels": []
+ },
+ {
+ "group": "Sex",
+ "labels": [
+ "Female",
+ "Male"
+ ]
+ },
+ {
+ "group": "Race and ethnicity",
+ "labels": [
+ "Hispanic or Latino, any race",
+ "White, non-Hispanic",
+ "Black or African American, non-Hispanic",
+ "All other races, non-Hispanic"
+ ]
+ },
+ {
+ "group": "Country of birth",
+ "labels": [
+ "United States",
+ "Other countries"
+ ]
+ },
+ {
+ "group": "Age",
+ "labels": [
+ "60–69",
+ "70–79",
+ "80–89",
+ "90 or older"
+ ]
+ },
+ {
+ "group": "Marital status",
+ "labels": [
+ "Married",
+ "Divorced",
+ "Widowed",
+ "Never married"
+ ]
+ },
+ {
+ "group": "Highest education level",
+ "labels": [
+ "Graduate",
+ "Bachelor",
+ "Associate",
+ "High school",
+ "Less than high school"
+ ]
+ },
+ {
+ "group": "Current-law poverty status",
+ "labels": [
+ "Above poverty",
+ "In poverty"
+ ]
+ },
+ {
+ "group": "Current-law household income quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Current-law benefit type",
+ "labels": [
+ "Retired worker only",
+ "Widow(er) (includes dually entitled)",
+ "Spousal (includes dually entitled)",
+ "Disabled worker only"
+ ]
+ }
+ ]
+ },
+ "9": {
+ "caption": "Projected Effects of Proposal on Household Income in 2070 POPULATION: Current-law beneficiaries aged 60 or older (characteristics)",
+ "columns": [
+ "Characteristic",
+ "Percent of population with an—",
+ "Percent change in household income at the—",
+ "Income decrease",
+ "Income increase",
+ "10th %ile",
+ "Median",
+ "90th %ile"
+ ],
+ "groups": [
+ {
+ "group": "Total",
+ "labels": []
+ },
+ {
+ "group": "Sex",
+ "labels": [
+ "Female",
+ "Male"
+ ]
+ },
+ {
+ "group": "Race and ethnicity",
+ "labels": [
+ "Hispanic or Latino, any race",
+ "White, non-Hispanic",
+ "Black or African American, non-Hispanic",
+ "All other races, non-Hispanic"
+ ]
+ },
+ {
+ "group": "Country of birth",
+ "labels": [
+ "United States",
+ "Other countries"
+ ]
+ },
+ {
+ "group": "Age",
+ "labels": [
+ "60–69",
+ "70–79",
+ "80–89",
+ "90 or older"
+ ]
+ },
+ {
+ "group": "Marital status",
+ "labels": [
+ "Married",
+ "Divorced",
+ "Widowed",
+ "Never married"
+ ]
+ },
+ {
+ "group": "Highest education level",
+ "labels": [
+ "Graduate",
+ "Bachelor",
+ "Associate",
+ "High school",
+ "Less than high school"
+ ]
+ },
+ {
+ "group": "Current-law poverty status",
+ "labels": [
+ "Above poverty",
+ "In poverty"
+ ]
+ },
+ {
+ "group": "Current-law household income quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Current-law benefit type",
+ "labels": [
+ "Retired worker only",
+ "Widow(er) (includes dually entitled)",
+ "Spousal (includes dually entitled)",
+ "Disabled worker only"
+ ]
+ }
+ ]
+ },
+ "10": {
+ "caption": "Projected Effects of Proposal on Official Poverty Measure in 2030 POPULATION: Current-law beneficiaries aged 60 or older (characteristics)",
+ "columns": [
+ "Characteristic",
+ "Official poverty rate",
+ "Number of population in poverty (in thousands)",
+ "Percent change in the number in poverty",
+ "Under current law",
+ "With proposal",
+ "Under current law",
+ "With proposal",
+ "Change"
+ ],
+ "groups": [
+ {
+ "group": "Total",
+ "labels": []
+ },
+ {
+ "group": "Sex",
+ "labels": [
+ "Female",
+ "Male"
+ ]
+ },
+ {
+ "group": "Race and ethnicity",
+ "labels": [
+ "Hispanic or Latino, any race",
+ "White, non-Hispanic",
+ "Black or African American, non-Hispanic",
+ "All other races, non-Hispanic"
+ ]
+ },
+ {
+ "group": "Country of birth",
+ "labels": [
+ "United States",
+ "Other countries"
+ ]
+ },
+ {
+ "group": "Age",
+ "labels": [
+ "60–69",
+ "70–79",
+ "80–89",
+ "90 or older"
+ ]
+ },
+ {
+ "group": "Marital status",
+ "labels": [
+ "Married",
+ "Divorced",
+ "Widowed",
+ "Never married"
+ ]
+ },
+ {
+ "group": "Highest education level",
+ "labels": [
+ "Graduate",
+ "Bachelor",
+ "Associate",
+ "High school",
+ "Less than high school"
+ ]
+ },
+ {
+ "group": "Current-law poverty status",
+ "labels": [
+ "Above poverty",
+ "In poverty"
+ ]
+ },
+ {
+ "group": "Current-law benefit type",
+ "labels": [
+ "Retired worker only",
+ "Widow(er) (includes dually entitled)",
+ "Spousal (includes dually entitled)",
+ "Disabled worker only"
+ ]
+ }
+ ]
+ },
+ "11": {
+ "caption": "Projected Effects of Proposal on Official Poverty Measure in 2050 POPULATION: Current-law beneficiaries aged 60 or older (characteristics)",
+ "columns": [
+ "Characteristic",
+ "Official poverty rate",
+ "Number of population in poverty (in thousands)",
+ "Percent change in the number in poverty",
+ "Under current law",
+ "With proposal",
+ "Under current law",
+ "With proposal",
+ "Change"
+ ],
+ "groups": [
+ {
+ "group": "Total",
+ "labels": []
+ },
+ {
+ "group": "Sex",
+ "labels": [
+ "Female",
+ "Male"
+ ]
+ },
+ {
+ "group": "Race and ethnicity",
+ "labels": [
+ "Hispanic or Latino, any race",
+ "White, non-Hispanic",
+ "Black or African American, non-Hispanic",
+ "All other races, non-Hispanic"
+ ]
+ },
+ {
+ "group": "Country of birth",
+ "labels": [
+ "United States",
+ "Other countries"
+ ]
+ },
+ {
+ "group": "Age",
+ "labels": [
+ "60–69",
+ "70–79",
+ "80–89",
+ "90 or older"
+ ]
+ },
+ {
+ "group": "Marital status",
+ "labels": [
+ "Married",
+ "Divorced",
+ "Widowed",
+ "Never married"
+ ]
+ },
+ {
+ "group": "Highest education level",
+ "labels": [
+ "Graduate",
+ "Bachelor",
+ "Associate",
+ "High school",
+ "Less than high school"
+ ]
+ },
+ {
+ "group": "Current-law poverty status",
+ "labels": [
+ "Above poverty",
+ "In poverty"
+ ]
+ },
+ {
+ "group": "Current-law benefit type",
+ "labels": [
+ "Retired worker only",
+ "Widow(er) (includes dually entitled)",
+ "Spousal (includes dually entitled)",
+ "Disabled worker only"
+ ]
+ }
+ ]
+ },
+ "12": {
+ "caption": "Projected Effects of Proposal on Official Poverty Measure in 2070 POPULATION: Current-law beneficiaries aged 60 or older (characteristics)",
+ "columns": [
+ "Characteristic",
+ "Official poverty rate",
+ "Number of population in poverty (in thousands)",
+ "Percent change in the number in poverty",
+ "Under current law",
+ "With proposal",
+ "Under current law",
+ "With proposal",
+ "Change"
+ ],
+ "groups": [
+ {
+ "group": "Total",
+ "labels": []
+ },
+ {
+ "group": "Sex",
+ "labels": [
+ "Female",
+ "Male"
+ ]
+ },
+ {
+ "group": "Race and ethnicity",
+ "labels": [
+ "Hispanic or Latino, any race",
+ "White, non-Hispanic",
+ "Black or African American, non-Hispanic",
+ "All other races, non-Hispanic"
+ ]
+ },
+ {
+ "group": "Country of birth",
+ "labels": [
+ "United States",
+ "Other countries"
+ ]
+ },
+ {
+ "group": "Age",
+ "labels": [
+ "60–69",
+ "70–79",
+ "80–89",
+ "90 or older"
+ ]
+ },
+ {
+ "group": "Marital status",
+ "labels": [
+ "Married",
+ "Divorced",
+ "Widowed",
+ "Never married"
+ ]
+ },
+ {
+ "group": "Highest education level",
+ "labels": [
+ "Graduate",
+ "Bachelor",
+ "Associate",
+ "High school",
+ "Less than high school"
+ ]
+ },
+ {
+ "group": "Current-law poverty status",
+ "labels": [
+ "Above poverty",
+ "In poverty"
+ ]
+ },
+ {
+ "group": "Current-law benefit type",
+ "labels": [
+ "Retired worker only",
+ "Widow(er) (includes dually entitled)",
+ "Spousal (includes dually entitled)",
+ "Disabled worker only"
+ ]
+ }
+ ]
+ },
+ "13": {
+ "caption": "Projected Effects of Proposal on Benefit/Tax Ratios POPULATION: Workers born 1960–1969 with a benefit/tax ratio (characteristics)",
+ "columns": [
+ "Characteristic",
+ "Percent of population with a—",
+ "Percent change in benefit/tax ratio at the—",
+ "Benefit/tax ratio under current law at the—",
+ "Benefit/tax ratio with proposal at the—",
+ "Ratio decrease",
+ "Ratio increase",
+ "10th %ile",
+ "Median",
+ "90th %ile",
+ "10th %ile",
+ "Median",
+ "90th %ile",
+ "10th %ile",
+ "Median",
+ "90th %ile"
+ ],
+ "groups": [
+ {
+ "group": "Total",
+ "labels": []
+ },
+ {
+ "group": "Sex",
+ "labels": [
+ "Female",
+ "Male"
+ ]
+ },
+ {
+ "group": "Race and ethnicity",
+ "labels": [
+ "Hispanic or Latino, any race",
+ "White, non-Hispanic",
+ "Black or African American, non-Hispanic",
+ "All other races, non-Hispanic"
+ ]
+ },
+ {
+ "group": "Country of birth",
+ "labels": [
+ "United States",
+ "Other countries"
+ ]
+ },
+ {
+ "group": "Highest education level",
+ "labels": [
+ "Graduate",
+ "Bachelor",
+ "Associate",
+ "High school",
+ "Less than high school"
+ ]
+ },
+ {
+ "group": "Current-law initial AIME quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Lifetime payroll tax quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Lifetime payroll tax quintile (shared)",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ }
+ ]
+ },
+ "14": {
+ "caption": "Projected Effects of Proposal on Benefit/Tax Ratios POPULATION: Workers born 1980–1989 with a benefit/tax ratio (characteristics)",
+ "columns": [
+ "Characteristic",
+ "Percent of population with a—",
+ "Percent change in benefit/tax ratio at the—",
+ "Benefit/tax ratio under current law at the—",
+ "Benefit/tax ratio with proposal at the—",
+ "Ratio decrease",
+ "Ratio increase",
+ "10th %ile",
+ "Median",
+ "90th %ile",
+ "10th %ile",
+ "Median",
+ "90th %ile",
+ "10th %ile",
+ "Median",
+ "90th %ile"
+ ],
+ "groups": [
+ {
+ "group": "Total",
+ "labels": []
+ },
+ {
+ "group": "Sex",
+ "labels": [
+ "Female",
+ "Male"
+ ]
+ },
+ {
+ "group": "Race and ethnicity",
+ "labels": [
+ "Hispanic or Latino, any race",
+ "White, non-Hispanic",
+ "Black or African American, non-Hispanic",
+ "All other races, non-Hispanic"
+ ]
+ },
+ {
+ "group": "Country of birth",
+ "labels": [
+ "United States",
+ "Other countries"
+ ]
+ },
+ {
+ "group": "Highest education level",
+ "labels": [
+ "Graduate",
+ "Bachelor",
+ "Associate",
+ "High school",
+ "Less than high school"
+ ]
+ },
+ {
+ "group": "Current-law initial AIME quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Lifetime payroll tax quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Lifetime payroll tax quintile (shared)",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ }
+ ]
+ },
+ "15": {
+ "caption": "Projected Effects of Proposal on Benefit/Tax Ratios POPULATION: Workers born 2000–2009 with a benefit/tax ratio (characteristics)",
+ "columns": [
+ "Characteristic",
+ "Percent of population with a—",
+ "Percent change in benefit/tax ratio at the—",
+ "Benefit/tax ratio under current law at the—",
+ "Benefit/tax ratio with proposal at the—",
+ "Ratio decrease",
+ "Ratio increase",
+ "10th %ile",
+ "Median",
+ "90th %ile",
+ "10th %ile",
+ "Median",
+ "90th %ile",
+ "10th %ile",
+ "Median",
+ "90th %ile"
+ ],
+ "groups": [
+ {
+ "group": "Total",
+ "labels": []
+ },
+ {
+ "group": "Sex",
+ "labels": [
+ "Female",
+ "Male"
+ ]
+ },
+ {
+ "group": "Race and ethnicity",
+ "labels": [
+ "Hispanic or Latino, any race",
+ "White, non-Hispanic",
+ "Black or African American, non-Hispanic",
+ "All other races, non-Hispanic"
+ ]
+ },
+ {
+ "group": "Country of birth",
+ "labels": [
+ "United States",
+ "Other countries"
+ ]
+ },
+ {
+ "group": "Highest education level",
+ "labels": [
+ "Graduate",
+ "Bachelor",
+ "Associate",
+ "High school",
+ "Less than high school"
+ ]
+ },
+ {
+ "group": "Current-law initial AIME quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Lifetime payroll tax quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Lifetime payroll tax quintile (shared)",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ }
+ ]
+ },
+ "16": {
+ "caption": "Projected Effects of Proposal on Benefit/Tax Ratios POPULATION: Workers born 2020–2029 with a benefit/tax ratio (characteristics)",
+ "columns": [
+ "Characteristic",
+ "Percent of population with a—",
+ "Percent change in benefit/tax ratio at the—",
+ "Benefit/tax ratio under current law at the—",
+ "Benefit/tax ratio with proposal at the—",
+ "Ratio decrease",
+ "Ratio increase",
+ "10th %ile",
+ "Median",
+ "90th %ile",
+ "10th %ile",
+ "Median",
+ "90th %ile",
+ "10th %ile",
+ "Median",
+ "90th %ile"
+ ],
+ "groups": [
+ {
+ "group": "Total",
+ "labels": []
+ },
+ {
+ "group": "Sex",
+ "labels": [
+ "Female",
+ "Male"
+ ]
+ },
+ {
+ "group": "Race and ethnicity",
+ "labels": [
+ "Hispanic or Latino, any race",
+ "White, non-Hispanic",
+ "Black or African American, non-Hispanic",
+ "All other races, non-Hispanic"
+ ]
+ },
+ {
+ "group": "Country of birth",
+ "labels": [
+ "United States",
+ "Other countries"
+ ]
+ },
+ {
+ "group": "Highest education level",
+ "labels": [
+ "Graduate",
+ "Bachelor",
+ "Associate",
+ "High school",
+ "Less than high school"
+ ]
+ },
+ {
+ "group": "Current-law initial AIME quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Lifetime payroll tax quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Lifetime payroll tax quintile (shared)",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ }
+ ]
+ },
+ "17": {
+ "caption": "Projected Effects of Proposal on Initial Replacement Rates POPULATION: Current-law beneficiaries born 1960–1969 with a replacement rate (characteristics)",
+ "columns": [
+ "Characteristic",
+ "Percent of population with a—",
+ "Percent change in initial replacement rate at the—",
+ "Initial replacement rate under current law at the—",
+ "Initial replacement rate with proposal at the—",
+ "Rate decrease",
+ "Rate increase",
+ "10th %ile",
+ "Median",
+ "90th %ile",
+ "10th %ile",
+ "Median",
+ "90th %ile",
+ "10th %ile",
+ "Median",
+ "90th %ile"
+ ],
+ "groups": [
+ {
+ "group": "Total",
+ "labels": []
+ },
+ {
+ "group": "Sex",
+ "labels": [
+ "Female",
+ "Male"
+ ]
+ },
+ {
+ "group": "Race and ethnicity",
+ "labels": [
+ "Hispanic or Latino, any race",
+ "White, non-Hispanic",
+ "Black or African American, non-Hispanic",
+ "All other races, non-Hispanic"
+ ]
+ },
+ {
+ "group": "Country of birth",
+ "labels": [
+ "United States",
+ "Other countries"
+ ]
+ },
+ {
+ "group": "Highest education level",
+ "labels": [
+ "Graduate",
+ "Bachelor",
+ "Associate",
+ "High school",
+ "Less than high school"
+ ]
+ },
+ {
+ "group": "Current-law initial AIME quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Lifetime payroll tax quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Lifetime payroll tax quintile (shared)",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ }
+ ]
+ },
+ "18": {
+ "caption": "Projected Effects of Proposal on Initial Replacement Rates POPULATION: Current-law beneficiaries born 1980–1989 with a replacement rate (characteristics)",
+ "columns": [
+ "Characteristic",
+ "Percent of population with a—",
+ "Percent change in initial replacement rate at the—",
+ "Initial replacement rate under current law at the—",
+ "Initial replacement rate with proposal at the—",
+ "Rate decrease",
+ "Rate increase",
+ "10th %ile",
+ "Median",
+ "90th %ile",
+ "10th %ile",
+ "Median",
+ "90th %ile",
+ "10th %ile",
+ "Median",
+ "90th %ile"
+ ],
+ "groups": [
+ {
+ "group": "Total",
+ "labels": []
+ },
+ {
+ "group": "Sex",
+ "labels": [
+ "Female",
+ "Male"
+ ]
+ },
+ {
+ "group": "Race and ethnicity",
+ "labels": [
+ "Hispanic or Latino, any race",
+ "White, non-Hispanic",
+ "Black or African American, non-Hispanic",
+ "All other races, non-Hispanic"
+ ]
+ },
+ {
+ "group": "Country of birth",
+ "labels": [
+ "United States",
+ "Other countries"
+ ]
+ },
+ {
+ "group": "Highest education level",
+ "labels": [
+ "Graduate",
+ "Bachelor",
+ "Associate",
+ "High school",
+ "Less than high school"
+ ]
+ },
+ {
+ "group": "Current-law initial AIME quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Lifetime payroll tax quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Lifetime payroll tax quintile (shared)",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ }
+ ]
+ },
+ "19": {
+ "caption": "Projected Effects of Proposal on Initial Replacement Rates POPULATION: Current-law beneficiaries born 2000–2009 with a replacement rate (characteristics)",
+ "columns": [
+ "Characteristic",
+ "Percent of population with a—",
+ "Percent change in initial replacement rate at the—",
+ "Initial replacement rate under current law at the—",
+ "Initial replacement rate with proposal at the—",
+ "Rate decrease",
+ "Rate increase",
+ "10th %ile",
+ "Median",
+ "90th %ile",
+ "10th %ile",
+ "Median",
+ "90th %ile",
+ "10th %ile",
+ "Median",
+ "90th %ile"
+ ],
+ "groups": [
+ {
+ "group": "Total",
+ "labels": []
+ },
+ {
+ "group": "Sex",
+ "labels": [
+ "Female",
+ "Male"
+ ]
+ },
+ {
+ "group": "Race and ethnicity",
+ "labels": [
+ "Hispanic or Latino, any race",
+ "White, non-Hispanic",
+ "Black or African American, non-Hispanic",
+ "All other races, non-Hispanic"
+ ]
+ },
+ {
+ "group": "Country of birth",
+ "labels": [
+ "United States",
+ "Other countries"
+ ]
+ },
+ {
+ "group": "Highest education level",
+ "labels": [
+ "Graduate",
+ "Bachelor",
+ "Associate",
+ "High school",
+ "Less than high school"
+ ]
+ },
+ {
+ "group": "Current-law initial AIME quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Lifetime payroll tax quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Lifetime payroll tax quintile (shared)",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ }
+ ]
+ },
+ "20": {
+ "caption": "Projected Effects of Proposal on Initial Replacement Rates POPULATION: Current-law beneficiaries born 2020–2029 with a replacement rate (characteristics)",
+ "columns": [
+ "Characteristic",
+ "Percent of population with a—",
+ "Percent change in initial replacement rate at the—",
+ "Initial replacement rate under current law at the—",
+ "Initial replacement rate with proposal at the—",
+ "Rate decrease",
+ "Rate increase",
+ "10th %ile",
+ "Median",
+ "90th %ile",
+ "10th %ile",
+ "Median",
+ "90th %ile",
+ "10th %ile",
+ "Median",
+ "90th %ile"
+ ],
+ "groups": [
+ {
+ "group": "Total",
+ "labels": []
+ },
+ {
+ "group": "Sex",
+ "labels": [
+ "Female",
+ "Male"
+ ]
+ },
+ {
+ "group": "Race and ethnicity",
+ "labels": [
+ "Hispanic or Latino, any race",
+ "White, non-Hispanic",
+ "Black or African American, non-Hispanic",
+ "All other races, non-Hispanic"
+ ]
+ },
+ {
+ "group": "Country of birth",
+ "labels": [
+ "United States",
+ "Other countries"
+ ]
+ },
+ {
+ "group": "Highest education level",
+ "labels": [
+ "Graduate",
+ "Bachelor",
+ "Associate",
+ "High school",
+ "Less than high school"
+ ]
+ },
+ {
+ "group": "Current-law initial AIME quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Lifetime payroll tax quintile",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ },
+ {
+ "group": "Lifetime payroll tax quintile (shared)",
+ "labels": [
+ "Highest",
+ "Second highest",
+ "Middle",
+ "Second lowest",
+ "Lowest"
+ ]
+ }
+ ]
+ }
+ }
+}
\ No newline at end of file
diff --git a/data/external/ssa_mint8_payroll_option_row_labels.source.provenance.json b/data/external/ssa_mint8_payroll_option_row_labels.source.provenance.json
new file mode 100644
index 00000000..3c109d61
--- /dev/null
+++ b/data/external/ssa_mint8_payroll_option_row_labels.source.provenance.json
@@ -0,0 +1,11 @@
+{
+ "schema_version": "external_source_provenance.v1",
+ "committed_source_file": "data/external/ssa_mint8_payroll_option_row_labels.json",
+ "source_url": "https://www.ssa.gov/policy/docs/projections/policy-options/increase-payroll-tax-rate.html",
+ "document": "SSA MINT 8 policy-option tables, 'Increase payroll tax rate' — row and column labels only",
+ "fetch_method": "Built by the MINT-categories lane of orchestrating session 95606380 from the Internet Archive capture http://web.archive.org/web/20260519190449id_/https://www.ssa.gov/policy/docs/projections/policy-options/increase-payroll-tax-rate.html (raw page SHA-256 3cfb6eb636cad6d062eb22734c533fb349a0bc4761014eee8f2ccb0eb6e28d8c, recorded inside the file) with a parser that emitted table captions, header cells and row labels and never an ordinary data cell. Copied byte-identically from the lane's label file.",
+ "source_sha256": "23fbfbc8dbc14b06144a83bf0583ea0a6c89505638f3f7dc5fd6b34a3a4ac650",
+ "acquired_by": "MINT-categories lane (workflow wf_35b496b0-085), 2026-10-01; committed by the G1 group-attributes builder",
+ "consumed_by": "data/external/group_category_schemes_v1.json (scheme mint8 attribute and lifetime-quintile row labels; annual and cohort table row-group orders)",
+ "note": "LABELS ONLY. No MINT data cell is in this file. The restricted-files list allows one payroll-tax option table to be read for its row labels only; this is that table's label record."
+}
diff --git a/data/external/ssa_mint8_table_user_guide.source.html b/data/external/ssa_mint8_table_user_guide.source.html
new file mode 100644
index 00000000..889f4d81
--- /dev/null
+++ b/data/external/ssa_mint8_table_user_guide.source.html
@@ -0,0 +1,633 @@
+
+
+
+
+
+
+Table User Guide - Modeling Income in the Near Term (MINT) 8
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+ An official website of the United States government
+ Here's how you know
+
+
+
+
Official websites use .gov A .gov website belongs to an official government organization in the United States.
+
+
Secure .gov websites use HTTPS A lock (
+
+ ) or https:// means you’ve safely connected to the .gov website. Share sensitive information only on official, secure websites.
+
The first four sets of tables provide results for the analysis years, 2030, 2050, and 2070, while the last two provide results for four birth cohorts: 1960–1969,1980–1989,2000–2009, and 2020–2029.
+
Each set of tables and what they show are discussed below.
The first two columns show the percent of the population with a benefit decrease or increase, and the next three columns show the percent change in individual Social Security benefits at three percentiles.
The first five columns are the same as the Social Security benefits tables except tax changes are shown rather than benefit changes. Columns 6–8 show the dollar amount changes in individual Social Security payroll taxes at the same percentiles as the percentage change columns.
+
Household Income
+
These tables have the same structure and population as the Social Security benefits tables except the effects on household income are shown rather than individual Social Security benefits.
The first two columns show the poverty rate with and without (current law) the proposed option. The next three columns show the number in poverty (expressed in thousands) with and without (current law) the proposed option and the difference between them. The final column shows the percent change in the number in poverty with the proposed change. This percent change is calculated by dividing the change in thousands in poverty (5th column) by the thousands in poverty without the proposal (3rd column).
+
CORRECT interpretation: the number of people in poverty would decline by 2 percent. INCORRECT interpretation: the poverty rate would decline by 2 percent.
+
Benefit/Tax Ratios
+
These tables show the projected changes in the benefit/tax ratios for workers born in four birth cohorts: 1960–1969,1980–1989,2000–2009, and 2020–2029 who have a tax record from which to calculate a benefit/tax ratio. We excluded those who paid zero payroll taxes over their lifetime and, therefore, could not have a benefit/tax ratio calculated.
+
The first five columns are similar to the Social Security benefits and taxes paid tables except the percent of the population with a benefit/tax ratio decrease and increase, and the percent change in the benefit/tax ratio at three percentiles are shown.
+
The last six columns show the distribution of benefit/tax ratios with and without (current law) the proposed option at three percentiles.
+
The benefit/tax ratio is a money's worth measure that is the lifetime present value of benefits divided by the lifetime present value of payroll taxes. The ratio represents how much in benefits an individual received for every dollar of payroll taxes paid.
+
How to interpret the ratios:
+
+
0% means that an individual received no benefits despite paying payroll taxes (or $0.00 in benefits for every $1.00 of taxes).
+
50% means that an individual received half as much benefits as was paid in taxes (or $0.50 in benefits for every $1.00 of taxes).
+
100% means that an individual received the same amount of benefits as was paid in taxes (or $1.00 in benefits for every $1.00 of taxes).
+
1,350% means that an individual received over 13 times the amount of benefits as was paid in taxes (or $13.50 in benefits for every $1.00 of taxes).
+
+
The present value of benefits includes all Social Security benefits the individual received, regardless of earnings record or type of benefit. The present value of payroll taxes includes all the payroll taxes that the individual paid over a lifetime. We use the Social Security Trust Fund interest rate to adjust benefits and taxes to their present values at age 62.
+
Initial Replacement Rates
+
These tables are the same as the benefit/tax ratio tables except initial replacement rates are shown for current-law beneficiaries born in four birth cohorts: 1960–1969,1980–1989,2000–2009, and 2020–2029. Only beneficiaries with both income and benefit records from which to calculate an initial replacement rate are included in the population. Beneficiaries with zero average indexed monthly earnings (AIME) or zero benefit at claiming (due to the earnings test or other fixed-dollar reductions) are excluded because no replacement took place. When no benefit is received, the initial replacement of earnings may take place in a later year or never, if the beneficiary dies before a benefit is paid.
+
The initial replacement rate represents how much of the initial AIME is replaced by the initial total monthly Social Security benefit. It is calculated by dividing the initial monthly benefit by the initial AIME.
+
How to interpret the initial replacement rate values:
+
+
20% means that the initial benefit replaced one-fifth of lifetime earnings (or $0.20 in monthly benefits for every $1.00 of AIME).
+
100% means that the initial benefit replaced all lifetime earnings (or $1.00 in monthly benefits for every $1.00 of AIME).
+
150% means that the initial benefit replaced one and a half times lifetime earnings (or $1.50 in monthly benefits for every $1.00 of AIME).
+
+
The initial monthly benefit is the total individual Social Security benefit received at the person's claiming age, including any spousal, survivor, or disability benefits received. We calculate the replacement rate at claiming age, regardless of what type of benefits the beneficiary claimed.
+
Interpreting the Profile of Beneficiaries by Race & Ethnicity Tables
The tables also include the same information for four race and ethnicity groupings:
+
+
Hispanic or Latino, any race
+
White, non-Hispanic
+
Black or African American, non-Hispanic
+
All other races, non-Hispanic
+
+
Each set of tables and what they show are discussed below.
+
Social Security Benefits
+
These tables show the projected distribution of individual monthly Social Security benefits at three percentiles. The benefit amount is the total monthly benefit an individual would receive, regardless of the type of benefit or earnings record it came from.
+
Poverty Rates and Numbers
+
These tables show the poverty rate and number in poverty under the Official Poverty Measure and the Supplemental Poverty Measure.
+
The first and third columns show the rates and numbers under the Official Poverty Measure while the second and fourth columns show the same information under the Supplemental Poverty Measure. Further information about the Official and Supplemental Poverty Measures and how they relate to the aged, is available in this paper.
The last five columns show the projected mean share of household income from five sources:
+
+
Social Security benefits, which include benefits for the individual, spouse, and any children.
+
Annuitized asset income, which includes income from defined contribution plans (such as 401(k) accounts) and personal savings.
+
Defined benefit pension income, which includes the individual's and any spouse's defined benefit pension income.
+
All earnings, including covered earnings (from which Social Security taxes are withheld) and non-covered earnings (no Social Security taxes withheld) of the individual and his or her spouse.
+
Coresident income, which is the income of non-spousal coresidents in the household.
+
+
Rows may not sum to 100 percent because minor sources of income are excluded.
+
Total Earnings
+
These tables show the projected distribution of annual individual total earnings at three percentiles. Total earnings includes both covered and non-covered wages.
+
Household Wealth
+
These tables show the projected distribution of household wealth at three percentiles. Wealth includes retirement account balances and savings, but excludes household equity.
+
Health Status and Costs
+
These tables show projected health status and expenses. The first two columns cover the percent and number (in thousands) of beneficiaries who are projected to self-report fair or poor health. The third column shows family median annual health insurance premiums, and the fourth column shows family median annual out-of-pocket health expenses.
+
Interpreting the Profile of Taxpayers by Race & Ethnicity Tables
These tables show the projected distribution of annual individual Social Security taxes paid at three percentiles. Both the employer and employee shares of the Social Security portion of FICA (Federal Insurance Contributions Act) taxes are included.
+
Covered Earnings
+
These tables show the projected distribution of annual individual covered earnings at three percentiles. Covered earnings are wages from work that is subject to the Social Security payroll tax. The covered earnings in these tables are not capped at the taxable maximum.
Interpreting the Population Characteristics Tables
+
Each projection table links (at the top) to the corresponding population characteristics table. From there, you can use the tabs on the right side to view each population analyzed across the MINT projections including:
We use 10-year birth cohorts to increase the sample size to a point where the characteristic subgroups could be examined.
+
The three columns in every population characteristics table are:
+
+
Unweighted sample: The number of people in the sample for each row. We use it to verify that we comply with disclosure avoidance policy (see “Sample Size Restrictions”).
+
Population (in thousands): The weighted population in thousands. A value of 71,500 means 71 million, 500 thousand.
+
Share of population: The percentage of the population in that particular characteristic subgroup. The Total row at the top of the table is always 100%. The percentages for the each subgroup underneath should add to 100%. For instance, under Country of Birth, United States may be 84% and Other Countries would be 16%, which adds to 100%.
+
+
CORRECT interpretation: 40% of the population is female, 60% of the population is male INCORRECT interpretation: 40% of females are in this population, 60% of males are in this population
Total: Refers to the total population of the table.
+
Sex: Female or Male.
+
Race/Ethnicity: We list “Hispanic or Latino, any race” first; the rest of the groups (White, Black or African American, and All other races) are non-Hispanic. Additional racial or ethnic identifications are not covered because they are not in the datasets used to build the MINT8 model.
+
Country of Birth: We differentiate between the United States and other countries.
+
Age:
+
+
The beneficiary population includes those aged 60 or older because 60 is the earliest eligibility age for any aged benefits under current law.
+
The taxpayer population includes those aged 31 or older because 31 is the earliest age in MINT for household income and poverty information.
+
+
Marital Status: Refers to the marital status in the year of analysis only. An individual's marital status can change in the future and may have been different in the past.
+
Highest Education Level: Reported number of years of education.
+
+
Graduate means more than 16 years of education,
+
Bachelor means 16 years of education,
+
Associate means 14–15 years of education,
+
High school means 12–13 years of education, and
+
Less than high school means less than 12 years of education.
+
+
Current-Law Poverty Status: Indicates whether the person is in a household that has income above (“above poverty”) or below (“in poverty”) the official poverty line under current law. The household income used for the official poverty measure is the same as the household income used in our results except for how asset income is counted. The official poverty measure of asset income only includes dividend income, interest income, and rental income (non-annuitized) as reported on income tax returns. In contrast, we include annuitized asset income from all household wealth held in defined contribution plans (such as 401(k) accounts) and personal savings in that year. We add the annuitized asset income to account for the expected spend-down of assets in retirement. The asset income value used for official poverty calculations generally produces a substantially lower asset income value than the household income measure.
+
Current-Law Household Income Quintile: Represents an individual's annual household income under current law, including:
+
+
household earnings;
+
asset income (annuitized), which includes income from defined contribution plans (such as 401(k) accounts) and personal savings;
+
defined benefit pensions;
+
means-tested income;
+
non-means-tested income;
+
Social Security;
+
Supplemental Security Income; and
+
non-spousal co-residents' income.
+
+
We calculate the income quintiles for each year (e.g., 2030, 2050, or 2070) for the population analyzed, determine the dollar thresholds for each income quintile, and assign each beneficiary to the appropriate quintile. The dollar ranges are available upon request.
+
Current-Law Benefit Type: Some Social Security benefits are based on one's own work, while others are based on the work of a current, divorced, or deceased spouse. The current-law benefit type refers to one of the following benefit types received in the specified analysis year:
+
+
Retired-worker only: receives only a retired-worker benefit based on his or her earnings record.
+
Widow(er) (includes dually entitled): receives a survivor benefit (may or may not also receive a lower worker benefit from his or her own earnings record, known as dually entitled).
+
Spousal (includes dually entitled): receives a spousal benefit (may or may not also receive a lower worker benefit from his or her own earnings record, known as dually entitled).
+
Disabled-worker only: receives a disabled-worker benefit on his or her earnings record and is under the full retirement age (FRA). Disabled workers convert to retired workers at FRA.
+
+
Our results do not show different benefit types a beneficiary might receive under a policy option/proposal or in a different year under current law.
+
Current-Law Payroll Taxes Quintile: Represents an individual's annual Social Security payroll taxes under current law. We calculate the payroll tax quintiles for each analysis year's population of payroll taxpayers aged 31 or older.
+
Current-Law Initial AIME Quintile: Represents an individual's average indexed monthly earnings (AIME) under current law at age 62, the earliest eligibility age for retired-worker benefits. We calculate the AIME quintiles for each birth cohort. The dollar ranges are available upon request.
+
Lifetime Payroll Tax Quintile: Represents the present value of an individual's current-law payroll taxes at age 62. We calculate the payroll tax quintiles for each birth cohort. The dollar ranges are available upon request.
+
Lifetime Payroll Tax Quintile (Shared): Represents the present value of an individual's current-law payroll taxes at age 62. For married couples, the payroll taxes paid while married are shared equally between them. For never-married individuals, this is the same as the lifetime payroll tax. In any year where an individual is not married, we count only their individual payroll taxes. We calculate the quintiles for each birth cohort. The dollar ranges are available upon request.
+
+
Measures—Column Headings
+
+
Threshold for Categorization in the “Decrease” or “Increase” Groups (“Percent of Population with a—[decrease or increase]” Columns)
+
We categorize individuals as having a “decrease” in the amount being analyzed (benefits, taxes, income, etc.) when a proposal would reduce the analyzed quantity by 1% or more. Individuals are categorized as having an “increase” when a proposal would raise the analyzed quantity by 1% or more. We consider individuals with differences between −1% and 1% to be unaffected.
+
For example, consider two individuals with benefits under an option/proposal that are lower than benefits without the proposal—individual A has a benefit decrease of 0.8% while individual B has a decrease of 1.6%. We would consider individual A unaffected and not categorize him or her in the “decrease” or “increase” columns. However, we would categorize individual B as having a decrease.
+
Thus, some beneficiaries who are technically “affected” by a proposal, but who will still receive essentially the same benefit amount (or pay essentially the same taxes, etc.) are considered to be unaffected in our projections.
+
Percent Change Values (“Percent Change in [item being analyzed] at the—[three percentiles]” Columns)
+
Understanding how to interpret the distribution of percent changes is critical to understanding the results correctly.
+
The formula we use to calculate the percent change for each individual is:
From this distribution of individual percent changes, ranked from high to low, we calculate the 10th percentile, median, and 90th percentile values (see below).
+
CORRECT interpretation: a −5% median indicates that this is the median of the distribution of individual percent changes (half the individuals have a percent change that is higher and half have a percent change that is lower). INCORRECT interpretation: a −5% median indicates that the median amount under the option is 5% less than the median amount under current law.
+
10th, Median, and 90th Percentile Values
+
The percentiles provide a picture of the distribution of policy option effects or beneficiaries/taxpayers financial status distributed from lowest to highest. The example below is for a percent change in benefits, but also applies to distributions of other policy options, dollar amounts, initial replacement rates, household income levels, etc.
+
+
+
Table example
+
+
+
+
+
+
Percent change in Social Security benefits at the—
+
+
+
10th percentile
+
Median
+
90th percentile
+
+
+
+
+
Total
+
2%
+
4%
+
21%
+
+
+
+
+
+
+
+
+
+
10th percentile: “2%” means that 10 percent of the population has a benefit change of less than 2 percent, while 90 percent have a benefit change of more than 2 percent.
+
Median: “4%” means that 50 percent of the population has a benefit change of less than 4 percent, while 50 percent have a benefit change of more than 4 percent.
+
90th percentile: “21%” means that 90 percent of the population has a benefit change of less than 21 percent, while 10 percent have a benefit change of more than 21 percent.
+
CORRECT interpretations include:
+
+
10% of this population has a benefit change of less than 2%.
+
10% of this population has a benefit change of more than 21%.
+
40% of this population has a benefit change of between 2% to 4%.
+
40% of this population has a benefit change of between 4% to 21%.
+
50% of this population has a benefit change of less than 4%.
+
50% of this population has a benefit change of more than 4%.
+
+
+
+
Additional Notes
+
Dollar Amounts
+
All dollar amounts are presented in today's dollars, meaning that they are in real dollars (inflation-adjusted) for the year the table is produced. If a table is run in 2021, the dollars are in 2021 dollars and the table will note that in the column label. Tables run in 2024 will be in 2024 dollars, and so on.
+
Sample Size Restrictions
+
To maintain the privacy of survey respondents, our tables have built-in disclosure avoidance protections that suppress an entire characteristic subgroup if the sample size for any row in that subgroup is less than 100 individuals. This minimum sample size protects privacy so that we can show the 10th and 90th percentiles based on a sample size of at least 10.
+
For example, if there are only 82 widow(er)s in the Widowed row of the Marital Status subgroup, the entire Marital Status subgroup is removed from the table. We remove the entire subgroup to avoid secondary or tertiary disclosure issues. The subgroups below Marital Status in this example would automatically move up the page. A “short” table can reveal at a glance that at least one subgroup has been removed.
+
Columns that show percent of the population with a decrease or increase in whatever value is being shown (e.g. benefits, payroll tax, household income) have disclosure restrictions on the numerator's sample size as well as the denominator. The numerator must have either zero cases or meet a minimum numerator threshold of 10 for the table to display it. The table would suppress any characteristic subgroup that has a percentage based on a numerator of 1–9 cases. (MINT results are weighted, but for this example, everything is presented in unweighted sample sizes.)
+
There is an exception for low numerator situations where the table would show a 0% for any numerator from zero up to and including the minimum numerator threshold. If a particular percentage was based on seven records with a denominator of 30,000, it would produce a percentage of 0.02%. By only showing percentages in single digits, we would display this result as 0%, which would be the same value displayed for any numerator from 0–149.
+
This exception is important because there are policy options where 27,000/30,000 beneficiaries (90%) would receive a benefit increase while 8/30,000 receive a decrease. Without the exception, the very small decrease numerator would suppress a number of subgroups for both the increased and the decreased results, which limits the presentable results more than is necessary for disclosure avoidance.
+
1 All the characteristic subgroups are not in every table. Birth cohort tables (benefit/tax ratios and initial replacement rates) do not have age or marital status breakouts. Characteristic subgroups can also drop out of tables because of sample size restrictions (see “Sample Size Restrictions” for details).
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
\ No newline at end of file
diff --git a/data/external/ssa_mint8_table_user_guide.source.provenance.json b/data/external/ssa_mint8_table_user_guide.source.provenance.json
new file mode 100644
index 00000000..c5c069a3
--- /dev/null
+++ b/data/external/ssa_mint8_table_user_guide.source.provenance.json
@@ -0,0 +1,14 @@
+{
+ "schema_version": "external_source_provenance.v1",
+ "committed_source_file": "data/external/ssa_mint8_table_user_guide.source.html",
+ "source_url": "https://www.ssa.gov/policy/docs/projections/user-guide.html",
+ "document": "SSA Office of Research, Evaluation, and Statistics — Table User Guide—Modeling Income in the Near Term (MINT) 8",
+ "date_certified": "2025-10-01 (the page's DCTERMS:dateCertified meta tag)",
+ "section_of_record": "Definitions—Table Rows and Columns > Characteristic Subgroups—Table Rows (race/ethnicity, country of birth, education and lifetime quintile dimensions); Explanation of Table Types > Benefit/Tax Ratios (present-value interest rate).",
+ "fetch_method": "Exact response body of the Internet Archive capture https://web.archive.org/web/20260419231110id_/https://www.ssa.gov/policy/docs/projections/user-guide.html (direct requests to ssa.gov returned an Akamai 'Access Denied' page). The body's SHA-1 in base32 is N2TYAIU2LPKHETR3M5H4E6MUZLL4KFVO, equal to the CDX digest of capture 20260419231110 (CDX query run 2026-10-01).",
+ "source_sha256": "278d5d19c1b50f1d354db1ada515288af563c035a67eb16fb971d25700fb94e9",
+ "source_length_bytes": 73926,
+ "acquired_by": "MINT-categories lane of orchestrating session 95606380 (Claude Code, Opus 5.5), 2026-10-01; locator verified and committed by the G1 group-attributes builder",
+ "consumed_by": "data/external/group_category_schemes_v1.json (scheme mint8 attribute mappings, full table profiles and cohort lifetime-earnings dimension definitions)",
+ "note": "A definitions page only: it holds no MINT result. The Table User Guide is cleared for builders under the NASI follow-ups entry of the evidence directory's restricted-files list ('A lane may read the MINT8 Table User Guide and methodology pages')."
+}
diff --git a/docs/design/nasi_g1_group_attributes.md b/docs/design/nasi_g1_group_attributes.md
new file mode 100644
index 00000000..76941118
--- /dev/null
+++ b/docs/design/nasi_g1_group_attributes.md
@@ -0,0 +1,522 @@
+# G1: cohort-side group attributes
+
+G1 adds a separate `person_id` side frame for the 2009/2011 projection cohorts, age-67 observations, and the Track M 2023 universe. It preserves the existing cohort frames and their pinned digests. The package computes person attributes; it does not produce reform outcomes, group shares, or new blind-test cells.
+
+## Adjudication and resolution
+
+Race and Hispanic origin use every asked race mention of present heads/reference persons and wives/spouses/partners in family waves 1985–2023. Sequence 1–20 identifies people present at interview; relationships 10 and 20/22 identify the two roles. The individual roster retains OFUMs. Race and nativity for a person never in an eligible role remain explicitly unknown, with availability counts in provenance.
+
+Race frames are wave-specific: 1985–1989 has two mentions and no Latino race code; 1990–1993 adds race codes 5 (Latino origin) and 6 (color other than black or white), with 8 meaning more than two races. From 1994, race code 8 means DK. There are three race mentions in 1994–1996 and four from 1997. Waves 1997–2003 ask no Spanish-descent item: a Latino race mention establishes Hispanic origin, while its absence leaves origin unknown. From 2005, code 5 means Native Hawaiian/Pacific Islander and the Spanish-descent question returns. Direct Spanish-origin answers take precedence over Latino race mentions when that question is asked.
+
+For the 1994–1996 wife Spanish-origin variables ER3880, ER6750, and ER8996, the captured codebook describes 0 only as no wife in the family unit. It supplies no non-Hispanic meaning for a present wife. Such reports have `hispanic=None` and `hispanic_basis="undocumented_zero_meaning"`; documentation frequencies are never used to infer an omitted meaning.
+
+Static race/ethnicity uses the most recent complete head/spouse report in the full window; Hispanic reports are complete regardless of race, while non-Hispanic reports need a known full race set. Conflicts remain visible in report and distinct-reading counts. If no complete report exists, the latest known Hispanic-origin answer is retained separately. Nativity uses the latest classifiable 2013–2023 report. Sources after an anchor are explicitly counted; no modal fallback or imputation is used.
+
+Education uses the latest known individual-file answer at or before `max(anchor_waves)`, including OFUMs aged 16 or older. Codes 1–16 give completed grades and 17 means at least some postgraduate work. Individual code 0 is Inap.; it establishes zero years only when a present head/spouse also has the family completed-education recode 0, whose text explicitly documents completed no grades. Otherwise it remains inapplicable. Missing and undocumented codes remain distinct.
+
+Nativity is first available for both roles in 2013. The state/year-came pair identifies a U.S. state when state is 1–56 and year-came is 0, a U.S. territory when both codes are 0, and a foreign country when state is 0 and year-came records a year or an applicable nonzero response. Inconsistent skip pairs and unknown reports do not decide birthplace. U.S. territory births remain unresolved in MINT because the guide does not define their placement. Earlier immigrant-supplement and grew-up items are not substituted.
+
+Each selected variable must match its exact whitespace-normalized SPSS label. Every observed code must belong to its captured codebook or SAS-format domain. Invalid integer types, undocumented codes, ambiguous label changes, duplicate role joins, missing family rows, and changed pinned documentation are refused.
+
+The pure builder enforces the same wave-specific education and family-item domains on supplied frames, including source role and question availability. The reader partitions the long individual roster once by wave before selecting head/spouse rows; an independent synthetic row-join oracle checks that this preserves the earlier whole-frame selection semantics.
+
+## Schemes and lifetime dimensions
+
+`group_category_schemes_v1.json` records MINT four-way race/ethnicity and five education bands, report four-way race/ethnicity, and unresolved report education definitions. The existing repository records the three Report education labels but does not supply numerical definitions; G1 leaves all three assignments unavailable rather than imposing the suggested 12–15 boundary. The MINT convention placing non-Hispanic multiple-race reports in All other races is an explicit builder assumption pending registration. Alternate report schemes can replace this rule without changing the reader.
+
+Lifetime measures (initial AIME at 62 and lifetime payroll-tax present values, own and shared) live in `estimates/lifetime_measures.py`, documented in `docs/design/lifetime_measures.md`. This package supplies only the person attributes and the category schemes those measures are tabulated under.
+
+## Provenance and source pins
+
+`load_group_attribute_inputs` performs all PSID reads inside the cohort file-audit context. Provenance seals the opened files, raw input frames, requested identifiers, schemes, and final attribute frame with SHA-256. Frozen wrappers carry the frames; builders verify recorded input seals, and lifetime quintile assignment refuses a changed frame. Label/codebook inspection and captured-domain regeneration read documentation only. No restricted comparator, policy-option data table, Urban results source, or outcome was read for this package.
+
+| Capture | SHA-256 | Locator |
+|---|---|---|
+| `data/external/psid_group_attribute_codebook_values_v1.json` | `09fce5627b0271a68eadb2e8748e33fe1ff3e6b0e025bfef9575bb2804ebef63` | Per-wave family variable, exact label, one-based PDF page and value table; documentation counts removed |
+| `data/external/group_category_schemes_v1.json` | `da15a940d85dc1b8ea49480ad5ae5d0c4179ff916571bc44c9d901e22c6772c0` | Per-scheme categories, education bands, explicit assumptions and unresolved definitions |
+| `data/external/ssa_mint8_table_user_guide.source.html` | `278d5d19c1b50f1d354db1ada515288af563c035a67eb16fb971d25700fb94e9` | Definitions—Table Rows and Columns > Characteristic Subgroups—Table Rows |
+| `data/external/ssa_mint8_payroll_option_row_labels.json` | `23fbfbc8dbc14b06144a83bf0583ea0a6c89505638f3f7dc5fd6b34a3a4ac650` | Tables 1–3 (benefits), 7–9 (household income), 10–12 (official poverty), 13–16 (benefit/tax ratio), and 17–20 (initial replacement rate): captions and row labels only; no data cells. |
+
+The source paths below are relative to the staged PSID root. Each family inventory row gives the one-based PDF page locating that variable; the captured JSON preserves exact value/range texts. Family SAS formats, when captured, have an additional SHA-256 pin in that JSON. Individual domains and role descriptions are verified in `IND2023ER_formats.sas`, with its actual file SHA-256 included in the runtime file audit.
+
+| Wave | Family codebook source | PDF SHA-256 |
+|---|---|---|
+| 1985 | `family/1985/FAM1985_codebook.pdf` | `ef9cdf2eafccf167b5ab0495cab085877dce4787ccb9a2c01fdf4a868d13909a` |
+| 1986 | `family/1986/FAM1986_codebook.pdf` | `6483a31953e3f969b8757c97779ff6f2c45734b54a2806e8fd506db72c220112` |
+| 1987 | `family/1987/FAM1987_codebook.pdf` | `86065d5cb127766a27037735ee428d1cec5afeeac554bfc9cc2cabfebd300056` |
+| 1988 | `family/1988/FAM1988_codebook.pdf` | `4e2a75c05ee4b2ca9d44a31063096567f45c97285127a124d2d98a913f772c8f` |
+| 1989 | `family/1989/FAM1989_codebook.pdf` | `1dff0473a8df0b7564f1716fd01ef67d276920105d17a986f6f167de5e1a049a` |
+| 1990 | `family/1990/FAM1990_codebook.pdf` | `4e6e0d3c22eedc50137b3d8a8b33e949064681b5bf42275d51da4b349734b32f` |
+| 1991 | `family/1991/FAM1991_codebook.pdf` | `cb25d86256acb143cbc7b6f400702dc6a20787af9476dfbe9fc277ccecca424c` |
+| 1992 | `family/1992/FAM1992_codebook.pdf` | `249b4c1a9a333792df66511c4fc63ae2ac51a28bb764f38aff4bc04a6d389d52` |
+| 1993 | `family/1993/fam1993_codebook.pdf` | `a4659bad3b0d2a4ec0040f9ad63994060d1d98c5924327853dfc7899ac172411` |
+| 1994 | `family/1994/FAM1994ER_codebook_public.pdf` | `ecf8667f3b6d2a5ac7e312ecd9761a00ae7c1d2db6bee2f8af049e495ff1cf56` |
+| 1995 | `family/1995/FAM1995ER_codebook_public.pdf` | `c4a161e935244487d26d7a581ff68b1c7f844c8032fe756eda82c035f51cb985` |
+| 1996 | `family/1996/FAM1996ER_codebook_public.pdf` | `aa52222228e4ecc3d7309c39c61929a7603951615c6a80b3ba3423e7f3ba8300` |
+| 1997 | `family/1997/FAM1997ER_codebook_public.pdf` | `9716779d9b1102fb5be611e620525759a6cb8676714ba9722a34da8c7287bb3d` |
+| 1999 | `family/1999/FAM1999ER_codebook.pdf` | `5accefc4c1b50b3447b1a674ac2750de94568c538873f9bab34c2049c80c6107` |
+| 2001 | `family/2001/FAM2001ER_codebook.pdf` | `13ae45ce24125c2ac30c15e62f1efeddc7842dc924b3415fe5aaa058666029a0` |
+| 2003 | `family/2003/FAM2003ER_codebook.pdf` | `5a69ee0e605aec1305a23931330da97d75dd21c32d183b8d67c3b6597f23a201` |
+| 2005 | `family/2005/FAM2005ER_codebook.pdf` | `eec8de3c04ac88ffce0495622f0cbe0ac621780b201ce6bd0b4b54344a3a6243` |
+| 2007 | `family/2007/FAM2007ER_codebook.pdf` | `d92bc05151cf1b77513e4520578402305d356a669c7998f10f91ad48ac66c89e` |
+| 2009 | `family/2009/FAM2009ER_codebook.pdf` | `2639c0cad9d5b6e1a08eae57c578d9698a249ed331475b79df525a0caa7c4c6b` |
+| 2011 | `family/2011/FAM2011ER_codebook.pdf` | `00c0517569b94efcc7fc6594529a263feb9e3bb9a88efff3becf5b8aaeebdd3c` |
+| 2013 | `family/2013/FAM2013ER_codebook.pdf` | `40cc2de6af801b94b93518f7252e24c7c95098e713fc97f9ead1af74c4ce25d6` |
+| 2015 | `family/2015/FAM2015ER_codebook.pdf` | `7dcb71f823e309a286b904531410243d58cde48b8f50079401856de9ac762238` |
+| 2017 | `family/2017/FAM2017ER_codebook.pdf` | `19bfba9f7a3fecc82a18711466e2b2a5f00b49c9d3db0f9f8587795fad2bf8c1` |
+| 2019 | `family/2019/fam2019er_codebook.pdf` | `ce637e37e0f5e306dd7f10455e85069b56c9b13d5c138d3eb4b9c8d6a8452406` |
+| 2021 | `family/2021/FAM2021ER_codebook.pdf` | `e3b88ef394952fa05b9b98eb7f05bf9a99896840515942562aa37399e1bcd462` |
+| 2023 | `family/2023/FAM2023ER_codebook.pdf` | `85a4078ebb79146d023df3d9d4e2d2ef6017eced849656ca39f4b740cb5a5688` |
+
+## Public API
+
+```python
+load_group_attributes(person_ids: Iterable[int], *, anchor_waves: Iterable[int],
+ psid_dir: Path | None = None) -> GroupAttributes
+load_group_attribute_inputs(*, psid_dir: Path | None = None) -> GroupAttributeInputs
+build_group_attributes(inputs: GroupAttributeInputs, person_ids: Iterable[int], *,
+ anchor_waves: Iterable[int],
+ schemes: Mapping[str, Any] | None = None) -> GroupAttributes
+```
+
+`GroupAttributes(frame, provenance)` supplies the required scheme columns plus source wave, role, variables, mentions, conflict counts, and availability reasons. `GroupAttributeInputs(reports, education, universe, provenance)` supports pure synthetic or previously audited inputs. Join side frames by `person_id`; age-67 observations may legitimately repeat a person.
+
+## Tests and integration follow-ups
+
+| Test module | Tier | Coverage |
+|---|---|---|
+| `tests/data/test_group_attributes_psid.py` | artifact | INVENTED fixed-width products; exact labels; domains; role joins; ambiguous and unknown answers; education/nativity properties; differential domain representations; captured table identity; codebook extraction and documentation pins |
+| `tests/cohorts/test_group_attributes.py` | artifact | Synthetic resolution, conflicts, unknown preservation, joins, anchor education cutoff, seals and property tests; reads committed schemes |
+| `tests/cohorts/test_group_category_schemes.py` | artifact | Captured source/label identity and scheme mappings, unresolved definitions, boundaries and property tests |
+| `tests/cohorts/test_group_attributes_integration.py` | integration_psid | Staged labels, domains, pins and coverage for 2009/2011 cohorts, age-67 observations and Track M 2023; skips when sources are absent |
+
+Run only targeted tests with the main venv and one pytest process at a time. Host integration checks availability totals and identifier coverage only; it computes no policy outcome or distribution by group. A later #42 registration is required before any real group breakdown. The integrator must update `tests/tier_counts.json` and the source-exclusion/artifact-test inventory in `scripts/first_estimates_birth_evidence.py` and `tests/estimates/test_birth_evidence_artifact.py`; G1 does not edit those shared files or `pyproject.toml`.
+
+## Validation in the assigned workspace
+
+All pytest commands used `PYTHONPATH=src /Users/maxghenis/PolicyEngine/microcosm-dynamics/.venv/bin/python -m pytest`, with one process at a time:
+
+| Arguments | Result |
+|---|---|
+| `tests/data/test_group_attributes_psid.py tests/cohorts/test_group_attributes.py tests/cohorts/test_group_category_schemes.py -q -x` | 151 passed in 358.48 seconds before final validation/performance additions |
+| `tests/cohorts/test_group_attributes.py -q` | Final: 73 passed in 1425.64 seconds |
+| `tests/data/test_group_attributes_psid.py -q` | Final: 39 passed in 386.66 seconds |
+| `tests/cohorts/test_group_attributes.py tests/cohorts/test_group_attributes_integration.py -q -x` | Interrupted after 8296.62 seconds to limit shared-host load; 48 synthetic cases passed, no host population case completed |
+
+The final unchanged scheme/lifetime tests and final cohort/reader tests account for 183 distinct passing synthetic cases. Black at 79 columns and Ruff passed for all nine Python files, including rechecks after the final source/test edits. The captured codebook table regenerated exactly from staged documentation. Host population coverage remains unverified; no outcome or group-share pipeline was run.
+
+The branch remains `nasi/g1-group-attributes` at `75cd35245f1a0c12e7a23b048a7d126c74ad8ecb`. All package files remain uncommitted: Git staging was denied when creating `/Users/maxghenis/PolicyEngine/microcosm-dynamics/.git/worktrees/nasi-g1/index.lock`, which lies outside the writable workspace. Existing tracked files were preserved. Nothing was pushed and no PR was opened.
+
+## Exact individual-file inventory
+
+All variables below are in `ind2023er/IND2023ER.sps`; every label was verified against the staged file. Wave-independent identifiers are ER30001 — 1968 INTERVIEW NUMBER and ER30002 — PERSON NUMBER 68. Sequence and relationship domains come from the assigned SAS VALUE blocks; education domains have their wave-specific DK/NA codes.
+
+| Wave | Item | Variable | Exact SPSS label |
+|---|---|---|---|
+| 1985 | interview | `ER30463` | 1985 INTERVIEW NUMBER |
+| 1985 | sequence | `ER30464` | SEQUENCE NUMBER 85 |
+| 1985 | relationship | `ER30465` | RELATIONSHIP TO HEAD 85 |
+| 1985 | education | `ER30478` | COMPLETED EDUCATION 85 |
+| 1986 | interview | `ER30498` | 1986 INTERVIEW NUMBER |
+| 1986 | sequence | `ER30499` | SEQUENCE NUMBER 86 |
+| 1986 | relationship | `ER30500` | RELATIONSHIP TO HEAD 86 |
+| 1986 | education | `ER30513` | COMPLETED EDUCATION 86 |
+| 1987 | interview | `ER30535` | 1987 INTERVIEW NUMBER |
+| 1987 | sequence | `ER30536` | SEQUENCE NUMBER 87 |
+| 1987 | relationship | `ER30537` | RELATIONSHIP TO HEAD 87 |
+| 1987 | education | `ER30549` | COMPLETED EDUCATION 87 |
+| 1988 | interview | `ER30570` | 1988 INTERVIEW NUMBER |
+| 1988 | sequence | `ER30571` | SEQUENCE NUMBER 88 |
+| 1988 | relationship | `ER30572` | RELATION TO HEAD 88 |
+| 1988 | education | `ER30584` | COMPLETED EDUC-IND 88 |
+| 1989 | interview | `ER30606` | 1989 INTERVIEW NUMBER |
+| 1989 | sequence | `ER30607` | SEQUENCE NUMBER 89 |
+| 1989 | relationship | `ER30608` | RELATION TO HEAD 89 |
+| 1989 | education | `ER30620` | COMPLETED EDUC-IND 89 |
+| 1990 | interview | `ER30642` | 1990 INTERVIEW NUMBER |
+| 1990 | sequence | `ER30643` | SEQUENCE NUMBER 90 |
+| 1990 | relationship | `ER30644` | RELATION TO HEAD 90 |
+| 1990 | education | `ER30657` | COMPLETED EDUC-IND 90 |
+| 1991 | interview | `ER30689` | 1991 INTERVIEW NUMBER |
+| 1991 | sequence | `ER30690` | SEQUENCE NUMBER 91 |
+| 1991 | relationship | `ER30691` | RELATION TO HEAD 91 |
+| 1991 | education | `ER30703` | COMPLETED EDUC-IND 91 |
+| 1992 | interview | `ER30733` | 1992 INTERVIEW NUMBER |
+| 1992 | sequence | `ER30734` | SEQUENCE NUMBER 92 |
+| 1992 | relationship | `ER30735` | RELATION TO HEAD 92 |
+| 1992 | education | `ER30748` | COMPLETED EDUCATION 92 |
+| 1993 | interview | `ER30806` | 1993 INTERVIEW NUMBER |
+| 1993 | sequence | `ER30807` | SEQUENCE NUMBER 93 |
+| 1993 | relationship | `ER30808` | RELATION TO HEAD 93 |
+| 1993 | education | `ER30820` | YRS COMPLETED EDUCATION 93 |
+| 1994 | interview | `ER33101` | 1994 INTERVIEW NUMBER |
+| 1994 | sequence | `ER33102` | SEQUENCE NUMBER 94 |
+| 1994 | relationship | `ER33103` | RELATION TO HEAD 94 |
+| 1994 | education | `ER33115` | YRS COMPLETED EDUC 94 |
+| 1995 | interview | `ER33201` | 1995 INTERVIEW NUMBER |
+| 1995 | sequence | `ER33202` | SEQUENCE NUMBER 95 |
+| 1995 | relationship | `ER33203` | RELATION TO HEAD 95 |
+| 1995 | education | `ER33215` | YEARS COMPLETED EDUCATION 95 |
+| 1996 | interview | `ER33301` | 1996 INTERVIEW NUMBER |
+| 1996 | sequence | `ER33302` | SEQUENCE NUMBER 96 |
+| 1996 | relationship | `ER33303` | RELATION TO HEAD 96 |
+| 1996 | education | `ER33315` | YEARS COMPLETED EDUCATION 96 |
+| 1997 | interview | `ER33401` | 1997 INTERVIEW NUMBER |
+| 1997 | sequence | `ER33402` | SEQUENCE NUMBER 97 |
+| 1997 | relationship | `ER33403` | RELATION TO HEAD 97 |
+| 1997 | education | `ER33415` | YEARS COMPLETED EDUCATION 97 |
+| 1999 | interview | `ER33501` | 1999 INTERVIEW NUMBER |
+| 1999 | sequence | `ER33502` | SEQUENCE NUMBER 99 |
+| 1999 | relationship | `ER33503` | RELATION TO HEAD 99 |
+| 1999 | education | `ER33516` | YEARS COMPLETED EDUCATION 99 |
+| 2001 | interview | `ER33601` | 2001 INTERVIEW NUMBER |
+| 2001 | sequence | `ER33602` | SEQUENCE NUMBER 01 |
+| 2001 | relationship | `ER33603` | RELATION TO HEAD 01 |
+| 2001 | education | `ER33616` | YEARS COMPLETED EDUCATION 01 |
+| 2003 | interview | `ER33701` | 2003 INTERVIEW NUMBER |
+| 2003 | sequence | `ER33702` | SEQUENCE NUMBER 03 |
+| 2003 | relationship | `ER33703` | RELATION TO HEAD 03 |
+| 2003 | education | `ER33716` | YEARS COMPLETED EDUCATION 03 |
+| 2005 | interview | `ER33801` | 2005 INTERVIEW NUMBER |
+| 2005 | sequence | `ER33802` | SEQUENCE NUMBER 05 |
+| 2005 | relationship | `ER33803` | RELATION TO HEAD 05 |
+| 2005 | education | `ER33817` | YEARS COMPLETED EDUCATION 05 |
+| 2007 | interview | `ER33901` | 2007 INTERVIEW NUMBER |
+| 2007 | sequence | `ER33902` | SEQUENCE NUMBER 07 |
+| 2007 | relationship | `ER33903` | RELATION TO HEAD 07 |
+| 2007 | education | `ER33917` | YEARS COMPLETED EDUCATION 07 |
+| 2009 | interview | `ER34001` | 2009 INTERVIEW NUMBER |
+| 2009 | sequence | `ER34002` | SEQUENCE NUMBER 09 |
+| 2009 | relationship | `ER34003` | RELATION TO HEAD 09 |
+| 2009 | education | `ER34020` | YEARS COMPLETED EDUCATION 09 |
+| 2011 | interview | `ER34101` | 2011 INTERVIEW NUMBER |
+| 2011 | sequence | `ER34102` | SEQUENCE NUMBER 11 |
+| 2011 | relationship | `ER34103` | RELATION TO HEAD 11 |
+| 2011 | education | `ER34119` | YEARS COMPLETED EDUCATION 11 |
+| 2013 | interview | `ER34201` | 2013 INTERVIEW NUMBER |
+| 2013 | sequence | `ER34202` | SEQUENCE NUMBER 13 |
+| 2013 | relationship | `ER34203` | RELATION TO HEAD 13 |
+| 2013 | education | `ER34230` | YEARS COMPLETED EDUCATION 13 |
+| 2015 | interview | `ER34301` | 2015 INTERVIEW NUMBER |
+| 2015 | sequence | `ER34302` | SEQUENCE NUMBER 15 |
+| 2015 | relationship | `ER34303` | RELATION TO HEAD 15 |
+| 2015 | education | `ER34349` | YEARS COMPLETED EDUCATION 15 |
+| 2017 | interview | `ER34501` | 2017 INTERVIEW NUMBER |
+| 2017 | sequence | `ER34502` | SEQUENCE NUMBER 17 |
+| 2017 | relationship | `ER34503` | RELATION TO REFERENCE PERSON 17 |
+| 2017 | education | `ER34548` | YEARS COMPLETED EDUCATION 17 |
+| 2019 | interview | `ER34701` | 2019 INTERVIEW NUMBER |
+| 2019 | sequence | `ER34702` | SEQUENCE NUMBER 19 |
+| 2019 | relationship | `ER34703` | RELATION TO REFERENCE PERSON 19 |
+| 2019 | education | `ER34752` | YEARS COMPLETED EDUCATION 19 |
+| 2021 | interview | `ER34901` | 2021 INTERVIEW NUMBER |
+| 2021 | sequence | `ER34902` | SEQUENCE NUMBER 21 |
+| 2021 | relationship | `ER34903` | RELATION TO REFERENCE PERSON 21 |
+| 2021 | education | `ER34952` | YEARS COMPLETED EDUCATION 21 |
+| 2023 | interview | `ER35101` | 2023 INTERVIEW NUMBER |
+| 2023 | sequence | `ER35102` | SEQUENCE NUMBER 23 |
+| 2023 | relationship | `ER35103` | RELATION TO REFERENCE PERSON 23 |
+| 2023 | education | `ER35152` | YEARS COMPLETED EDUCATION 23 |
+
+## Exact family-file inventory
+
+Wave directories are `family//`; interview numbers join the individual roster to family records. Race mention numbers are one-based and preserve question order. Missing items in a wave are genuinely unasked and are not filled with another variable. Every listed label was verified against its staged SPSS setup file.
+
+| Wave | Role | Item | Variable | Exact SPSS label | PDF page |
+|---|---|---|---|---|---|
+| 1985 | family | interview | `V11102` | 1985 INTERVIEW NUMBER | — |
+| 1985 | head | hispanic | `V11937` | G31 SPANISH DESCENT-HEAD | 277 |
+| 1985 | head | race mention 1 | `V11938` | G32 RACE OF HEAD (1 MEN) | 277 |
+| 1985 | head | race mention 2 | `V11939` | G32 RACE OF HEAD (2 MEN) | 278 |
+| 1985 | spouse | hispanic | `V12292` | N31 SPANISH DESCENT-WIFE | 413 |
+| 1985 | spouse | race mention 1 | `V12293` | N32 RACE OF WIFE (1 MEN) | 413 |
+| 1985 | spouse | race mention 2 | `V12294` | N32 RACE OF WIFE (2 MEN) | 414 |
+| 1986 | family | interview | `V12502` | 1986 INTERVIEW NUMBER | — |
+| 1986 | head | hispanic | `V13564` | L31 SPANISH DESCENT HD | 376 |
+| 1986 | head | race mention 1 | `V13565` | L32 RACE OF HEAD 1 | 376 |
+| 1986 | head | race mention 2 | `V13566` | L32 RACE OF HEAD 2 | 377 |
+| 1986 | spouse | hispanic | `V13499` | K18 SPANISH DESCENT WF | 344 |
+| 1986 | spouse | race mention 1 | `V13500` | K19 RACE OF WIFE 1 | 345 |
+| 1986 | spouse | race mention 2 | `V13501` | K19 RACE OF WIFE 2 | 345 |
+| 1987 | family | interview | `V13702` | 1987 INTERVIEW NUMBER | — |
+| 1987 | head | hispanic | `V14611` | L31 SPANISH DESCENT HD | 326 |
+| 1987 | head | race mention 1 | `V14612` | L32 RACE OF HEAD 1 | 326 |
+| 1987 | head | race mention 2 | `V14613` | L32 RACE OF HEAD 2 | 327 |
+| 1987 | spouse | hispanic | `V14546` | K18 SPANISH DESCENT WF | 296 |
+| 1987 | spouse | race mention 1 | `V14547` | K19 RACE OF WIFE 1 | 296 |
+| 1987 | spouse | race mention 2 | `V14548` | K19 RACE OF WIFE 2 | 297 |
+| 1988 | family | interview | `V14802` | 1988 INTERVIEW NUMBER | — |
+| 1988 | head | hispanic | `V16085` | L31 SPANISH DESCENT HD | 437 |
+| 1988 | head | race mention 1 | `V16086` | L32 RACE OF HEAD 1 | 438 |
+| 1988 | head | race mention 2 | `V16087` | L32 RACE OF HEAD 2 | 438 |
+| 1988 | spouse | hispanic | `V16020` | K18 SPANISH DESCENT WF | 406 |
+| 1988 | spouse | race mention 1 | `V16021` | K19 RACE OF WIFE 1 | 406 |
+| 1988 | spouse | race mention 2 | `V16022` | K19 RACE OF WIFE 2 | 407 |
+| 1989 | family | interview | `V16302` | 1989 INTERVIEW NUMBER | — |
+| 1989 | head | hispanic | `V17482` | L31 SPANISH DESCENT HD | 389 |
+| 1989 | head | race mention 1 | `V17483` | L32 RACE OF HEAD 1 | 389 |
+| 1989 | head | race mention 2 | `V17484` | L32 RACE OF HEAD 2 | 390 |
+| 1989 | spouse | hispanic | `V17417` | K18 SPANISH DESCENT WF | 358 |
+| 1989 | spouse | race mention 1 | `V17418` | K19 RACE OF WIFE 1 | 359 |
+| 1989 | spouse | race mention 2 | `V17419` | K19 RACE OF WIFE 2 | 359 |
+| 1990 | family | interview | `V17702` | 1990 INTERVEW NUMBER | — |
+| 1990 | head | hispanic | `V18813` | M31 SPANISH DESCENT HD | 355 |
+| 1990 | head | race mention 1 | `V18814` | M32 RACE OF HEAD 1 | 355 |
+| 1990 | head | race mention 2 | `V18815` | M32 RACE OF HEAD 2 | 356 |
+| 1990 | spouse | hispanic | `V18748` | L18 SPANISH DESCENT WF | 326 |
+| 1990 | spouse | race mention 1 | `V18749` | L19 RACE OF WIFE 1 | 327 |
+| 1990 | spouse | race mention 2 | `V18750` | L19 RACE OF WIFE 2 | 327 |
+| 1991 | family | interview | `V19002` | 1991 INTERVIEW NUMBER | — |
+| 1991 | head | hispanic | `V20113` | L31 SPANISH DESCENT HD | 351 |
+| 1991 | head | race mention 1 | `V20114` | L32 RACE OF HEAD 1 | 351 |
+| 1991 | head | race mention 2 | `V20115` | L32 RACE OF HEAD 2 | 352 |
+| 1991 | spouse | hispanic | `V20048` | K18 SPANISH DESCENT WF | 326 |
+| 1991 | spouse | race mention 1 | `V20049` | K19 RACE OF WIFE 1 | 327 |
+| 1991 | spouse | race mention 2 | `V20050` | K19 RACE OF WIFE 2 | 327 |
+| 1992 | family | interview | `V20302` | 1992 INTERVIEW NUMBER | — |
+| 1992 | head | hispanic | `V21419` | M31 SPANISH DESCENT HD | 353 |
+| 1992 | head | race mention 1 | `V21420` | M32 RACE OF HEAD 1 | 354 |
+| 1992 | head | race mention 2 | `V21421` | M32 RACE OF HEAD 2 | 354 |
+| 1992 | spouse | hispanic | `V21354` | L18 SPANISH DESCENT WF | 329 |
+| 1992 | spouse | race mention 1 | `V21355` | L19 RACE OF WIFE 1 | 329 |
+| 1992 | spouse | race mention 2 | `V21356` | L19 RACE OF WIFE 2 | 329 |
+| 1993 | family | interview | `V21602` | 1993 INTERVIEW NUMBER | — |
+| 1993 | head | hispanic | `V23275` | L31 WTR HD OF SPANISH DESCENT | 432 |
+| 1993 | head | race mention 1 | `V23276` | L32 RACE OF HD-1ST MENTION | 432 |
+| 1993 | head | race mention 2 | `V23277` | L32 RACE OF HD-2ND MENTION | 433 |
+| 1993 | head | completed_education | `V23333` | COMPLETED ED-HD 1993 | 459 |
+| 1993 | spouse | hispanic | `V23211` | K18 WTR WF OF SPANISH DESCENT | 409 |
+| 1993 | spouse | race mention 1 | `V23212` | K19 RACE OF WF-1ST MENTION | 409 |
+| 1993 | spouse | race mention 2 | `V23213` | K19 RACE OF WF-2ND MENTION | 410 |
+| 1993 | spouse | completed_education | `V23334` | COMPLETED ED-WF 1993 | 460 |
+| 1994 | family | interview | `ER2002` | 1994 INTERVIEW # | — |
+| 1994 | head | hispanic | `ER3941` | L31 SPANISH DESCENT 1 HD | 566 |
+| 1994 | head | race mention 1 | `ER3944` | L32 RACE OF HEAD 1 | 567 |
+| 1994 | head | race mention 2 | `ER3945` | L32 RACE OF HEAD 2 | 568 |
+| 1994 | head | race mention 3 | `ER3946` | L32 RACE OF HEAD 3 | 568 |
+| 1994 | head | completed_education | `ER4158` | COMPLETED ED-HD | 669 |
+| 1994 | spouse | hispanic | `ER3880` | K18 SPANISH DESCENT 1 WF | 545 |
+| 1994 | spouse | race mention 1 | `ER3883` | K19 RACE OF WIFE 1 | 546 |
+| 1994 | spouse | race mention 2 | `ER3884` | K19 RACE OF WIFE 2 | 546 |
+| 1994 | spouse | race mention 3 | `ER3885` | K19 RACE OF WIFE 3 | 547 |
+| 1994 | spouse | completed_education | `ER4159` | COMPLETED ED-WF | 669 |
+| 1995 | family | interview | `ER5002` | 1995 INTERVIEW # | — |
+| 1995 | head | hispanic | `ER6811` | L31 SPANISH DESCENT 1 HD | 517 |
+| 1995 | head | race mention 1 | `ER6814` | L32 RACE OF HEAD 1 | 518 |
+| 1995 | head | race mention 2 | `ER6815` | L32 RACE OF HEAD 2 | 519 |
+| 1995 | head | race mention 3 | `ER6816` | L32 RACE OF HEAD 3 | 519 |
+| 1995 | head | completed_education | `ER6998` | COMPLETED ED-HD | 612 |
+| 1995 | spouse | hispanic | `ER6750` | K18 SPANISH DESCENT 1 WF | 496 |
+| 1995 | spouse | race mention 1 | `ER6753` | K19 RACE OF WIFE 1 | 497 |
+| 1995 | spouse | race mention 2 | `ER6754` | K19 RACE OF WIFE 2 | 498 |
+| 1995 | spouse | race mention 3 | `ER6755` | K19 RACE OF WIFE 3 | 498 |
+| 1995 | spouse | completed_education | `ER6999` | COMPLETED ED-WF | 612 |
+| 1996 | family | interview | `ER7002` | 1996 INTERVIEW # | — |
+| 1996 | head | hispanic | `ER9057` | L31 SPANISH DESCENT 1 HD | 625 |
+| 1996 | head | race mention 1 | `ER9060` | L32 RACE OF HEAD 1 | 626 |
+| 1996 | head | race mention 2 | `ER9061` | L32 RACE OF HEAD 2 | 627 |
+| 1996 | head | race mention 3 | `ER9062` | L32 RACE OF HEAD 3 | 627 |
+| 1996 | head | completed_education | `ER9249` | COMPLETED ED-HD | 719 |
+| 1996 | spouse | hispanic | `ER8996` | K18 SPANISH DESCENT 1 WF | 604 |
+| 1996 | spouse | race mention 1 | `ER8999` | K19 RACE OF WIFE 1 | 605 |
+| 1996 | spouse | race mention 2 | `ER9000` | K19 RACE OF WIFE 2 | 605 |
+| 1996 | spouse | race mention 3 | `ER9001` | K19 RACE OF WIFE 3 | 606 |
+| 1996 | spouse | completed_education | `ER9250` | COMPLETED ED-WF | 720 |
+| 1997 | family | interview | `ER10002` | 1997 INTERVIEW # | — |
+| 1997 | head | race mention 1 | `ER11848` | L40/95 RACE OF HEAD 1 | 504 |
+| 1997 | head | race mention 2 | `ER11849` | L40/95 RACE OF HEAD 2 | 504 |
+| 1997 | head | race mention 3 | `ER11850` | L40/95 RACE OF HEAD 3 | 504 |
+| 1997 | head | race mention 4 | `ER11851` | L40/95 RACE OF HEAD 4 | 505 |
+| 1997 | head | completed_education | `ER12222` | COMPLETED ED-HD | 660 |
+| 1997 | spouse | race mention 1 | `ER11760` | K34/87 RACE OF WIFE 1 | 470 |
+| 1997 | spouse | race mention 2 | `ER11761` | K34/87 RACE OF WIFE 2 | 471 |
+| 1997 | spouse | race mention 3 | `ER11762` | K34/87 RACE OF WIFE 3 | 471 |
+| 1997 | spouse | race mention 4 | `ER11763` | K34/87 RACE OF WIFE 4 | 471 |
+| 1997 | spouse | completed_education | `ER12223` | COMPLETED ED-WF | 660 |
+| 1999 | family | interview | `ER13002` | 1999 FAMILY INTERVIEW (ID) NUMBER | — |
+| 1999 | head | race mention 1 | `ER15928` | L40/95 RACE OF HEAD 1 | 858 |
+| 1999 | head | race mention 2 | `ER15929` | L40/95 RACE OF HEAD 2 | 858 |
+| 1999 | head | race mention 3 | `ER15930` | L40/95 RACE OF HEAD 3 | 858 |
+| 1999 | head | race mention 4 | `ER15931` | L40/95 RACE OF HEAD 4 | 859 |
+| 1999 | head | completed_education | `ER16516` | COMPLETED ED-HD | 1079 |
+| 1999 | spouse | race mention 1 | `ER15836` | K34/87 RACE OF WIFE 1 | 818 |
+| 1999 | spouse | race mention 2 | `ER15837` | K34/87 RACE OF WIFE 2 | 819 |
+| 1999 | spouse | race mention 3 | `ER15838` | K34/87 RACE OF WIFE 3 | 819 |
+| 1999 | spouse | race mention 4 | `ER15839` | K34/87 RACE OF WIFE 4 | 819 |
+| 1999 | spouse | completed_education | `ER16517` | COMPLETED ED-WF | 1079 |
+| 2001 | family | interview | `ER17002` | 2001 FAMILY INTERVIEW (ID) NUMBER | — |
+| 2001 | head | race mention 1 | `ER19989` | L40/95 RACE OF HEAD 1 | 860 |
+| 2001 | head | race mention 2 | `ER19990` | L40/95 RACE OF HEAD 2 | 861 |
+| 2001 | head | race mention 3 | `ER19991` | L40/95 RACE OF HEAD 3 | 861 |
+| 2001 | head | race mention 4 | `ER19992` | L40/95 RACE OF HEAD 4 | 862 |
+| 2001 | head | completed_education | `ER20457` | COMPLETED ED-HD | 1037 |
+| 2001 | spouse | race mention 1 | `ER19897` | K34/87 RACE OF WIFE 1 | 821 |
+| 2001 | spouse | race mention 2 | `ER19898` | K34/87 RACE OF WIFE 2 | 822 |
+| 2001 | spouse | race mention 3 | `ER19899` | K34/87 RACE OF WIFE 3 | 822 |
+| 2001 | spouse | race mention 4 | `ER19900` | K34/87 RACE OF WIFE 4 | 822 |
+| 2001 | spouse | completed_education | `ER20458` | COMPLETED ED-WF | 1038 |
+| 2003 | family | interview | `ER21002` | 2003 FAMILY INTERVIEW (ID) NUMBER | — |
+| 2003 | head | race mention 1 | `ER23426` | L40/95 RACE OF HEAD 1 | 714 |
+| 2003 | head | race mention 2 | `ER23427` | L40/95 RACE OF HEAD 2 | 714 |
+| 2003 | head | race mention 3 | `ER23428` | L40/95 RACE OF HEAD 3 | 714 |
+| 2003 | head | race mention 4 | `ER23429` | L40/95 RACE OF HEAD 4 | 715 |
+| 2003 | head | completed_education | `ER24148` | COMPLETED ED-HD | 1004 |
+| 2003 | spouse | race mention 1 | `ER23334` | K34/87 RACE OF WIFE 1 | 673 |
+| 2003 | spouse | race mention 2 | `ER23335` | K34/87 RACE OF WIFE 2 | 674 |
+| 2003 | spouse | race mention 3 | `ER23336` | K34/87 RACE OF WIFE 3 | 674 |
+| 2003 | spouse | race mention 4 | `ER23337` | K34/87 RACE OF WIFE 4 | 674 |
+| 2003 | spouse | completed_education | `ER24149` | COMPLETED ED-WF | 1005 |
+| 2005 | family | interview | `ER25002` | 2005 FAMILY INTERVIEW (ID) NUMBER | — |
+| 2005 | head | hispanic | `ER27392` | L39A SPANISH DESCENT-HEAD | 708 |
+| 2005 | head | race mention 1 | `ER27393` | L40 RACE OF HEAD-MENTION 1 | 708 |
+| 2005 | head | race mention 2 | `ER27394` | L40 RACE OF HEAD-MENTION 2 | 708 |
+| 2005 | head | race mention 3 | `ER27395` | L40 RACE OF HEAD-MENTION 3 | 709 |
+| 2005 | head | race mention 4 | `ER27396` | L40 RACE OF HEAD-MENTION 4 | 709 |
+| 2005 | head | completed_education | `ER28047` | COMPLETED ED-HD | 961 |
+| 2005 | spouse | hispanic | `ER27296` | K33A SPANISH DESCENT-WIFE | 665 |
+| 2005 | spouse | race mention 1 | `ER27297` | K34 RACE OF WIFE-MENTION 1 | 666 |
+| 2005 | spouse | race mention 2 | `ER27298` | K34 RACE OF WIFE-MENTION 2 | 666 |
+| 2005 | spouse | race mention 3 | `ER27299` | K34 RACE OF WIFE-MENTION 3 | 666 |
+| 2005 | spouse | race mention 4 | `ER27300` | K34 RACE OF WIFE-MENTION 4 | 667 |
+| 2005 | spouse | completed_education | `ER28048` | COMPLETED ED-WF | 962 |
+| 2007 | family | interview | `ER36002` | 2007 FAMILY INTERVIEW (ID) NUMBER | — |
+| 2007 | head | hispanic | `ER40564` | L39 SPANISH DESCENT-HEAD | 1399 |
+| 2007 | head | race mention 1 | `ER40565` | L40 RACE OF HEAD-MENTION 1 | 1399 |
+| 2007 | head | race mention 2 | `ER40566` | L40 RACE OF HEAD-MENTION 2 | 1399 |
+| 2007 | head | race mention 3 | `ER40567` | L40 RACE OF HEAD-MENTION 3 | 1400 |
+| 2007 | head | race mention 4 | `ER40568` | L40 RACE OF HEAD-MENTION 4 | 1400 |
+| 2007 | head | completed_education | `ER41037` | COMPLETED ED-HD | 1590 |
+| 2007 | spouse | hispanic | `ER40471` | K39 SPANISH DESCENT-WIFE | 1357 |
+| 2007 | spouse | race mention 1 | `ER40472` | K40 RACE OF WIFE-MENTION 1 | 1358 |
+| 2007 | spouse | race mention 2 | `ER40473` | K40 RACE OF WIFE-MENTION 2 | 1358 |
+| 2007 | spouse | race mention 3 | `ER40474` | K40 RACE OF WIFE-MENTION 3 | 1358 |
+| 2007 | spouse | race mention 4 | `ER40475` | K40 RACE OF WIFE-MENTION 4 | 1359 |
+| 2007 | spouse | completed_education | `ER41038` | COMPLETED ED-WF | 1591 |
+| 2009 | family | interview | `ER42002` | 2009 FAMILY INTERVIEW (ID) NUMBER | — |
+| 2009 | head | hispanic | `ER46542` | L39 SPANISH DESCENT-HEAD | 1500 |
+| 2009 | head | race mention 1 | `ER46543` | L40 RACE OF HEAD-MENTION 1 | 1501 |
+| 2009 | head | race mention 2 | `ER46544` | L40 RACE OF HEAD-MENTION 2 | 1501 |
+| 2009 | head | race mention 3 | `ER46545` | L40 RACE OF HEAD-MENTION 3 | 1501 |
+| 2009 | head | race mention 4 | `ER46546` | L40 RACE OF HEAD-MENTION 4 | 1502 |
+| 2009 | head | completed_education | `ER46981` | COMPLETED ED-HD | 1671 |
+| 2009 | spouse | hispanic | `ER46448` | K39 SPANISH DESCENT-WIFE | 1458 |
+| 2009 | spouse | race mention 1 | `ER46449` | K40 RACE OF WIFE-MENTION 1 | 1459 |
+| 2009 | spouse | race mention 2 | `ER46450` | K40 RACE OF WIFE-MENTION 2 | 1459 |
+| 2009 | spouse | race mention 3 | `ER46451` | K40 RACE OF WIFE-MENTION 3 | 1459 |
+| 2009 | spouse | race mention 4 | `ER46452` | K40 RACE OF WIFE-MENTION 4 | 1460 |
+| 2009 | spouse | completed_education | `ER46982` | COMPLETED ED-WF | 1672 |
+| 2011 | family | interview | `ER47302` | 2011 FAMILY INTERVIEW (ID) NUMBER | — |
+| 2011 | head | hispanic | `ER51903` | L39 SPANISH DESCENT-HEAD | 1655 |
+| 2011 | head | race mention 1 | `ER51904` | L40 RACE OF HEAD-MENTION 1 | 1655 |
+| 2011 | head | race mention 2 | `ER51905` | L40 RACE OF HEAD-MENTION 2 | 1656 |
+| 2011 | head | race mention 3 | `ER51906` | L40 RACE OF HEAD-MENTION 3 | 1656 |
+| 2011 | head | race mention 4 | `ER51907` | L40 RACE OF HEAD-MENTION 4 | 1656 |
+| 2011 | head | completed_education | `ER52405` | COMPLETED ED-HD | 1861 |
+| 2011 | spouse | hispanic | `ER51809` | K39 SPANISH DESCENT-WIFE | 1610 |
+| 2011 | spouse | race mention 1 | `ER51810` | K40 RACE OF WIFE-MENTION 1 | 1610 |
+| 2011 | spouse | race mention 2 | `ER51811` | K40 RACE OF WIFE-MENTION 2 | 1611 |
+| 2011 | spouse | race mention 3 | `ER51812` | K40 RACE OF WIFE-MENTION 3 | 1611 |
+| 2011 | spouse | race mention 4 | `ER51813` | K40 RACE OF WIFE-MENTION 4 | 1611 |
+| 2011 | spouse | completed_education | `ER52406` | COMPLETED ED-WF | 1862 |
+| 2013 | family | interview | `ER53002` | 2013 FAMILY INTERVIEW (ID) NUMBER | — |
+| 2013 | head | hispanic | `ER57658` | L39 SPANISH DESCENT-HEAD | 1677 |
+| 2013 | head | race mention 1 | `ER57659` | L40 RACE OF HEAD-MENTION 1 | 1677 |
+| 2013 | head | race mention 2 | `ER57660` | L40 RACE OF HEAD-MENTION 2 | 1678 |
+| 2013 | head | race mention 3 | `ER57661` | L40 RACE OF HEAD-MENTION 3 | 1678 |
+| 2013 | head | race mention 4 | `ER57662` | L40 RACE OF HEAD-MENTION 4 | 1678 |
+| 2013 | head | completed_education | `ER58223` | COMPLETED ED-HD | 1876 |
+| 2013 | head | birth_state | `ER57651` | L33 STATE HEAD WAS BORN | 1675 |
+| 2013 | head | year_came | `ER57652` | L33YR YEAR CAME TO UNITED STATES-HD | 1675 |
+| 2013 | spouse | hispanic | `ER57548` | K39 SPANISH DESCENT-WIFE | 1620 |
+| 2013 | spouse | race mention 1 | `ER57549` | K40 RACE OF WIFE-MENTION 1 | 1621 |
+| 2013 | spouse | race mention 2 | `ER57550` | K40 RACE OF WIFE-MENTION 2 | 1621 |
+| 2013 | spouse | race mention 3 | `ER57551` | K40 RACE OF WIFE-MENTION 3 | 1621 |
+| 2013 | spouse | race mention 4 | `ER57552` | K40 RACE OF WIFE-MENTION 4 | 1622 |
+| 2013 | spouse | completed_education | `ER58224` | COMPLETED ED-WF | 1877 |
+| 2013 | spouse | birth_state | `ER57541` | K33 STATE WIFE WAS BORN | 1618 |
+| 2013 | spouse | year_came | `ER57542` | K33YR YEAR CAME TO UNITED STATES-WF | 1619 |
+| 2015 | family | interview | `ER60002` | 2015 FAMILY INTERVIEW (ID) NUMBER | — |
+| 2015 | head | hispanic | `ER64809` | L39 SPANISH DESCENT-HEAD | 1771 |
+| 2015 | head | race mention 1 | `ER64810` | L40 RACE OF HEAD-MENTION 1 | 1771 |
+| 2015 | head | race mention 2 | `ER64811` | L40 RACE OF HEAD-MENTION 2 | 1772 |
+| 2015 | head | race mention 3 | `ER64812` | L40 RACE OF HEAD-MENTION 3 | 1772 |
+| 2015 | head | race mention 4 | `ER64813` | L40 RACE OF HEAD-MENTION 4 | 1772 |
+| 2015 | head | completed_education | `ER65459` | COMPLETED ED-HD | 2017 |
+| 2015 | head | birth_state | `ER64802` | L33 STATE HEAD WAS BORN | 1769 |
+| 2015 | head | year_came | `ER64803` | L33YR YEAR CAME TO UNITED STATES-HD | 1769 |
+| 2015 | spouse | hispanic | `ER64670` | K39 SPANISH DESCENT-SPOUSE | 1679 |
+| 2015 | spouse | race mention 1 | `ER64671` | K40 RACE OF SPOUSE-MENTION 1 | 1679 |
+| 2015 | spouse | race mention 2 | `ER64672` | K40 RACE OF SPOUSE-MENTION 2 | 1680 |
+| 2015 | spouse | race mention 3 | `ER64673` | K40 RACE OF SPOUSE-MENTION 3 | 1680 |
+| 2015 | spouse | race mention 4 | `ER64674` | K40 RACE OF SPOUSE-MENTION 4 | 1680 |
+| 2015 | spouse | completed_education | `ER65460` | COMPLETED ED-SP | 2018 |
+| 2015 | spouse | birth_state | `ER64663` | K33 STATE SPOUSE WAS BORN | 1677 |
+| 2015 | spouse | year_came | `ER64664` | K33YR YEAR CAME TO UNITED STATES-SP | 1677 |
+| 2017 | family | interview | `ER66002` | 2017 FAMILY INTERVIEW (ID) NUMBER | — |
+| 2017 | head | hispanic | `ER70881` | L39 SPANISH DESCENT-RP | 1801 |
+| 2017 | head | race mention 1 | `ER70882` | L40 RACE OF REFERENCE PERSON-MENTION 1 | 1802 |
+| 2017 | head | race mention 2 | `ER70883` | L40 RACE OF REFERENCE PERSON-MENTION 2 | 1802 |
+| 2017 | head | race mention 3 | `ER70884` | L40 RACE OF REFERENCE PERSON-MENTION 3 | 1802 |
+| 2017 | head | race mention 4 | `ER70885` | L40 RACE OF REFERENCE PERSON-MENTION 4 | 1803 |
+| 2017 | head | completed_education | `ER71538` | COMPLETED ED-RP | 2071 |
+| 2017 | head | birth_state | `ER70874` | L33 STATE REFERENCE PERSON WAS BORN | 1799 |
+| 2017 | head | year_came | `ER70875` | L33YR YEAR CAME TO UNITED STATES-RP | 1799 |
+| 2017 | spouse | hispanic | `ER70743` | K39 SPANISH DESCENT-SPOUSE | 1709 |
+| 2017 | spouse | race mention 1 | `ER70744` | K40 RACE OF SPOUSE-MENTION 1 | 1710 |
+| 2017 | spouse | race mention 2 | `ER70745` | K40 RACE OF SPOUSE-MENTION 2 | 1710 |
+| 2017 | spouse | race mention 3 | `ER70746` | K40 RACE OF SPOUSE-MENTION 3 | 1710 |
+| 2017 | spouse | race mention 4 | `ER70747` | K40 RACE OF SPOUSE-MENTION 4 | 1711 |
+| 2017 | spouse | completed_education | `ER71539` | COMPLETED ED-SP | 2072 |
+| 2017 | spouse | birth_state | `ER70736` | K33 STATE SPOUSE WAS BORN | 1708 |
+| 2017 | spouse | year_came | `ER70737` | K33YR YEAR CAME TO UNITED STATES-SP | 1708 |
+| 2019 | family | interview | `ER72002` | 2019 FAMILY INTERVIEW (ID) NUMBER | — |
+| 2019 | head | hispanic | `ER76896` | L39 SPANISH DESCENT-RP | 1792 |
+| 2019 | head | race mention 1 | `ER76897` | L40 RACE OF REFERENCE PERSON-MENTION 1 | 1793 |
+| 2019 | head | race mention 2 | `ER76898` | L40 RACE OF REFERENCE PERSON-MENTION 2 | 1793 |
+| 2019 | head | race mention 3 | `ER76899` | L40 RACE OF REFERENCE PERSON-MENTION 3 | 1793 |
+| 2019 | head | race mention 4 | `ER76900` | L40 RACE OF REFERENCE PERSON-MENTION 4 | 1794 |
+| 2019 | head | completed_education | `ER77599` | COMPLETED ED-RP | 2056 |
+| 2019 | head | birth_state | `ER76889` | L33 STATE REFERENCE PERSON WAS BORN | 1791 |
+| 2019 | head | year_came | `ER76890` | L33YR YEAR CAME TO UNITED STATES-RP | 1791 |
+| 2019 | spouse | hispanic | `ER76751` | K39 SPANISH DESCENT-SPOUSE | 1707 |
+| 2019 | spouse | race mention 1 | `ER76752` | K40 RACE OF SPOUSE-MENTION 1 | 1708 |
+| 2019 | spouse | race mention 2 | `ER76753` | K40 RACE OF SPOUSE-MENTION 2 | 1708 |
+| 2019 | spouse | race mention 3 | `ER76754` | K40 RACE OF SPOUSE-MENTION 3 | 1708 |
+| 2019 | spouse | race mention 4 | `ER76755` | K40 RACE OF SPOUSE-MENTION 4 | 1709 |
+| 2019 | spouse | completed_education | `ER77600` | COMPLETED ED-SP | 2057 |
+| 2019 | spouse | birth_state | `ER76744` | K33 STATE SPOUSE WAS BORN | 1706 |
+| 2019 | spouse | year_came | `ER76745` | K33YR YEAR CAME TO UNITED STATES-SP | 1706 |
+| 2021 | family | interview | `ER78002` | 2021 FAMILY INTERVIEW (ID) NUMBER | — |
+| 2021 | head | hispanic | `ER81143` | L39 SPANISH DESCENT-RP | 1123 |
+| 2021 | head | race mention 1 | `ER81144` | L40 RACE OF REFERENCE PERSON-MENTION 1 | 1124 |
+| 2021 | head | race mention 2 | `ER81145` | L40 RACE OF REFERENCE PERSON-MENTION 2 | 1124 |
+| 2021 | head | race mention 3 | `ER81146` | L40 RACE OF REFERENCE PERSON-MENTION 3 | 1124 |
+| 2021 | head | race mention 4 | `ER81147` | L40 RACE OF REFERENCE PERSON-MENTION 4 | 1125 |
+| 2021 | head | completed_education | `ER81926` | COMPLETED ED-RP | 1420 |
+| 2021 | head | birth_state | `ER81136` | L33 STATE REFERENCE PERSON WAS BORN | 1122 |
+| 2021 | head | year_came | `ER81137` | L33YR YEAR CAME TO UNITED STATES-RP | 1122 |
+| 2021 | spouse | hispanic | `ER81016` | K39 SPANISH DESCENT-SPOUSE | 1050 |
+| 2021 | spouse | race mention 1 | `ER81017` | K40 RACE OF SPOUSE-MENTION 1 | 1050 |
+| 2021 | spouse | race mention 2 | `ER81018` | K40 RACE OF SPOUSE-MENTION 2 | 1050 |
+| 2021 | spouse | race mention 3 | `ER81019` | K40 RACE OF SPOUSE-MENTION 3 | 1051 |
+| 2021 | spouse | race mention 4 | `ER81020` | K40 RACE OF SPOUSE-MENTION 4 | 1051 |
+| 2021 | spouse | completed_education | `ER81927` | COMPLETED ED-SP | 1421 |
+| 2021 | spouse | birth_state | `ER81009` | K33 STATE SPOUSE WAS BORN | 1048 |
+| 2021 | spouse | year_came | `ER81010` | K33YR YEAR CAME TO UNITED STATES-SP | 1048 |
+| 2023 | family | interview | `ER82002` | 2023 FAMILY INTERVIEW (ID) NUMBER | — |
+| 2023 | head | hispanic | `ER85120` | L39 SPANISH DESCENT-RP | 1114 |
+| 2023 | head | race mention 1 | `ER85121` | L40 RACE OF REFERENCE PERSON-MENTION 1 | 1115 |
+| 2023 | head | race mention 2 | `ER85122` | L40 RACE OF REFERENCE PERSON-MENTION 2 | 1115 |
+| 2023 | head | race mention 3 | `ER85123` | L40 RACE OF REFERENCE PERSON-MENTION 3 | 1115 |
+| 2023 | head | race mention 4 | `ER85124` | L40 RACE OF REFERENCE PERSON-MENTION 4 | 1116 |
+| 2023 | head | completed_education | `ER85780` | COMPLETED ED-RP | 1338 |
+| 2023 | head | birth_state | `ER85113` | L33 STATE REFERENCE PERSON WAS BORN | 1113 |
+| 2023 | head | year_came | `ER85114` | L33YR YEAR CAME TO UNITED STATES-RP | 1113 |
+| 2023 | spouse | hispanic | `ER84993` | K39 SPANISH DESCENT-SPOUSE | 1041 |
+| 2023 | spouse | race mention 1 | `ER84994` | K40 RACE OF SPOUSE-MENTION 1 | 1041 |
+| 2023 | spouse | race mention 2 | `ER84995` | K40 RACE OF SPOUSE-MENTION 2 | 1041 |
+| 2023 | spouse | race mention 3 | `ER84996` | K40 RACE OF SPOUSE-MENTION 3 | 1042 |
+| 2023 | spouse | race mention 4 | `ER84997` | K40 RACE OF SPOUSE-MENTION 4 | 1042 |
+| 2023 | spouse | completed_education | `ER85781` | COMPLETED ED-SP | 1339 |
+| 2023 | spouse | birth_state | `ER84986` | K33 STATE SPOUSE WAS BORN | 1039 |
+| 2023 | spouse | year_came | `ER84987` | K33YR YEAR CAME TO UNITED STATES-SP | 1039 |
diff --git a/scripts/build_group_attribute_codebook_values.py b/scripts/build_group_attribute_codebook_values.py
new file mode 100644
index 00000000..9ca9922f
--- /dev/null
+++ b/scripts/build_group_attribute_codebook_values.py
@@ -0,0 +1,214 @@
+"""Build the codebook value tables behind the group-attribute readers.
+
+For every family-file variable that
+:mod:`populace_dynamics.data.group_attributes_psid` reads (race and
+Spanish descent of head and wife 1985-2023, COMPLETED ED-HD/WF 1993-2023,
+state born and year came to the U.S. 2013-2023), this script extracts the
+"Value/Range Code" table from the wave's staged codebook PDF and writes
+``data/external/psid_group_attribute_codebook_values_v1.json``: each
+variable's label, the codebook page it sits on, and its documented codes
+and ranges with their text, plus the codebook PDF's path and SHA-256.
+
+It reads documentation only (codebook PDFs through ``pdftotext -layout``,
+and the label blocks of the ``.sps`` setup files); it opens no PSID data
+file. The readers load the table, pinned by SHA-256, to check every
+observed code; a test holds every documented code to a meaning in the
+module's code frames.
+
+Usage::
+
+ python scripts/build_group_attribute_codebook_values.py [--psid-dir D]
+"""
+
+from __future__ import annotations
+
+import argparse
+import hashlib
+import json
+import re
+import subprocess
+import sys
+from pathlib import Path
+
+from populace_dynamics.data import family, psid
+from populace_dynamics.data import group_attributes_psid as gap
+
+OUT = gap.CODEBOOK_VALUES_PATH
+
+_HEADER = re.compile(r'^(V\d+|ER\d+[A-Z0-9]*)\s+"([^"]*)"')
+_ROW = re.compile(
+ r"^\s*(?:[\d,]+|-)\s+(?:[\d.]+|-)\s+"
+ r"(-?[\d,]+(?:\s*-\s*-?[\d,]+)?)\s+(\S.*)$"
+)
+_SKIP = re.compile(
+ r"(^\s*Page \d+ of \d+\s*$)|(^\s*Filename\s*=)"
+ r"|(PANEL STUDY OF INCOME DYNAMICS)",
+ re.IGNORECASE,
+)
+
+
+def _sha256(path: Path) -> str:
+ return hashlib.sha256(path.read_bytes()).hexdigest()
+
+
+def _codebook_pdf(wave: int, root: Path) -> Path:
+ base = root / "family" / str(wave)
+ hits = sorted(
+ p
+ for p in base.iterdir()
+ if re.search(r"codebook.*\.pdf$", p.name, re.IGNORECASE)
+ )
+ if len(hits) != 1:
+ raise SystemExit(
+ f"wave {wave}: expected one codebook PDF in {base}, found "
+ f"{[p.name for p in hits]}"
+ )
+ return hits[0]
+
+
+def _pdftotext_version() -> str:
+ run = subprocess.run(
+ ["pdftotext", "-v"], capture_output=True, text=True, check=False
+ )
+ return (run.stderr or run.stdout).splitlines()[0].strip()
+
+
+def _parse_code(token: str) -> dict[str, object]:
+ token = token.replace(",", "").replace(" ", "")
+ match = re.fullmatch(r"(-?\d+)-(-?\d+)", token)
+ if match:
+ return {"range": [int(match.group(1)), int(match.group(2))]}
+ return {"code": int(token)}
+
+
+def extract_entry(text: str, variable: str) -> dict[str, object]:
+ """The label, page and value table of ``variable`` in codebook text.
+
+ ``text`` is ``pdftotext -layout`` output, whose pages are separated
+ by form feeds; the table is read from the variable's header to the
+ next variable header, across page breaks, skipping running headers
+ and footers and joining continuation lines to the row above.
+ """
+
+ pages = text.split("\f")
+ for page_number, page in enumerate(pages, start=1):
+ lines = page.splitlines()
+ for i, line in enumerate(lines):
+ header = _HEADER.match(line)
+ if not header or header.group(1) != variable:
+ continue
+ label = " ".join(header.group(2).split())
+ values: list[dict[str, object]] = []
+ in_table = False
+ rest = lines[i + 1 :] + [
+ ln
+ for later in pages[page_number:]
+ for ln in later.splitlines()
+ ]
+ for ln in rest:
+ if _HEADER.match(ln):
+ break
+ if "Value/Range Code" in ln:
+ in_table = True
+ continue
+ if not in_table or not ln.strip() or _SKIP.search(ln):
+ continue
+ row = _ROW.match(ln)
+ if row:
+ entry = _parse_code(row.group(1))
+ entry["text"] = " ".join(row.group(2).split())
+ values.append(entry)
+ elif values:
+ values[-1][
+ "text"
+ ] = f"{values[-1]['text']} {' '.join(ln.split())}"
+ if not values:
+ raise SystemExit(f"{variable}: empty value table")
+ return {"label": label, "page": page_number, "values": values}
+ raise SystemExit(f"{variable}: header not found in the codebook text")
+
+
+def build(root: Path) -> dict[str, object]:
+ """The full value table for every adjudicated family wave."""
+
+ out: dict[str, object] = {}
+ for wave, items in sorted(gap.FAMILY_ITEMS.items()):
+ pdf = _codebook_pdf(wave, root)
+ text = subprocess.run(
+ ["pdftotext", "-layout", str(pdf), "-"],
+ capture_output=True,
+ text=True,
+ check=True,
+ ).stdout
+ sps_path, _ = family._family_paths(wave, root)
+ sps_labels = psid.parse_sps_labels(sps_path)
+ variables: dict[str, object] = {}
+ for var, want in items.labels().items():
+ if var == items.interview[0]:
+ continue
+ entry = extract_entry(text, var)
+ if " ".join(want.split()) != entry["label"]:
+ raise SystemExit(
+ f"wave {wave} {var}: codebook label {entry['label']!r} "
+ f"differs from the adjudicated {want!r}"
+ )
+ if " ".join(sps_labels.get(var, "").split()) != entry["label"]:
+ raise SystemExit(
+ f"wave {wave} {var}: the .sps label differs from the "
+ "codebook's"
+ )
+ variables[var] = entry
+ formats = sorted((root / "family" / str(wave)).glob("*_formats.sas"))
+ if len(formats) > 1:
+ raise SystemExit(
+ f"wave {wave}: multiple SAS formats files in the source "
+ "documentation; adjudicate one before building"
+ )
+ out[str(wave)] = {
+ "codebook": {
+ "path": str(pdf.relative_to(root)),
+ "sha256": _sha256(pdf),
+ },
+ "formats_sas": (
+ {
+ "path": str(formats[0].relative_to(root)),
+ "sha256": _sha256(formats[0]),
+ }
+ if len(formats) == 1
+ else None
+ ),
+ "variables": variables,
+ }
+ return {
+ "schema_version": "psid_group_attribute_codebook.v1",
+ "generated_by": "scripts/build_group_attribute_codebook_values.py",
+ "extraction": (
+ f"{_pdftotext_version()} -layout; value rows are the codebook's "
+ "'Count / % / Value/Range Code / Value/Range Text' table with "
+ "continuation lines joined (counts and percents dropped); page "
+ "is the 1-based PDF page of the variable header"
+ ),
+ "note": (
+ "Documentation only: labels, value codes and their text from "
+ "the staged PSID family codebook PDFs. No data value or count "
+ "is recorded."
+ ),
+ "family": out,
+ }
+
+
+def main(argv: list[str] | None = None) -> int:
+ parser = argparse.ArgumentParser(description=__doc__)
+ parser.add_argument("--psid-dir", type=Path, default=None)
+ parser.add_argument("--out", type=Path, default=OUT)
+ args = parser.parse_args(argv)
+ root = psid._resolve_data_dir(args.psid_dir)
+ table = build(root)
+ payload = json.dumps(table, indent=1, sort_keys=True) + "\n"
+ args.out.write_text(payload)
+ print(args.out, hashlib.sha256(payload.encode()).hexdigest())
+ return 0
+
+
+if __name__ == "__main__":
+ sys.exit(main())
diff --git a/src/populace_dynamics/cohorts/group_attributes.py b/src/populace_dynamics/cohorts/group_attributes.py
new file mode 100644
index 00000000..8d4aecf8
--- /dev/null
+++ b/src/populace_dynamics/cohorts/group_attributes.py
@@ -0,0 +1,1102 @@
+"""Group attributes for the MINT breakdowns: a side frame by person_id.
+
+The four blind tests' populations (the psid2010 cohorts of the 2011 and
+2009 anchor waves, the age-67 observations of exercise 2, and Track M's
+2023 universe) carry sex, age and marital status but not race and
+ethnicity, education or country of birth. This module adds those three
+as a separate frame keyed by ``person_id``, with its own file audit and
+SHA-256 seal, so the existing cohorts, their frames and the digests the
+committed runs pin are untouched (the U2 "neither edits nor extends"
+precedent; the readers are
+:mod:`populace_dynamics.data.group_attributes_psid`).
+
+Resolution rules (``RULES_VERSION``), decided 2026-10-01:
+
+* **Race and Hispanic origin** are fixed person attributes, resolved from
+ every staged head/wife report 1985-2023 whatever the anchor. Each
+ report is harmonized through its wave's code frame
+ (:func:`~populace_dynamics.data.group_attributes_psid.race_ethnicity_report`).
+ A report is *complete* when it is Hispanic, or not Hispanic with every
+ race mention known. The person's value is their **most recent complete
+ report** (the latest self-identification; a person is head or wife at
+ most once per wave). ``race_ethnicity_n_distinct`` counts the distinct
+ four-way readings among their complete reports, so conflicts are
+ visible. ``hispanic`` comes from the selected report; a person with no
+ complete report takes it from their most recent report that shows it.
+* **Country of birth** is likewise fixed and comes from the most recent
+ head/wife report 2013-2023 that classifies it (United States, U.S.
+ territory or foreign country); DK/NA and the few pairs that contradict
+ the skip pattern (``inconsistent``) never decide it.
+* **Education** can change, so it is resolved **as of the population's
+ last anchor wave** (``education_cutoff_wave = max(anchor_waves)``):
+ the most recent reported years of schooling at or before it, from the
+ individual-file series (heads, wives and OFUMs aged 16 or older), with
+ "completed no grades" (0 years) for a head or wife whose family-file
+ code says so.
+* **No value is guessed.** A person never head or wife in the window has
+ ``race_ethnicity_status == "never_head_or_spouse"`` (OFUMs; the PSID
+ asks race only of heads and wives), and likewise for country of
+ birth; other unknowns name their reason. Every requested person gets
+ exactly one row, and the provenance counts each status.
+
+Report categories come from the data-driven schemes in
+``data/external/group_category_schemes_v1.json`` (SSA's MINT 8 Table
+User Guide; Butrica and Uccello 2004's rows as this repository records
+them): each scheme column has a ``_status`` of ``assigned``,
+``attribute_unknown`` or ``unresolved:`` for values the source
+does not place (a U.S.-territory birth for MINT 8; the Report's education
+rows, whose numerical definitions are absent from the cleared sources).
+
+This module computes attributes only. It opens no outcome and computes
+no group share; tabulation code receives its frame.
+"""
+
+from __future__ import annotations
+
+import hashlib
+import json
+from collections.abc import Iterable, Mapping
+from copy import deepcopy
+from dataclasses import dataclass, field
+from numbers import Integral
+from pathlib import Path
+from typing import Any
+
+import numpy as np
+import pandas as pd
+
+from populace_dynamics.cohorts import psid2010
+from populace_dynamics.data import group_attributes_psid as gap
+from populace_dynamics.data import psid
+
+__all__ = [
+ "RULES_VERSION",
+ "SCHEMES_PATH",
+ "SCHEMES_SHA256",
+ "FRAME_COLUMNS",
+ "GroupAttributeInputs",
+ "GroupAttributes",
+ "load_schemes",
+ "load_group_attribute_inputs",
+ "input_frames_sha256",
+ "build_group_attributes",
+ "load_group_attributes",
+ "content_sha256",
+ "apply_category_scheme",
+]
+
+RULES_VERSION = "g1-rules-1"
+
+_REPO_ROOT = Path(__file__).resolve().parents[3]
+SCHEMES_PATH = (
+ _REPO_ROOT / "data" / "external" / "group_category_schemes_v1.json"
+)
+#: SHA-256 of the committed scheme file; a different file is refused.
+SCHEMES_SHA256 = (
+ "da15a940d85dc1b8ea49480ad5ae5d0c4179ff916571bc44c9d901e22c6772c0"
+)
+
+KNOWN = "known"
+NEVER_HEAD_OR_SPOUSE = "never_head_or_spouse"
+HISPANIC_ORIGIN_NOT_ASKED = "hispanic_origin_not_asked"
+DK_NA_REFUSED = "dk_na_refused"
+UNDOCUMENTED_ZERO_MEANING = "undocumented_zero_meaning"
+INCONSISTENT_ONLY = "inconsistent"
+NO_REPORT_BY_CUTOFF = "no_report_by_cutoff"
+ASSIGNED = "assigned"
+ATTRIBUTE_UNKNOWN = "attribute_unknown"
+
+#: The four-way race/ethnicity codes every race scheme maps.
+RACE_ETHNICITY_CODES: tuple[str, ...] = (
+ "hispanic",
+ "white_non_hispanic",
+ "black_non_hispanic",
+ "other_non_hispanic",
+)
+#: The country-of-birth values a country scheme maps.
+COUNTRY_OF_BIRTH_CODES: tuple[str, ...] = (
+ gap.UNITED_STATES,
+ gap.US_TERRITORY,
+ gap.FOREIGN_COUNTRY,
+)
+_EDUCATION_DOMAIN = (0, gap.EDUCATION_YEARS_RANGE[1])
+
+_ATTRIBUTE_COLUMNS: dict[str, str] = {
+ "person_id": "int64",
+ "hispanic": "boolean",
+ "race_mentions": "string",
+ "race_ethnicity_status": "string",
+ "race_ethnicity_source_wave": "Int64",
+ "race_ethnicity_source_role": "string",
+ "race_ethnicity_source_variables": "string",
+ "race_ethnicity_source_mentions": "string",
+ "race_ethnicity_hispanic_basis": "string",
+ "race_ethnicity_n_reports": "Int64",
+ "race_ethnicity_n_distinct": "Int64",
+ "hispanic_source_wave": "Int64",
+ "hispanic_source_role": "string",
+ "hispanic_source_variables": "string",
+ "hispanic_source_mentions": "string",
+ "education_years": "Int64",
+ "education_status": "string",
+ "education_source_wave": "Int64",
+ "education_source_variables": "string",
+ "education_source_role": "string",
+ "education_source_mentions": "string",
+ "education_n_reports": "Int64",
+ "education_n_distinct": "Int64",
+ "country_of_birth": "string",
+ "country_of_birth_status": "string",
+ "country_of_birth_source_wave": "Int64",
+ "country_of_birth_source_role": "string",
+ "country_of_birth_source_variables": "string",
+ "country_of_birth_source_mentions": "string",
+ "country_of_birth_n_reports": "Int64",
+ "country_of_birth_n_distinct": "Int64",
+}
+#: The scheme columns of the default schemes, each followed by its
+#: ``_status`` column, in frame order.
+_DEFAULT_SCHEME_COLUMNS: tuple[str, ...] = (
+ "race_ethnicity_mint8",
+ "education_mint8",
+ "country_of_birth_mint8",
+ "race_ethnicity_report4",
+ "education_report3",
+)
+FRAME_COLUMNS: tuple[str, ...] = (
+ *_ATTRIBUTE_COLUMNS,
+ *(
+ name
+ for column in _DEFAULT_SCHEME_COLUMNS
+ for name in (column, f"{column}_status")
+ ),
+)
+#: The attribute each scheme dimension reads.
+_DIMENSION_INPUT = {
+ "race_ethnicity": "race_ethnicity",
+ "education": "education_years",
+ "country_of_birth": "country_of_birth",
+}
+
+
+# --------------------------------------------------------------------------
+# Schemes
+# --------------------------------------------------------------------------
+def load_schemes(path: Path | None = None) -> dict[str, Any]:
+ """Load and validate the category schemes, refusing a changed file.
+
+ With ``path`` given (tests), the SHA-256 pin is not applied, but the
+ file must still pass :func:`_validate_schemes`.
+ """
+
+ target = SCHEMES_PATH if path is None else Path(path)
+ raw = target.read_bytes()
+ if path is None:
+ digest = hashlib.sha256(raw).hexdigest()
+ if digest != SCHEMES_SHA256:
+ raise ValueError(
+ f"{target} has SHA-256 {digest}, expected the pinned "
+ f"{SCHEMES_SHA256}"
+ )
+ data = json.loads(raw)
+ if data.get("schema_version") != "group_category_schemes.v1":
+ raise ValueError(f"{target} is not a v1 scheme file")
+ _validate_schemes(data)
+ if path is None:
+ for source in data["sources"].values():
+ if "committed_file" not in source:
+ continue
+ source_path = _REPO_ROOT / source["committed_file"]
+ digest = hashlib.sha256(source_path.read_bytes()).hexdigest()
+ if digest != source["sha256"]:
+ raise ValueError(f"scheme source {source_path} changed")
+ return data
+
+
+def _validate_schemes(data: Mapping[str, Any]) -> None:
+ """Refuse a scheme that leaves a value neither placed nor unresolved."""
+
+ columns: set[str] = set()
+ for scheme_id, scheme in data["schemes"].items():
+ for dimension, spec in scheme["dimensions"].items():
+ where = f"scheme {scheme_id} {dimension}"
+ if dimension not in _DIMENSION_INPUT:
+ raise ValueError(f"{where}: unknown dimension")
+ column = spec["column"]
+ new_columns = {column, f"{column}_status"}
+ if new_columns & (columns | set(_ATTRIBUTE_COLUMNS)):
+ raise ValueError(f"{where}: column {column} is taken")
+ columns.update(new_columns)
+ if dimension == "education":
+ covered: list[int] = []
+ for band in [*spec["bands"], *spec["unresolved_bands"]]:
+ hi = (
+ _EDUCATION_DOMAIN[1]
+ if band["max"] is None
+ else band["max"]
+ )
+ covered.extend(range(band["min"], hi + 1))
+ want = list(
+ range(_EDUCATION_DOMAIN[0], _EDUCATION_DOMAIN[1] + 1)
+ )
+ if sorted(covered) != want:
+ raise ValueError(
+ f"{where}: bands must cover each of years "
+ f"{want[0]}-{want[-1]} exactly once"
+ )
+ continue
+ codes = (
+ RACE_ETHNICITY_CODES
+ if dimension == "race_ethnicity"
+ else COUNTRY_OF_BIRTH_CODES
+ )
+ placed = set(spec["categories"]) | set(spec.get("unresolved", {}))
+ if placed != set(codes) or set(spec["categories"]) & set(
+ spec.get("unresolved", {})
+ ):
+ raise ValueError(
+ f"{where}: categories and unresolved must partition "
+ f"{codes}"
+ )
+ if dimension == "race_ethnicity" and spec[
+ "multiple_races"
+ ] not in (*RACE_ETHNICITY_CODES, "first_mention"):
+ raise ValueError(f"{where}: bad multiple_races rule")
+
+
+def _race_code(
+ hispanic: Any, races: str | None, multiple_races: str
+) -> str | None:
+ """The four-way code of one person's selected report, or None."""
+
+ if hispanic is pd.NA or hispanic is None or pd.isna(hispanic):
+ return None
+ if not isinstance(hispanic, (bool, np.bool_)):
+ raise ValueError("hispanic must be a boolean or unknown")
+ if bool(hispanic):
+ return "hispanic"
+ if races is None or races is pd.NA or pd.isna(races):
+ return None
+ mentions = races.split("|")
+ if any(mention not in gap.RACE_CATEGORIES for mention in mentions):
+ raise ValueError("race_mentions contains undocumented categories")
+ if len(mentions) > 1 or gap.MORE_THAN_TWO in mentions:
+ if multiple_races != "first_mention":
+ return multiple_races
+ mentions = mentions[:1]
+ single = mentions[0]
+ if single == gap.WHITE:
+ return "white_non_hispanic"
+ if single == gap.BLACK:
+ return "black_non_hispanic"
+ return "other_non_hispanic"
+
+
+def apply_category_scheme(
+ frame: pd.DataFrame, dimension: str, spec: Mapping[str, Any]
+) -> tuple[pd.Series, pd.Series]:
+ """Map one attribute to a scheme's labels: ``(labels, status)``.
+
+ ``frame`` holds the attribute columns (``hispanic`` and
+ ``race_mentions`` for race and ethnicity, ``education_years``,
+ ``country_of_birth``). A known value the scheme places gets its label
+ and ``assigned``; one it leaves unresolved gets NA and
+ ``unresolved:``; an unknown attribute gets NA and
+ ``attribute_unknown``.
+ """
+
+ n = len(frame)
+ labels: list[Any] = [pd.NA] * n
+ status: list[str] = [ATTRIBUTE_UNKNOWN] * n
+ if dimension == "race_ethnicity":
+ values = [
+ _race_code(h, r, spec["multiple_races"])
+ for h, r in zip(
+ frame["hispanic"], frame["race_mentions"], strict=True
+ )
+ ]
+ for i, code in enumerate(values):
+ if code is None:
+ continue
+ if code in spec["categories"]:
+ labels[i], status[i] = spec["categories"][code], ASSIGNED
+ elif code in spec.get("unresolved", {}):
+ status[i] = f"unresolved:{code}"
+ else:
+ raise ValueError(f"race and ethnicity {code!r} is not mapped")
+ elif dimension == "country_of_birth":
+ unresolved = spec.get("unresolved", {})
+ for i, value in enumerate(frame["country_of_birth"]):
+ if value is pd.NA or pd.isna(value):
+ continue
+ if value in spec["categories"]:
+ labels[i], status[i] = spec["categories"][value], ASSIGNED
+ elif value in unresolved:
+ status[i] = f"unresolved:{value}"
+ else:
+ raise ValueError(f"country of birth {value!r} is not mapped")
+ elif dimension == "education":
+ bands = [
+ (b["min"], b["max"], b["label"], None) for b in spec["bands"]
+ ] + [
+ (b["min"], b["max"], None, b["status"])
+ for b in spec["unresolved_bands"]
+ ]
+ for i, years in enumerate(frame["education_years"]):
+ if years is pd.NA or pd.isna(years):
+ continue
+ if isinstance(years, (bool, np.bool_)) or not isinstance(
+ years, Integral
+ ):
+ raise ValueError("education_years must be integers")
+ y = int(years)
+ if not _EDUCATION_DOMAIN[0] <= y <= _EDUCATION_DOMAIN[1]:
+ raise ValueError(f"education_years {y} is undocumented")
+ hits = [
+ b for b in bands if b[0] <= y and (b[1] is None or y <= b[1])
+ ]
+ if len(hits) != 1:
+ raise ValueError(f"{y} years fall in {len(hits)} bands")
+ _, _, label, reason = hits[0]
+ if label is None:
+ status[i] = f"unresolved:{reason}"
+ else:
+ labels[i], status[i] = label, ASSIGNED
+ else:
+ raise ValueError(f"unknown dimension {dimension!r}")
+ return (
+ pd.Series(labels, index=frame.index, dtype="string"),
+ pd.Series(status, index=frame.index, dtype="string"),
+ )
+
+
+# --------------------------------------------------------------------------
+# Inputs
+# --------------------------------------------------------------------------
+@dataclass(frozen=True)
+class GroupAttributeInputs:
+ """Raw, documented codes the pure builder consumes.
+
+ ``reports``: one row per present head or spouse per wave
+ (:func:`~populace_dynamics.data.group_attributes_psid.head_spouse_reports`).
+ ``education``: person-waves with an individual education code other
+ than 0, plus present heads and spouses with code 0 (who may have
+ completed no grades), with their role's family COMPLETED ED code.
+ ``universe``: every person in the individual file (sorted).
+ """
+
+ reports: pd.DataFrame
+ education: pd.DataFrame
+ universe: pd.Series
+ provenance: Mapping[str, Any] = field(default_factory=dict)
+
+
+def _education_rows(
+ individual: pd.DataFrame, reports: pd.DataFrame
+) -> pd.DataFrame:
+ roles = reports[["person_id", "wave", "role", "family_education_code"]]
+ merged = individual[["person_id", "wave", "education_code"]].merge(
+ roles, on=["person_id", "wave"], how="left"
+ )
+ keep = (merged["education_code"] != 0) | merged["role"].notna()
+ out = merged.loc[keep].copy()
+ out["role"] = out["role"].astype("string")
+ out["family_education_code"] = out["family_education_code"].astype("Int64")
+ return out.sort_values(["person_id", "wave"]).reset_index(drop=True)
+
+
+def _frame_digest(digest: Any, name: str, frame: pd.DataFrame) -> None:
+ digest.update(f"{name}\n".encode())
+ digest.update(json.dumps([str(c) for c in frame.columns]).encode())
+ digest.update(json.dumps([str(t) for t in frame.dtypes]).encode())
+ digest.update(frame.to_csv(index=False, lineterminator="\n").encode())
+
+
+def input_frames_sha256(inputs: GroupAttributeInputs) -> str:
+ """SHA-256 of the inputs' frames (reports, education, universe)."""
+
+ digest = hashlib.sha256()
+ _frame_digest(digest, "reports", inputs.reports)
+ _frame_digest(digest, "education", inputs.education)
+ _frame_digest(digest, "universe", inputs.universe.to_frame("person_id"))
+ return digest.hexdigest()
+
+
+def _mapping_sha256(values: Mapping[str, str]) -> str:
+ encoded = (
+ json.dumps(dict(values), sort_keys=True, separators=(",", ":")) + "\n"
+ ).encode()
+ return hashlib.sha256(encoded).hexdigest()
+
+
+def load_group_attribute_inputs(
+ *, psid_dir: Path | None = None
+) -> GroupAttributeInputs:
+ """Read every input from the staged PSID, recording file hashes.
+
+ Inside :func:`populace_dynamics.cohorts.psid2010.record_files_read`:
+ the codebook PDFs' SHA-256 against the value table's pins, the
+ individual file's roster and education series (labels and formats
+ verified), and each 1985-2023 family file's race, Spanish-descent,
+ education and birthplace items (labels and codes verified).
+ """
+
+ root = psid._resolve_data_dir(psid_dir)
+ codebook = gap.load_codebook_values()
+ with psid2010.record_files_read(root) as files:
+ pins = gap.verify_codebook_pins(
+ gap.RACE_WAVES, data_dir=psid_dir, codebook=codebook
+ )
+ individual = gap.read_individual_items(data_dir=psid_dir)
+ family_items = {
+ wave: gap.read_family_items(
+ wave, data_dir=psid_dir, codebook=codebook
+ )
+ for wave in gap.RACE_WAVES
+ }
+ reports = gap.head_spouse_reports(individual, family_items)
+ education = _education_rows(individual, reports)
+ universe = pd.Series(
+ np.sort(individual["person_id"].unique()), name="person_id"
+ )
+ inputs = GroupAttributeInputs(
+ reports=reports,
+ education=education,
+ universe=universe,
+ )
+ provenance = {
+ "kind": "psid_files",
+ "psid_files_sha256": dict(files),
+ "psid_files_bundle_sha256": _mapping_sha256(files),
+ "codebook_values_sha256": gap.CODEBOOK_VALUES_SHA256,
+ "codebook_pdf_sha256": {str(w): d for w, d in pins.items()},
+ "input_frames_sha256": input_frames_sha256(inputs),
+ }
+ return GroupAttributeInputs(
+ reports=reports,
+ education=education,
+ universe=universe,
+ provenance=provenance,
+ )
+
+
+# --------------------------------------------------------------------------
+# Resolution
+# --------------------------------------------------------------------------
+def _race_report_columns(reports: pd.DataFrame) -> pd.DataFrame:
+ """Harmonize every head/spouse report (cached by code tuple)."""
+
+ cache: dict[tuple, gap.RaceEthnicityReport] = {}
+ out = []
+ code_columns = ["race_code_1", "race_code_2", "race_code_3", "race_code_4"]
+ for row in reports[
+ ["wave", "role", "hispanic_code", *code_columns]
+ ].itertuples(index=False):
+ wave = int(row[0])
+ role = str(row[1])
+ hisp = None if pd.isna(row[2]) else int(row[2])
+ codes = tuple(None if pd.isna(c) else int(c) for c in row[3:])
+ key = (wave, role, hisp, codes)
+ if key not in cache:
+ cache[key] = gap.race_ethnicity_report(
+ wave, hisp, codes, role=role
+ )
+ out.append(cache[key])
+ frame = reports[["person_id", "wave", "role"]].copy()
+ frame["complete"] = pd.array([r.complete for r in out], dtype="bool")
+ frame["hispanic"] = pd.array([r.hispanic for r in out], dtype="boolean")
+ frame["races"] = pd.array(
+ [None if r.races is None else "|".join(r.races) for r in out],
+ dtype="string",
+ )
+ frame["basis"] = pd.array([r.hispanic_basis for r in out], dtype="string")
+ frame["reading"] = pd.array(
+ [
+ (
+ None
+ if not r.complete
+ else _race_code(
+ r.hispanic, "|".join(r.races or ()), "other_non_hispanic"
+ )
+ )
+ for r in out
+ ],
+ dtype="string",
+ )
+ variables = []
+ mentions = []
+ hispanic_variables = []
+ hispanic_mentions = []
+ for row, report in zip(
+ reports[["wave", "role", *code_columns]].itertuples(index=False),
+ out,
+ strict=True,
+ ):
+ items = gap.FAMILY_ITEMS[int(row[0])]
+ role = str(row[1])
+ names = []
+ used_mentions = []
+ hisp_names = []
+ hisp_mentions = []
+ if items.asks_hispanic_origin:
+ names.append(items.hispanic[role][0])
+ hisp_names.append(items.hispanic[role][0])
+ hisp_mentions.append("direct_question")
+ for i, ((var, _), code) in enumerate(
+ zip(items.race[role], row[2:], strict=False), start=1
+ ):
+ if pd.isna(code):
+ continue
+ meaning = gap.race_mention_meaning(int(row[0]), i, int(code))
+ if meaning != gap.NO_FURTHER_MENTION:
+ names.append(var)
+ used_mentions.append(str(i))
+ if (
+ report.hispanic_basis == "latino_race_mention"
+ and meaning == gap.LATINO_ORIGIN
+ ):
+ hisp_names.append(var)
+ hisp_mentions.append(str(i))
+ variables.append("|".join(names))
+ mentions.append("|".join(used_mentions))
+ hispanic_variables.append("|".join(hisp_names))
+ hispanic_mentions.append("|".join(hisp_mentions))
+ frame["variables"] = pd.array(variables, dtype="string")
+ frame["mentions"] = pd.array(mentions, dtype="string")
+ frame["hispanic_variables"] = pd.array(hispanic_variables, dtype="string")
+ frame["hispanic_mentions"] = pd.array(hispanic_mentions, dtype="string")
+ return frame
+
+
+def _latest(frame: pd.DataFrame) -> pd.DataFrame:
+ """Each person's most recent row (a person has one row per wave)."""
+
+ ordered = frame.sort_values(["person_id", "wave"])
+ return (
+ ordered.groupby("person_id", sort=True).tail(1).set_index("person_id")
+ )
+
+
+def _resolve_race(persons: pd.Index, reports: pd.DataFrame) -> pd.DataFrame:
+ harmonized = _race_report_columns(reports)
+ complete = harmonized[harmonized["complete"]]
+ selected = _latest(complete)
+ out = pd.DataFrame(index=persons)
+ out["hispanic"] = selected["hispanic"].reindex(persons)
+ out["race_mentions"] = selected["races"].reindex(persons)
+ out["race_ethnicity_source_wave"] = selected["wave"].reindex(persons)
+ out["race_ethnicity_source_role"] = selected["role"].reindex(persons)
+ out["race_ethnicity_source_variables"] = selected["variables"].reindex(
+ persons
+ )
+ out["race_ethnicity_source_mentions"] = selected["mentions"].reindex(
+ persons
+ )
+ out["race_ethnicity_hispanic_basis"] = selected["basis"].reindex(persons)
+ out["race_ethnicity_n_reports"] = (
+ complete.groupby("person_id").size().reindex(persons, fill_value=0)
+ )
+ out["race_ethnicity_n_distinct"] = (
+ complete.groupby("person_id")["reading"]
+ .nunique()
+ .reindex(persons, fill_value=0)
+ )
+ # Hispanic origin for persons with no complete report: their most
+ # recent report that shows it.
+ shown = harmonized[harmonized["hispanic"].notna()]
+ fallback = _latest(shown)
+ out["hispanic_source_wave"] = out["race_ethnicity_source_wave"]
+ for suffix, source in (
+ ("role", "role"),
+ ("variables", "hispanic_variables"),
+ ("mentions", "hispanic_mentions"),
+ ):
+ out[f"hispanic_source_{suffix}"] = selected[source].reindex(persons)
+ lacking = persons[out["hispanic"].isna().to_numpy()]
+ out.loc[lacking, "hispanic"] = fallback["hispanic"].reindex(lacking)
+ out.loc[lacking, "hispanic_source_wave"] = fallback["wave"].reindex(
+ lacking
+ )
+ for suffix, source in (
+ ("role", "role"),
+ ("variables", "hispanic_variables"),
+ ("mentions", "hispanic_mentions"),
+ ):
+ out.loc[lacking, f"hispanic_source_{suffix}"] = fallback[
+ source
+ ].reindex(lacking)
+ # Status.
+ has_any = (
+ harmonized.groupby("person_id").size().reindex(persons, fill_value=0)
+ )
+ asked = harmonized[harmonized["basis"] != "not_asked"]
+ has_asked = (
+ asked.groupby("person_id").size().reindex(persons, fill_value=0)
+ )
+ status = pd.Series(DK_NA_REFUSED, index=persons, dtype="string")
+ status[has_any == 0] = NEVER_HEAD_OR_SPOUSE
+ status[(has_any > 0) & (has_asked == 0)] = HISPANIC_ORIGIN_NOT_ASKED
+ latest_report = _latest(harmonized)
+ ambiguous = (
+ latest_report["basis"].reindex(persons).eq(UNDOCUMENTED_ZERO_MEANING)
+ )
+ status[ambiguous] = UNDOCUMENTED_ZERO_MEANING
+ status[out["race_ethnicity_n_reports"] > 0] = KNOWN
+ out["race_ethnicity_status"] = status
+ return out
+
+
+def _resolve_country(persons: pd.Index, reports: pd.DataFrame) -> pd.DataFrame:
+ window = reports[reports["wave"].isin(gap.BIRTHPLACE_WAVES)]
+ classes = [
+ gap.birthplace_class(int(s), int(y))
+ for s, y in zip(
+ window["birth_state_code"], window["year_came_code"], strict=True
+ )
+ ]
+ frame = window[["person_id", "wave", "role"]].copy()
+ frame["class"] = pd.array(classes, dtype="string")
+ frame["variables"] = pd.array(
+ [
+ f"{gap.FAMILY_ITEMS[int(w)].birth_state[r][0]}|"
+ f"{gap.FAMILY_ITEMS[int(w)].year_came[r][0]}"
+ for w, r in zip(frame["wave"], frame["role"], strict=True)
+ ],
+ dtype="string",
+ )
+ valid = frame[frame["class"].isin(COUNTRY_OF_BIRTH_CODES)]
+ selected = _latest(valid)
+ out = pd.DataFrame(index=persons)
+ out["country_of_birth"] = selected["class"].reindex(persons)
+ out["country_of_birth_source_wave"] = selected["wave"].reindex(persons)
+ out["country_of_birth_source_role"] = selected["role"].reindex(persons)
+ out["country_of_birth_source_variables"] = selected["variables"].reindex(
+ persons
+ )
+ out["country_of_birth_source_mentions"] = pd.Series(
+ "birth_state|year_came", index=persons, dtype="string"
+ ).where(out["country_of_birth_source_wave"].notna())
+ out["country_of_birth_n_reports"] = (
+ valid.groupby("person_id").size().reindex(persons, fill_value=0)
+ )
+ out["country_of_birth_n_distinct"] = (
+ valid.groupby("person_id")["class"]
+ .nunique()
+ .reindex(persons, fill_value=0)
+ )
+ has_any = frame.groupby("person_id").size().reindex(persons, fill_value=0)
+ has_missing = (
+ frame[frame["class"] == gap.MISSING]
+ .groupby("person_id")
+ .size()
+ .reindex(persons, fill_value=0)
+ )
+ status = pd.Series(INCONSISTENT_ONLY, index=persons, dtype="string")
+ status[has_missing > 0] = DK_NA_REFUSED
+ status[has_any == 0] = NEVER_HEAD_OR_SPOUSE
+ status[out["country_of_birth_n_reports"] > 0] = KNOWN
+ out["country_of_birth_status"] = status
+ return out
+
+
+def _resolve_education(
+ persons: pd.Index, education: pd.DataFrame, cutoff: int
+) -> pd.DataFrame:
+ window = education[education["wave"] <= cutoff]
+ cache: dict[tuple, tuple[int | None, str]] = {}
+ years: list[Any] = []
+ statuses: list[str] = []
+ for code, fam in zip(
+ window["education_code"], window["family_education_code"], strict=True
+ ):
+ key = (int(code), None if pd.isna(fam) else int(fam))
+ if key not in cache:
+ cache[key] = gap.education_report(*key)
+ y, s = cache[key]
+ years.append(pd.NA if y is None else y)
+ statuses.append(s)
+ frame = window[["person_id", "wave"]].copy()
+ frame["years"] = pd.array(years, dtype="Int64")
+ frame["status"] = pd.array(statuses, dtype="string")
+ frame["variables"] = pd.array(
+ [
+ gap.INDIVIDUAL_ITEMS[int(w)].education[0]
+ + (
+ f"|{gap.FAMILY_ITEMS[int(w)].completed_education[str(r)][0]}"
+ if s == gap.EDUCATION_NO_GRADES
+ else ""
+ )
+ for w, r, s in zip(
+ frame["wave"],
+ window["role"].fillna(""),
+ statuses,
+ strict=True,
+ )
+ ],
+ dtype="string",
+ )
+ valid = frame[
+ frame["status"].isin([gap.EDUCATION_REPORTED, gap.EDUCATION_NO_GRADES])
+ ]
+ selected = _latest(valid)
+ out = pd.DataFrame(index=persons)
+ out["education_years"] = selected["years"].reindex(persons)
+ out["education_source_wave"] = selected["wave"].reindex(persons)
+ out["education_source_variables"] = selected["variables"].reindex(persons)
+ frame["role"] = window["role"].fillna("individual")
+ selected = _latest(frame.loc[valid.index])
+ out["education_source_role"] = selected["role"].reindex(persons)
+ out["education_source_mentions"] = (
+ selected["status"]
+ .map(
+ {
+ gap.EDUCATION_REPORTED: "individual",
+ gap.EDUCATION_NO_GRADES: "individual|family_recode",
+ }
+ )
+ .reindex(persons)
+ )
+ out["education_n_reports"] = (
+ valid.groupby("person_id").size().reindex(persons, fill_value=0)
+ )
+ out["education_n_distinct"] = (
+ valid.groupby("person_id")["years"]
+ .nunique()
+ .reindex(persons, fill_value=0)
+ )
+ has_missing = (
+ frame[frame["status"] == gap.EDUCATION_MISSING]
+ .groupby("person_id")
+ .size()
+ .reindex(persons, fill_value=0)
+ )
+ status = pd.Series(NO_REPORT_BY_CUTOFF, index=persons, dtype="string")
+ status[has_missing > 0] = DK_NA_REFUSED
+ status[out["education_n_reports"] > 0] = KNOWN
+ out["education_status"] = status
+ return out
+
+
+@dataclass(frozen=True)
+class GroupAttributes:
+ """The side frame (:data:`FRAME_COLUMNS` order, sorted by person_id)
+ and its provenance (rules, windows, status counts, input and file
+ digests, and the frame's ``content_sha256``)."""
+
+ frame: pd.DataFrame
+ provenance: Mapping[str, Any]
+
+
+def content_sha256(frame: pd.DataFrame) -> str:
+ """SHA-256 of the frame's columns, dtypes and CSV bytes."""
+
+ digest = hashlib.sha256()
+ _frame_digest(digest, "group_attributes", frame)
+ return digest.hexdigest()
+
+
+def _person_index(person_ids: Iterable[int], universe: pd.Series) -> pd.Index:
+ ids = pd.Series(list(person_ids), dtype="object")
+ if ids.empty:
+ raise ValueError("no person_ids given")
+ if any(
+ isinstance(p, (bool, np.bool_)) or not isinstance(p, Integral)
+ for p in ids
+ ):
+ raise ValueError("person_ids must be integers")
+ try:
+ ids = ids.astype("int64")
+ except (TypeError, ValueError, OverflowError) as exc:
+ raise ValueError("person_ids must be int64 integers") from exc
+ if ids.duplicated().any():
+ raise ValueError(
+ f"{int(ids.duplicated().sum())} duplicate person_ids; pass each "
+ "person once (the frame is keyed by person_id)"
+ )
+ unknown = sorted(set(ids) - set(universe))
+ if unknown:
+ raise ValueError(
+ f"{len(unknown)} person_ids are not in the individual file "
+ f"(first {unknown[:5]})"
+ )
+ return pd.Index(np.sort(ids.to_numpy()), name="person_id")
+
+
+def _anchor_waves(anchor_waves: Iterable[int]) -> tuple[int, ...]:
+ values = tuple(anchor_waves)
+ if any(
+ isinstance(w, (bool, np.bool_)) or not isinstance(w, Integral)
+ for w in values
+ ):
+ raise ValueError("anchor_waves must be integers")
+ waves = tuple(sorted({int(w) for w in values}))
+ if not waves:
+ raise ValueError("anchor_waves is empty")
+ bad = [w for w in waves if w not in gap.INDIVIDUAL_WAVES]
+ if bad:
+ raise ValueError(
+ f"anchor waves {bad} are not PSID waves "
+ f"{gap.INDIVIDUAL_WAVES[0]}-{gap.INDIVIDUAL_WAVES[-1]}"
+ )
+ return waves
+
+
+def _validate_inputs(inputs: GroupAttributeInputs) -> None:
+ """Require the documented person-wave schema before harmonization."""
+
+ if inputs.universe.duplicated().any():
+ raise ValueError("inputs.universe has duplicate person_ids")
+ for name, required in (
+ ("reports", {"person_id", "wave", "role", *gap.REPORT_CODE_COLUMNS}),
+ (
+ "education",
+ {
+ "person_id",
+ "wave",
+ "role",
+ "education_code",
+ "family_education_code",
+ },
+ ),
+ ):
+ frame = getattr(inputs, name)
+ if not required.issubset(frame.columns):
+ raise ValueError(f"inputs.{name} lacks required columns")
+ for column in ("person_id", "wave"):
+ if any(
+ isinstance(v, (bool, np.bool_)) or not isinstance(v, Integral)
+ for v in frame[column]
+ ):
+ raise ValueError(f"inputs.{name}.{column} must be integers")
+ if not frame["wave"].isin(gap.INDIVIDUAL_WAVES).all():
+ raise ValueError(f"inputs.{name} has unadjudicated waves")
+ if not frame["person_id"].isin(inputs.universe).all():
+ raise ValueError(f"inputs.{name} has persons outside the universe")
+ roles = frame["role"] if name == "reports" else frame["role"].dropna()
+ if not roles.isin(gap.ROLES).all():
+ raise ValueError(f"inputs.{name} has undocumented roles")
+ if frame.duplicated(["person_id", "wave"]).any():
+ raise ValueError(
+ f"inputs.{name} has more than one row for a person-wave"
+ )
+ code_columns = (
+ gap.REPORT_CODE_COLUMNS
+ if name == "reports"
+ else ("education_code", "family_education_code")
+ )
+ for column in code_columns:
+ if any(
+ isinstance(v, (bool, np.bool_)) or not isinstance(v, Integral)
+ for v in frame[column].dropna()
+ ):
+ raise ValueError(f"inputs.{name}.{column} must be integers")
+ if name == "education" and frame["education_code"].isna().any():
+ raise ValueError("inputs.education has absent education codes")
+ _validate_attribute_code_domains(inputs)
+
+
+def _validate_attribute_code_domains(inputs: GroupAttributeInputs) -> None:
+ """Use reader authority for supplied education and birthplace codes.
+
+ The individual education frame is the reader's adjudicated domain,
+ verified against IND2023ER_formats.sas at each real-data load. Family
+ domains come from the pinned per-variable codebook captures. Grouped
+ checks apply even to reports outside the requested population/cutoff.
+ """
+
+ education = inputs.education
+ supplied_family = education["family_education_code"].notna()
+ if (supplied_family & education["role"].isna()).any():
+ raise ValueError(
+ "inputs.education.family_education_code requires a head or "
+ "spouse role"
+ )
+ codebook = (
+ gap.load_codebook_values()
+ if not inputs.reports.empty or supplied_family.any()
+ else None
+ )
+
+ def family_code_domain(
+ wave: int, role: str, column: str, variable: str, part: pd.DataFrame
+ ) -> None:
+ gap._check_codes(
+ part[column],
+ gap.documented_domain(codebook, wave, variable),
+ context=f"inputs.{column} ({wave} {role} {variable})",
+ )
+
+ for wave, part in education.groupby("wave", sort=False):
+ wave = int(wave)
+ variable = gap.INDIVIDUAL_ITEMS[wave].education[0]
+ gap._check_codes(
+ part["education_code"],
+ gap._expected_education_domain(wave),
+ context=f"inputs.education.education_code ({wave} {variable})",
+ )
+ supplied = part[part["family_education_code"].notna()]
+ if supplied.empty:
+ continue
+ items = gap.FAMILY_ITEMS[wave]
+ if not items.completed_education:
+ raise ValueError(
+ f"inputs.education.family_education_code: not asked in {wave}"
+ )
+ for role, role_rows in supplied.groupby("role", sort=False):
+ family_code_domain(
+ wave,
+ str(role),
+ "family_education_code",
+ items.completed_education[str(role)][0],
+ role_rows,
+ )
+
+ for (wave, role), part in inputs.reports.groupby(
+ ["wave", "role"], sort=False
+ ):
+ wave, role = int(wave), str(role)
+ items = gap.FAMILY_ITEMS[wave]
+ supplied = part[part["family_education_code"].notna()]
+ if not supplied.empty:
+ if not items.completed_education:
+ raise ValueError(
+ "inputs.reports.family_education_code: not asked in "
+ f"{wave}"
+ )
+ family_code_domain(
+ wave,
+ role,
+ "family_education_code",
+ items.completed_education[role][0],
+ supplied,
+ )
+ for column, mapping in (
+ ("birth_state_code", items.birth_state),
+ ("year_came_code", items.year_came),
+ ):
+ if not mapping:
+ if part[column].notna().any():
+ raise ValueError(
+ f"inputs.reports.{column}: not asked in {wave}"
+ )
+ continue
+ family_code_domain(wave, role, column, mapping[role][0], part)
+
+
+def build_group_attributes(
+ inputs: GroupAttributeInputs,
+ person_ids: Iterable[int],
+ *,
+ anchor_waves: Iterable[int],
+ schemes: Mapping[str, Any] | None = None,
+) -> GroupAttributes:
+ """Resolve the attributes and scheme categories of ``person_ids``.
+
+ Reads no PSID files: consumes ``inputs`` and category/code documentation.
+ ``anchor_waves`` are the population's
+ anchor waves (``(2011,)``, ``(2009,)``, the age-67 observation waves,
+ ``(2023,)``); education is resolved as of their maximum. Refuses
+ duplicate or unknown person ids and non-PSID anchor waves.
+ """
+
+ waves = _anchor_waves(anchor_waves)
+ cutoff = max(waves)
+ schemes = load_schemes() if schemes is None else schemes
+ _validate_schemes(schemes)
+ _validate_inputs(inputs)
+ persons = _person_index(person_ids, inputs.universe)
+ actual_input_digest = input_frames_sha256(inputs)
+ recorded_input_digest = inputs.provenance.get("input_frames_sha256")
+ if (
+ recorded_input_digest is not None
+ and recorded_input_digest != actual_input_digest
+ ):
+ raise ValueError("input frames changed after their provenance seal")
+ wanted = set(persons)
+ reports = inputs.reports[inputs.reports["person_id"].isin(wanted)]
+ education = inputs.education[inputs.education["person_id"].isin(wanted)]
+
+ race = _resolve_race(persons, reports)
+ edu = _resolve_education(persons, education, cutoff)
+ country = _resolve_country(persons, reports)
+ frame = pd.concat([race, edu, country], axis=1).reset_index()
+ for column, dtype in _ATTRIBUTE_COLUMNS.items():
+ frame[column] = frame[column].astype(dtype)
+ scheme_columns: list[str] = []
+ for scheme in schemes["schemes"].values():
+ for dimension, spec in scheme["dimensions"].items():
+ labels, status = apply_category_scheme(frame, dimension, spec)
+ frame[spec["column"]] = labels
+ frame[f"{spec['column']}_status"] = status
+ scheme_columns.extend([spec["column"], f"{spec['column']}_status"])
+ frame = frame[[*_ATTRIBUTE_COLUMNS, *scheme_columns]]
+
+ def _counts(column: str) -> dict[str, int]:
+ return {
+ str(k): int(v)
+ for k, v in frame[column].value_counts(dropna=False).items()
+ }
+
+ provenance = {
+ "builder": (
+ "populace_dynamics.cohorts.group_attributes."
+ "build_group_attributes"
+ ),
+ "rules_version": RULES_VERSION,
+ "anchor_waves": list(waves),
+ "education_cutoff_wave": cutoff,
+ "windows": {
+ "race_ethnicity": list(gap.RACE_WAVES),
+ "country_of_birth": list(gap.BIRTHPLACE_WAVES),
+ "education": [w for w in gap.INDIVIDUAL_WAVES if w <= cutoff],
+ },
+ "n_persons": len(frame),
+ "person_ids_sha256": hashlib.sha256(
+ ",".join(str(p) for p in frame["person_id"]).encode()
+ ).hexdigest(),
+ "status_counts": {
+ column: _counts(column)
+ for column in (
+ "race_ethnicity_status",
+ "education_status",
+ "country_of_birth_status",
+ *(c for c in scheme_columns if c.endswith("_status")),
+ )
+ },
+ "sourced_after_cutoff": {
+ "race_ethnicity": int(
+ (frame["race_ethnicity_source_wave"] > cutoff).sum()
+ ),
+ "country_of_birth": int(
+ (frame["country_of_birth_source_wave"] > cutoff).sum()
+ ),
+ },
+ "inputs": deepcopy(dict(inputs.provenance)),
+ "input_frames_sha256": actual_input_digest,
+ "schemes_sha256": hashlib.sha256(
+ json.dumps(schemes, sort_keys=True).encode()
+ ).hexdigest(),
+ "content_sha256": content_sha256(frame),
+ }
+ return GroupAttributes(frame=frame, provenance=provenance)
+
+
+def load_group_attributes(
+ person_ids: Iterable[int],
+ *,
+ anchor_waves: Iterable[int],
+ psid_dir: Path | None = None,
+) -> GroupAttributes:
+ """Read the staged PSID and build the side frame for ``person_ids``.
+
+ ``anchor_waves`` as in :func:`build_group_attributes`; ``psid_dir``
+ defaults to ``POPULACE_DYNAMICS_PSID_DIR`` or the staged default.
+ """
+
+ waves = _anchor_waves(anchor_waves)
+ ids = list(person_ids)
+ # Reject malformed requests before any expensive PSID file read. The
+ # actual universe check follows the audited read.
+ _person_index(ids, pd.Series(ids, dtype="object"))
+ inputs = load_group_attribute_inputs(psid_dir=psid_dir)
+ return build_group_attributes(inputs, ids, anchor_waves=waves)
diff --git a/src/populace_dynamics/data/group_attributes_psid.py b/src/populace_dynamics/data/group_attributes_psid.py
new file mode 100644
index 00000000..4d520f91
--- /dev/null
+++ b/src/populace_dynamics/data/group_attributes_psid.py
@@ -0,0 +1,2000 @@
+"""Label-verified PSID readers for MINT group attributes.
+
+This module reads, from the staged PSID, the three person attributes the
+MINT group breakdowns need and the existing cohorts do not carry: race and
+Hispanic origin, years of completed education, and country of birth. It
+returns raw, documented codes and the pure functions that give each code
+its meaning. Choosing among a person's reports across waves, and mapping
+them to report categories, is cohort code
+(:mod:`populace_dynamics.cohorts.group_attributes`), which is the only
+caller of these readers.
+
+Where each attribute lives in the PSID (adjudicated 2026-10-01 from the
+staged label files, the codebook PDFs and the formats files):
+
+* **Race and Hispanic origin** are asked of the family's head (reference
+ person from 2017) and wife ("Wife", spouse or partner from 2015) in the
+ per-wave family files, never of other family-unit members (OFUMs). The
+ instrument changed over time, so each wave carries a code frame
+ (:data:`RACE_ERAS`):
+
+ - 1985-1989: a Spanish-descent question (Mexican, Mexican American,
+ Chicano, Puerto Rican, Cuban, combination, other Spanish) and up to
+ two race mentions (white, black, American Indian/Aleut/Eskimo,
+ Asian/Pacific Islander, other; code 8 "more than two mentions").
+ - 1990-1993: the same, with race codes 5 "mentions Latino origin or
+ descent" and 6 "mentions color other than black or white".
+ - 1994-2003: three (1994-1996) or four (1997-2003) race mentions with
+ the 1990 codes, except that code 8 is "DK" rather than "more than two
+ mentions". **1997-2003 ask no Spanish-descent question**, so a
+ report from those waves shows Hispanic origin only through a race
+ mention of Latino origin and never shows its absence.
+ - 2005-2023: a Spanish-descent question (no "combination" code) and up
+ to four race mentions with the OMB 1997 codes: white; black, African
+ American or Negro; American Indian or Alaska Native; Asian; Native
+ Hawaiian or Pacific Islander; other. **Code 5 here means Native
+ Hawaiian or Pacific Islander; in 1990-2003 it meant Latino origin**,
+ which is why every code is read through its wave's frame.
+
+ Before 1985 the family files carry one "RACE" item for the head only,
+ with a Spanish-American race category and no wife item; those waves
+ are not read (see the cohort module for what that costs).
+
+* **Years of completed education** come from the cross-year individual
+ file's per-wave "YEARS COMPLETED EDUCATION" series (1985-2023 labels
+ vary: "COMPLETED EDUCATION", "COMPLETED EDUC-IND", "YRS COMPLETED
+ EDUC"), which covers heads, wives and OFUMs aged 16 or older: 1-16 is
+ the highest grade completed and 17 "at least some post-graduate work";
+ a GED without college is 12. Its code 0 is documented only as Inap.,
+ and therefore does not by itself establish zero schooling. The family
+ file's COMPLETED ED-HD/WF recode (1993 on) explicitly documents
+ "completed no grades of school": a present head or spouse with both
+ codes 0 has 0 years of schooling. An OFUM with individual code 0 is
+ left unreported.
+
+* **Country of birth** is asked of heads and wives from 2013: "L33/K33
+ STATE WAS BORN" (FIPS state 1-56; 0 "U.S. territory or foreign
+ country"; 99 DK/NA/refused) with "L33YR/K33YR YEAR CAME TO UNITED
+ STATES" (0 "born in the United States or U.S. territory"). The pair
+ separates a U.S. state, a U.S. territory and a foreign country. The
+ 1997/1999 immigrant-supplement items ("M50 CKPT HEAD BORN IN US",
+ individual "ES1 STATE WHERE BORN") cover only new or immigrant-sample
+ heads and wives, and "STATE GREW UP" is not a birthplace, so
+ none of them is read.
+
+Every variable is checked against its exact label (whitespace
+normalized) before any fixed-width read, and every observed code against
+the codes its wave documents: the codebook value tables committed in
+``data/external/psid_group_attribute_codebook_values_v1.json`` (built by
+``scripts/build_group_attribute_codebook_values.py`` from the staged
+codebook PDFs, whose SHA-256 the table pins), cross-checked against the
+2021 and 2023 family formats files and ``IND2023ER_formats.sas`` where a
+format block exists. An undocumented code raises.
+"""
+
+from __future__ import annotations
+
+import hashlib
+import json
+import re
+from collections.abc import Iterable, Mapping
+from dataclasses import dataclass, field
+from functools import cache
+from numbers import Integral
+from pathlib import Path
+from typing import Any
+
+import pandas as pd
+
+from populace_dynamics.data import disability, family, psid
+
+__all__ = [
+ "ROLES",
+ "RACE_WAVES",
+ "INDIVIDUAL_WAVES",
+ "FAMILY_EDUCATION_WAVES",
+ "BIRTHPLACE_WAVES",
+ "RACE_ERAS",
+ "FamilyWaveItems",
+ "IndividualWaveItems",
+ "FAMILY_ITEMS",
+ "INDIVIDUAL_ITEMS",
+ "CodeDomain",
+ "CODEBOOK_VALUES_PATH",
+ "CODEBOOK_VALUES_SHA256",
+ "load_codebook_values",
+ "documented_domain",
+ "parse_sas_value_domains",
+ "read_individual_items",
+ "read_family_items",
+ "verify_codebook_pins",
+ "head_spouse_reports",
+ "hispanic_meaning",
+ "race_mention_meaning",
+ "RaceEthnicityReport",
+ "race_ethnicity_report",
+ "birthplace_class",
+ "education_report",
+]
+
+#: The two family-file roles whose items this module reads. "head" is the
+#: head (reference person from 2017); "spouse" is the wife/"wife" (spouse
+#: or partner from 2015): individual-file relationship 20 (legal wife or
+#: spouse) or 22 (cohabiting partner of a year or more).
+ROLES: tuple[str, ...] = ("head", "spouse")
+
+#: Waves whose family files carry the race items read here: annual
+#: 1985-1997, biennial 1999-2023.
+RACE_WAVES: tuple[int, ...] = (*range(1985, 1998), *range(1999, 2024, 2))
+#: Individual-file waves read for the roster and education (the same).
+INDIVIDUAL_WAVES: tuple[int, ...] = RACE_WAVES
+#: Waves whose family files carry COMPLETED ED-HD/WF in years (1993 on;
+#: 1985-1992 carry a bracketed "EDUCATION HEAD" code instead).
+FAMILY_EDUCATION_WAVES: tuple[int, ...] = tuple(
+ w for w in RACE_WAVES if w >= 1993
+)
+#: Waves whose family files ask where the head and wife were born.
+BIRTHPLACE_WAVES: tuple[int, ...] = tuple(w for w in RACE_WAVES if w >= 2013)
+
+#: Individual-file relationship codes (1983 on) for the two roles; their
+#: labels are verified against IND2023ER_formats.sas at read time.
+HEAD_RELATIONSHIP = 10
+SPOUSE_RELATIONSHIPS: tuple[int, ...] = (20, 22)
+#: Sequence numbers of persons in the family at the interview.
+IN_FAMILY_SEQUENCE: tuple[int, int] = (1, 20)
+
+_PERSON_VARS: dict[str, str] = {
+ "ER30001": "1968 INTERVIEW NUMBER",
+ "ER30002": "PERSON NUMBER 68",
+}
+
+# --------------------------------------------------------------------------
+# Code frames and their meanings
+# --------------------------------------------------------------------------
+RACE_ERA_1985 = "1985-1989"
+RACE_ERA_1990 = "1990-1993"
+RACE_ERA_1994 = "1994-2003"
+RACE_ERA_2005 = "2005-2023"
+RACE_ERAS: tuple[str, ...] = (
+ RACE_ERA_1985,
+ RACE_ERA_1990,
+ RACE_ERA_1994,
+ RACE_ERA_2005,
+)
+
+# Race-mention meanings (harmonized categories).
+WHITE = "white"
+BLACK = "black"
+AMERICAN_INDIAN = "american_indian_alaska_native"
+ASIAN_PACIFIC = "asian_pacific_islander"
+ASIAN = "asian"
+PACIFIC_ISLANDER = "native_hawaiian_pacific_islander"
+OTHER_RACE = "other_race"
+OTHER_COLOR = "color_other_than_black_or_white"
+MORE_THAN_TWO = "more_than_two_races"
+LATINO_ORIGIN = "latino_origin_mention"
+NO_FURTHER_MENTION = "no_further_mention"
+MISSING = "missing"
+
+#: The race categories (a person's race set is drawn from these).
+#: ``more_than_two_races`` marks a 1985-1993 "more than two mentions"
+#: response: at least three races, the rest not recorded.
+RACE_CATEGORIES: tuple[str, ...] = (
+ WHITE,
+ BLACK,
+ AMERICAN_INDIAN,
+ ASIAN_PACIFIC,
+ ASIAN,
+ PACIFIC_ISLANDER,
+ OTHER_RACE,
+ OTHER_COLOR,
+ MORE_THAN_TWO,
+)
+
+#: Code -> meaning per race code frame, from the codebook value tables
+#: (``data/external/psid_group_attribute_codebook_values_v1.json``; a test
+#: holds every documented code to a meaning whose text agrees). Code 0 on
+#: a second or later mention is "no further mention"; code 0 on the first
+#: mention is never a valid answer (for a wife it is "no wife in FU", for
+#: the 2005/2007 head a documented "wild code") and reads as missing.
+RACE_CODE_MEANINGS: dict[str, dict[int, str]] = {
+ RACE_ERA_1985: {
+ 1: WHITE,
+ 2: BLACK,
+ 3: AMERICAN_INDIAN,
+ 4: ASIAN_PACIFIC,
+ 7: OTHER_RACE,
+ 8: MORE_THAN_TWO,
+ 9: MISSING,
+ 0: NO_FURTHER_MENTION,
+ },
+ RACE_ERA_1990: {
+ 1: WHITE,
+ 2: BLACK,
+ 3: AMERICAN_INDIAN,
+ 4: ASIAN_PACIFIC,
+ 5: LATINO_ORIGIN,
+ 6: OTHER_COLOR,
+ 7: OTHER_RACE,
+ 8: MORE_THAN_TWO,
+ 9: MISSING,
+ 0: NO_FURTHER_MENTION,
+ },
+ RACE_ERA_1994: {
+ 1: WHITE,
+ 2: BLACK,
+ 3: AMERICAN_INDIAN,
+ 4: ASIAN_PACIFIC,
+ 5: LATINO_ORIGIN,
+ 6: OTHER_COLOR,
+ 7: OTHER_RACE,
+ 8: MISSING,
+ 9: MISSING,
+ 0: NO_FURTHER_MENTION,
+ },
+ RACE_ERA_2005: {
+ 1: WHITE,
+ 2: BLACK,
+ 3: AMERICAN_INDIAN,
+ 4: ASIAN,
+ 5: PACIFIC_ISLANDER,
+ 7: OTHER_RACE,
+ 8: MISSING,
+ 9: MISSING,
+ 0: NO_FURTHER_MENTION,
+ },
+}
+
+HISPANIC = "hispanic"
+NOT_HISPANIC = "not_hispanic"
+#: Spanish-descent code -> meaning (every wave that asks it). 1-7 are the
+#: origins (6 "combination; more than one mention" before 2005); 0 is
+#: "not Spanish, Hispanic or Latino". For a wife the codebook's 0 also
+#: covers "no wife in FU". In 1994-1996 its wife-variable text names
+#: only that case; a present spouse's 0 therefore has an undocumented
+#: meaning and remains unknown (:func:`race_ethnicity_report`).
+HISPANIC_CODE_MEANINGS: dict[int, str] = {
+ 0: NOT_HISPANIC,
+ 1: HISPANIC,
+ 2: HISPANIC,
+ 3: HISPANIC,
+ 4: HISPANIC,
+ 5: HISPANIC,
+ 6: HISPANIC,
+ 7: HISPANIC,
+ 8: MISSING,
+ 9: MISSING,
+}
+
+# Birthplace classes.
+UNITED_STATES = "united_states"
+US_TERRITORY = "us_territory"
+FOREIGN_COUNTRY = "foreign_country"
+INCONSISTENT = "inconsistent"
+BIRTHPLACE_CLASSES: tuple[str, ...] = (
+ UNITED_STATES,
+ US_TERRITORY,
+ FOREIGN_COUNTRY,
+ MISSING,
+ INCONSISTENT,
+)
+#: "STATE WAS BORN": FIPS state codes, "U.S. territory or foreign
+#: country", and DK/NA/refused.
+BIRTH_STATE_RANGE: tuple[int, int] = (1, 56)
+BIRTH_STATE_NOT_A_STATE = 0
+BIRTH_STATE_MISSING = 99
+#: "YEAR CAME TO UNITED STATES": 0 is "born in the United States or U.S.
+#: territory"; 9997 "not living in the United States"; 9998/9999 DK/NA.
+YEAR_CAME_BORN_IN_US = 0
+YEAR_CAME_NOT_LIVING_IN_US = 9997
+YEAR_CAME_MISSING: tuple[int, ...] = (9998, 9999)
+
+# Education report statuses.
+EDUCATION_REPORTED = "reported"
+EDUCATION_NO_GRADES = "no_grades"
+EDUCATION_MISSING = "missing"
+EDUCATION_INAPPLICABLE = "inapplicable"
+#: Individual-file "YEARS COMPLETED EDUCATION" codes.
+EDUCATION_YEARS_RANGE: tuple[int, int] = (1, 17)
+EDUCATION_DK_NA: tuple[int, ...] = (98, 99)
+#: Family-file COMPLETED ED-HD/WF code for "completed no grades".
+FAMILY_EDUCATION_NO_GRADES = 0
+FAMILY_EDUCATION_MISSING = 99
+
+
+# --------------------------------------------------------------------------
+# Adjudicated per-wave tables
+# --------------------------------------------------------------------------
+@dataclass(frozen=True)
+class FamilyWaveItems:
+ """One family-file wave's label-verified group-attribute variables.
+
+ Every item maps a role (``"head"``, ``"spouse"``) to ``(variable,
+ exact label)``; ``race`` maps each role to its mentions in order. An
+ empty mapping means the wave does not carry the item (no
+ Spanish-descent question 1997-2003; no years-of-education recode
+ before 1993; no birthplace before 2013). ``race_era`` names the race
+ code frame (:data:`RACE_CODE_MEANINGS`).
+ """
+
+ wave: int
+ race_era: str
+ interview: tuple[str, str]
+ race: Mapping[str, tuple[tuple[str, str], ...]]
+ hispanic: Mapping[str, tuple[str, str]] = field(default_factory=dict)
+ completed_education: Mapping[str, tuple[str, str]] = field(
+ default_factory=dict
+ )
+ birth_state: Mapping[str, tuple[str, str]] = field(default_factory=dict)
+ year_came: Mapping[str, tuple[str, str]] = field(default_factory=dict)
+
+ @property
+ def asks_hispanic_origin(self) -> bool:
+ return bool(self.hispanic)
+
+ def labels(self) -> dict[str, str]:
+ """Every variable this wave reads, with its exact label."""
+ out = {self.interview[0]: self.interview[1]}
+ for item in (
+ self.hispanic,
+ self.completed_education,
+ self.birth_state,
+ self.year_came,
+ ):
+ out.update({var: label for var, label in item.values()})
+ for mentions in self.race.values():
+ out.update({var: label for var, label in mentions})
+ return out
+
+ def role_columns(self, role: str) -> dict[str, str]:
+ """``{variable: generic column}`` for one role's items."""
+ out: dict[str, str] = {}
+ if self.hispanic:
+ out[self.hispanic[role][0]] = "hispanic_code"
+ for i, (var, _) in enumerate(self.race[role], start=1):
+ out[var] = f"race_code_{i}"
+ if self.completed_education:
+ out[self.completed_education[role][0]] = "family_education_code"
+ if self.birth_state:
+ out[self.birth_state[role][0]] = "birth_state_code"
+ if self.year_came:
+ out[self.year_came[role][0]] = "year_came_code"
+ return out
+
+
+@dataclass(frozen=True)
+class IndividualWaveItems:
+ """One wave's label-verified individual-file roster and education."""
+
+ wave: int
+ interview: tuple[str, str]
+ sequence: tuple[str, str]
+ relationship: tuple[str, str]
+ education: tuple[str, str]
+
+ def labels(self) -> dict[str, str]:
+ return {
+ var: label
+ for var, label in (
+ self.interview,
+ self.sequence,
+ self.relationship,
+ self.education,
+ )
+ }
+
+
+#: The adjudicated family-file variables, wave by wave. Built 2026-10-01
+#: by matching each concept's labels in every staged FAM[ER].sps
+#: (one match per role, mention and wave; the 1990 interview label is
+#: PSID's own "1990 INTERVEW NUMBER") and verified at read time under
+#: whitespace normalization.
+FAMILY_ITEMS: dict[int, FamilyWaveItems] = {
+ 1985: FamilyWaveItems(
+ wave=1985,
+ race_era=RACE_ERA_1985,
+ interview=("V11102", "1985 INTERVIEW NUMBER"),
+ hispanic={
+ "head": ("V11937", "G31 SPANISH DESCENT-HEAD"),
+ "spouse": ("V12292", "N31 SPANISH DESCENT-WIFE"),
+ },
+ race={
+ "head": (
+ ("V11938", "G32 RACE OF HEAD (1 MEN)"),
+ ("V11939", "G32 RACE OF HEAD (2 MEN)"),
+ ),
+ "spouse": (
+ ("V12293", "N32 RACE OF WIFE (1 MEN)"),
+ ("V12294", "N32 RACE OF WIFE (2 MEN)"),
+ ),
+ },
+ ),
+ 1986: FamilyWaveItems(
+ wave=1986,
+ race_era=RACE_ERA_1985,
+ interview=("V12502", "1986 INTERVIEW NUMBER"),
+ hispanic={
+ "head": ("V13564", "L31 SPANISH DESCENT HD"),
+ "spouse": ("V13499", "K18 SPANISH DESCENT WF"),
+ },
+ race={
+ "head": (
+ ("V13565", "L32 RACE OF HEAD 1"),
+ ("V13566", "L32 RACE OF HEAD 2"),
+ ),
+ "spouse": (
+ ("V13500", "K19 RACE OF WIFE 1"),
+ ("V13501", "K19 RACE OF WIFE 2"),
+ ),
+ },
+ ),
+ 1987: FamilyWaveItems(
+ wave=1987,
+ race_era=RACE_ERA_1985,
+ interview=("V13702", "1987 INTERVIEW NUMBER"),
+ hispanic={
+ "head": ("V14611", "L31 SPANISH DESCENT HD"),
+ "spouse": ("V14546", "K18 SPANISH DESCENT WF"),
+ },
+ race={
+ "head": (
+ ("V14612", "L32 RACE OF HEAD 1"),
+ ("V14613", "L32 RACE OF HEAD 2"),
+ ),
+ "spouse": (
+ ("V14547", "K19 RACE OF WIFE 1"),
+ ("V14548", "K19 RACE OF WIFE 2"),
+ ),
+ },
+ ),
+ 1988: FamilyWaveItems(
+ wave=1988,
+ race_era=RACE_ERA_1985,
+ interview=("V14802", "1988 INTERVIEW NUMBER"),
+ hispanic={
+ "head": ("V16085", "L31 SPANISH DESCENT HD"),
+ "spouse": ("V16020", "K18 SPANISH DESCENT WF"),
+ },
+ race={
+ "head": (
+ ("V16086", "L32 RACE OF HEAD 1"),
+ ("V16087", "L32 RACE OF HEAD 2"),
+ ),
+ "spouse": (
+ ("V16021", "K19 RACE OF WIFE 1"),
+ ("V16022", "K19 RACE OF WIFE 2"),
+ ),
+ },
+ ),
+ 1989: FamilyWaveItems(
+ wave=1989,
+ race_era=RACE_ERA_1985,
+ interview=("V16302", "1989 INTERVIEW NUMBER"),
+ hispanic={
+ "head": ("V17482", "L31 SPANISH DESCENT HD"),
+ "spouse": ("V17417", "K18 SPANISH DESCENT WF"),
+ },
+ race={
+ "head": (
+ ("V17483", "L32 RACE OF HEAD 1"),
+ ("V17484", "L32 RACE OF HEAD 2"),
+ ),
+ "spouse": (
+ ("V17418", "K19 RACE OF WIFE 1"),
+ ("V17419", "K19 RACE OF WIFE 2"),
+ ),
+ },
+ ),
+ 1990: FamilyWaveItems(
+ wave=1990,
+ race_era=RACE_ERA_1990,
+ interview=("V17702", "1990 INTERVEW NUMBER"),
+ hispanic={
+ "head": ("V18813", "M31 SPANISH DESCENT HD"),
+ "spouse": ("V18748", "L18 SPANISH DESCENT WF"),
+ },
+ race={
+ "head": (
+ ("V18814", "M32 RACE OF HEAD 1"),
+ ("V18815", "M32 RACE OF HEAD 2"),
+ ),
+ "spouse": (
+ ("V18749", "L19 RACE OF WIFE 1"),
+ ("V18750", "L19 RACE OF WIFE 2"),
+ ),
+ },
+ ),
+ 1991: FamilyWaveItems(
+ wave=1991,
+ race_era=RACE_ERA_1990,
+ interview=("V19002", "1991 INTERVIEW NUMBER"),
+ hispanic={
+ "head": ("V20113", "L31 SPANISH DESCENT HD"),
+ "spouse": ("V20048", "K18 SPANISH DESCENT WF"),
+ },
+ race={
+ "head": (
+ ("V20114", "L32 RACE OF HEAD 1"),
+ ("V20115", "L32 RACE OF HEAD 2"),
+ ),
+ "spouse": (
+ ("V20049", "K19 RACE OF WIFE 1"),
+ ("V20050", "K19 RACE OF WIFE 2"),
+ ),
+ },
+ ),
+ 1992: FamilyWaveItems(
+ wave=1992,
+ race_era=RACE_ERA_1990,
+ interview=("V20302", "1992 INTERVIEW NUMBER"),
+ hispanic={
+ "head": ("V21419", "M31 SPANISH DESCENT HD"),
+ "spouse": ("V21354", "L18 SPANISH DESCENT WF"),
+ },
+ race={
+ "head": (
+ ("V21420", "M32 RACE OF HEAD 1"),
+ ("V21421", "M32 RACE OF HEAD 2"),
+ ),
+ "spouse": (
+ ("V21355", "L19 RACE OF WIFE 1"),
+ ("V21356", "L19 RACE OF WIFE 2"),
+ ),
+ },
+ ),
+ 1993: FamilyWaveItems(
+ wave=1993,
+ race_era=RACE_ERA_1990,
+ interview=("V21602", "1993 INTERVIEW NUMBER"),
+ hispanic={
+ "head": ("V23275", "L31 WTR HD OF SPANISH DESCENT"),
+ "spouse": ("V23211", "K18 WTR WF OF SPANISH DESCENT"),
+ },
+ race={
+ "head": (
+ ("V23276", "L32 RACE OF HD-1ST MENTION"),
+ ("V23277", "L32 RACE OF HD-2ND MENTION"),
+ ),
+ "spouse": (
+ ("V23212", "K19 RACE OF WF-1ST MENTION"),
+ ("V23213", "K19 RACE OF WF-2ND MENTION"),
+ ),
+ },
+ completed_education={
+ "head": ("V23333", "COMPLETED ED-HD 1993"),
+ "spouse": ("V23334", "COMPLETED ED-WF 1993"),
+ },
+ ),
+ 1994: FamilyWaveItems(
+ wave=1994,
+ race_era=RACE_ERA_1994,
+ interview=("ER2002", "1994 INTERVIEW #"),
+ hispanic={
+ "head": ("ER3941", "L31 SPANISH DESCENT 1 HD"),
+ "spouse": ("ER3880", "K18 SPANISH DESCENT 1 WF"),
+ },
+ race={
+ "head": (
+ ("ER3944", "L32 RACE OF HEAD 1"),
+ ("ER3945", "L32 RACE OF HEAD 2"),
+ ("ER3946", "L32 RACE OF HEAD 3"),
+ ),
+ "spouse": (
+ ("ER3883", "K19 RACE OF WIFE 1"),
+ ("ER3884", "K19 RACE OF WIFE 2"),
+ ("ER3885", "K19 RACE OF WIFE 3"),
+ ),
+ },
+ completed_education={
+ "head": ("ER4158", "COMPLETED ED-HD"),
+ "spouse": ("ER4159", "COMPLETED ED-WF"),
+ },
+ ),
+ 1995: FamilyWaveItems(
+ wave=1995,
+ race_era=RACE_ERA_1994,
+ interview=("ER5002", "1995 INTERVIEW #"),
+ hispanic={
+ "head": ("ER6811", "L31 SPANISH DESCENT 1 HD"),
+ "spouse": ("ER6750", "K18 SPANISH DESCENT 1 WF"),
+ },
+ race={
+ "head": (
+ ("ER6814", "L32 RACE OF HEAD 1"),
+ ("ER6815", "L32 RACE OF HEAD 2"),
+ ("ER6816", "L32 RACE OF HEAD 3"),
+ ),
+ "spouse": (
+ ("ER6753", "K19 RACE OF WIFE 1"),
+ ("ER6754", "K19 RACE OF WIFE 2"),
+ ("ER6755", "K19 RACE OF WIFE 3"),
+ ),
+ },
+ completed_education={
+ "head": ("ER6998", "COMPLETED ED-HD"),
+ "spouse": ("ER6999", "COMPLETED ED-WF"),
+ },
+ ),
+ 1996: FamilyWaveItems(
+ wave=1996,
+ race_era=RACE_ERA_1994,
+ interview=("ER7002", "1996 INTERVIEW #"),
+ hispanic={
+ "head": ("ER9057", "L31 SPANISH DESCENT 1 HD"),
+ "spouse": ("ER8996", "K18 SPANISH DESCENT 1 WF"),
+ },
+ race={
+ "head": (
+ ("ER9060", "L32 RACE OF HEAD 1"),
+ ("ER9061", "L32 RACE OF HEAD 2"),
+ ("ER9062", "L32 RACE OF HEAD 3"),
+ ),
+ "spouse": (
+ ("ER8999", "K19 RACE OF WIFE 1"),
+ ("ER9000", "K19 RACE OF WIFE 2"),
+ ("ER9001", "K19 RACE OF WIFE 3"),
+ ),
+ },
+ completed_education={
+ "head": ("ER9249", "COMPLETED ED-HD"),
+ "spouse": ("ER9250", "COMPLETED ED-WF"),
+ },
+ ),
+ 1997: FamilyWaveItems(
+ wave=1997,
+ race_era=RACE_ERA_1994,
+ interview=("ER10002", "1997 INTERVIEW #"),
+ race={
+ "head": (
+ ("ER11848", "L40/95 RACE OF HEAD 1"),
+ ("ER11849", "L40/95 RACE OF HEAD 2"),
+ ("ER11850", "L40/95 RACE OF HEAD 3"),
+ ("ER11851", "L40/95 RACE OF HEAD 4"),
+ ),
+ "spouse": (
+ ("ER11760", "K34/87 RACE OF WIFE 1"),
+ ("ER11761", "K34/87 RACE OF WIFE 2"),
+ ("ER11762", "K34/87 RACE OF WIFE 3"),
+ ("ER11763", "K34/87 RACE OF WIFE 4"),
+ ),
+ },
+ completed_education={
+ "head": ("ER12222", "COMPLETED ED-HD"),
+ "spouse": ("ER12223", "COMPLETED ED-WF"),
+ },
+ ),
+ 1999: FamilyWaveItems(
+ wave=1999,
+ race_era=RACE_ERA_1994,
+ interview=("ER13002", "1999 FAMILY INTERVIEW (ID) NUMBER"),
+ race={
+ "head": (
+ ("ER15928", "L40/95 RACE OF HEAD 1"),
+ ("ER15929", "L40/95 RACE OF HEAD 2"),
+ ("ER15930", "L40/95 RACE OF HEAD 3"),
+ ("ER15931", "L40/95 RACE OF HEAD 4"),
+ ),
+ "spouse": (
+ ("ER15836", "K34/87 RACE OF WIFE 1"),
+ ("ER15837", "K34/87 RACE OF WIFE 2"),
+ ("ER15838", "K34/87 RACE OF WIFE 3"),
+ ("ER15839", "K34/87 RACE OF WIFE 4"),
+ ),
+ },
+ completed_education={
+ "head": ("ER16516", "COMPLETED ED-HD"),
+ "spouse": ("ER16517", "COMPLETED ED-WF"),
+ },
+ ),
+ 2001: FamilyWaveItems(
+ wave=2001,
+ race_era=RACE_ERA_1994,
+ interview=("ER17002", "2001 FAMILY INTERVIEW (ID) NUMBER"),
+ race={
+ "head": (
+ ("ER19989", "L40/95 RACE OF HEAD 1"),
+ ("ER19990", "L40/95 RACE OF HEAD 2"),
+ ("ER19991", "L40/95 RACE OF HEAD 3"),
+ ("ER19992", "L40/95 RACE OF HEAD 4"),
+ ),
+ "spouse": (
+ ("ER19897", "K34/87 RACE OF WIFE 1"),
+ ("ER19898", "K34/87 RACE OF WIFE 2"),
+ ("ER19899", "K34/87 RACE OF WIFE 3"),
+ ("ER19900", "K34/87 RACE OF WIFE 4"),
+ ),
+ },
+ completed_education={
+ "head": ("ER20457", "COMPLETED ED-HD"),
+ "spouse": ("ER20458", "COMPLETED ED-WF"),
+ },
+ ),
+ 2003: FamilyWaveItems(
+ wave=2003,
+ race_era=RACE_ERA_1994,
+ interview=("ER21002", "2003 FAMILY INTERVIEW (ID) NUMBER"),
+ race={
+ "head": (
+ ("ER23426", "L40/95 RACE OF HEAD 1"),
+ ("ER23427", "L40/95 RACE OF HEAD 2"),
+ ("ER23428", "L40/95 RACE OF HEAD 3"),
+ ("ER23429", "L40/95 RACE OF HEAD 4"),
+ ),
+ "spouse": (
+ ("ER23334", "K34/87 RACE OF WIFE 1"),
+ ("ER23335", "K34/87 RACE OF WIFE 2"),
+ ("ER23336", "K34/87 RACE OF WIFE 3"),
+ ("ER23337", "K34/87 RACE OF WIFE 4"),
+ ),
+ },
+ completed_education={
+ "head": ("ER24148", "COMPLETED ED-HD"),
+ "spouse": ("ER24149", "COMPLETED ED-WF"),
+ },
+ ),
+ 2005: FamilyWaveItems(
+ wave=2005,
+ race_era=RACE_ERA_2005,
+ interview=("ER25002", "2005 FAMILY INTERVIEW (ID) NUMBER"),
+ hispanic={
+ "head": ("ER27392", "L39A SPANISH DESCENT-HEAD"),
+ "spouse": ("ER27296", "K33A SPANISH DESCENT-WIFE"),
+ },
+ race={
+ "head": (
+ ("ER27393", "L40 RACE OF HEAD-MENTION 1"),
+ ("ER27394", "L40 RACE OF HEAD-MENTION 2"),
+ ("ER27395", "L40 RACE OF HEAD-MENTION 3"),
+ ("ER27396", "L40 RACE OF HEAD-MENTION 4"),
+ ),
+ "spouse": (
+ ("ER27297", "K34 RACE OF WIFE-MENTION 1"),
+ ("ER27298", "K34 RACE OF WIFE-MENTION 2"),
+ ("ER27299", "K34 RACE OF WIFE-MENTION 3"),
+ ("ER27300", "K34 RACE OF WIFE-MENTION 4"),
+ ),
+ },
+ completed_education={
+ "head": ("ER28047", "COMPLETED ED-HD"),
+ "spouse": ("ER28048", "COMPLETED ED-WF"),
+ },
+ ),
+ 2007: FamilyWaveItems(
+ wave=2007,
+ race_era=RACE_ERA_2005,
+ interview=("ER36002", "2007 FAMILY INTERVIEW (ID) NUMBER"),
+ hispanic={
+ "head": ("ER40564", "L39 SPANISH DESCENT-HEAD"),
+ "spouse": ("ER40471", "K39 SPANISH DESCENT-WIFE"),
+ },
+ race={
+ "head": (
+ ("ER40565", "L40 RACE OF HEAD-MENTION 1"),
+ ("ER40566", "L40 RACE OF HEAD-MENTION 2"),
+ ("ER40567", "L40 RACE OF HEAD-MENTION 3"),
+ ("ER40568", "L40 RACE OF HEAD-MENTION 4"),
+ ),
+ "spouse": (
+ ("ER40472", "K40 RACE OF WIFE-MENTION 1"),
+ ("ER40473", "K40 RACE OF WIFE-MENTION 2"),
+ ("ER40474", "K40 RACE OF WIFE-MENTION 3"),
+ ("ER40475", "K40 RACE OF WIFE-MENTION 4"),
+ ),
+ },
+ completed_education={
+ "head": ("ER41037", "COMPLETED ED-HD"),
+ "spouse": ("ER41038", "COMPLETED ED-WF"),
+ },
+ ),
+ 2009: FamilyWaveItems(
+ wave=2009,
+ race_era=RACE_ERA_2005,
+ interview=("ER42002", "2009 FAMILY INTERVIEW (ID) NUMBER"),
+ hispanic={
+ "head": ("ER46542", "L39 SPANISH DESCENT-HEAD"),
+ "spouse": ("ER46448", "K39 SPANISH DESCENT-WIFE"),
+ },
+ race={
+ "head": (
+ ("ER46543", "L40 RACE OF HEAD-MENTION 1"),
+ ("ER46544", "L40 RACE OF HEAD-MENTION 2"),
+ ("ER46545", "L40 RACE OF HEAD-MENTION 3"),
+ ("ER46546", "L40 RACE OF HEAD-MENTION 4"),
+ ),
+ "spouse": (
+ ("ER46449", "K40 RACE OF WIFE-MENTION 1"),
+ ("ER46450", "K40 RACE OF WIFE-MENTION 2"),
+ ("ER46451", "K40 RACE OF WIFE-MENTION 3"),
+ ("ER46452", "K40 RACE OF WIFE-MENTION 4"),
+ ),
+ },
+ completed_education={
+ "head": ("ER46981", "COMPLETED ED-HD"),
+ "spouse": ("ER46982", "COMPLETED ED-WF"),
+ },
+ ),
+ 2011: FamilyWaveItems(
+ wave=2011,
+ race_era=RACE_ERA_2005,
+ interview=("ER47302", "2011 FAMILY INTERVIEW (ID) NUMBER"),
+ hispanic={
+ "head": ("ER51903", "L39 SPANISH DESCENT-HEAD"),
+ "spouse": ("ER51809", "K39 SPANISH DESCENT-WIFE"),
+ },
+ race={
+ "head": (
+ ("ER51904", "L40 RACE OF HEAD-MENTION 1"),
+ ("ER51905", "L40 RACE OF HEAD-MENTION 2"),
+ ("ER51906", "L40 RACE OF HEAD-MENTION 3"),
+ ("ER51907", "L40 RACE OF HEAD-MENTION 4"),
+ ),
+ "spouse": (
+ ("ER51810", "K40 RACE OF WIFE-MENTION 1"),
+ ("ER51811", "K40 RACE OF WIFE-MENTION 2"),
+ ("ER51812", "K40 RACE OF WIFE-MENTION 3"),
+ ("ER51813", "K40 RACE OF WIFE-MENTION 4"),
+ ),
+ },
+ completed_education={
+ "head": ("ER52405", "COMPLETED ED-HD"),
+ "spouse": ("ER52406", "COMPLETED ED-WF"),
+ },
+ ),
+ 2013: FamilyWaveItems(
+ wave=2013,
+ race_era=RACE_ERA_2005,
+ interview=("ER53002", "2013 FAMILY INTERVIEW (ID) NUMBER"),
+ hispanic={
+ "head": ("ER57658", "L39 SPANISH DESCENT-HEAD"),
+ "spouse": ("ER57548", "K39 SPANISH DESCENT-WIFE"),
+ },
+ race={
+ "head": (
+ ("ER57659", "L40 RACE OF HEAD-MENTION 1"),
+ ("ER57660", "L40 RACE OF HEAD-MENTION 2"),
+ ("ER57661", "L40 RACE OF HEAD-MENTION 3"),
+ ("ER57662", "L40 RACE OF HEAD-MENTION 4"),
+ ),
+ "spouse": (
+ ("ER57549", "K40 RACE OF WIFE-MENTION 1"),
+ ("ER57550", "K40 RACE OF WIFE-MENTION 2"),
+ ("ER57551", "K40 RACE OF WIFE-MENTION 3"),
+ ("ER57552", "K40 RACE OF WIFE-MENTION 4"),
+ ),
+ },
+ completed_education={
+ "head": ("ER58223", "COMPLETED ED-HD"),
+ "spouse": ("ER58224", "COMPLETED ED-WF"),
+ },
+ birth_state={
+ "head": ("ER57651", "L33 STATE HEAD WAS BORN"),
+ "spouse": ("ER57541", "K33 STATE WIFE WAS BORN"),
+ },
+ year_came={
+ "head": ("ER57652", "L33YR YEAR CAME TO UNITED STATES-HD"),
+ "spouse": ("ER57542", "K33YR YEAR CAME TO UNITED STATES-WF"),
+ },
+ ),
+ 2015: FamilyWaveItems(
+ wave=2015,
+ race_era=RACE_ERA_2005,
+ interview=("ER60002", "2015 FAMILY INTERVIEW (ID) NUMBER"),
+ hispanic={
+ "head": ("ER64809", "L39 SPANISH DESCENT-HEAD"),
+ "spouse": ("ER64670", "K39 SPANISH DESCENT-SPOUSE"),
+ },
+ race={
+ "head": (
+ ("ER64810", "L40 RACE OF HEAD-MENTION 1"),
+ ("ER64811", "L40 RACE OF HEAD-MENTION 2"),
+ ("ER64812", "L40 RACE OF HEAD-MENTION 3"),
+ ("ER64813", "L40 RACE OF HEAD-MENTION 4"),
+ ),
+ "spouse": (
+ ("ER64671", "K40 RACE OF SPOUSE-MENTION 1"),
+ ("ER64672", "K40 RACE OF SPOUSE-MENTION 2"),
+ ("ER64673", "K40 RACE OF SPOUSE-MENTION 3"),
+ ("ER64674", "K40 RACE OF SPOUSE-MENTION 4"),
+ ),
+ },
+ completed_education={
+ "head": ("ER65459", "COMPLETED ED-HD"),
+ "spouse": ("ER65460", "COMPLETED ED-SP"),
+ },
+ birth_state={
+ "head": ("ER64802", "L33 STATE HEAD WAS BORN"),
+ "spouse": ("ER64663", "K33 STATE SPOUSE WAS BORN"),
+ },
+ year_came={
+ "head": ("ER64803", "L33YR YEAR CAME TO UNITED STATES-HD"),
+ "spouse": ("ER64664", "K33YR YEAR CAME TO UNITED STATES-SP"),
+ },
+ ),
+ 2017: FamilyWaveItems(
+ wave=2017,
+ race_era=RACE_ERA_2005,
+ interview=("ER66002", "2017 FAMILY INTERVIEW (ID) NUMBER"),
+ hispanic={
+ "head": ("ER70881", "L39 SPANISH DESCENT-RP"),
+ "spouse": ("ER70743", "K39 SPANISH DESCENT-SPOUSE"),
+ },
+ race={
+ "head": (
+ ("ER70882", "L40 RACE OF REFERENCE PERSON-MENTION 1"),
+ ("ER70883", "L40 RACE OF REFERENCE PERSON-MENTION 2"),
+ ("ER70884", "L40 RACE OF REFERENCE PERSON-MENTION 3"),
+ ("ER70885", "L40 RACE OF REFERENCE PERSON-MENTION 4"),
+ ),
+ "spouse": (
+ ("ER70744", "K40 RACE OF SPOUSE-MENTION 1"),
+ ("ER70745", "K40 RACE OF SPOUSE-MENTION 2"),
+ ("ER70746", "K40 RACE OF SPOUSE-MENTION 3"),
+ ("ER70747", "K40 RACE OF SPOUSE-MENTION 4"),
+ ),
+ },
+ completed_education={
+ "head": ("ER71538", "COMPLETED ED-RP"),
+ "spouse": ("ER71539", "COMPLETED ED-SP"),
+ },
+ birth_state={
+ "head": ("ER70874", "L33 STATE REFERENCE PERSON WAS BORN"),
+ "spouse": ("ER70736", "K33 STATE SPOUSE WAS BORN"),
+ },
+ year_came={
+ "head": ("ER70875", "L33YR YEAR CAME TO UNITED STATES-RP"),
+ "spouse": ("ER70737", "K33YR YEAR CAME TO UNITED STATES-SP"),
+ },
+ ),
+ 2019: FamilyWaveItems(
+ wave=2019,
+ race_era=RACE_ERA_2005,
+ interview=("ER72002", "2019 FAMILY INTERVIEW (ID) NUMBER"),
+ hispanic={
+ "head": ("ER76896", "L39 SPANISH DESCENT-RP"),
+ "spouse": ("ER76751", "K39 SPANISH DESCENT-SPOUSE"),
+ },
+ race={
+ "head": (
+ ("ER76897", "L40 RACE OF REFERENCE PERSON-MENTION 1"),
+ ("ER76898", "L40 RACE OF REFERENCE PERSON-MENTION 2"),
+ ("ER76899", "L40 RACE OF REFERENCE PERSON-MENTION 3"),
+ ("ER76900", "L40 RACE OF REFERENCE PERSON-MENTION 4"),
+ ),
+ "spouse": (
+ ("ER76752", "K40 RACE OF SPOUSE-MENTION 1"),
+ ("ER76753", "K40 RACE OF SPOUSE-MENTION 2"),
+ ("ER76754", "K40 RACE OF SPOUSE-MENTION 3"),
+ ("ER76755", "K40 RACE OF SPOUSE-MENTION 4"),
+ ),
+ },
+ completed_education={
+ "head": ("ER77599", "COMPLETED ED-RP"),
+ "spouse": ("ER77600", "COMPLETED ED-SP"),
+ },
+ birth_state={
+ "head": ("ER76889", "L33 STATE REFERENCE PERSON WAS BORN"),
+ "spouse": ("ER76744", "K33 STATE SPOUSE WAS BORN"),
+ },
+ year_came={
+ "head": ("ER76890", "L33YR YEAR CAME TO UNITED STATES-RP"),
+ "spouse": ("ER76745", "K33YR YEAR CAME TO UNITED STATES-SP"),
+ },
+ ),
+ 2021: FamilyWaveItems(
+ wave=2021,
+ race_era=RACE_ERA_2005,
+ interview=("ER78002", "2021 FAMILY INTERVIEW (ID) NUMBER"),
+ hispanic={
+ "head": ("ER81143", "L39 SPANISH DESCENT-RP"),
+ "spouse": ("ER81016", "K39 SPANISH DESCENT-SPOUSE"),
+ },
+ race={
+ "head": (
+ ("ER81144", "L40 RACE OF REFERENCE PERSON-MENTION 1"),
+ ("ER81145", "L40 RACE OF REFERENCE PERSON-MENTION 2"),
+ ("ER81146", "L40 RACE OF REFERENCE PERSON-MENTION 3"),
+ ("ER81147", "L40 RACE OF REFERENCE PERSON-MENTION 4"),
+ ),
+ "spouse": (
+ ("ER81017", "K40 RACE OF SPOUSE-MENTION 1"),
+ ("ER81018", "K40 RACE OF SPOUSE-MENTION 2"),
+ ("ER81019", "K40 RACE OF SPOUSE-MENTION 3"),
+ ("ER81020", "K40 RACE OF SPOUSE-MENTION 4"),
+ ),
+ },
+ completed_education={
+ "head": ("ER81926", "COMPLETED ED-RP"),
+ "spouse": ("ER81927", "COMPLETED ED-SP"),
+ },
+ birth_state={
+ "head": ("ER81136", "L33 STATE REFERENCE PERSON WAS BORN"),
+ "spouse": ("ER81009", "K33 STATE SPOUSE WAS BORN"),
+ },
+ year_came={
+ "head": ("ER81137", "L33YR YEAR CAME TO UNITED STATES-RP"),
+ "spouse": ("ER81010", "K33YR YEAR CAME TO UNITED STATES-SP"),
+ },
+ ),
+ 2023: FamilyWaveItems(
+ wave=2023,
+ race_era=RACE_ERA_2005,
+ interview=("ER82002", "2023 FAMILY INTERVIEW (ID) NUMBER"),
+ hispanic={
+ "head": ("ER85120", "L39 SPANISH DESCENT-RP"),
+ "spouse": ("ER84993", "K39 SPANISH DESCENT-SPOUSE"),
+ },
+ race={
+ "head": (
+ ("ER85121", "L40 RACE OF REFERENCE PERSON-MENTION 1"),
+ ("ER85122", "L40 RACE OF REFERENCE PERSON-MENTION 2"),
+ ("ER85123", "L40 RACE OF REFERENCE PERSON-MENTION 3"),
+ ("ER85124", "L40 RACE OF REFERENCE PERSON-MENTION 4"),
+ ),
+ "spouse": (
+ ("ER84994", "K40 RACE OF SPOUSE-MENTION 1"),
+ ("ER84995", "K40 RACE OF SPOUSE-MENTION 2"),
+ ("ER84996", "K40 RACE OF SPOUSE-MENTION 3"),
+ ("ER84997", "K40 RACE OF SPOUSE-MENTION 4"),
+ ),
+ },
+ completed_education={
+ "head": ("ER85780", "COMPLETED ED-RP"),
+ "spouse": ("ER85781", "COMPLETED ED-SP"),
+ },
+ birth_state={
+ "head": ("ER85113", "L33 STATE REFERENCE PERSON WAS BORN"),
+ "spouse": ("ER84986", "K33 STATE SPOUSE WAS BORN"),
+ },
+ year_came={
+ "head": ("ER85114", "L33YR YEAR CAME TO UNITED STATES-RP"),
+ "spouse": ("ER84987", "K33YR YEAR CAME TO UNITED STATES-SP"),
+ },
+ ),
+}
+
+
+#: The adjudicated individual-file roster and education variables, wave
+#: by wave (ind2023er; one exact label match per item and wave, 2026-10-01).
+INDIVIDUAL_ITEMS: dict[int, IndividualWaveItems] = {
+ 1985: IndividualWaveItems(
+ wave=1985,
+ interview=("ER30463", "1985 INTERVIEW NUMBER"),
+ sequence=("ER30464", "SEQUENCE NUMBER 85"),
+ relationship=("ER30465", "RELATIONSHIP TO HEAD 85"),
+ education=("ER30478", "COMPLETED EDUCATION 85"),
+ ),
+ 1986: IndividualWaveItems(
+ wave=1986,
+ interview=("ER30498", "1986 INTERVIEW NUMBER"),
+ sequence=("ER30499", "SEQUENCE NUMBER 86"),
+ relationship=("ER30500", "RELATIONSHIP TO HEAD 86"),
+ education=("ER30513", "COMPLETED EDUCATION 86"),
+ ),
+ 1987: IndividualWaveItems(
+ wave=1987,
+ interview=("ER30535", "1987 INTERVIEW NUMBER"),
+ sequence=("ER30536", "SEQUENCE NUMBER 87"),
+ relationship=("ER30537", "RELATIONSHIP TO HEAD 87"),
+ education=("ER30549", "COMPLETED EDUCATION 87"),
+ ),
+ 1988: IndividualWaveItems(
+ wave=1988,
+ interview=("ER30570", "1988 INTERVIEW NUMBER"),
+ sequence=("ER30571", "SEQUENCE NUMBER 88"),
+ relationship=("ER30572", "RELATION TO HEAD 88"),
+ education=("ER30584", "COMPLETED EDUC-IND 88"),
+ ),
+ 1989: IndividualWaveItems(
+ wave=1989,
+ interview=("ER30606", "1989 INTERVIEW NUMBER"),
+ sequence=("ER30607", "SEQUENCE NUMBER 89"),
+ relationship=("ER30608", "RELATION TO HEAD 89"),
+ education=("ER30620", "COMPLETED EDUC-IND 89"),
+ ),
+ 1990: IndividualWaveItems(
+ wave=1990,
+ interview=("ER30642", "1990 INTERVIEW NUMBER"),
+ sequence=("ER30643", "SEQUENCE NUMBER 90"),
+ relationship=("ER30644", "RELATION TO HEAD 90"),
+ education=("ER30657", "COMPLETED EDUC-IND 90"),
+ ),
+ 1991: IndividualWaveItems(
+ wave=1991,
+ interview=("ER30689", "1991 INTERVIEW NUMBER"),
+ sequence=("ER30690", "SEQUENCE NUMBER 91"),
+ relationship=("ER30691", "RELATION TO HEAD 91"),
+ education=("ER30703", "COMPLETED EDUC-IND 91"),
+ ),
+ 1992: IndividualWaveItems(
+ wave=1992,
+ interview=("ER30733", "1992 INTERVIEW NUMBER"),
+ sequence=("ER30734", "SEQUENCE NUMBER 92"),
+ relationship=("ER30735", "RELATION TO HEAD 92"),
+ education=("ER30748", "COMPLETED EDUCATION 92"),
+ ),
+ 1993: IndividualWaveItems(
+ wave=1993,
+ interview=("ER30806", "1993 INTERVIEW NUMBER"),
+ sequence=("ER30807", "SEQUENCE NUMBER 93"),
+ relationship=("ER30808", "RELATION TO HEAD 93"),
+ education=("ER30820", "YRS COMPLETED EDUCATION 93"),
+ ),
+ 1994: IndividualWaveItems(
+ wave=1994,
+ interview=("ER33101", "1994 INTERVIEW NUMBER"),
+ sequence=("ER33102", "SEQUENCE NUMBER 94"),
+ relationship=("ER33103", "RELATION TO HEAD 94"),
+ education=("ER33115", "YRS COMPLETED EDUC 94"),
+ ),
+ 1995: IndividualWaveItems(
+ wave=1995,
+ interview=("ER33201", "1995 INTERVIEW NUMBER"),
+ sequence=("ER33202", "SEQUENCE NUMBER 95"),
+ relationship=("ER33203", "RELATION TO HEAD 95"),
+ education=("ER33215", "YEARS COMPLETED EDUCATION 95"),
+ ),
+ 1996: IndividualWaveItems(
+ wave=1996,
+ interview=("ER33301", "1996 INTERVIEW NUMBER"),
+ sequence=("ER33302", "SEQUENCE NUMBER 96"),
+ relationship=("ER33303", "RELATION TO HEAD 96"),
+ education=("ER33315", "YEARS COMPLETED EDUCATION 96"),
+ ),
+ 1997: IndividualWaveItems(
+ wave=1997,
+ interview=("ER33401", "1997 INTERVIEW NUMBER"),
+ sequence=("ER33402", "SEQUENCE NUMBER 97"),
+ relationship=("ER33403", "RELATION TO HEAD 97"),
+ education=("ER33415", "YEARS COMPLETED EDUCATION 97"),
+ ),
+ 1999: IndividualWaveItems(
+ wave=1999,
+ interview=("ER33501", "1999 INTERVIEW NUMBER"),
+ sequence=("ER33502", "SEQUENCE NUMBER 99"),
+ relationship=("ER33503", "RELATION TO HEAD 99"),
+ education=("ER33516", "YEARS COMPLETED EDUCATION 99"),
+ ),
+ 2001: IndividualWaveItems(
+ wave=2001,
+ interview=("ER33601", "2001 INTERVIEW NUMBER"),
+ sequence=("ER33602", "SEQUENCE NUMBER 01"),
+ relationship=("ER33603", "RELATION TO HEAD 01"),
+ education=("ER33616", "YEARS COMPLETED EDUCATION 01"),
+ ),
+ 2003: IndividualWaveItems(
+ wave=2003,
+ interview=("ER33701", "2003 INTERVIEW NUMBER"),
+ sequence=("ER33702", "SEQUENCE NUMBER 03"),
+ relationship=("ER33703", "RELATION TO HEAD 03"),
+ education=("ER33716", "YEARS COMPLETED EDUCATION 03"),
+ ),
+ 2005: IndividualWaveItems(
+ wave=2005,
+ interview=("ER33801", "2005 INTERVIEW NUMBER"),
+ sequence=("ER33802", "SEQUENCE NUMBER 05"),
+ relationship=("ER33803", "RELATION TO HEAD 05"),
+ education=("ER33817", "YEARS COMPLETED EDUCATION 05"),
+ ),
+ 2007: IndividualWaveItems(
+ wave=2007,
+ interview=("ER33901", "2007 INTERVIEW NUMBER"),
+ sequence=("ER33902", "SEQUENCE NUMBER 07"),
+ relationship=("ER33903", "RELATION TO HEAD 07"),
+ education=("ER33917", "YEARS COMPLETED EDUCATION 07"),
+ ),
+ 2009: IndividualWaveItems(
+ wave=2009,
+ interview=("ER34001", "2009 INTERVIEW NUMBER"),
+ sequence=("ER34002", "SEQUENCE NUMBER 09"),
+ relationship=("ER34003", "RELATION TO HEAD 09"),
+ education=("ER34020", "YEARS COMPLETED EDUCATION 09"),
+ ),
+ 2011: IndividualWaveItems(
+ wave=2011,
+ interview=("ER34101", "2011 INTERVIEW NUMBER"),
+ sequence=("ER34102", "SEQUENCE NUMBER 11"),
+ relationship=("ER34103", "RELATION TO HEAD 11"),
+ education=("ER34119", "YEARS COMPLETED EDUCATION 11"),
+ ),
+ 2013: IndividualWaveItems(
+ wave=2013,
+ interview=("ER34201", "2013 INTERVIEW NUMBER"),
+ sequence=("ER34202", "SEQUENCE NUMBER 13"),
+ relationship=("ER34203", "RELATION TO HEAD 13"),
+ education=("ER34230", "YEARS COMPLETED EDUCATION 13"),
+ ),
+ 2015: IndividualWaveItems(
+ wave=2015,
+ interview=("ER34301", "2015 INTERVIEW NUMBER"),
+ sequence=("ER34302", "SEQUENCE NUMBER 15"),
+ relationship=("ER34303", "RELATION TO HEAD 15"),
+ education=("ER34349", "YEARS COMPLETED EDUCATION 15"),
+ ),
+ 2017: IndividualWaveItems(
+ wave=2017,
+ interview=("ER34501", "2017 INTERVIEW NUMBER"),
+ sequence=("ER34502", "SEQUENCE NUMBER 17"),
+ relationship=("ER34503", "RELATION TO REFERENCE PERSON 17"),
+ education=("ER34548", "YEARS COMPLETED EDUCATION 17"),
+ ),
+ 2019: IndividualWaveItems(
+ wave=2019,
+ interview=("ER34701", "2019 INTERVIEW NUMBER"),
+ sequence=("ER34702", "SEQUENCE NUMBER 19"),
+ relationship=("ER34703", "RELATION TO REFERENCE PERSON 19"),
+ education=("ER34752", "YEARS COMPLETED EDUCATION 19"),
+ ),
+ 2021: IndividualWaveItems(
+ wave=2021,
+ interview=("ER34901", "2021 INTERVIEW NUMBER"),
+ sequence=("ER34902", "SEQUENCE NUMBER 21"),
+ relationship=("ER34903", "RELATION TO REFERENCE PERSON 21"),
+ education=("ER34952", "YEARS COMPLETED EDUCATION 21"),
+ ),
+ 2023: IndividualWaveItems(
+ wave=2023,
+ interview=("ER35101", "2023 INTERVIEW NUMBER"),
+ sequence=("ER35102", "SEQUENCE NUMBER 23"),
+ relationship=("ER35103", "RELATION TO REFERENCE PERSON 23"),
+ education=("ER35152", "YEARS COMPLETED EDUCATION 23"),
+ ),
+}
+
+
+def _relationship_prefixes(wave: int) -> dict[int, tuple[str, ...]]:
+ """Accepted label prefixes of the role codes in a wave's format."""
+
+ return {
+ HEAD_RELATIONSHIP: (
+ f"head in {wave}",
+ f"reference person in {wave}",
+ ),
+ 20: (f"legal wife in {wave}", f"legal spouse in {wave}"),
+ 22: ('"wife"--', "partner--"),
+ }
+
+
+#: The individual-file education code frame per wave, as
+#: IND2023ER_formats.sas documents it (1-17 the grade completed, 98 DK and
+#: 99 NA where present, 0 Inap.); verified at read time.
+def _expected_education_domain(wave: int) -> CodeDomain:
+ singles = {0, 99} if (wave <= 1993 or wave >= 2013) else {0, 98, 99}
+ return CodeDomain(codes=frozenset(singles), ranges=((1, 17),))
+
+
+# --------------------------------------------------------------------------
+# Documented code domains
+# --------------------------------------------------------------------------
+@dataclass(frozen=True)
+class CodeDomain:
+ """The codes a variable documents: single codes plus closed ranges."""
+
+ codes: frozenset[int]
+ ranges: tuple[tuple[int, int], ...] = ()
+
+ def __post_init__(self) -> None:
+ for code in self.codes:
+ _integer_code(code)
+ for lo, hi in self.ranges:
+ _integer_code(lo)
+ _integer_code(hi)
+ if lo > hi:
+ raise ValueError(f"code-domain range {lo}-{hi} is reversed")
+
+ def contains(self, value: int) -> bool:
+ value = _integer_code(value)
+ if value in self.codes:
+ return True
+ return any(lo <= value <= hi for lo, hi in self.ranges)
+
+ def undocumented(self, values: Iterable[Any]) -> list[int]:
+ """Sorted distinct values the domain does not document."""
+ seen = {_integer_code(v) for v in values if not pd.isna(v)}
+ return sorted(v for v in seen if not self.contains(v))
+
+ def normalized(self) -> tuple[frozenset[int], tuple[tuple[int, int], ...]]:
+ """Merge adjacent singles into ranges for a canonical comparison."""
+ values = set(self.codes)
+ for lo, hi in self.ranges:
+ values.update(range(lo, hi + 1))
+ return frozenset(values), ()
+
+ def same_as(self, other: CodeDomain) -> bool:
+ return self.normalized() == other.normalized()
+
+ @classmethod
+ def from_values(cls, entries: Iterable[Mapping[str, Any]]) -> CodeDomain:
+ codes: set[int] = set()
+ ranges: list[tuple[int, int]] = []
+ for entry in entries:
+ if "range" in entry:
+ lo, hi = entry["range"]
+ ranges.append((int(lo), int(hi)))
+ else:
+ codes.add(int(entry["code"]))
+ return cls(codes=frozenset(codes), ranges=tuple(sorted(ranges)))
+
+
+def _integer_code(value: Any) -> int:
+ """Refuse noninteger codes before a coercion can change their meaning."""
+ if isinstance(value, bool) or not isinstance(value, Integral):
+ raise ValueError(f"PSID code {value!r} is not a finite integer")
+ return int(value)
+
+
+_REPO_ROOT = Path(__file__).resolve().parents[3]
+#: The committed codebook value tables (generator:
+#: ``scripts/build_group_attribute_codebook_values.py``).
+CODEBOOK_VALUES_PATH = (
+ _REPO_ROOT
+ / "data"
+ / "external"
+ / "psid_group_attribute_codebook_values_v1.json"
+)
+#: SHA-256 of the committed table; a different file is refused.
+CODEBOOK_VALUES_SHA256 = (
+ "09fce5627b0271a68eadb2e8748e33fe1ff3e6b0e025bfef9575bb2804ebef63"
+)
+
+
+def load_codebook_values(path: Path | None = None) -> dict[str, Any]:
+ """Load the committed codebook value tables, refusing a changed file.
+
+ With ``path`` given (tests), the SHA-256 pin is not applied.
+ """
+
+ target = CODEBOOK_VALUES_PATH if path is None else Path(path)
+ raw = target.read_bytes()
+ if path is None:
+ digest = hashlib.sha256(raw).hexdigest()
+ if digest != CODEBOOK_VALUES_SHA256:
+ raise ValueError(
+ f"{target} has SHA-256 {digest}, expected the pinned "
+ f"{CODEBOOK_VALUES_SHA256}; regenerate the table with "
+ "scripts/build_group_attribute_codebook_values.py and "
+ "re-adjudicate before changing the pin."
+ )
+ data = json.loads(raw)
+ if data.get("schema_version") != "psid_group_attribute_codebook.v1":
+ raise ValueError(f"{target} is not a v1 codebook value table")
+ return data
+
+
+def documented_domain(
+ codebook: Mapping[str, Any], wave: int, variable: str
+) -> CodeDomain:
+ """The codes ``variable`` documents in ``wave``'s family codebook."""
+
+ try:
+ entry = codebook["family"][str(wave)]["variables"][variable]
+ except KeyError as exc:
+ raise KeyError(
+ f"wave {wave} variable {variable} has no codebook value table"
+ ) from exc
+ return CodeDomain.from_values(entry["values"])
+
+
+@cache
+def _semantic_codebook() -> dict[str, Any]:
+ """Pinned documentation shared by pure semantic domain checks."""
+ return load_codebook_values()
+
+
+@cache
+def _semantic_domain(wave: int, variable: str) -> CodeDomain:
+ return documented_domain(_semantic_codebook(), wave, variable)
+
+
+def _require_semantic_code(wave: int, variable: str, code: Any) -> int:
+ value = _integer_code(code)
+ if not _semantic_domain(wave, variable).contains(value):
+ raise ValueError(
+ f"undocumented code {value} for wave {wave} {variable}"
+ )
+ return value
+
+
+_SAS_VALUE_HEADER = re.compile(r"^\s*VALUE\s+(\S+)\s*$")
+_SAS_VALUE_LINE = re.compile(r"^\s*(-?\d+)(?:\s*-\s*(-?\d+))?\s*=\s*'")
+_SAS_BLOCK_END = re.compile(r"^\s*;\s*$")
+
+
+def parse_sas_value_domains(path: str | Path) -> dict[str, CodeDomain]:
+ """``{format name: CodeDomain}`` from a PSID SAS formats file.
+
+ Unlike :func:`populace_dynamics.data.disability.parse_sas_value_labels`
+ this keeps ranges (``1 - 17 = '...'``), which the education and
+ birthplace items use; continuation lines of long labels are skipped,
+ and each block ends at a line holding only ``;``.
+ """
+
+ path = Path(path)
+ if not path.is_file():
+ raise FileNotFoundError(f"PSID SAS formats file not found: {path}")
+ out: dict[str, CodeDomain] = {}
+ name: str | None = None
+ codes: set[int] = set()
+ ranges: list[tuple[int, int]] = []
+ for line in path.read_text(errors="replace").splitlines():
+ header = _SAS_VALUE_HEADER.match(line)
+ if header:
+ if name is not None or header.group(1) in out:
+ raise ValueError(
+ f"Malformed or duplicate VALUE block in {path}"
+ )
+ name, codes, ranges = header.group(1), set(), []
+ continue
+ if name is None:
+ continue
+ if _SAS_BLOCK_END.match(line):
+ out[name] = CodeDomain(
+ codes=frozenset(codes), ranges=tuple(sorted(ranges))
+ )
+ name = None
+ continue
+ value = _SAS_VALUE_LINE.match(line)
+ if value:
+ lo = int(value.group(1))
+ if value.group(2) is None:
+ codes.add(lo)
+ else:
+ ranges.append((lo, int(value.group(2))))
+ if name is not None:
+ raise ValueError(f"Unterminated VALUE block {name} in {path}")
+ if not out:
+ raise ValueError(f"No VALUE block found in {path}")
+ return out
+
+
+def _format_domain(
+ domains: Mapping[str, CodeDomain],
+ assignments: Mapping[str, str],
+ variable: str,
+) -> CodeDomain | None:
+ return domains.get(assignments.get(variable, f"{variable}F"))
+
+
+def _sas_value_labels(path: Path, fmt: str) -> dict[int, str]:
+ """Single-code labels of one SAS format block (for prefix checks)."""
+
+ labels: dict[int, str] = {}
+ inside = False
+ for line in path.read_text(errors="replace").splitlines():
+ header = _SAS_VALUE_HEADER.match(line)
+ if header:
+ inside = header.group(1) == fmt
+ continue
+ if not inside:
+ continue
+ if _SAS_BLOCK_END.match(line):
+ break
+ match = re.match(r"^\s*(-?\d+)\s*=\s*'((?:[^']|'')*)'", line)
+ if match:
+ labels[int(match.group(1))] = match.group(2).replace("''", "'")
+ return labels
+
+
+def _check_codes(
+ values: pd.Series, domain: CodeDomain, *, context: str
+) -> None:
+ if values.isna().any():
+ raise ValueError(f"{context}: blank codes are not documented")
+ bad = domain.undocumented(values.unique())
+ if bad:
+ raise ValueError(
+ f"{context}: codes {bad[:10]} are not documented "
+ f"(documented: singles {sorted(domain.codes)}, ranges "
+ f"{list(domain.ranges)}). The release may have changed; "
+ "re-adjudicate before reading."
+ )
+
+
+# --------------------------------------------------------------------------
+# Readers
+# --------------------------------------------------------------------------
+def _individual_formats_path(data_dir: Path | None) -> Path:
+ return disability.employment_status_formats_path(data_dir)
+
+
+def verify_individual_formats(
+ *,
+ data_dir: Path | None = None,
+ waves: tuple[int, ...] = INDIVIDUAL_WAVES,
+) -> dict[int, dict[str, str]]:
+ """Check each wave's education and relationship formats.
+
+ The education format must document exactly the expected frame
+ (:func:`_expected_education_domain`); relationship codes 10, 20 and
+ 22 must carry the head, legal wife/spouse and partner labels.
+ Returns ``{wave: {variable: format}}`` for the verified items.
+ """
+
+ path = _individual_formats_path(data_dir)
+ domains = parse_sas_value_domains(path)
+ assignments = disability.parse_sas_format_assignments(path)
+ out: dict[int, dict[str, str]] = {}
+ for wave in waves:
+ items = INDIVIDUAL_ITEMS[wave]
+ edu_var = items.education[0]
+ edu_fmt = assignments.get(edu_var, f"{edu_var}F")
+ documented = domains.get(edu_fmt)
+ expected = _expected_education_domain(wave)
+ if documented is None or not documented.same_as(expected):
+ raise ValueError(
+ f"wave {wave}: {edu_var} format {edu_fmt} documents "
+ f"{documented}, expected {expected}"
+ )
+ rel_var = items.relationship[0]
+ rel_fmt = assignments.get(rel_var, f"{rel_var}F")
+ rel_labels = _sas_value_labels(path, rel_fmt)
+ for code, prefixes in _relationship_prefixes(wave).items():
+ label = " ".join(rel_labels.get(code, "").lower().split())
+ if not label.startswith(prefixes):
+ raise ValueError(
+ f"wave {wave}: relationship code {code} in {rel_fmt} "
+ f"is {label!r}, expected a label starting with one of "
+ f"{prefixes}"
+ )
+ out[wave] = {edu_var: edu_fmt, rel_var: rel_fmt}
+ return out
+
+
+def read_individual_items(
+ *,
+ data_dir: Path | None = None,
+ waves: tuple[int, ...] = INDIVIDUAL_WAVES,
+ nrows: int | None = None,
+) -> pd.DataFrame:
+ """Read the per-wave roster and education codes, long by wave.
+
+ Columns: ``person_id``, ``wave``, ``interview``, ``sequence``,
+ ``relationship``, ``education_code`` (raw, verified against the
+ wave's documented frame). Every person appears once per wave in
+ ``waves``; presence filtering is the caller's.
+ """
+
+ unknown = [w for w in waves if w not in INDIVIDUAL_ITEMS]
+ if unknown:
+ raise ValueError(f"waves {unknown} have no adjudicated layout")
+ if not waves or len(set(waves)) != len(waves):
+ raise ValueError("waves must be nonempty and distinct")
+ labels = psid.parse_sps_labels(
+ psid.product_sps_path("ind2023er", data_dir)
+ )
+ expected = dict(_PERSON_VARS)
+ for wave in waves:
+ expected.update(INDIVIDUAL_ITEMS[wave].labels())
+ psid.verify_labels(labels, expected, context="ind2023er group attributes")
+ verify_individual_formats(data_dir=data_dir, waves=tuple(waves))
+ formats_path = _individual_formats_path(data_dir)
+ format_domains = parse_sas_value_domains(formats_path)
+ format_assignments = disability.parse_sas_format_assignments(formats_path)
+ raw = psid.read_psid(
+ "ind2023er", columns=list(expected), data_dir=data_dir, nrows=nrows
+ )
+ person_id = raw["ER30001"].astype("int64") * 1000 + raw["ER30002"].astype(
+ "int64"
+ )
+ frames = []
+ for wave in waves:
+ items = INDIVIDUAL_ITEMS[wave]
+ for var, _ in (items.sequence, items.relationship):
+ domain = _format_domain(format_domains, format_assignments, var)
+ if domain is None:
+ raise ValueError(f"wave {wave}: {var} has no format domain")
+ _check_codes(raw[var], domain, context=f"ind2023er {var} ({wave})")
+ _check_codes(
+ raw[items.education[0]],
+ _expected_education_domain(wave),
+ context=f"ind2023er {items.education[0]} ({wave})",
+ )
+ frames.append(
+ pd.DataFrame(
+ {
+ "person_id": person_id,
+ "wave": wave,
+ "interview": raw[items.interview[0]].astype("int64"),
+ "sequence": raw[items.sequence[0]].astype("int64"),
+ "relationship": raw[items.relationship[0]].astype("int64"),
+ "education_code": raw[items.education[0]].astype("int64"),
+ }
+ )
+ )
+ return pd.concat(frames, ignore_index=True)
+
+
+def _family_formats_path(wave: int, data_dir: Path | None) -> Path | None:
+ base = psid._resolve_data_dir(data_dir) / "family" / str(wave)
+ hits = sorted(base.glob("*_formats.sas"))
+ if len(hits) > 1:
+ raise ValueError(f"family {wave}: multiple SAS formats files found")
+ return hits[0] if hits else None
+
+
+def read_family_items(
+ wave: int,
+ *,
+ data_dir: Path | None = None,
+ codebook: Mapping[str, Any] | None = None,
+ nrows: int | None = None,
+) -> pd.DataFrame:
+ """Read one family wave's group-attribute codes, one row per family.
+
+ Columns: ``interview`` plus, for each role, the wave's items under
+ ```` names (the raw PSID names; :func:`head_spouse_reports`
+ maps them to generic columns). Every label is verified first, and
+ every observed code must be documented by the wave's codebook value
+ table; where a formats file exists (2021, 2023), each item's format
+ block, when it has one, must document the same codes.
+ """
+
+ if wave not in FAMILY_ITEMS:
+ raise ValueError(f"wave {wave} has no adjudicated family layout")
+ codebook = load_codebook_values() if codebook is None else codebook
+ items = FAMILY_ITEMS[wave]
+ sps_path, txt_path = family._family_paths(wave, data_dir)
+ labels = psid.parse_sps_labels(sps_path)
+ expected = items.labels()
+ psid.verify_labels(labels, expected, context=f"family {wave}")
+ formats_path = _family_formats_path(wave, data_dir)
+ if formats_path is not None:
+ domains = parse_sas_value_domains(formats_path)
+ assignments = disability.parse_sas_format_assignments(formats_path)
+ names = list(expected)
+ layout = psid.parse_sps_layout(sps_path).set_index("name")
+ colspecs = [
+ (int(layout.loc[n, "start"]) - 1, int(layout.loc[n, "end"]))
+ for n in names
+ ]
+ raw = pd.read_fwf(
+ txt_path, colspecs=colspecs, names=names, header=None, nrows=nrows
+ )
+ interview_var = items.interview[0]
+ for var in names:
+ if var == interview_var:
+ continue
+ domain = documented_domain(codebook, wave, var)
+ if formats_path is not None:
+ fmt_domain = _format_domain(domains, assignments, var)
+ if fmt_domain is not None and not fmt_domain.same_as(domain):
+ raise ValueError(
+ f"family {wave} {var}: the formats file documents "
+ f"{fmt_domain}, the codebook {domain}"
+ )
+ _check_codes(raw[var], domain, context=f"family {wave} {var}")
+ out = raw.rename(columns={interview_var: "interview"})
+ if out["interview"].duplicated().any():
+ raise ValueError(f"family {wave}: duplicate interview numbers")
+ return out.astype("int64")
+
+
+def verify_codebook_pins(
+ waves: Iterable[int],
+ *,
+ data_dir: Path | None = None,
+ codebook: Mapping[str, Any] | None = None,
+) -> dict[int, str]:
+ """Check each wave's staged codebook PDF against the table's pin.
+
+ The code meanings were adjudicated from these exact PDFs; a changed
+ PDF means the documentation moved and the table must be rebuilt.
+ Returns ``{wave: sha256}``.
+ """
+
+ codebook = load_codebook_values() if codebook is None else codebook
+ root = psid._resolve_data_dir(data_dir)
+ out: dict[int, str] = {}
+ for wave in waves:
+ source = codebook["family"][str(wave)]
+ entry = source["codebook"]
+ path = root / entry["path"]
+ if not path.is_file():
+ raise FileNotFoundError(
+ f"codebook {path} is not staged; the code meanings for "
+ f"wave {wave} cannot be checked against their source"
+ )
+ digest = hashlib.sha256(path.read_bytes()).hexdigest()
+ if digest != entry["sha256"]:
+ raise ValueError(
+ f"codebook {path} has SHA-256 {digest}, the value table "
+ f"pins {entry['sha256']}"
+ )
+ out[wave] = digest
+ formats = source.get("formats_sas")
+ if formats is not None:
+ format_path = root / formats["path"]
+ if not format_path.is_file():
+ raise FileNotFoundError(
+ f"pinned formats file missing: {format_path}"
+ )
+ format_digest = hashlib.sha256(
+ format_path.read_bytes()
+ ).hexdigest()
+ if format_digest != formats["sha256"]:
+ raise ValueError(
+ f"formats {format_path} has SHA-256 {format_digest}, "
+ f"the value table pins {formats['sha256']}"
+ )
+ return out
+
+
+#: Generic columns of a head/spouse report, in order.
+REPORT_CODE_COLUMNS: tuple[str, ...] = (
+ "hispanic_code",
+ "race_code_1",
+ "race_code_2",
+ "race_code_3",
+ "race_code_4",
+ "family_education_code",
+ "birth_state_code",
+ "year_came_code",
+)
+
+
+def head_spouse_reports(
+ individual: pd.DataFrame, family_items: Mapping[int, pd.DataFrame]
+) -> pd.DataFrame:
+ """Attach each wave's family codes to its present head and spouse.
+
+ A person reports in a wave when present (sequence 1-20) with
+ relationship 10 (head) or 20/22 (spouse); the family row joins on the
+ wave's interview number. Columns: ``person_id``, ``wave``, ``role``
+ and :data:`REPORT_CODE_COLUMNS` (nullable ``Int64``; NA where the
+ wave does not carry the item). Refuses a present head or spouse
+ whose family row is missing, and a family with two heads or two
+ spouses.
+ """
+
+ lo, hi = IN_FAMILY_SEQUENCE
+ # The long roster contains a full person frame for each wave. Partition
+ # it once so presence masks scan each wave only, rather than rescanning
+ # all person-waves for every family file. Positional indices preserve
+ # the input's row order and labels, including nonconsecutive indices.
+ wave_positions = individual.groupby("wave", sort=False).indices
+ frames = []
+ for wave in sorted(family_items):
+ items = FAMILY_ITEMS[wave]
+ positions = wave_positions.get(wave)
+ wave_people = (
+ individual.iloc[positions]
+ if positions is not None
+ else individual.iloc[0:0]
+ )
+ people = wave_people[wave_people["sequence"].between(lo, hi)]
+ roles = pd.Series(pd.NA, index=people.index, dtype="string")
+ roles[people["relationship"] == HEAD_RELATIONSHIP] = "head"
+ roles[people["relationship"].isin(SPOUSE_RELATIONSHIPS)] = "spouse"
+ people = people.assign(role=roles).dropna(subset=["role"])
+ dup = people.duplicated(["interview", "role"], keep=False)
+ if dup.any():
+ raise ValueError(
+ f"wave {wave}: {int(dup.sum())} present persons share a "
+ "family's head or spouse role"
+ )
+ fam = family_items[wave]
+ missing = set(people["interview"]) - set(fam["interview"])
+ if missing:
+ raise ValueError(
+ f"wave {wave}: {len(missing)} present heads/spouses have "
+ "no family-file row"
+ )
+ merged = people.merge(fam, on="interview", how="left")
+ for role in ROLES:
+ part = merged[merged["role"] == role]
+ columns = items.role_columns(role)
+ report = pd.DataFrame(
+ {
+ "person_id": part["person_id"].astype("int64"),
+ "wave": wave,
+ "role": role,
+ }
+ )
+ for column in REPORT_CODE_COLUMNS:
+ report[column] = pd.array([pd.NA] * len(part), dtype="Int64")
+ for var, column in columns.items():
+ report[column] = part[var].astype("Int64").to_numpy()
+ frames.append(report)
+ if not frames:
+ return pd.DataFrame(
+ columns=["person_id", "wave", "role", *REPORT_CODE_COLUMNS]
+ )
+ out = pd.concat(frames, ignore_index=True)
+ out["role"] = out["role"].astype("string")
+ for column in REPORT_CODE_COLUMNS:
+ out[column] = out[column].astype("Int64")
+ return out.sort_values(["person_id", "wave"]).reset_index(drop=True)
+
+
+# --------------------------------------------------------------------------
+# Meanings (pure)
+# --------------------------------------------------------------------------
+def hispanic_meaning(code: int) -> str:
+ """Meaning of a Spanish-descent code (:data:`HISPANIC_CODE_MEANINGS`)."""
+
+ try:
+ return HISPANIC_CODE_MEANINGS[_integer_code(code)]
+ except KeyError as exc:
+ raise ValueError(f"undocumented Spanish-descent code {code}") from exc
+
+
+def race_mention_meaning(wave: int, mention: int, code: int) -> str:
+ """Meaning of race ``mention`` (1-based) code ``code`` in ``wave``."""
+
+ era = FAMILY_ITEMS[wave].race_era
+ if not isinstance(mention, Integral) or isinstance(mention, bool):
+ raise ValueError(f"race mention {mention!r} is not an integer")
+ if not 1 <= mention <= len(FAMILY_ITEMS[wave].race["head"]):
+ raise ValueError(f"wave {wave}: undocumented race mention {mention}")
+ try:
+ meaning = RACE_CODE_MEANINGS[era][_integer_code(code)]
+ except KeyError as exc:
+ raise ValueError(
+ f"undocumented race code {code} (wave {wave}, era {era})"
+ ) from exc
+ if mention == 1 and meaning == NO_FURTHER_MENTION:
+ return MISSING
+ return meaning
+
+
+@dataclass(frozen=True)
+class RaceEthnicityReport:
+ """One head/spouse report's race and Hispanic origin, harmonized.
+
+ ``hispanic`` is True/False, or None when the report cannot tell.
+ ``races`` is the ordered, distinct race categories mentioned (Latino
+ origin mentions excluded; they are an ethnicity answer to the race
+ question), or None when any mention is missing or none names a race.
+ ``hispanic_basis`` is ``"direct_question"`` (a Spanish-descent item),
+ ``"latino_race_mention"`` (1997-2003: a race mention of Latino
+ origin), ``"not_asked"`` (1997-2003 with no such mention) or
+ ``"missing"`` (DK/NA to the direct question), or
+ ``"undocumented_zero_meaning"`` (1994-1996 present spouse code 0).
+ """
+
+ hispanic: bool | None
+ races: tuple[str, ...] | None
+ hispanic_basis: str
+
+ @property
+ def complete(self) -> bool:
+ """Whether the report fixes a four-way race/ethnicity group."""
+ return self.hispanic is True or (
+ self.hispanic is False and self.races is not None
+ )
+
+
+def race_ethnicity_report(
+ wave: int,
+ hispanic_code: int | None,
+ race_codes: Iterable[int | None],
+ *,
+ role: str = "head",
+) -> RaceEthnicityReport:
+ """Harmonize one report's Spanish-descent and race-mention codes.
+
+ ``hispanic_code`` must be None exactly when the wave asks no
+ Spanish-descent question (1997-2003). ``race_codes`` are the wave's
+ mentions in order (trailing None for mentions the wave lacks).
+ Rules: where the wave asks the direct question, Hispanic origin is
+ that answer alone; in 1997-2003 a Latino-origin race mention makes
+ the report Hispanic and its absence leaves Hispanic origin unknown.
+ ``role`` distinguishes a present spouse's ambiguous 1994-1996
+ Spanish-descent code 0: its codebook text names only "no wife", so
+ the report leaves ethnicity unknown with basis
+ ``undocumented_zero_meaning``. A missing (DK/NA/wild) mention at any
+ position leaves the race set
+ unknown, since the person's full set of races is then not known.
+ """
+
+ items = FAMILY_ITEMS[wave]
+ if role not in ROLES:
+ raise ValueError(f"undocumented family role {role!r}")
+ if items.asks_hispanic_origin != (hispanic_code is not None):
+ raise ValueError(
+ f"wave {wave}: Spanish-descent code {hispanic_code!r} given "
+ f"where asks_hispanic_origin={items.asks_hispanic_origin}"
+ )
+ codes = tuple(race_codes)
+ mentions = items.race[role]
+ if not len(mentions) <= len(codes) <= 4:
+ raise ValueError(
+ f"wave {wave}: expected {len(mentions)} race codes with "
+ "optional unasked trailing mentions"
+ )
+ if any(
+ code is not None and not pd.isna(code)
+ for code in codes[len(mentions) :]
+ ):
+ raise ValueError(
+ f"wave {wave}: codes supplied for unasked race mentions"
+ )
+ meanings = []
+ for mention, (variable, _) in enumerate(mentions, start=1):
+ code = _require_semantic_code(wave, variable, codes[mention - 1])
+ meanings.append(race_mention_meaning(wave, mention, code))
+ latino = LATINO_ORIGIN in meanings
+ races: tuple[str, ...] | None
+ if MISSING in meanings:
+ races = None
+ else:
+ ordered = [
+ m for m in meanings if m not in (NO_FURTHER_MENTION, LATINO_ORIGIN)
+ ]
+ races = tuple(dict.fromkeys(ordered)) or None
+ if items.asks_hispanic_origin:
+ hispanic_code = _require_semantic_code(
+ wave, items.hispanic[role][0], hispanic_code
+ )
+ meaning = hispanic_meaning(hispanic_code)
+ if role == "spouse" and 1994 <= wave <= 1996 and hispanic_code == 0:
+ hispanic, basis = None, "undocumented_zero_meaning"
+ else:
+ hispanic = {HISPANIC: True, NOT_HISPANIC: False}.get(meaning)
+ basis = "missing" if hispanic is None else "direct_question"
+ elif latino:
+ hispanic, basis = True, "latino_race_mention"
+ else:
+ hispanic, basis = None, "not_asked"
+ return RaceEthnicityReport(
+ hispanic=hispanic, races=races, hispanic_basis=basis
+ )
+
+
+def birthplace_class(birth_state_code: int, year_came_code: int) -> str:
+ """Classify one report's birthplace pair (:data:`BIRTHPLACE_CLASSES`).
+
+ A FIPS state with year-came 0 is the United States; "territory or
+ foreign country" with year-came 0 ("born in the United States or U.S.
+ territory") is a U.S. territory, and with any other year-came code
+ (a year, "not living in the U.S." or DK/NA, all asked only of persons
+ born outside the U.S. and its territories) a foreign country. DK/NA
+ on the state is missing. Any other pairing contradicts the skip
+ pattern and is ``inconsistent``.
+ """
+
+ state, came = _integer_code(birth_state_code), _integer_code(
+ year_came_code
+ )
+ if came not in (0, 9997, 9998, 9999) and not 1901 <= came <= 2023:
+ raise ValueError(f"undocumented year-came code {came}")
+ lo, hi = BIRTH_STATE_RANGE
+ born_in_us_or_territory = came == YEAR_CAME_BORN_IN_US
+ if lo <= state <= hi:
+ return UNITED_STATES if born_in_us_or_territory else INCONSISTENT
+ if state == BIRTH_STATE_NOT_A_STATE:
+ return US_TERRITORY if born_in_us_or_territory else FOREIGN_COUNTRY
+ if state == BIRTH_STATE_MISSING:
+ return MISSING if born_in_us_or_territory else INCONSISTENT
+ raise ValueError(f"undocumented birth-state code {state}")
+
+
+def education_report(
+ individual_code: int, family_code: int | None = None
+) -> tuple[int | None, str]:
+ """Years of schooling and status from one person-wave's codes.
+
+ ``family_code`` is the family COMPLETED ED code of the person's role
+ when the person is that wave's head or spouse (1993 on), else None.
+ Returns ``(years, status)``: 1-17 are ``reported``; individual 0 with
+ family 0 is ``no_grades`` (0 years); 98/99 are ``missing``; any other
+ individual 0 is ``inapplicable``.
+ """
+
+ code = _integer_code(individual_code)
+ if family_code is not None and not pd.isna(family_code):
+ family_code = _integer_code(family_code)
+ if family_code not in (*range(18), FAMILY_EDUCATION_MISSING):
+ raise ValueError(
+ f"undocumented family education code {family_code}"
+ )
+ lo, hi = EDUCATION_YEARS_RANGE
+ if lo <= code <= hi:
+ return code, EDUCATION_REPORTED
+ if code in EDUCATION_DK_NA:
+ return None, EDUCATION_MISSING
+ if code != 0:
+ raise ValueError(f"undocumented education code {code}")
+ if family_code is not None and not pd.isna(family_code):
+ if int(family_code) == FAMILY_EDUCATION_NO_GRADES:
+ return 0, EDUCATION_NO_GRADES
+ return None, EDUCATION_INAPPLICABLE
diff --git a/tests/cohorts/test_group_attributes.py b/tests/cohorts/test_group_attributes.py
new file mode 100644
index 00000000..c797b647
--- /dev/null
+++ b/tests/cohorts/test_group_attributes.py
@@ -0,0 +1,792 @@
+"""Cohort-side group attributes on INVENTED persons and fixed-width files.
+
+No observation in this module comes from the PSID. The differential
+oracle implements latest-complete report selection with ordinary Python
+records, independently of the pandas builder. The invariants are exact
+requested-person coverage, order-independent resolution, education at the
+anchor cutoff, explicit unknowns, and independently reproducible seals.
+The committed category schemes under ``"data" / "external"`` are read,
+so the module belongs to the artifact tier.
+"""
+
+from __future__ import annotations
+
+import copy
+import dataclasses
+import hashlib
+import json
+from pathlib import Path
+
+import numpy as np
+import pandas as pd
+import pytest
+from hypothesis import given, settings
+from hypothesis import strategies as st
+from pandas.testing import assert_frame_equal
+
+from populace_dynamics.cohorts import group_attributes as cohort
+from populace_dynamics.data import group_attributes_psid as gap
+from tests.data.psid_fixtures import write_product
+
+_REPORT_COLUMNS = ["person_id", "wave", "role", *gap.REPORT_CODE_COLUMNS]
+_EDUCATION_COLUMNS = [
+ "person_id",
+ "wave",
+ "education_code",
+ "role",
+ "family_education_code",
+]
+
+
+def _report(
+ person_id: int,
+ wave: int,
+ *,
+ hispanic: int | None = 0,
+ race: tuple[int, ...] = (1,),
+ role: str = "head",
+ birth_state: int = 1,
+ year_came: int = 0,
+) -> dict:
+ """One INVENTED report, with every actual mention explicitly coded."""
+
+ items = gap.FAMILY_ITEMS[wave]
+ row = dict.fromkeys(_REPORT_COLUMNS, pd.NA)
+ row.update(person_id=person_id, wave=wave, role=role)
+ row["hispanic_code"] = hispanic if items.asks_hispanic_origin else pd.NA
+ for mention in range(1, len(items.race[role]) + 1):
+ row[f"race_code_{mention}"] = (
+ race[mention - 1] if mention <= len(race) else 0
+ )
+ if wave in gap.FAMILY_EDUCATION_WAVES:
+ row["family_education_code"] = 12
+ if wave in gap.BIRTHPLACE_WAVES:
+ row["birth_state_code"] = birth_state
+ row["year_came_code"] = year_came
+ return row
+
+
+def _education(
+ person_id: int,
+ wave: int,
+ years: int,
+ *,
+ role: str | None = None,
+ family_code: int | None = None,
+) -> dict:
+ return {
+ "person_id": person_id,
+ "wave": wave,
+ "education_code": years,
+ "role": pd.NA if role is None else role,
+ "family_education_code": (
+ pd.NA if family_code is None else family_code
+ ),
+ }
+
+
+def _inputs(
+ reports: list[dict] | None = None,
+ education: list[dict] | None = None,
+ *,
+ universe: tuple[int, ...] = (1001, 1002, 1003),
+) -> cohort.GroupAttributeInputs:
+ race = pd.DataFrame(reports or [], columns=_REPORT_COLUMNS)
+ edu = pd.DataFrame(education or [], columns=_EDUCATION_COLUMNS)
+ for frame in (race, edu):
+ for column in frame:
+ frame[column] = frame[column].astype(
+ "string" if column == "role" else "Int64"
+ )
+ return cohort.GroupAttributeInputs(
+ reports=race,
+ education=edu,
+ universe=pd.Series(universe, name="person_id", dtype="int64"),
+ provenance={"kind": "INVENTED", "fixture": "group attributes"},
+ )
+
+
+def _build(inputs, person_ids=(1001, 1002, 1003), waves=(2011,)):
+ return cohort.build_group_attributes(
+ inputs, person_ids, anchor_waves=waves
+ )
+
+
+def test_latest_complete_report_wins_and_unknown_person_is_retained():
+ inputs = _inputs(
+ [
+ _report(1001, 2005, race=(1,)),
+ _report(1001, 2011, race=(2,)),
+ _report(1001, 2023, hispanic=9, race=(1,)),
+ _report(1002, 2023, hispanic=1, race=(9,)),
+ ],
+ [_education(1001, 2011, 12), _education(1003, 2011, 14)],
+ )
+ result = _build(inputs, [1003, 1002, 1001])
+ frame = result.frame.set_index("person_id")
+ assert result.frame.person_id.tolist() == [1001, 1002, 1003]
+ assert tuple(result.frame.columns) == cohort.FRAME_COLUMNS
+ assert frame.loc[1001, "race_ethnicity_mint8"] == (
+ "Black or African American, non-Hispanic"
+ )
+ assert frame.loc[1001, "race_ethnicity_source_wave"] == 2011
+ assert frame.loc[1001, "race_ethnicity_n_reports"] == 2
+ assert frame.loc[1001, "race_ethnicity_n_distinct"] == 2
+ assert frame.loc[1002, "race_ethnicity_mint8"] == (
+ "Hispanic or Latino, any race"
+ )
+ assert frame.loc[1002, "race_ethnicity_source_wave"] == 2023
+ assert frame.loc[1003, "race_ethnicity_status"] == ("never_head_or_spouse")
+ assert pd.isna(frame.loc[1003, "race_ethnicity_mint8"])
+ assert frame.loc[1003, "education_mint8"] == "Associate"
+ assert result.provenance["n_persons"] == 3
+ assert result.provenance["sourced_after_cutoff"]["race_ethnicity"] == 1
+ assert result.provenance["status_counts"]["race_ethnicity_status"] == {
+ "known": 2,
+ "never_head_or_spouse": 1,
+ }
+
+
+def test_all_never_head_spouse_still_get_rows_and_status_counts():
+ result = _build(_inputs())
+ assert result.frame.person_id.tolist() == [1001, 1002, 1003]
+ for column in ("race_ethnicity_status", "country_of_birth_status"):
+ assert result.frame[column].tolist() == ["never_head_or_spouse"] * 3
+ assert result.provenance["status_counts"][column] == {
+ "never_head_or_spouse": 3
+ }
+ assert (
+ result.frame.education_status.tolist() == ["no_report_by_cutoff"] * 3
+ )
+ for column in (
+ "race_ethnicity_mint8",
+ "education_mint8",
+ "education_report3",
+ "country_of_birth_mint8",
+ ):
+ assert result.frame[column].isna().all()
+ assert (
+ result.frame[f"{column}_status"].tolist()
+ == ["attribute_unknown"] * 3
+ )
+
+
+def test_hispanic_fallback_does_not_invent_race_or_non_hispanic_origin():
+ frame = _build(
+ _inputs(
+ [
+ _report(1001, 2011, hispanic=0, race=(9,)),
+ _report(1001, 2023, hispanic=9, race=(1,)),
+ _report(1002, 1999, hispanic=None, race=(1,)),
+ _report(1003, 1999, hispanic=None, race=(5,)),
+ ]
+ )
+ ).frame.set_index("person_id")
+ assert not frame.loc[1001, "hispanic"]
+ assert frame.loc[1001, "hispanic_source_wave"] == 2011
+ assert frame.loc[1001, "race_ethnicity_status"] == "dk_na_refused"
+ assert pd.isna(frame.loc[1001, "race_ethnicity_mint8"])
+ assert pd.isna(frame.loc[1002, "hispanic"])
+ assert frame.loc[1002, "race_ethnicity_status"] == (
+ "hispanic_origin_not_asked"
+ )
+ assert frame.loc[1003, "hispanic"]
+ assert frame.loc[1003, "race_ethnicity_hispanic_basis"] == (
+ "latino_race_mention"
+ )
+ assert frame.loc[1003, "race_ethnicity_report4"] == "Hispanic"
+
+
+def test_distinct_race_count_uses_four_way_semantics_and_not_mention_order():
+ frame = _build(
+ _inputs(
+ [
+ _report(1001, 2005, race=(1, 2)),
+ _report(1001, 2011, race=(2, 1)),
+ _report(1001, 2023, race=(4,)),
+ ]
+ )
+ ).frame.set_index("person_id")
+ assert frame.loc[1001, "race_ethnicity_n_reports"] == 3
+ assert frame.loc[1001, "race_ethnicity_n_distinct"] == 1
+ assert frame.loc[1001, "race_ethnicity_mint8"] == (
+ "All other races, non-Hispanic"
+ )
+
+
+def test_undocumented_spouse_zero_does_not_override_known_ethnicity():
+ frame = _build(
+ _inputs(
+ [
+ _report(1001, 1993, hispanic=1, role="spouse"),
+ _report(1001, 1994, hispanic=0, role="spouse"),
+ _report(1002, 1996, hispanic=0, role="spouse"),
+ ]
+ )
+ ).frame.set_index("person_id")
+ assert frame.loc[1001, "race_ethnicity_source_wave"] == 1993
+ assert frame.loc[1001, "race_ethnicity_report4"] == "Hispanic"
+ assert pd.isna(frame.loc[1002, "race_ethnicity_report4"])
+ assert pd.isna(frame.loc[1002, "hispanic"])
+ assert frame.loc[1002, "race_ethnicity_n_reports"] == 0
+
+
+def test_education_uses_latest_known_by_cutoff_and_family_zero_adjudication():
+ inputs = _inputs(
+ education=[
+ _education(1001, 2009, 12),
+ _education(1001, 2011, 99),
+ _education(1001, 2023, 17),
+ _education(1002, 2011, 0, role="head", family_code=0),
+ _education(1003, 2011, 0),
+ ]
+ )
+ frame = _build(inputs, waves=(2011, 2009, 2011)).frame.set_index(
+ "person_id"
+ )
+ assert frame.loc[1001, "education_years"] == 12
+ assert frame.loc[1001, "education_source_wave"] == 2009
+ assert frame.loc[1001, "education_n_reports"] == 1
+ assert frame.loc[1002, "education_years"] == 0
+ assert frame.loc[1002, "education_source_variables"] == (
+ gap.INDIVIDUAL_ITEMS[2011].education[0]
+ + "|"
+ + gap.FAMILY_ITEMS[2011].completed_education["head"][0]
+ )
+ assert pd.isna(frame.loc[1003, "education_years"])
+ assert frame.loc[1003, "education_status"] == "no_report_by_cutoff"
+ later = _build(inputs, waves=(2023,)).frame.set_index("person_id")
+ assert later.loc[1001, "education_years"] == 17
+ assert later.loc[1001, "education_n_distinct"] == 2
+
+
+def test_country_latest_valid_pair_wins_and_territory_stays_unresolved():
+ frame = _build(
+ _inputs(
+ [
+ _report(1001, 2013, birth_state=1),
+ _report(1001, 2015, birth_state=0, year_came=1980),
+ _report(1001, 2023, birth_state=1, year_came=1980),
+ _report(1002, 2023, birth_state=0),
+ _report(1003, 2023, birth_state=99),
+ ]
+ )
+ ).frame.set_index("person_id")
+ assert frame.loc[1001, "country_of_birth"] == "foreign_country"
+ assert frame.loc[1001, "country_of_birth_source_wave"] == 2015
+ assert frame.loc[1001, "country_of_birth_n_reports"] == 2
+ assert frame.loc[1001, "country_of_birth_n_distinct"] == 2
+ assert frame.loc[1001, "country_of_birth_mint8"] == "Other countries"
+ assert frame.loc[1002, "country_of_birth_status"] == "known"
+ assert pd.isna(frame.loc[1002, "country_of_birth_mint8"])
+ assert frame.loc[1002, "country_of_birth_mint8_status"] == (
+ "unresolved:us_territory"
+ )
+ assert frame.loc[1003, "country_of_birth_status"] == "dk_na_refused"
+
+
+@pytest.mark.parametrize(
+ ("years", "mint_label"),
+ [
+ (0, "Less than high school"),
+ (11, "Less than high school"),
+ (12, "High school"),
+ (13, "High school"),
+ (14, "Associate"),
+ (15, "Associate"),
+ (16, "Bachelor"),
+ (17, "Graduate"),
+ ],
+)
+def test_mint_boundaries_and_unrecorded_report_definitions(years, mint_label):
+ inputs = _inputs(
+ education=[
+ _education(1001, 2011, years, role="head", family_code=years)
+ ]
+ )
+ row = _build(inputs, (1001,)).frame.iloc[0]
+ assert row.education_mint8 == mint_label
+ assert row.education_mint8_status == "assigned"
+ assert pd.isna(row.education_report3)
+ assert row.education_report3_status == (
+ "unresolved:definition_not_recorded"
+ )
+
+
+@given(
+ st.dictionaries(
+ st.sampled_from((2005, 2007, 2009, 2011, 2013, 2023)),
+ st.tuples(st.sampled_from((0, 1, 9)), st.sampled_from((1, 2, 4, 9))),
+ min_size=1,
+ )
+)
+@settings(max_examples=40, deadline=None)
+def test_latest_complete_selection_matches_independent_record_oracle(codes):
+ reports = [
+ _report(1001, wave, hispanic=hispanic, race=(race,))
+ for wave, (hispanic, race) in codes.items()
+ ]
+ row = _build(_inputs(reports), (1001,)).frame.iloc[0]
+ # This oracle deliberately does not call the reader's harmonization,
+ # cohort helper, pandas grouping, or data-driven category mapper.
+ complete = []
+ for wave, (hispanic, race) in sorted(codes.items()):
+ if hispanic == 1:
+ complete.append((wave, "Hispanic or Latino, any race"))
+ elif hispanic == 0 and race != 9:
+ label = {
+ 1: "White, non-Hispanic",
+ 2: "Black or African American, non-Hispanic",
+ 4: "All other races, non-Hispanic",
+ }[race]
+ complete.append((wave, label))
+ assert row.race_ethnicity_n_reports == len(complete)
+ assert row.race_ethnicity_n_distinct == len({r[1] for r in complete})
+ if complete:
+ assert row.race_ethnicity_source_wave == complete[-1][0]
+ assert row.race_ethnicity_mint8 == complete[-1][1]
+ else:
+ assert pd.isna(row.race_ethnicity_source_wave)
+ assert pd.isna(row.race_ethnicity_mint8)
+
+
+@given(
+ st.dictionaries(
+ st.sampled_from((2005, 2007, 2009, 2011, 2013, 2023)),
+ st.sampled_from((0, 1, 11, 12, 14, 16, 17, 99)),
+ min_size=1,
+ ),
+ st.sampled_from((2005, 2009, 2011, 2023)),
+)
+@settings(max_examples=40, deadline=None)
+def test_education_latest_known_matches_independent_cutoff_oracle(
+ codes, cutoff
+):
+ education = [_education(1001, wave, code) for wave, code in codes.items()]
+ row = _build(_inputs(education=education), (1001,), (cutoff,)).frame.iloc[
+ 0
+ ]
+ known = sorted(
+ (wave, code)
+ for wave, code in codes.items()
+ if wave <= cutoff and 1 <= code <= 17
+ )
+ assert row.education_n_reports == len(known)
+ assert row.education_n_distinct == len({r[1] for r in known})
+ if known:
+ assert row.education_source_wave == known[-1][0]
+ assert row.education_years == known[-1][1]
+ assert row.education_source_wave <= cutoff
+ else:
+ assert pd.isna(row.education_source_wave)
+ assert pd.isna(row.education_years)
+
+
+@given(st.permutations((1001, 1002, 1003)))
+@settings(deadline=None)
+def test_requested_person_order_does_not_change_frame_or_provenance(order):
+ inputs = _inputs(
+ [_report(1001, 2011), _report(1002, 2023, hispanic=1)],
+ [_education(1003, 2011, 13)],
+ )
+ expected = _build(inputs)
+ reordered = dataclasses.replace(
+ inputs,
+ reports=inputs.reports.iloc[::-1].reset_index(drop=True),
+ education=inputs.education.iloc[::-1].reset_index(drop=True),
+ )
+ actual = _build(reordered, order)
+ assert_frame_equal(actual.frame, expected.frame)
+ assert {
+ k: v
+ for k, v in actual.provenance.items()
+ if k != "input_frames_sha256"
+ } == {
+ k: v
+ for k, v in expected.provenance.items()
+ if k != "input_frames_sha256"
+ }
+
+
+@pytest.mark.parametrize("ids", [(), (1001, 1001), (9999,)])
+def test_empty_duplicate_and_absent_requested_ids_are_refused(ids):
+ with pytest.raises(ValueError):
+ _build(_inputs(), ids)
+
+
+@pytest.mark.parametrize("bad", [1001.5, 1001.0, "1001", True, None])
+def test_requested_ids_require_exact_integer_type(bad):
+ with pytest.raises(ValueError, match="integer"):
+ _build(_inputs(universe=(1, 1001)), (bad,))
+
+
+def test_numpy_integer_identifiers_are_accepted():
+ result = _build(_inputs(), (np.int64(1001),), (np.int64(2011),))
+ assert result.frame.person_id.tolist() == [1001]
+ assert json.loads(json.dumps(result.provenance))["anchor_waves"] == [2011]
+
+
+@pytest.mark.parametrize("waves", [(), (2010,), (2011.5,), ("2011",), (True,)])
+def test_missing_unsupported_or_coerced_anchor_waves_are_refused(waves):
+ with pytest.raises(ValueError):
+ _build(_inputs(), waves=waves)
+
+
+@pytest.mark.parametrize("name", ["reports", "education"])
+def test_duplicate_person_wave_input_is_refused(name):
+ inputs = _inputs([_report(1001, 2011)], [_education(1001, 2011, 12)])
+ duplicated = pd.concat([getattr(inputs, name)] * 2, ignore_index=True)
+ with pytest.raises(ValueError, match="more than one row"):
+ _build(dataclasses.replace(inputs, **{name: duplicated}))
+
+
+def test_builder_is_pure_and_independent_seals_bind_values_types_and_columns():
+ inputs = _inputs([_report(1001, 2011)], [_education(1001, 2011, 12)])
+ before = copy.deepcopy(inputs)
+ digest = cohort.input_frames_sha256(inputs)
+ result = _build(inputs)
+ assert_frame_equal(inputs.reports, before.reports)
+ assert_frame_equal(inputs.education, before.education)
+ assert inputs.universe.equals(before.universe)
+ assert inputs.provenance == before.provenance
+ assert cohort.input_frames_sha256(inputs) == digest
+ assert result.provenance["inputs"] == inputs.provenance
+ assert (
+ result.provenance["person_ids_sha256"]
+ == hashlib.sha256(b"1001,1002,1003").hexdigest()
+ )
+ assert result.provenance["content_sha256"] == cohort.content_sha256(
+ result.frame
+ )
+ assert result.provenance["content_sha256"] == _independent_frame_digest(
+ "group_attributes", result.frame
+ )
+ changed = result.frame.copy()
+ changed.loc[0, "education_years"] = 16
+ assert (
+ cohort.content_sha256(changed) != result.provenance["content_sha256"]
+ )
+ assert cohort.content_sha256(result.frame.iloc[:, ::-1]) != (
+ result.provenance["content_sha256"]
+ )
+ changed = result.frame.astype({"education_years": "object"})
+ assert (
+ cohort.content_sha256(changed) != result.provenance["content_sha256"]
+ )
+ changed_inputs = dataclasses.replace(
+ inputs, education=inputs.education.assign(education_code=16)
+ )
+ assert cohort.input_frames_sha256(changed_inputs) != digest
+
+
+def _independent_frame_digest(name: str, frame: pd.DataFrame) -> str:
+ digest = hashlib.sha256()
+ digest.update(f"{name}\n".encode())
+ digest.update(json.dumps([str(c) for c in frame.columns]).encode())
+ digest.update(json.dumps([str(d) for d in frame.dtypes]).encode())
+ digest.update(frame.to_csv(index=False, lineterminator="\n").encode())
+ return digest.hexdigest()
+
+
+def _write_loader_fixture(root: Path) -> dict:
+ """INVENTED one-wave source, including a pinned fake source document."""
+
+ wave = 2011
+ individual = gap.INDIVIDUAL_ITEMS[wave]
+ fields = [
+ ("ER30001", 4, "1968 INTERVIEW NUMBER", [1, 1, 1]),
+ ("ER30002", 3, "PERSON NUMBER 68", [1, 2, 3]),
+ ]
+ for item, values in (
+ (individual.interview, [11, 11, 11]),
+ (individual.sequence, [1, 2, 3]),
+ (individual.relationship, [10, 20, 30]),
+ (individual.education, [0, 16, 13]),
+ ):
+ fields.append((*item[:1], 4, item[1], values))
+ write_product(root / "ind2023er", "IND2023ER.sps", "IND2023ER.txt", fields)
+ formats = root / "ind2023er" / "IND2023ER_formats.sas"
+ formats.write_text(
+ f"VALUE {individual.education[0]}F\n"
+ "0 = 'Inap.'\n1 - 17 = 'INVENTED grades'\n"
+ "98 = 'DK'\n99 = 'NA'\n;\n"
+ f"VALUE {individual.relationship[0]}F\n"
+ "10 = 'Head in 2011'\n20 = 'Legal wife in 2011'\n"
+ "22 = '\"Wife\"--partner'\n30 = 'Child'\n;\n"
+ f"VALUE {individual.sequence[0]}F\n"
+ "1 - 20 = 'INVENTED present family sequence'\n;\n"
+ )
+ family = gap.FAMILY_ITEMS[wave]
+ family_values = {family.interview[0]: 11}
+ for role, first_race in (("head", 1), ("spouse", 2)):
+ family_values[family.hispanic[role][0]] = 0
+ for mention, (variable, _) in enumerate(family.race[role], start=1):
+ family_values[variable] = first_race if mention == 1 else 0
+ family_values[family.completed_education[role][0]] = (
+ 0 if role == "head" else 16
+ )
+ fields = [
+ (var, 4, label, [family_values[var]])
+ for var, label in family.labels().items()
+ ]
+ write_product(
+ root / "family" / "2011", "FAM2011ER.sps", "FAM2011ER.txt", fields
+ )
+ documentation = root / "documentation"
+ documentation.mkdir()
+ fake_document = documentation / "INVENTED-codebook.txt"
+ fake_document.write_text("INVENTED source document for audit tests.\n")
+ return {
+ "schema_version": "psid_group_attribute_codebook.v1",
+ "family": {
+ "2011": {
+ "codebook": {
+ "path": str(fake_document.relative_to(root)),
+ "sha256": hashlib.sha256(
+ fake_document.read_bytes()
+ ).hexdigest(),
+ },
+ "variables": {
+ var: {"values": [{"code": value, "label": "INVENTED"}]}
+ for var, value in family_values.items()
+ if var != family.interview[0]
+ },
+ }
+ },
+ }
+
+
+def test_fixed_width_cohort_loader_audits_sources_and_retains_ofum(
+ tmp_path, monkeypatch
+):
+ codebook = _write_loader_fixture(tmp_path)
+ original_reader = gap.read_individual_items
+ monkeypatch.setattr(gap, "load_codebook_values", lambda: codebook)
+ monkeypatch.setattr(gap, "RACE_WAVES", (2011,))
+ monkeypatch.setattr(
+ gap,
+ "read_individual_items",
+ lambda *, data_dir: original_reader(data_dir=data_dir, waves=(2011,)),
+ )
+ inputs = cohort.load_group_attribute_inputs(psid_dir=tmp_path)
+ result = _build(inputs)
+ frame = result.frame.set_index("person_id")
+ assert frame.loc[1001, "education_years"] == 0
+ assert frame.loc[1002, "race_ethnicity_report4"] == "Black, non-hispanic"
+ assert frame.loc[1003, "race_ethnicity_status"] == "never_head_or_spouse"
+ assert frame.loc[1003, "education_years"] == 13
+ assert inputs.provenance["input_frames_sha256"] == (
+ cohort.input_frames_sha256(inputs)
+ )
+ expected_files = {
+ "documentation/INVENTED-codebook.txt",
+ "family/2011/FAM2011ER.sps",
+ "family/2011/FAM2011ER.txt",
+ "ind2023er/IND2023ER.sps",
+ "ind2023er/IND2023ER.txt",
+ "ind2023er/IND2023ER_formats.sas",
+ }
+ files = inputs.provenance["psid_files_sha256"]
+ assert set(files) == expected_files
+ for path, digest in files.items():
+ assert (
+ digest
+ == hashlib.sha256((tmp_path / path).read_bytes()).hexdigest()
+ )
+ bundle = (
+ json.dumps(files, sort_keys=True, separators=(",", ":")) + "\n"
+ ).encode()
+ assert (
+ inputs.provenance["psid_files_bundle_sha256"]
+ == hashlib.sha256(bundle).hexdigest()
+ )
+
+
+def test_custom_education_scheme_and_unresolved_bands_are_data_driven():
+ spec = {
+ "bands": [{"min": 0, "max": 11, "label": "INVENTED lower"}],
+ "unresolved_bands": [
+ {"min": 12, "max": None, "status": "INVENTED unsupported"}
+ ],
+ }
+ frame = pd.DataFrame(
+ {"education_years": pd.array([11, 12, pd.NA], dtype="Int64")},
+ index=[9, 4, 7],
+ )
+ labels, statuses = cohort.apply_category_scheme(frame, "education", spec)
+ assert labels.index.tolist() == [9, 4, 7]
+ assert labels.iloc[0] == "INVENTED lower"
+ assert labels.iloc[1:].isna().all()
+ assert statuses.tolist() == [
+ "assigned",
+ "unresolved:INVENTED unsupported",
+ "attribute_unknown",
+ ]
+
+
+@pytest.mark.parametrize("years", [12.5, 12.0, "12", True, -1, 18])
+def test_category_mapping_refuses_nonintegral_or_unsupported_schooling(years):
+ spec = {
+ "bands": [{"min": 0, "max": 17, "label": "INVENTED supported"}],
+ "unresolved_bands": [],
+ }
+ with pytest.raises(ValueError):
+ cohort.apply_category_scheme(
+ pd.DataFrame({"education_years": [years]}), "education", spec
+ )
+
+
+def test_every_attribute_traces_the_selected_source_and_mention():
+ row = _build(
+ _inputs(
+ [
+ _report(
+ 1001, 2023, role="spouse", birth_state=0, year_came=1999
+ )
+ ],
+ [_education(1001, 2011, 0, role="head", family_code=0)],
+ ),
+ (1001,),
+ ).frame.iloc[0]
+ items = gap.FAMILY_ITEMS[2023]
+ assert row.race_ethnicity_source_role == "spouse"
+ assert row.race_ethnicity_source_mentions == "1"
+ assert row.hispanic_source_variables == items.hispanic["spouse"][0]
+ assert row.hispanic_source_mentions == "direct_question"
+ assert row.hispanic_source_wave == 2023
+ assert row.education_source_role == "head"
+ assert row.education_source_mentions == "individual|family_recode"
+ assert row.country_of_birth_source_role == "spouse"
+ assert row.country_of_birth_source_mentions == "birth_state|year_came"
+ assert row.country_of_birth_source_variables == (
+ items.birth_state["spouse"][0] + "|" + items.year_came["spouse"][0]
+ )
+
+
+def test_mutated_audited_inputs_are_refused():
+ inputs = _inputs(education=[_education(1001, 2011, 12)])
+ sealed = dataclasses.replace(
+ inputs,
+ provenance={"input_frames_sha256": cohort.input_frames_sha256(inputs)},
+ )
+ sealed.education.loc[0, "education_code"] = 16
+ with pytest.raises(
+ ValueError, match="changed after their provenance seal"
+ ):
+ _build(sealed)
+
+
+@pytest.mark.parametrize(
+ "ids,waves", [([1001.5], [2011]), ([1001], [2010]), ([], [2011])]
+)
+def test_loader_refuses_malformed_request_before_opening_psid(
+ monkeypatch, ids, waves
+):
+ def forbidden(**kwargs):
+ raise AssertionError("PSID opened for a malformed request")
+
+ monkeypatch.setattr(cohort, "load_group_attribute_inputs", forbidden)
+ with pytest.raises(ValueError):
+ cohort.load_group_attributes(ids, anchor_waves=waves)
+
+
+@pytest.mark.parametrize(
+ "wave,admitted",
+ [(1993, False), (1994, True), (2011, True), (2013, False), (2023, False)],
+)
+def test_supplied_education_codes_follow_reader_wave_domain(wave, admitted):
+ inputs = _inputs(education=[_education(1001, wave, 98)])
+ if admitted:
+ row = _build(inputs, (1001,), (wave,)).frame.iloc[0]
+ assert pd.isna(row.education_years)
+ assert row.education_status == "dk_na_refused"
+ else:
+ with pytest.raises(ValueError, match="not documented"):
+ _build(inputs, (1001,), (wave,))
+
+
+@pytest.mark.parametrize(
+ "wave,role,code,reason",
+ [
+ (2023, None, 0, "requires a head or spouse role"),
+ (1992, "head", 0, "not asked"),
+ (2023, "head", 18, "not documented"),
+ (2023, "spouse", 98, "not documented"),
+ ],
+)
+def test_supplied_family_education_requires_role_wave_and_domain(
+ wave, role, code, reason
+):
+ inputs = _inputs(
+ education=[_education(1001, wave, 0, role=role, family_code=code)]
+ )
+ with pytest.raises(ValueError, match=reason):
+ _build(inputs, (1001,), (wave,))
+
+
+@pytest.mark.parametrize("wave,code", [(1992, 0), (2023, 98)])
+def test_supplied_report_family_education_requires_asked_documented_code(
+ wave, code
+):
+ report = _report(1001, wave)
+ report["family_education_code"] = code
+ with pytest.raises(ValueError, match="not asked|not documented"):
+ _build(_inputs([report]), (1001,), (wave,))
+
+
+@pytest.mark.parametrize(
+ "wave,role,year,admitted",
+ [
+ (2013, "head", 1900, False),
+ (2013, "head", 2014, False),
+ (2013, "head", 9998, True),
+ (2023, "head", 1901, True),
+ (2023, "spouse", 1901, False),
+ (2023, "spouse", 1920, True),
+ (2023, "head", 9998, False),
+ ],
+)
+def test_supplied_birthplace_codes_follow_wave_and_role_domain(
+ wave, role, year, admitted
+):
+ inputs = _inputs(
+ [_report(1001, wave, role=role, birth_state=0, year_came=year)]
+ )
+ if admitted:
+ row = _build(inputs, (1001,), (wave,)).frame.iloc[0]
+ assert row.country_of_birth == "foreign_country"
+ else:
+ with pytest.raises(ValueError, match="not documented"):
+ _build(inputs, (1001,), (wave,))
+
+
+@pytest.mark.parametrize(
+ "column,code",
+ [
+ ("birth_state_code", 57),
+ ("year_came_code", -1),
+ ("birth_state_code", pd.NA),
+ ("year_came_code", pd.NA),
+ ],
+)
+def test_supplied_birthplace_pairs_refuse_undocumented_or_blank_codes(
+ column, code
+):
+ report = _report(1001, 2023)
+ report[column] = code
+ with pytest.raises(ValueError, match="not documented"):
+ _build(_inputs([report]), (1001,), (2023,))
+
+
+@pytest.mark.parametrize("column", ["birth_state_code", "year_came_code"])
+def test_supplied_birthplace_codes_are_refused_before_question_exists(column):
+ report = _report(1001, 2011)
+ report[column] = 0
+ with pytest.raises(ValueError, match="not asked"):
+ _build(_inputs([report]), (1001,))
+
+
+def test_supplied_domains_are_checked_even_after_education_cutoff():
+ inputs = _inputs(education=[_education(1001, 2023, 98)])
+ with pytest.raises(ValueError, match="not documented"):
+ _build(inputs, (1001,), (2011,))
diff --git a/tests/cohorts/test_group_attributes_integration.py b/tests/cohorts/test_group_attributes_integration.py
new file mode 100644
index 00000000..8ae0b84d
--- /dev/null
+++ b/tests/cohorts/test_group_attributes_integration.py
@@ -0,0 +1,182 @@
+"""Staged-PSID side-frame coverage for the four blind-test populations.
+
+Skipped without the products under ``~/PolicyEngine/psid-data``. The
+official cohort readers and selectors determine the population; these
+tests run no projection, benefit rule, poverty measure or tabulation.
+Only identifiers, labels, documented code domains, source seals and
+attribute availability are checked. No group distribution is computed
+or pinned. Bulky cohort inputs are released between populations.
+"""
+
+from __future__ import annotations
+
+import gc
+from collections.abc import Iterable
+from pathlib import Path
+from types import SimpleNamespace
+
+import pandas as pd
+import pytest
+
+from populace_dynamics.cohorts import age67, group_attributes, psid2010
+from populace_dynamics.data import family, family_income
+from populace_dynamics.data import group_attributes_psid as gap
+from populace_dynamics.min_benefit_track_m import cohort as track_m
+from populace_dynamics.min_benefit_track_m import structure
+
+REAL_DATA = Path("~/PolicyEngine/psid-data").expanduser()
+_CODEBOOKS = gap.load_codebook_values()["family"]
+_NEEDED = (
+ REAL_DATA / "ind2023er" / "IND2023ER.txt",
+ REAL_DATA / "ind2023er" / "IND2023ER.sps",
+ REAL_DATA / "ind2023er" / "IND2023ER_formats.sas",
+ REAL_DATA / "mh85_23" / "MH85_23.txt",
+ REAL_DATA / "mh85_23" / "MH85_23.sps",
+ *(
+ REAL_DATA / "family" / str(wave)
+ for wave in set(family.FAMILY_WAVES) | set(gap.RACE_WAVES)
+ ),
+ *(
+ REAL_DATA / _CODEBOOKS[str(wave)]["codebook"]["path"]
+ for wave in gap.RACE_WAVES
+ ),
+ *(
+ REAL_DATA / "wealth" / str(wave) / name
+ for wave, pins in family_income.WEALTH_SUPPLEMENT_SHA256.items()
+ for name in pins
+ ),
+)
+needs_real_psid = pytest.mark.skipif(
+ not all(path.exists() for path in _NEEDED),
+ reason="staged PSID products, family codebooks or wealth absent",
+)
+
+
+@pytest.fixture(scope="module")
+def attribute_inputs():
+ # This loader verifies every selected .sps label, individual-format
+ # domain, family-codebook domain and staged-codebook SHA-256 before
+ # returning data. Its audit records the files those checks opened.
+ inputs = group_attributes.load_group_attribute_inputs(psid_dir=REAL_DATA)
+ provenance = inputs.provenance
+ assert provenance["kind"] == "psid_files"
+ assert inputs.universe.is_unique
+ assert not inputs.reports.duplicated(["person_id", "wave"]).any()
+ assert not inputs.education.duplicated(["person_id", "wave"]).any()
+ assert provenance["psid_files_sha256"]
+ assert all(
+ len(digest) == 64
+ for digest in provenance["psid_files_sha256"].values()
+ )
+ assert set(provenance["codebook_pdf_sha256"]) == {
+ str(wave) for wave in gap.RACE_WAVES
+ }
+ assert provenance["input_frames_sha256"] == (
+ group_attributes.input_frames_sha256(inputs)
+ )
+ yield inputs
+ del inputs
+ gc.collect()
+
+
+def _assert_coverage(
+ inputs: group_attributes.GroupAttributeInputs,
+ roster: pd.DataFrame,
+ anchor_waves: Iterable[int],
+) -> None:
+ """Check an identifier-only roster's total join and availability."""
+
+ assert len(roster) > 0
+ ids = sorted(set(int(pid) for pid in roster["person_id"]))
+ before = group_attributes.input_frames_sha256(inputs)
+ built = group_attributes.build_group_attributes(
+ inputs, ids, anchor_waves=anchor_waves
+ )
+ frame = built.frame
+ assert len(frame) == len(ids)
+ assert frame["person_id"].is_unique
+ assert set(frame["person_id"]) == set(ids)
+ assert built.provenance["n_persons"] == len(ids)
+ assert built.provenance["content_sha256"] == (
+ group_attributes.content_sha256(frame)
+ )
+ assert group_attributes.input_frames_sha256(inputs) == before
+ joined = roster.merge(frame, on="person_id", validate="many_to_one")
+ assert len(joined) == len(roster)
+ pd.testing.assert_series_equal(
+ joined["person_id"], roster["person_id"].reset_index(drop=True)
+ )
+
+ # Unknown availability is explicit and counted for every attribute.
+ # These are status totals, never population counts by group label.
+ for column, counts in built.provenance["status_counts"].items():
+ assert sum(counts.values()) == len(frame)
+ assert frame[column].notna().all()
+ for value, status in (
+ ("education_years", "education_status"),
+ ("country_of_birth", "country_of_birth_status"),
+ ("race_ethnicity_mint8", "race_ethnicity_status"),
+ ):
+ assert (
+ frame[value].notna() == frame[status].eq(group_attributes.KNOWN)
+ ).all()
+ for source in (
+ "race_ethnicity_source_wave",
+ "education_source_wave",
+ "country_of_birth_source_wave",
+ ):
+ assert frame[source].dropna().isin(gap.RACE_WAVES).all()
+ assert (
+ frame["education_source_wave"].dropna()
+ <= built.provenance["education_cutoff_wave"]
+ ).all()
+
+
+@needs_real_psid
+@pytest.mark.parametrize("anchor_wave", (2009, 2011))
+def test_projection_cohort_coverage(attribute_inputs, anchor_wave):
+ inputs = psid2010.load_psid2010_inputs(
+ data_dir=REAL_DATA, anchor_wave=anchor_wave
+ )
+ built = psid2010.build_psid2010_cohort(
+ inputs, psid2010.Psid2010CohortSpec(anchor_wave=anchor_wave)
+ )
+ roster = built.persons[["person_id"]].copy()
+ del built, inputs
+ gc.collect()
+ _assert_coverage(attribute_inputs, roster, (anchor_wave,))
+
+
+@needs_real_psid
+def test_age67_observation_coverage(attribute_inputs):
+ inputs = age67.load_age67_inputs(data_dir=REAL_DATA)
+ # U0 is the primary, U1 includes all 1936-45 births, and U0-F is the
+ # registered fallback. Repeated observations join the same side row.
+ rosters = []
+ for row in age67.ROWS:
+ built = age67.build_age67_cohort(inputs, age67.Age67Spec(row=row))
+ rosters.append(built.observations[["person_id", "wave"]].copy())
+ del built
+ del inputs
+ gc.collect()
+ for roster in rosters:
+ _assert_coverage(attribute_inputs, roster, age67.WAVES)
+
+
+@needs_real_psid
+def test_track_m_universe_coverage(attribute_inputs):
+ inputs = structure.load_structure_inputs(data_dir=REAL_DATA)
+ # Reuse the exact official selector (cohort._universe), whose only
+ # inputs are these three frames. This does not call build_cohort or
+ # load receipt histories, worker records, eligibility or benefits.
+ minimal = SimpleNamespace(
+ anchor=inputs.anchor,
+ earnings=inputs.observed_earnings,
+ structure_inputs=SimpleNamespace(
+ marriage_history=inputs.marriage_history
+ ),
+ )
+ roster = track_m._universe(minimal)[["person_id"]].copy()
+ del minimal, inputs
+ gc.collect()
+ _assert_coverage(attribute_inputs, roster, (2023,))
diff --git a/tests/cohorts/test_group_category_schemes.py b/tests/cohorts/test_group_category_schemes.py
new file mode 100644
index 00000000..ca943de6
--- /dev/null
+++ b/tests/cohorts/test_group_category_schemes.py
@@ -0,0 +1,296 @@
+"""Pins and definitions of the committed group-category scheme artifact.
+
+Only the cleared guide and its labels-only capture under ``data/external``
+are read. No policy-option page or comparator is opened. Altered captures,
+malformed schemes and person attributes are explicitly INVENTED fixtures.
+Invariants: committed sources match their seals, row labels and order
+match the capture, education bands partition their documented support,
+output/status column names do not collide, and an unresolved definition
+produces an explicit status rather than an inferred category.
+"""
+
+from __future__ import annotations
+
+import copy
+import hashlib
+import html
+import json
+import re
+from pathlib import Path
+
+import pandas as pd
+import pytest
+from hypothesis import given, settings
+from hypothesis import strategies as st
+
+from populace_dynamics.cohorts import group_attributes as cohort
+from populace_dynamics.data import group_attributes_psid as gap
+
+_ROOT = Path(__file__).resolve().parents[2]
+_EXTERNAL = _ROOT / "data" / "external"
+_PROFILE_TABLES = {
+ "annual_benefits": (1, 2, 3),
+ "annual_household_income": (7, 8, 9),
+ "annual_official_poverty": (10, 11, 12),
+ "cohort_benefit_tax_ratio": (13, 14, 15, 16),
+ "cohort_initial_replacement_rate": (17, 18, 19, 20),
+}
+_SOURCE_PROVENANCE = {
+ "mint8_user_guide": "ssa_mint8_table_user_guide.source.provenance.json",
+ "mint8_row_labels": (
+ "ssa_mint8_payroll_option_row_labels.source.provenance.json"
+ ),
+}
+
+
+@pytest.fixture(scope="module")
+def schemes():
+ return cohort.load_schemes()
+
+
+@pytest.fixture(scope="module")
+def captured_rows(schemes):
+ source = schemes["sources"]["mint8_row_labels"]["committed_file"]
+ return json.loads((_ROOT / source).read_text())
+
+
+def _plain_text(value: str) -> str:
+ return " ".join(html.unescape(re.sub(r"<[^>]+>", " ", value)).split())
+
+
+def _paragraph(guide: str, identifier: str) -> str:
+ match = re.search(
+ rf'
]*\bid="{re.escape(identifier)}"[^>]*>(.*?)
',
+ guide,
+ flags=re.DOTALL,
+ )
+ assert match is not None
+ return _plain_text(match.group(1))
+
+
+def _invented_scheme(schemes: dict) -> dict:
+ altered = copy.deepcopy(schemes)
+ altered["note"] = "INVENTED altered scheme for validation tests"
+ return altered
+
+
+def test_scheme_and_source_seals(schemes):
+ assert hashlib.sha256(cohort.SCHEMES_PATH.read_bytes()).hexdigest() == (
+ cohort.SCHEMES_SHA256
+ )
+ for source_id, provenance_file in _SOURCE_PROVENANCE.items():
+ source = schemes["sources"][source_id]
+ provenance = json.loads((_EXTERNAL / provenance_file).read_text())
+ raw = (_ROOT / source["committed_file"]).read_bytes()
+ digest = hashlib.sha256(raw).hexdigest()
+ assert digest == source["sha256"] == provenance["source_sha256"]
+ assert provenance["committed_source_file"] == source["committed_file"]
+ assert provenance["source_url"] == source["url"]
+ assert source["locator"]
+ guide = _ROOT / schemes["sources"]["mint8_user_guide"]["committed_file"]
+ assert (
+ 'name="DCTERMS:dateCertified" content="2025-10-01"'
+ in guide.read_text()
+ )
+
+
+@pytest.mark.parametrize("profile_id", tuple(_PROFILE_TABLES))
+def test_profile_row_groups_match_every_captured_table(
+ schemes, captured_rows, profile_id
+):
+ profiles = schemes["schemes"]["mint8"]["table_profiles"]
+ assert set(profiles) == set(_PROFILE_TABLES)
+ profile = profiles[profile_id]
+ assert profile["source"] == "mint8_row_labels"
+ table_ids = _PROFILE_TABLES[profile_id]
+ assert profile["locator"] == "tables " + ", ".join(map(str, table_ids))
+ for table_id in table_ids:
+ table = captured_rows["tables"][str(table_id)]
+ assert profile["row_groups"] == table["groups"]
+ if profile_id.startswith("annual_"):
+ assert profile["analysis_years"] == [2030, 2050, 2070]
+ for year, table_id in zip(
+ profile["analysis_years"], table_ids, strict=True
+ ):
+ caption = captured_rows["tables"][str(table_id)]["caption"]
+ assert str(year) in caption
+ assert profile["population"] in caption
+ else:
+ for (lower, upper), table_id in zip(
+ profile["birth_cohorts"], table_ids, strict=True
+ ):
+ caption = captured_rows["tables"][str(table_id)]["caption"]
+ assert f"{lower}–{upper}" in caption
+
+
+def test_lifetime_descriptors_match_captured_definitions_and_labels(
+ schemes, captured_rows
+):
+ mint = schemes["schemes"]["mint8"]
+ descriptors = mint["lifetime_earnings_dimensions"]
+ assert set(descriptors) == {
+ "initial_aime",
+ "payroll_tax_own",
+ "payroll_tax_shared",
+ }
+ guide_file = schemes["sources"]["mint8_user_guide"]["committed_file"]
+ guide = (_ROOT / guide_file).read_text()
+ paragraphs = {
+ "initial_aime": _paragraph(guide, "AIME"),
+ "payroll_tax_own": _paragraph(guide, "lifetime-tax"),
+ "payroll_tax_shared": _paragraph(guide, "lifetime-tax-shared"),
+ }
+ for dimension_id, descriptor in descriptors.items():
+ assert descriptor["source"] == "mint8_user_guide"
+ assert descriptor["labels_source"] == "mint8_row_labels"
+ assert descriptor["reference_age"] == 62
+ assert "at age 62" in paragraphs[dimension_id]
+ assert descriptor["quintile_population"] == "each birth cohort"
+ assert "for each birth cohort" in paragraphs[dimension_id]
+ for table_id in range(13, 21):
+ groups = captured_rows["tables"][str(table_id)]["groups"]
+ group = next(
+ row
+ for row in groups
+ if row["group"] == descriptor["group_label"]
+ )
+ assert descriptor["labels"] == group["labels"]
+
+ assert "under current law at age 62" in paragraphs["initial_aime"]
+ assert "present value" in paragraphs["payroll_tax_own"]
+ plain_guide = _plain_text(guide)
+ assert (
+ "We use the Social Security Trust Fund interest rate to adjust "
+ "benefits and taxes to their present values at age 62."
+ ) in plain_guide
+ for dimension in ("payroll_tax_own", "payroll_tax_shared"):
+ assert descriptors[dimension]["discount_rate"] == (
+ "Social Security Trust Fund interest rate"
+ )
+ shared = descriptors["payroll_tax_shared"]
+ assert shared["married_year_rule"] == (
+ "(own payroll taxes + spouse payroll taxes) / 2"
+ )
+ assert shared["unmarried_year_rule"] == "own payroll taxes"
+ assert (
+ "paid while married are shared equally"
+ in paragraphs["payroll_tax_shared"]
+ )
+ assert (
+ "For never-married individuals, this is the same"
+ in paragraphs["payroll_tax_shared"]
+ )
+ assert "not married, we count only their individual payroll taxes" in (
+ paragraphs["payroll_tax_shared"]
+ )
+
+
+def test_report_education_definitions_remain_unresolved(schemes):
+ spec = schemes["schemes"]["boomers2004"]["dimensions"]["education"]
+ assert not spec["bands"]
+ assert not spec["assumptions"]
+ assert spec["row_labels"] == [
+ "High school dropout",
+ "High school graduate",
+ "College graduate",
+ ]
+ # INVENTED persons span the documented years-of-education support.
+ frame = pd.DataFrame(
+ {"education_years": pd.array([*range(18), pd.NA], dtype="Int64")}
+ )
+ labels, statuses = cohort.apply_category_scheme(frame, "education", spec)
+ assert labels.isna().all()
+ assert statuses.iloc[:-1].eq("unresolved:definition_not_recorded").all()
+ assert statuses.iloc[-1] == cohort.ATTRIBUTE_UNKNOWN
+
+
+@pytest.mark.parametrize("source_id", tuple(_SOURCE_PROVENANCE))
+def test_changed_capture_is_refused(schemes, monkeypatch, tmp_path, source_id):
+ for source in schemes["sources"].values():
+ if "committed_file" not in source:
+ continue
+ path = tmp_path / source["committed_file"]
+ path.parent.mkdir(parents=True, exist_ok=True)
+ path.write_bytes((_ROOT / source["committed_file"]).read_bytes())
+ altered = tmp_path / schemes["sources"][source_id]["committed_file"]
+ altered.write_bytes(altered.read_bytes() + b"\nINVENTED altered bytes\n")
+ monkeypatch.setattr(cohort, "_REPO_ROOT", tmp_path)
+ with pytest.raises(ValueError, match="scheme source .* changed"):
+ cohort.load_schemes()
+
+
+def test_changed_scheme_is_refused(schemes, monkeypatch, tmp_path):
+ path = tmp_path / "INVENTED_altered_scheme.json"
+ path.write_text(json.dumps(_invented_scheme(schemes)))
+ monkeypatch.setattr(cohort, "SCHEMES_PATH", path)
+ with pytest.raises(ValueError, match="expected the pinned"):
+ cohort.load_schemes()
+
+
+@pytest.mark.parametrize("defect", ("gap", "overlap"))
+@settings(max_examples=18, deadline=None)
+@given(year=st.integers(min_value=0, max_value=17))
+def test_education_support_must_be_partitioned(defect, year):
+ invented = _invented_scheme(cohort.load_schemes())
+ spec = invented["schemes"]["boomers2004"]["dimensions"]["education"]
+ spec["unresolved_bands"] = []
+ intervals = (
+ [(0, year - 1), (year + 1, 17)]
+ if defect == "gap"
+ else [(0, 17), (year, year)]
+ )
+ for lower, upper in intervals:
+ if lower <= upper:
+ spec["unresolved_bands"].append(
+ {
+ "min": lower,
+ "max": upper,
+ "status": "INVENTED_missing_definition",
+ "reason": "INVENTED malformed partition",
+ }
+ )
+ with pytest.raises(ValueError, match="bands must cover each"):
+ cohort._validate_schemes(invented)
+
+
+@pytest.mark.parametrize(
+ "column",
+ (
+ "race_ethnicity_mint8_status",
+ "education_years",
+ "education_source_wave",
+ ),
+)
+def test_output_and_status_columns_cannot_collide(schemes, column):
+ invented = _invented_scheme(schemes)
+ invented["schemes"]["mint8"]["dimensions"]["education"]["column"] = column
+ with pytest.raises(ValueError, match="column .* is taken"):
+ cohort._validate_schemes(invented)
+
+
+def test_race_dimension_can_explicitly_leave_a_known_code_unresolved(schemes):
+ invented = _invented_scheme(schemes)
+ spec = invented["schemes"]["mint8"]["dimensions"]["race_ethnicity"]
+ del spec["categories"]["white_non_hispanic"]
+ spec["unresolved"] = {
+ "white_non_hispanic": "INVENTED definition withheld in a test scheme"
+ }
+ cohort._validate_schemes(invented)
+ # INVENTED attribute rows exercise unresolved, assigned and unknown.
+ frame = pd.DataFrame(
+ {
+ "hispanic": pd.array([False, True, pd.NA], dtype="boolean"),
+ "race_mentions": pd.array(
+ [gap.WHITE, pd.NA, pd.NA], dtype="string"
+ ),
+ }
+ )
+ labels, statuses = cohort.apply_category_scheme(
+ frame, "race_ethnicity", spec
+ )
+ assert pd.isna(labels.iloc[0])
+ assert statuses.iloc[0] == "unresolved:white_non_hispanic"
+ assert labels.iloc[1] == spec["categories"]["hispanic"]
+ assert statuses.iloc[1] == cohort.ASSIGNED
+ assert pd.isna(labels.iloc[2])
+ assert statuses.iloc[2] == cohort.ATTRIBUTE_UNKNOWN
diff --git a/tests/data/test_group_attributes_psid.py b/tests/data/test_group_attributes_psid.py
new file mode 100644
index 00000000..b796bc60
--- /dev/null
+++ b/tests/data/test_group_attributes_psid.py
@@ -0,0 +1,507 @@
+"""INVENTED fixed-width fixtures and pinned documentation checks.
+
+The committed ``data/external`` table supplies code domains, never
+microdata. Fixture persons and all their answers are INVENTED. Invariants:
+labels and domains are checked before interpretation; role joins preserve
+OFUMs in the individual roster but attach family answers only to present
+heads/spouses; ambiguous or absent answers never become negative answers.
+"""
+
+from __future__ import annotations
+
+import hashlib
+import importlib.util
+from pathlib import Path
+
+import pandas as pd
+import pytest
+from hypothesis import given, settings
+from hypothesis import strategies as st
+
+from populace_dynamics.data import group_attributes_psid as gap
+from tests.data.psid_fixtures import write_product
+
+
+def _write_individual(root: Path, wave: int = 2023) -> Path:
+ items = gap.INDIVIDUAL_ITEMS[wave]
+ fields = [
+ ("ER30001", 5, "1968 INTERVIEW NUMBER", [1] * 5),
+ ("ER30002", 3, "PERSON NUMBER 68", [1, 2, 3, 4, 5]),
+ (*items.interview, [101, 101, 101, 202, 303]),
+ (*items.sequence, [1, 2, 3, 2, 81]),
+ (*items.relationship, [10, 20, 30, 22, 10]),
+ (*items.education, [17, 0, 99, 12, 0]),
+ ]
+ # The first two tuples already contain width; remaining tuples do not.
+ fields = [
+ field if len(field) == 4 else (field[0], 4, field[1], field[2])
+ for field in fields
+ ]
+ directory = root / "ind2023er"
+ write_product(directory, "IND2023ER.sps", "IND2023ER.txt", fields)
+ edu, rel, seq = (
+ items.education[0],
+ items.relationship[0],
+ items.sequence[0],
+ )
+ # Domains and role descriptions of this synthetic product are INVENTED
+ # documentation matching the reader's adjudicated frame.
+ dk = "98 = 'DK'\n" if 1994 <= wave <= 2011 else ""
+ formats = (
+ f"VALUE {edu}F\n1 - 17 = 'Grades'\n{dk}"
+ "99 = 'NA'\n0 = 'Inap.'\n;\n"
+ f"VALUE {rel}F\n10 = 'Head in {wave}'\n"
+ f"20 = 'Legal wife in {wave}'\n"
+ "22 = '\"Wife\"--cohabitor'\n30 = 'Child'\n0 = 'Inap.'\n;\n"
+ f"VALUE {seq}F\n1 - 20 = 'Present'\n"
+ "51 - 59 = 'Institution'\n71 - 89 = 'Absent'\n0 = 'Inap.'\n;\n"
+ )
+ target = directory / "IND2023ER_formats.sas"
+ target.write_text(formats)
+ return target
+
+
+def _write_family(root: Path, wave: int = 2023) -> Path:
+ items = gap.FAMILY_ITEMS[wave]
+ values = {variable: [0, 0] for variable in items.labels()}
+ values[items.interview[0]] = [101, 202]
+ for role in gap.ROLES:
+ values[items.race[role][0][0]] = [1, 2]
+ if items.completed_education:
+ values[items.completed_education[role][0]] = [0, 12]
+ if items.birth_state:
+ values[items.birth_state[role][0]] = [1, 0]
+ values[items.year_came[role][0]] = [0, 1999]
+ fields = [
+ (variable, 5, label, values[variable])
+ for variable, label in items.labels().items()
+ ]
+ directory = root / "family" / str(wave)
+ write_product(directory, "FAM.sps", "FAM.txt", fields)
+ return directory / "FAM.sps"
+
+
+def test_fixed_width_reader_and_role_join(tmp_path):
+ _write_individual(tmp_path)
+ _write_family(tmp_path)
+ individual = gap.read_individual_items(data_dir=tmp_path, waves=(2023,))
+ family = gap.read_family_items(2023, data_dir=tmp_path)
+ reports = gap.head_spouse_reports(individual, {2023: family})
+ assert len(individual) == 5
+ assert reports["person_id"].tolist() == [1001, 1002, 1004]
+ assert reports["role"].tolist() == ["head", "spouse", "spouse"]
+ assert reports["race_code_1"].tolist() == [1, 1, 2]
+ assert all(
+ reports[column].dtype == "Int64" for column in gap.REPORT_CODE_COLUMNS
+ )
+ # OFUM education remains available without a fabricated family role.
+ assert (
+ individual.loc[individual.person_id == 1003, "education_code"].item()
+ == 99
+ )
+ assert gap.education_report(
+ 0, reports.loc[1, "family_education_code"]
+ ) == (
+ 0,
+ "no_grades",
+ )
+
+
+@pytest.mark.parametrize("kind", ["individual", "family"])
+def test_label_mismatch_refused_before_data_read(tmp_path, monkeypatch, kind):
+ _write_individual(tmp_path)
+ family_sps = _write_family(tmp_path)
+ if kind == "individual":
+ target = tmp_path / "ind2023er" / "IND2023ER.sps"
+ target.write_text(
+ target.read_text().replace(
+ "YEARS COMPLETED EDUCATION 23", "WRONG LABEL"
+ )
+ )
+
+ def call():
+ return gap.read_individual_items(data_dir=tmp_path, waves=(2023,))
+
+ else:
+ family_sps.write_text(
+ family_sps.read_text().replace(
+ "L39 SPANISH DESCENT-RP", "WRONG LABEL"
+ )
+ )
+
+ def call():
+ return gap.read_family_items(2023, data_dir=tmp_path)
+
+ def forbidden(*args, **kwargs):
+ raise AssertionError("fixed-width data opened before label check")
+
+ monkeypatch.setattr(pd, "read_fwf", forbidden)
+ with pytest.raises(ValueError, match="label"):
+ call()
+
+
+@pytest.mark.parametrize(
+ "concept,code", [("education", 18), ("relationship", 11), ("sequence", 21)]
+)
+def test_individual_undocumented_codes_refused(tmp_path, concept, code):
+ _write_individual(tmp_path)
+ variable = getattr(gap.INDIVIDUAL_ITEMS[2023], concept)[0]
+ # Rewrite the INVENTED fixture through its explicit SPSS colspecs.
+ sps = tmp_path / "ind2023er" / "IND2023ER.sps"
+ from populace_dynamics.data import psid
+
+ spec = psid.parse_sps_layout(sps).set_index("name").loc[variable]
+ txt = sps.with_suffix(".txt")
+ lines = txt.read_text().splitlines()
+ start, end = int(spec.start) - 1, int(spec.end)
+ lines[0] = lines[0][:start] + f"{code:>{end - start}}" + lines[0][end:]
+ txt.write_text("\n".join(lines) + "\n")
+ with pytest.raises(ValueError, match="not documented"):
+ gap.read_individual_items(data_dir=tmp_path, waves=(2023,))
+
+
+def test_family_undocumented_code_and_duplicate_interview_refused(tmp_path):
+ sps = _write_family(tmp_path)
+ txt = sps.with_suffix(".txt")
+ rows = txt.read_text().splitlines()
+ rows[0] = rows[0][:5] + " 6" + rows[0][10:]
+ txt.write_text("\n".join(rows) + "\n")
+ with pytest.raises(ValueError, match="not documented"):
+ gap.read_family_items(2023, data_dir=tmp_path)
+ _write_family(tmp_path)
+ rows = txt.read_text().splitlines()
+ rows[1] = rows[0][:5] + rows[1][5:]
+ txt.write_text("\n".join(rows) + "\n")
+ with pytest.raises(ValueError, match="duplicate interview"):
+ gap.read_family_items(2023, data_dir=tmp_path)
+
+
+def test_family_formats_disagreement_refused(tmp_path):
+ _write_family(tmp_path)
+ variable = gap.FAMILY_ITEMS[2023].hispanic["head"][0]
+ (tmp_path / "family" / "2023" / "FAM_formats.sas").write_text(
+ f"VALUE {variable}F\n0 = 'Negative'\n1 = 'Positive'\n;\n"
+ )
+ with pytest.raises(ValueError, match="formats file documents"):
+ gap.read_family_items(2023, data_dir=tmp_path)
+
+
+def test_missing_family_and_duplicate_role_refused(tmp_path):
+ _write_individual(tmp_path)
+ _write_family(tmp_path)
+ individual = gap.read_individual_items(data_dir=tmp_path, waves=(2023,))
+ family = gap.read_family_items(2023, data_dir=tmp_path)
+ with pytest.raises(ValueError, match="no family-file row"):
+ gap.head_spouse_reports(individual, {2023: family.iloc[:1]})
+ individual.loc[individual.person_id == 1003, "relationship"] = 10
+ with pytest.raises(ValueError, match="share a family's head"):
+ gap.head_spouse_reports(individual, {2023: family})
+
+
+def _invented_multiwave_join_inputs(rows):
+ """An INVENTED interleaved roster and family answers for join tests."""
+
+ waves = (1985, 1993, 2005, 2023)
+ individual = pd.DataFrame(
+ [
+ {
+ "person_id": position + 1001,
+ "wave": wave,
+ "interview": position + 101,
+ "sequence": sequence,
+ "relationship": relationship,
+ "education_code": 12,
+ }
+ for position, (wave, relationship, sequence) in enumerate(rows)
+ ],
+ columns=[
+ "person_id",
+ "wave",
+ "interview",
+ "sequence",
+ "relationship",
+ "education_code",
+ ],
+ dtype="int64",
+ )
+ # The partition must preserve selection and alignment with these
+ # nonconsecutive, decreasing input labels.
+ individual.index = range(7 * len(individual), 0, -7)
+ families = {}
+ for wave in waves:
+ items = gap.FAMILY_ITEMS[wave]
+ interviews = individual.loc[
+ individual["wave"] == wave, "interview"
+ ].tolist()
+ values = {
+ variable: [0] * len(interviews) for variable in items.labels()
+ }
+ values[items.interview[0]] = interviews
+ for role in gap.ROLES:
+ values[items.race[role][0][0]] = [1] * len(interviews)
+ if items.completed_education:
+ values[items.completed_education[role][0]] = [12] * len(
+ interviews
+ )
+ if items.birth_state:
+ values[items.birth_state[role][0]] = [1] * len(interviews)
+ families[wave] = pd.DataFrame(values, dtype="int64").rename(
+ columns={items.interview[0]: "interview"}
+ )
+ return individual, families
+
+
+def _whole_frame_mask_reference(individual, families):
+ """The original whole-roster masks with independent row-based joins."""
+
+ result = []
+ for wave, family in sorted(families.items()):
+ selected = individual[
+ (individual["wave"] == wave)
+ & individual["sequence"].between(*gap.IN_FAMILY_SEQUENCE)
+ ]
+ family_by_interview = family.set_index("interview")
+ for person in selected.itertuples(index=False):
+ if person.relationship == gap.HEAD_RELATIONSHIP:
+ role = "head"
+ elif person.relationship in gap.SPOUSE_RELATIONSHIPS:
+ role = "spouse"
+ else:
+ continue
+ row = dict.fromkeys(gap.REPORT_CODE_COLUMNS, pd.NA)
+ row.update(person_id=person.person_id, wave=wave, role=role)
+ family_row = family_by_interview.loc[person.interview]
+ for variable, column in (
+ gap.FAMILY_ITEMS[wave].role_columns(role).items()
+ ):
+ row[column] = int(family_row[variable])
+ result.append(row)
+ frame = pd.DataFrame(
+ result,
+ columns=["person_id", "wave", "role", *gap.REPORT_CODE_COLUMNS],
+ )
+ frame = frame.astype(
+ {
+ "person_id": "int64",
+ "wave": "int64",
+ "role": "string",
+ **dict.fromkeys(gap.REPORT_CODE_COLUMNS, "Int64"),
+ }
+ )
+ return frame.sort_values(["person_id", "wave"]).reset_index(drop=True)
+
+
+@settings(max_examples=40, deadline=None)
+@given(
+ st.lists(
+ st.tuples(
+ st.sampled_from((1985, 1993, 2005, 2023)),
+ st.sampled_from((10, 20, 22, 30)),
+ st.sampled_from((1, 20, 51, 81)),
+ ),
+ max_size=24,
+ )
+)
+def test_wave_partition_differential_preserves_whole_frame_masks(rows):
+ individual, families = _invented_multiwave_join_inputs(rows)
+ expected = _whole_frame_mask_reference(individual, families)
+ actual = gap.head_spouse_reports(individual, families)
+ pd.testing.assert_frame_equal(actual, expected)
+
+
+def test_wave_partition_preserves_missing_wave_and_empty_role_outputs():
+ individual, families = _invented_multiwave_join_inputs(
+ [(2023, 10, 1), (2005, 30, 1), (1993, 20, 81)]
+ )
+ expected = _whole_frame_mask_reference(individual, families)
+ actual = gap.head_spouse_reports(individual, families)
+ pd.testing.assert_frame_equal(actual, expected)
+ assert actual["person_id"].tolist() == [1001]
+
+
+def test_domains_pinned_and_all_documented_race_codes_mean_something():
+ expected_path = (
+ Path(__file__).resolve().parents[2]
+ / "data"
+ / "external"
+ / "psid_group_attribute_codebook_values_v1.json"
+ )
+ assert gap.CODEBOOK_VALUES_PATH == expected_path
+ table = gap.load_codebook_values()
+ assert (
+ hashlib.sha256(gap.CODEBOOK_VALUES_PATH.read_bytes()).hexdigest()
+ == gap.CODEBOOK_VALUES_SHA256
+ )
+ for wave, items in gap.FAMILY_ITEMS.items():
+ for role in gap.ROLES:
+ for mention, (variable, label) in enumerate(items.race[role], 1):
+ entry = table["family"][str(wave)]["variables"][variable]
+ assert entry["label"] == " ".join(label.split())
+ domain = gap.documented_domain(table, wave, variable)
+ for code in domain.normalized()[0]:
+ assert gap.race_mention_meaning(wave, mention, code) in (
+ *gap.RACE_CATEGORIES,
+ gap.MISSING,
+ gap.LATINO_ORIGIN,
+ gap.NO_FURTHER_MENTION,
+ )
+
+
+@pytest.mark.parametrize("wave", [1994, 1995, 1996])
+def test_ambiguous_spouse_zero_remains_unknown(wave):
+ head = gap.race_ethnicity_report(wave, 0, (1, 0, 0), role="head")
+ spouse = gap.race_ethnicity_report(wave, 0, (1, 0, 0), role="spouse")
+ assert head.hispanic is False
+ assert spouse.hispanic is None
+ assert not spouse.complete
+ assert spouse.hispanic_basis == "undocumented_zero_meaning"
+
+
+def test_wave_specific_race_and_missing_origin():
+ assert gap.race_mention_meaning(1993, 1, 5) == gap.LATINO_ORIGIN
+ assert gap.race_mention_meaning(2005, 1, 5) == gap.PACIFIC_ISLANDER
+ assert gap.race_mention_meaning(1993, 2, 8) == gap.MORE_THAN_TWO
+ assert gap.race_mention_meaning(1994, 2, 8) == gap.MISSING
+ assert gap.race_ethnicity_report(1997, None, (1, 0, 0, 0)).hispanic is None
+ assert gap.race_ethnicity_report(1997, None, (5, 0, 0, 0)).hispanic is True
+ assert gap.race_ethnicity_report(2023, 1, (9, 0, 0, 0)).complete
+ assert gap.race_ethnicity_report(2023, 0, (9, 0, 0, 0)).races is None
+
+
+@pytest.mark.parametrize(
+ "codes",
+ [(1,), (1, None, 0, 0), (1, 6, 0, 0), (1, 9, 0, 0), (1, 0, 0, 0, 0)],
+)
+def test_incomplete_or_undocumented_mentions_refused(codes):
+ with pytest.raises(ValueError):
+ gap.race_ethnicity_report(2023, 0, codes)
+
+
+def test_unasked_origin_and_modern_undocumented_origin_refused():
+ with pytest.raises(ValueError, match="asks_hispanic_origin"):
+ gap.race_ethnicity_report(1997, 0, (1, 0, 0, 0))
+ with pytest.raises(ValueError, match="undocumented code"):
+ gap.race_ethnicity_report(2023, 6, (1, 0, 0, 0))
+
+
+@given(st.integers(1, 17))
+def test_education_codes_preserve_years_and_never_impute(code):
+ assert gap.education_report(code) == (code, "reported")
+ assert gap.education_report(code, 99) == (code, "reported")
+ assert gap.education_report(0) == (None, "inapplicable")
+ assert gap.education_report(0, 99) == (None, "inapplicable")
+ assert gap.education_report(0, 0) == (0, "no_grades")
+
+
+@pytest.mark.parametrize(
+ "value", [True, 1.0, 1.5, "1", float("nan"), float("inf")]
+)
+def test_numeric_codes_must_be_integers(value):
+ for call in (
+ lambda: gap.education_report(value),
+ lambda: gap.hispanic_meaning(value),
+ lambda: gap.race_mention_meaning(2023, 1, value),
+ lambda: gap.birthplace_class(value, 0),
+ ):
+ with pytest.raises(ValueError):
+ call()
+
+
+@pytest.mark.parametrize("code", [-1, 18, 98, 100])
+def test_undocumented_family_education_refused(code):
+ with pytest.raises(ValueError, match="family education"):
+ gap.education_report(12, code)
+
+
+@given(st.integers(1, 56), st.integers(1901, 2023))
+def test_birthplace_pair_respects_skip_pattern(state, came):
+ assert gap.birthplace_class(state, 0) == gap.UNITED_STATES
+ assert gap.birthplace_class(state, came) == gap.INCONSISTENT
+ assert gap.birthplace_class(0, 0) == gap.US_TERRITORY
+ assert gap.birthplace_class(0, came) == gap.FOREIGN_COUNTRY
+
+
+def test_birthplace_missing_and_invalid_year():
+ assert gap.birthplace_class(99, 0) == gap.MISSING
+ assert gap.birthplace_class(99, 1999) == gap.INCONSISTENT
+ assert gap.birthplace_class(0, 9999) == gap.FOREIGN_COUNTRY
+ with pytest.raises(ValueError, match="year-came"):
+ gap.birthplace_class(0, -1)
+
+
+@given(
+ st.sets(st.integers(-10, 10)),
+ st.lists(
+ st.tuples(st.integers(-10, 10), st.integers(-10, 10)), max_size=5
+ ),
+ st.integers(-15, 15),
+)
+def test_domain_representations_differential(singles, ranges, value):
+ intervals = tuple((min(lo, hi), max(lo, hi)) for lo, hi in ranges)
+ domain = gap.CodeDomain(frozenset(singles), intervals)
+ expected = value in singles or any(
+ lo <= value <= hi for lo, hi in intervals
+ )
+ assert domain.contains(value) is expected
+ expanded = gap.CodeDomain(domain.normalized()[0])
+ assert domain.same_as(expanded)
+ assert expanded.contains(value) is expected
+
+
+def test_sas_domains_preserve_ranges_and_reject_truncation(tmp_path):
+ path = tmp_path / "formats.sas"
+ path.write_text(
+ "VALUE EXAMPLE\n1 - 17 = 'Grades'\n99 = 'NA'\n0 = 'Inap.'\n;\n"
+ )
+ parsed = gap.parse_sas_value_domains(path)["EXAMPLE"]
+ assert parsed.same_as(gap.CodeDomain(frozenset({0, 99}), ((1, 17),)))
+ path.write_text(path.read_text().removesuffix(";\n"))
+ with pytest.raises(ValueError, match="Unterminated"):
+ gap.parse_sas_value_domains(path)
+
+
+def test_codebook_extraction_continuations_and_page_locator():
+ script = (
+ Path(__file__).resolve().parents[2]
+ / "scripts"
+ / "build_group_attribute_codebook_values.py"
+ )
+ spec = importlib.util.spec_from_file_location(
+ "group_codebook_builder", script
+ )
+ module = importlib.util.module_from_spec(spec)
+ spec.loader.exec_module(module)
+ text = (
+ 'ER00001 "INVENTED LABEL"\nQuestion?\n'
+ " Count % Value/Range Code Value/Range Text\n"
+ " 1 10.0 1 - 17 Grades\n"
+ " continuation\nPage 1 of 2\n\f"
+ "PANEL STUDY OF INCOME DYNAMICS\n"
+ " 1 10.0 99 NA\n 1 10.0 0 Inap.\n"
+ 'ER00002 "NEXT INVENTED LABEL"\n'
+ )
+ entry = module.extract_entry(text, "ER00001")
+ assert entry == {
+ "label": "INVENTED LABEL",
+ "page": 1,
+ "values": [
+ {"range": [1, 17], "text": "Grades continuation"},
+ {"code": 99, "text": "NA"},
+ {"code": 0, "text": "Inap."},
+ ],
+ }
+
+
+def test_source_pins_refuse_changed_documentation(tmp_path):
+ pdf = tmp_path / "INVENTED.pdf"
+ pdf.write_bytes(b"INVENTED source documentation")
+ sha = hashlib.sha256(pdf.read_bytes()).hexdigest()
+ table = {
+ "family": {"2023": {"codebook": {"path": pdf.name, "sha256": sha}}}
+ }
+ assert gap.verify_codebook_pins(
+ (2023,), data_dir=tmp_path, codebook=table
+ ) == {2023: sha}
+ pdf.write_bytes(b"changed INVENTED source documentation")
+ with pytest.raises(ValueError, match="pins"):
+ gap.verify_codebook_pins((2023,), data_dir=tmp_path, codebook=table)
From c65b1f7ca481307cf2058f7c6cae7c4b59e3ecf6 Mon Sep 17 00:00:00 2001
From: Max Ghenis
Date: Sat, 3 Oct 2026 07:30:59 -0400
Subject: [PATCH 05/10] Add MINT group breakdowns for the projection tests
(exercises 1 and 3)
group_breakdowns/cola.py and fra68.py re-execute each frozen registered
specification (projection, benefits, union rows) by composing the
existing modules read-only, and refuse before loading any group
attribute unless every committed cell and per-draw diagnostic matches
runs/replication_urban2010_cola_v1.json or
runs/replication_urban2010_fra68_v1.json. They then join G1's race,
education and nativity side frame and G2's lifetime measures by
person_id and tabulate with G3 for both the tests' own population and
MINT's beneficiaries aged 60 or older in 2030. Poverty status and
household income quintile are not computable in the projection (no
income) and are marked so.
scripts/run_projection_groups_registered.py refuses without an issue
#42 pointer, a clean HEAD equal to the registered commit and an absent
output; scripts/make_nasi_repro_venv.sh builds the environment the
committed artifacts recorded. Invented dry runs only; the large
invented result files are kept outside the repository and the
RESULTS.md summaries are committed.
Co-Authored-By: Claude Opus 5.5
---
.../cola/RESULTS.md | 7 +
.../fra68/RESULTS.md | 7 +
docs/design/nasi_projection_groups_posthoc.md | 262 ++++++++
scripts/make_nasi_repro_venv.sh | 146 +++++
scripts/projection_groups_dry_run.py | 221 +++++++
scripts/run_projection_groups_registered.py | 440 +++++++++++++
.../group_breakdowns/__init__.py | 1 +
.../group_breakdowns/cola.py | 248 +++++++
.../group_breakdowns/common.py | 616 ++++++++++++++++++
.../group_breakdowns/fra68.py | 392 +++++++++++
tests/group_breakdowns/test_cola.py | 179 +++++
tests/group_breakdowns/test_common.py | 402 ++++++++++++
tests/group_breakdowns/test_fra68.py | 121 ++++
tests/group_breakdowns/test_registration.py | 421 ++++++++++++
14 files changed, 3463 insertions(+)
create mode 100644 docs/analysis/nasi_group_breakdowns_invented_20261001/cola/RESULTS.md
create mode 100644 docs/analysis/nasi_group_breakdowns_invented_20261001/fra68/RESULTS.md
create mode 100644 docs/design/nasi_projection_groups_posthoc.md
create mode 100755 scripts/make_nasi_repro_venv.sh
create mode 100644 scripts/projection_groups_dry_run.py
create mode 100644 scripts/run_projection_groups_registered.py
create mode 100644 src/populace_dynamics/group_breakdowns/__init__.py
create mode 100644 src/populace_dynamics/group_breakdowns/cola.py
create mode 100644 src/populace_dynamics/group_breakdowns/common.py
create mode 100644 src/populace_dynamics/group_breakdowns/fra68.py
create mode 100644 tests/group_breakdowns/test_cola.py
create mode 100644 tests/group_breakdowns/test_common.py
create mode 100644 tests/group_breakdowns/test_fra68.py
create mode 100644 tests/group_breakdowns/test_registration.py
diff --git a/docs/analysis/nasi_group_breakdowns_invented_20261001/cola/RESULTS.md b/docs/analysis/nasi_group_breakdowns_invented_20261001/cola/RESULTS.md
new file mode 100644
index 00000000..0722105a
--- /dev/null
+++ b/docs/analysis/nasi_group_breakdowns_invented_20261001/cola/RESULTS.md
@@ -0,0 +1,7 @@
+# INVENTED DATA - NOT A COMPARISON
+
+Exercise: cola. Registered, one-shot, post hoc, not blind; report-only labels describe the intended registered artifact. This dry run is unregistered and uses invented data only.
+
+All 7 registered rows reproduced the frozen runner exactly before group loading. Both original waves and both population variants are in [result.json](result.json). G1 supplies attributes, G2 supplies lifetime measures, and G3 supplies labels, suppression flags and every group cell.
+
+No PSID file, real-data outcome or comparator value was used. Income/poverty are unavailable. Lifetime careers stop at the opening year, and interest/tax inputs here are invented.
diff --git a/docs/analysis/nasi_group_breakdowns_invented_20261001/fra68/RESULTS.md b/docs/analysis/nasi_group_breakdowns_invented_20261001/fra68/RESULTS.md
new file mode 100644
index 00000000..1fbeb97f
--- /dev/null
+++ b/docs/analysis/nasi_group_breakdowns_invented_20261001/fra68/RESULTS.md
@@ -0,0 +1,7 @@
+# INVENTED DATA - NOT A COMPARISON
+
+Exercise: fra68. Registered, one-shot, post hoc, not blind; report-only labels describe the intended registered artifact. This dry run is unregistered and uses invented data only.
+
+All 9 registered rows reproduced the frozen runner exactly before group loading. Both original waves and both population variants are in [result.json](result.json). G1 supplies attributes, G2 supplies lifetime measures, and G3 supplies labels, suppression flags and every group cell.
+
+No PSID file, real-data outcome or comparator value was used. Income/poverty are unavailable. Lifetime careers stop at the opening year, and interest/tax inputs here are invented.
diff --git a/docs/design/nasi_projection_groups_posthoc.md b/docs/design/nasi_projection_groups_posthoc.md
new file mode 100644
index 00000000..dafe4e08
--- /dev/null
+++ b/docs/design/nasi_projection_groups_posthoc.md
@@ -0,0 +1,262 @@
+# NASI projection group breakdowns: exercises 1 and 3
+
+Package G4a implements the 2026-10-01 NASI follow-up request for group
+breakdowns of the completed COLA and FRA-68 projection tests. This is a
+methodology specification for a new registration on issue #42. Every
+real-data output must carry the original test labels plus **registered,
+one-shot, post hoc, not blind** and **report-only**. The development outputs
+carry **INVENTED DATA - NOT A COMPARISON**. No comparator value is read, no
+acceptance gate is added, and no real-data group output was produced while
+building the package.
+
+The existing registered specifications remain the authorities for the
+projection and benefit mechanisms:
+
+- Exercise 1: `docs/design/urban2010_cola_comparison.md`,
+ `cola_track_a/runner.py`, and
+ `runs/replication_urban2010_cola_v1.json`.
+- Exercise 3: `docs/design/urban2010_fra68_comparison.md`,
+ `fra68_track/runner.py`, and
+ `runs/replication_urban2010_fra68_v1.json`.
+- G1: `cohorts/group_attributes.py` for person attributes and its pinned
+ category schemes and reader provenance.
+- G2: `estimates/lifetime_measures.py` for every lifetime measure.
+- G3: `estimates/group_breakdown.py` for every group assignment and cell,
+ including the MINT8 schemes, definitions, uncertainty and suppression.
+
+Paths to source modules above are relative to `src/populace_dynamics/`.
+The MINT8 labels and definitions come from the source captures already
+pinned by G3; this package does not capture another source or implement
+another lifetime measure. The three inherited User Guide captures
+`data/external/mint8_table_user_guide.source.html`,
+`data/external/ssa_mint8_table_user_guide.source.html`, and
+`data/external/ssa_mint8_user_guide_2026.source.html` remain intact for the
+integrator to consolidate.
+
+## Frozen replay and the reproduction gate
+
+`group_breakdowns.cola.reproduce_cola` replays R0-R6, and
+`group_breakdowns.fra68.reproduce_fra68` replays F0-F8. The populations are
+unchanged: the 2011-wave opening cohort for R0-R5/F0-F7, and the 2009-wave
+opening cohort for R6/F8. Each population is projected once per registered
+draw, using the frozen `_project_population`. All of its rows share that
+projection. Exercise 1 calls the frozen `StateLookups` and
+`reference_benefit_rows`; exercise 3 calls the frozen scenario calculators
+and `union_benefit_rows`. The frozen runners' guards, five-age-group
+statistics, draw diagnostics, counters, membership diagnostics, floors and
+provenance remain part of the replayed result.
+
+The replay returns a frozen `ProjectionReplay` containing the aggregate
+result, the already-computed benefit rows by registered row, and the final
+projected state by anchor wave. Retention adds no group attribute and
+computes no group cell. Its frames stay in memory; person rows are not
+serialized into the output artifact.
+
+`ProjectionReplay` also seals every retained benefit and state frame when
+constructed. The SHA-256 covers column order, dtypes and CSV values,
+including the component mappings. The reproduction gate refuses any later
+mutation before calling a loader or tabulator and records the retention
+seal on success. An aggregate match therefore cannot certify person rows
+that were altered after the frozen calculation.
+
+`common.verify_reproduction` compares **every field returned by the frozen
+runner** against the corresponding field in the committed parent
+artifact. The entry script's wrapper fields, such as wall-clock run times,
+are outside that runner result. Nested mappings, sequences and scalar
+values must match as canonical JSON, with **absolute tolerance 0 and
+relative tolerance 0**. This includes all five age cells and every draw's
+diagnostics. No floating summation tolerance is proposed in this version.
+The check records the verified fields, row manifest, comparison rule and
+SHA-256 of the matched canonical payload. A failure reports a field path
+without printing outcome values.
+
+On any mismatch, the adapter raises `ReproductionMismatch` before calling
+a G1 attribute loader, G2 lifetime measure, G3 group assignment or G3 cell
+tabulator. No artifact or environment sidecar is written. An invented test
+changes a committed-cell stand-in and verifies both that the attribute
+loader is never called and that neither output file is created. This
+ordering is binding even if the real-data replay fails because the host
+environment or frozen code no longer reproduces the original result.
+
+The frozen replay uses the original parent's registration pointer so that
+its result can match the original registration metadata exactly. The
+new issue #42 pointer authorizes the post hoc group calculation and is
+recorded separately on the new output. The old pointer alone does not
+authorize a real-data group run.
+
+## Person attributes and current-law benefit type
+
+After the reproduction gate succeeds, G1 supplies a side frame for every
+opening person, joined by `person_id`. Its labels are translated to the
+codes of G3's existing MINT8 dimensions; no category label is invented.
+The frame must retain each requested person exactly once and must match
+its recorded content seal when present. The output carries G1 provenance.
+
+Race and ethnicity and country of birth use G1's documented resolution of
+its head/spouse reports. Education uses G1's anchor-wave cutoff, separately
+for each population. Missing reports and unresolved placements, including
+G1's treatment of U.S.-territory births, remain unclassified. There is no
+imputation or guessed category.
+
+Sex and marital status come from the retained reference-year projected
+state, joined on `(draw, person_id)`. Age is reference year minus birth
+year. The projection carries opening divorced and never-married statuses
+forward and changes married to widowed when its existing linked-spouse
+mortality mechanism requires it. Its default opening specification treats
+separated people as married; the separate original separation flag is not
+retained in `TrackACohort`. Therefore this analysis **inherits that frozen
+married treatment**. A literal `separated` code, `unknown`, or
+`no_marriage_history` is unclassified. This is a disclosed projection
+convention; MINT8 does not supply a placement rule for separated people.
+
+Current-law benefit type uses the positive baseline amounts in the frozen
+`benefit_components` mapping:
+
+| Positive baseline components | MINT8 category |
+|---|---|
+| Aged or disabled widow component, with or without an own-worker component | Widow(er) (includes dually entitled) |
+| Spouse excess component, with or without an own-worker component | Spousal (includes dually entitled) |
+| Retired worker component alone | Retired worker only |
+| Disabled worker component alone | Disabled worker only |
+
+A baseline nonrecipient is unclassified for current-law benefit type.
+Concurrent spouse and widow components, unknown components, or concurrent
+retired and disabled worker components are refused. No alternative
+entitlement mechanism is inferred from the group labels.
+
+The closed projection does not calculate household income or an official
+poverty measure. Current-law poverty status, current-law household income
+quintiles, household income statistics and official poverty statistics are
+explicitly **not computed** with that reason. Their missing group inputs
+remain unclassified; they never become zero income or above-poverty codes.
+
+## Lifetime dimensions and their limitations
+
+The annual MINT8 beneficiary scheme has no lifetime-earnings dimension.
+This report appends G3's three MINT8 cohort-table dimensions: current-law
+initial AIME quintile, lifetime payroll tax quintile, and lifetime payroll
+tax quintile (shared). The result is labeled as a composite scheme, not
+presented as an unchanged SSA annual-table layout.
+
+Every measure calls G2 over the existing cohort careers, cut at the
+opening year: 2010 for the 2011 wave and 2008 for the 2009 wave. No earnings
+are projected or added. Initial AIME calls `initial_aime_at_62` with the
+explicit `mint8_initial_aime` convention: statutory computation years and
+a last-earnings-age cutoff of 61, also capped at the opening year. G2
+identifies younger persons' values as provisional through the last
+observed year and exposes truncated-history flags. These opening-career
+measures do not establish a fully observed lifetime history for younger
+members of the closed cohort. The output carries the supplied parameter
+revision, source/input/output hashes and G2 coverage flags.
+
+Own payroll-tax present value calls G2 with
+`TaxRateBasis.EMPLOYEE_EMPLOYER_PAID`: combined OASDI taxes actually paid,
+including the effective employee reductions in the captured credit years,
+rather than general-revenue reimbursements to the trust funds. G2 applies
+the contribution and benefit base and its explicit end-of-year timing,
+valuing each year's tax at age 62 with the captured trust-fund interest
+rates. The source loaders enforce their inherited SHA-256 pins. No future
+interest rate is invented or silently extended; a person needing a rate
+outside the captured coverage is not computed under
+`MissingRatePolicy.NOT_COMPUTED`.
+
+Shared taxes call the same G2 present-value function with an already-loaded
+marriage-episode frame. Taxes paid while married are shared equally using
+G2's stated convention; unmarried years retain own taxes. The adapter uses
+`MissingSpousePolicy.NOT_COMPUTED`. Missing marriage histories are not
+interpreted as never-married: those persons' shared values are excluded
+from classification and their unavailable count is recorded separately,
+without mutating G2's measure frame or provenance. If no episode loader is
+supplied, the shared dimension is not computed for that reason. Real-data
+history loading remains in cohort code; measure and tabulation code never
+open PSID files.
+
+Quintile thresholds are computed once per dimension and **ten-year birth
+cohort of the entire opening population**, with opening weights, by G3's
+weighted-percentile rule. Ties use that rule, including its midpoint at an
+exact cumulative-weight tie. The resulting category is fixed by
+`person_id` across registered rows, projection draws, scenarios and
+population variants. A survivor subset is not re-ranked. Missing measures
+remain unclassified. Labels retain G3's MINT8 order: Highest, Second
+highest, Middle, Second lowest, Lowest.
+
+## Populations, statistics and disclosure
+
+Each registered row has two separately labeled variants:
+
+| Variant | Population |
+|---|---|
+| Test population | Persons alive in the reference year with a positive frozen benefit in either scenario; the original row's scenario membership rules are retained |
+| MINT population | The row's current-law beneficiaries aged 60 or older in the reference year, with positive baseline benefit |
+
+For exercise 1, the frozen rows have positive benefits in both scenarios.
+Exercise 3 retains the union and its registered scenario-specific membership
+differences. The age threshold is an additional report population filter;
+it does not change the replayed five-age-group parent cells.
+
+G3 computes every reported cell. It reuses the existing A7 ratio of
+scenario means, the mean of individual ratios, registered draw dispersion
+and floor conventions. It also reports the MINT8 benefit statistics:
+percent with a decrease of at least 1 percent, percent with an increase of
+at least 1 percent, the unaffected remainder, and the weighted 10th,
+50th and 90th percentiles of each individual's percent change relative to
+current law. The MINT change statistics use current-law recipients with a
+positive selected baseline benefit. Empty or undefined statistics remain
+explicitly undefined.
+
+Within each variant, G3 computes the family-unit half split once on its
+full input frame and intersects that split with each group. It never
+re-splits a group's subset. G3 reports unweighted sample sizes, the
+under-100 SSA flag, the under-30 flag, numerator flags, uncertainty and
+unclassified counts. If any category in a dimension is below SSA's
+disclosure minimum, the **whole subgroup** is flagged suppressed. Cells
+remain in this research artifact with their flags; no small row is silently
+dropped.
+
+## Entry points, artifacts and permitted verification
+
+`scripts/run_projection_groups_registered.py` selects one exercise and
+requires a new issue #42 comment URL, a full 40-hex registered commit equal
+to HEAD, a clean tree, the ratified frozen specification, and a new output
+pair. It imports the Track A `_environment` resolver through `importlib`.
+The real input builder reproduces the frozen cohort/parameter construction;
+recorded PSID hashes and frozen output values are verified before groups.
+The entry script also pins this methodology specification.
+
+There is **one artifact per exercise**. This preserves independent parent
+identities, row manifests, source specifications and one-shot registrations
+and makes a failed reproduction of one exercise unable to certify the
+other. Each output is a new `runs/_groups_posthoc_v1.json`, with an
+exclusive-created `.env.json` sidecar. Both files are created only after
+the full replay and group computation succeed. The prior committed run
+artifacts are read-only and are never rewritten.
+
+`scripts/projection_groups_dry_run.py` runs only the inherited invented
+cohorts and supplied invented group reports, rates and interest factors.
+Its artifacts go under
+`docs/analysis/nasi_group_breakdowns_invented_20261001//` with the
+invented-data header. These runs verify mechanics and do not compare any
+model outcome to DYNASIM or SSA policy-option values.
+
+`scripts/make_nasi_repro_venv.sh` builds the reproduction environment from
+the versions recorded in the parents' environment sidecars, with an
+editable install. Reproduction requires the existing policyengine-us
+checkout at revision `a03e82e503`, selected by
+`POPULACE_DYNAMICS_PE_US_DIR`, and the original PSID file hashes. Building
+the environment is permitted during development; running a real-data
+entry point to produce group cells is deferred to the new registration.
+
+Development verification is targeted and invented-only: adapter
+comparisons with both frozen runners, property tests for row/draw
+selection, joins and classification invariants, the mismatch-before-loader
+and no-write test, and entry-point guard tests. Run only one pytest process
+at a time. Black at 79 columns and Ruff apply to the new Python files.
+
+No existing frozen engine, Track A, FRA68, SS or A7 module is edited. No
+new branch or worktree is created; changes stay uncommitted in the assigned
+`nasi/g4a-projection-groups` worktree for the orchestrator. The integrator
+alone updates `POST_REVIEW_SOURCE_EXCLUSIONS` in
+`scripts/first_estimates_birth_evidence.py`, its corresponding artifact
+test, `tests/tier_counts.json`, and any required development extras in
+`pyproject.toml`. New source modules and test modules with tiers are listed
+in the package's final report.
diff --git a/scripts/make_nasi_repro_venv.sh b/scripts/make_nasi_repro_venv.sh
new file mode 100755
index 00000000..d03d8ed5
--- /dev/null
+++ b/scripts/make_nasi_repro_venv.sh
@@ -0,0 +1,146 @@
+#!/usr/bin/env bash
+# Build only the reproduction environment; never run a PSID pipeline.
+# Pins are read from the bound exercise-1/3 model sidecars, not guessed.
+set -euo pipefail
+
+repo_root=$(CDPATH= cd -- "$(dirname -- "$0")/.." && pwd)
+default_checkout=/Users/maxghenis/PolicyEngine/microcosm-dynamics
+target_dir="$default_checkout/.venv-nasi-repro"
+source_python="$default_checkout/.venv/bin/python"
+
+usage() {
+ cat <<'EOF'
+Usage: scripts/make_nasi_repro_venv.sh [--target-dir PATH] [--source-python PATH]
+
+Defaults to the main checkout's .venv-nasi-repro. --target-dir supports a
+restricted-workspace build verification. Existing environments are refused.
+The script checks both frozen sidecars, installs their exact Python/numpy/
+pandas/scipy versions and this checkout editable, then verifies those versions.
+No microdata or projected outcome is read or computed.
+EOF
+}
+
+while (($#)); do
+ case "$1" in
+ --target-dir|--source-python)
+ if (($# < 2)); then
+ usage >&2
+ exit 2
+ fi
+ if [[ "$1" == --target-dir ]]; then
+ target_dir=$2
+ else
+ source_python=$2
+ fi
+ shift 2
+ ;;
+ --help|-h)
+ usage
+ exit 0
+ ;;
+ *)
+ usage >&2
+ exit 2
+ ;;
+ esac
+done
+
+if [[ -e "$target_dir" ]]; then
+ printf 'Refusing existing environment: %s\n' "$target_dir" >&2
+ exit 1
+fi
+command -v uv >/dev/null
+scratch_dir=$(mktemp -d "${TMPDIR:-/tmp}/nasi-repro.XXXXXX")
+trap 'rm -rf -- "$scratch_dir"' EXIT
+# uv's user cache and Python install directories are outside a worktree's
+# write sandbox. Keep downloads in the assigned temporary directory.
+export UV_CACHE_DIR="$scratch_dir/uv-cache"
+export UV_PYTHON_INSTALL_DIR="$scratch_dir/uv-python"
+
+"$source_python" - "$repo_root" "$scratch_dir" <<'PY'
+import hashlib
+import json
+import platform
+import sys
+from pathlib import Path
+
+root, scratch = map(Path, sys.argv[1:])
+pins = {
+ "cola": "270acf292682b8111133f9047f366c93173e713bc422cb7e108bd064ac33d53e",
+ "fra68": "9c768fff17bfd828d16bdca738f92d8487f082ec99cd94ee0c73959bb050746e",
+}
+environment_pins = {
+ "cola": "a79b54ec5d3aadaacee41e8877e5a985821fbd3bba82b97fe8fa962d8aa8e6f9",
+ "fra68": "8c04da45818cc88ece6693b12fd630b624c3c8d6843829ed92dec258af81fab6",
+}
+environments = []
+for exercise, expected_digest in pins.items():
+ artifact = root / "runs" / f"replication_urban2010_{exercise}_v1.json"
+ digest = hashlib.sha256(artifact.read_bytes()).hexdigest()
+ sidecar_path = artifact.with_suffix(".env.json")
+ if hashlib.sha256(sidecar_path.read_bytes()).hexdigest() != environment_pins[exercise]:
+ raise SystemExit(f"{exercise}: environment sidecar SHA-256 differs")
+ sidecar = json.loads(sidecar_path.read_text())
+ if digest != expected_digest or sidecar["artifact_sha256"] != digest:
+ raise SystemExit(f"{exercise}: parent/sidecar binding differs")
+ if sidecar["artifact"] != artifact.name:
+ raise SystemExit(f"{exercise}: sidecar names a different artifact")
+ environment = sidecar["environment"]
+ versions = {"python": environment["python"]}
+ for package in ("numpy", "pandas", "scipy"):
+ version = environment["packages"].get(package)
+ if not version:
+ raise SystemExit(f"{exercise}: missing {package} version")
+ versions[package] = version
+ environments.append(versions)
+if environments[0] != environments[1]:
+ raise SystemExit("Exercise-1 and exercise-3 environments differ")
+versions = environments[0]
+if platform.python_version() != versions["python"]:
+ raise SystemExit("--source-python must match the exact recorded Python")
+(scratch / "versions.json").write_text(json.dumps(versions))
+(scratch / "requirements.txt").write_text(
+ "".join(f"{key}=={versions[key]}\n" for key in ("numpy", "pandas", "scipy"))
+)
+print("Recorded versions: " + json.dumps(versions, sort_keys=True))
+PY
+
+# Reuse the exact interpreter already on the host rather than making the
+# resulting environment depend on a temporary downloaded Python tree.
+uv venv --python "$source_python" "$target_dir"
+uv pip install --python "$target_dir/bin/python" \
+ --requirements "$scratch_dir/requirements.txt" --editable "$repo_root"
+"$target_dir/bin/python" - "$scratch_dir/versions.json" "$repo_root" <<'PY'
+import importlib.metadata
+import json
+import platform
+import sys
+from pathlib import Path
+
+expected = json.loads(Path(sys.argv[1]).read_text())
+actual = {"python": platform.python_version()}
+for package in ("numpy", "pandas", "scipy"):
+ actual[package] = importlib.metadata.version(package)
+if actual != expected:
+ raise SystemExit(f"Version verification failed: {actual!r} != {expected!r}")
+import populace_dynamics
+
+source = Path(populace_dynamics.__file__).resolve()
+if not source.is_relative_to(Path(sys.argv[2]) / "src"):
+ raise SystemExit(f"Editable install points outside assigned source: {source}")
+print("Verified exact recorded versions and assigned editable source")
+PY
+
+parameter_checkout="$default_checkout/.claude/pe-us-a03e82e503"
+if [[ ! -d "$parameter_checkout" ]]; then
+ printf 'Missing existing parameter checkout: %s\n' "$parameter_checkout" >&2
+ exit 1
+fi
+parameter_revision=$(git -C "$parameter_checkout" rev-parse HEAD)
+if [[ "$parameter_revision" != a03e82e503* ]]; then
+ printf 'Parameter checkout revision differs: %s\n' "$parameter_revision" >&2
+ exit 1
+fi
+printf 'Built %s\n' "$target_dir"
+printf 'export POPULACE_DYNAMICS_PE_US_DIR=%q\n' "$parameter_checkout"
+printf 'Use %s/bin/python only after a new issue #42 registration.\n' "$target_dir"
diff --git a/scripts/projection_groups_dry_run.py b/scripts/projection_groups_dry_run.py
new file mode 100644
index 00000000..d07c4d7a
--- /dev/null
+++ b/scripts/projection_groups_dry_run.py
@@ -0,0 +1,221 @@
+"""INVENTED DATA - NOT A COMPARISON: exercise 1/3 group adapter dry run.
+
+All people, careers, rates, weights and group reports are invented. The
+inherited statutory FRA schedule supplies the mechanism, not observations.
+No PSID files, real parameters, model artifacts or comparator values load.
+G1 builds attributes, G2 builds lifetime measures and G3 builds every cell.
+"""
+
+from __future__ import annotations
+
+import argparse
+import hashlib
+import json
+import sys
+from dataclasses import replace
+from pathlib import Path
+
+import pandas as pd
+
+ROOT = Path(__file__).resolve().parents[1]
+if str(ROOT / "src") not in sys.path:
+ sys.path.insert(0, str(ROOT / "src"))
+
+from populace_dynamics.cohorts import group_attributes as g1 # noqa: E402
+from populace_dynamics.cohorts import psid2010 # noqa: E402
+from populace_dynamics.cola_track_a import ( # noqa: E402
+ TrackAConfig,
+ invented,
+ prepare_track_a_cohort,
+ run_track_a,
+)
+from populace_dynamics.data import group_attributes_psid as gap # noqa: E402
+from populace_dynamics.estimates import lifetime_measures as g2 # noqa: E402
+from populace_dynamics.fra68_track import FRA68Config, run_fra68 # noqa: E402
+from populace_dynamics.group_breakdowns.cola import ( # noqa: E402
+ reproduce_cola,
+)
+from populace_dynamics.group_breakdowns.common import ( # noqa: E402
+ INVENTED_HEADER,
+ LifetimeOptions,
+ run_group_breakdown,
+)
+from populace_dynamics.group_breakdowns.fra68 import ( # noqa: E402
+ reproduce_fra68,
+)
+from populace_dynamics.track_a_v2.invented import invented_inputs # noqa: E402
+
+
+def build_invented_inputs(seed: int = 7):
+ """Both original waves, with invented rates and their marriage histories."""
+ inputs = invented_inputs(seed)
+ raw_by_wave = {
+ wave: invented.invented_psid2010_inputs(
+ seed=seed, anchor_wave=wave, claiming_pmf=inputs.claiming_pmf
+ )
+ for wave in (2011, 2009)
+ }
+ secondary = prepare_track_a_cohort(
+ psid2010.build_psid2010_cohort(
+ raw_by_wave[2009], psid2010.Psid2010CohortSpec(anchor_wave=2009)
+ ),
+ data_provenance="invented",
+ config=TrackAConfig(),
+ )
+ mortality = replace(
+ inputs.population_mortality,
+ ratio_by_year={year: (1.0, 1.0) for year in range(2009, 2031)},
+ )
+ return (
+ replace(
+ inputs,
+ additional_cohorts=(secondary,),
+ population_mortality=mortality,
+ ),
+ raw_by_wave,
+ )
+
+
+def invented_attribute_loader(person_ids, *, anchor_waves):
+ """Pass INVENTED reports through the actual G1 cohort builder."""
+ wave = anchor_waves[0]
+ reports = []
+ education = []
+ for index, pid in enumerate(person_ids):
+ report = dict.fromkeys(gap.REPORT_CODE_COLUMNS, pd.NA)
+ report.update(person_id=pid, wave=wave, role="head")
+ report["hispanic_code"] = 1 if index % 4 == 0 else 0
+ for mention in range(1, len(gap.FAMILY_ITEMS[wave].race["head"]) + 1):
+ report[f"race_code_{mention}"] = (
+ (1, 2, 3)[index % 3] if mention == 1 else 0
+ )
+ if wave in gap.BIRTHPLACE_WAVES:
+ report["birth_state_code"] = 1
+ report["year_came_code"] = 0
+ if wave in gap.FAMILY_EDUCATION_WAVES:
+ report["family_education_code"] = 12
+ reports.append(report)
+ education.append(
+ {
+ "person_id": pid,
+ "wave": wave,
+ "education_code": (11, 12, 14, 16, 17)[index % 5],
+ "role": "head",
+ "family_education_code": pd.NA,
+ }
+ )
+ report_frame = pd.DataFrame(reports)
+ education_frame = pd.DataFrame(education)
+ for frame in (report_frame, education_frame):
+ for column in frame:
+ frame[column] = frame[column].astype(
+ "string" if column == "role" else "Int64"
+ )
+ supplied = g1.GroupAttributeInputs(
+ reports=report_frame,
+ education=education_frame,
+ universe=pd.Series(person_ids, dtype="int64", name="person_id"),
+ provenance={"kind": "INVENTED", "fixture": "projection group dry run"},
+ )
+ return g1.build_group_attributes(
+ supplied, person_ids, anchor_waves=anchor_waves
+ )
+
+
+def invented_lifetime_options(raw_by_wave) -> LifetimeOptions:
+ """INVENTED tax and interest values; no captured data values are used."""
+ rates = g2.OASDITaxRates(
+ combined_percent_by_year={},
+ open_ended_from=1937,
+ open_ended_combined_percent=12.4,
+ basis=g2.TaxRateBasis.EMPLOYEE_EMPLOYER_PAID,
+ provenance={"kind": "INVENTED"},
+ )
+ interest = g2.TrustFundInterestRates(
+ percent_by_year={year: 3.0 for year in range(1940, 2101)},
+ series="INVENTED",
+ provenance={"kind": "INVENTED"},
+ )
+
+ def episodes(wave):
+ history = raw_by_wave[wave].marriage_history
+ return psid2010._episodes_with_separation(history), tuple(
+ int(pid) for pid in history.person_id.unique()
+ )
+
+ return LifetimeOptions(
+ rates=rates, interest=interest, marriage_episode_loader=episodes
+ )
+
+
+def dry_run(exercise: str, *, draws: int = 2, seed: int = 7):
+ """Differential replay against a frozen runner on INVENTED inputs."""
+ inputs, raw = build_invented_inputs(seed)
+ if exercise == "cola":
+ config = TrackAConfig(draw_indices=tuple(range(draws)))
+ parent = run_track_a(inputs, config=config)
+ replay = reproduce_cola(inputs, config=config)
+ elif exercise == "fra68":
+ config = FRA68Config(draw_indices=tuple(range(draws)))
+ parent = run_fra68(inputs, config=config)
+ replay = reproduce_fra68(inputs, config=config)
+ else:
+ raise ValueError("exercise must be cola or fra68")
+ digest = hashlib.sha256(
+ json.dumps(parent, sort_keys=True, allow_nan=False).encode()
+ ).hexdigest()
+ return run_group_breakdown(
+ replay,
+ parent,
+ inputs=inputs,
+ config=config,
+ parent_sha256=digest,
+ attribute_loader=invented_attribute_loader,
+ lifetime_options=invented_lifetime_options(raw),
+ )
+
+
+def main(argv=None) -> int:
+ parser = argparse.ArgumentParser(description=__doc__.splitlines()[0])
+ parser.add_argument(
+ "--exercise", choices=("cola", "fra68", "both"), default="both"
+ )
+ parser.add_argument(
+ "--output-dir",
+ type=Path,
+ default=ROOT / "docs/analysis/nasi_group_breakdowns_invented_20261001",
+ )
+ parser.add_argument("--draws", type=int, default=2)
+ parser.add_argument("--seed", type=int, default=7)
+ args = parser.parse_args(argv)
+ if args.draws < 1:
+ parser.error("--draws must be positive")
+ exercises = (
+ ("cola", "fra68") if args.exercise == "both" else (args.exercise,)
+ )
+ for exercise in exercises:
+ result = dry_run(exercise, draws=args.draws, seed=args.seed)
+ encoded = json.dumps(result, indent=2, allow_nan=False) + "\n"
+ output = args.output_dir / exercise
+ output.mkdir(parents=True, exist_ok=True)
+ (output / "result.json").write_text(encoded)
+ (output / "RESULTS.md").write_text(
+ f"# {INVENTED_HEADER}\n\nExercise: {exercise}. "
+ "Registered, one-shot, post hoc, not blind; report-only labels "
+ "describe the intended registered artifact. This dry run is "
+ "unregistered and uses invented data only.\n\n"
+ f"All {len(result['rows'])} registered rows reproduced the "
+ "frozen runner exactly before group loading. Both original "
+ "waves and both population variants are in [result.json](result.json). "
+ "G1 supplies attributes, G2 supplies lifetime measures, and G3 "
+ "supplies labels, suppression flags and every group cell.\n\n"
+ "No PSID file, real-data outcome or comparator value was used. "
+ "Income/poverty are unavailable. Lifetime careers stop at the "
+ "opening year, and interest/tax inputs here are invented.\n"
+ )
+ print(output)
+ return 0
+
+
+if __name__ == "__main__":
+ raise SystemExit(main())
diff --git a/scripts/run_projection_groups_registered.py b/scripts/run_projection_groups_registered.py
new file mode 100644
index 00000000..43abc4a4
--- /dev/null
+++ b/scripts/run_projection_groups_registered.py
@@ -0,0 +1,440 @@
+"""Registered, one-shot post hoc group breakdowns of projection tests 1/3.
+
+This entry point composes the frozen registered builders and refuses before
+group attributes or cells unless every parent-artifact value reproduces.
+Each exercise has its own new artifact and environment sidecar because its
+registration, parent specification, and row manifest are independent. No
+real-data run is authorized without a new issue #42 comment and its clean
+registered commit. The v1 artifacts remain the reproduction references.
+
+Usage::
+
+ python scripts/run_projection_groups_registered.py \\
+ --exercise cola --registration-pointer \\
+ --registered-commit
+
+See scripts/run_track_a_registered.py and scripts/run_fra68_registered.py
+for the frozen builders, guards, provenance records and environment resolver.
+"""
+
+from __future__ import annotations
+
+import argparse
+import datetime
+import hashlib
+import importlib.util
+import json
+import re
+import subprocess
+import sys
+from pathlib import Path
+from typing import Any
+
+ROOT = Path(__file__).resolve().parents[1]
+if str(ROOT / "src") not in sys.path:
+ sys.path.insert(0, str(ROOT / "src"))
+
+REGISTRATION_POINTER = re.compile(
+ r"https://github\.com/PolicyEngine/microcosm-dynamics/issues/42"
+ r"#issuecomment-[0-9]+"
+)
+PARENT_SHA256 = {
+ "cola": (
+ "270acf292682b8111133f9047f366c93173e713bc422cb7e108bd064ac33d53e"
+ ),
+ "fra68": (
+ "9c768fff17bfd828d16bdca738f92d8487f082ec99cd94ee0c73959bb050746e"
+ ),
+}
+PARENT_ENVIRONMENT_SHA256 = {
+ "cola": (
+ "a79b54ec5d3aadaacee41e8877e5a985821fbd3bba82b97fe8fa962d8aa8e6f9"
+ ),
+ "fra68": (
+ "8c04da45818cc88ece6693b12fd630b624c3c8d6843829ed92dec258af81fab6"
+ ),
+}
+POSTHOC_LABEL = "registered, one-shot, post hoc, not blind"
+POSTHOC_SPECIFICATION_SHA256 = (
+ "690bcceb102e1ca928c8c9af0754739f433df00f29330294481051ab4ef7f4b9"
+)
+
+
+def _git(*args: str) -> str:
+ return subprocess.run(
+ ["git", "-C", str(ROOT), *args],
+ check=True,
+ capture_output=True,
+ text=True,
+ ).stdout.strip()
+
+
+def _sha256(path: Path) -> str:
+ return hashlib.sha256(path.read_bytes()).hexdigest()
+
+
+def _script_module(name: str) -> Any:
+ """Import the existing registered script, without executing its main."""
+ path = ROOT / "scripts" / f"{name}.py"
+ spec = importlib.util.spec_from_file_location(f"_groups_{name}", path)
+ if spec is None or spec.loader is None:
+ raise ValueError(f"cannot resolve frozen registered script {path}")
+ module = importlib.util.module_from_spec(spec)
+ spec.loader.exec_module(module)
+ return module
+
+
+def _check_specification(
+ exercise: str, specification: dict[str, Any] | None, config: Any
+) -> None:
+ if exercise == "cola":
+ legacy = _script_module("run_track_a_registered")
+ legacy.check_specification_ratified(
+ legacy.a1_parameter_block()
+ if specification is None
+ else specification
+ )
+ else:
+ legacy = _script_module("run_fra68_registered")
+ legacy.check_specification_for_registered_run(
+ (
+ legacy.e1_parameter_block()
+ if specification is None
+ else specification
+ ),
+ config or legacy.FRA68Config(),
+ )
+
+
+def preflight(
+ *,
+ exercise: str,
+ registration_pointer: str,
+ registered_commit: str,
+ output: Path,
+ git: Any = _git,
+ specification: dict[str, Any] | None = None,
+ config: Any = None,
+) -> dict[str, str]:
+ """Require registration, clean exact HEAD and a new output pair first."""
+ if exercise not in PARENT_SHA256:
+ raise ValueError("exercise must be cola or fra68")
+ if not REGISTRATION_POINTER.fullmatch(registration_pointer):
+ raise ValueError(
+ "registration pointer must be an issue #42 comment URL"
+ )
+ if not re.fullmatch(r"[0-9a-f]{40}", registered_commit):
+ raise ValueError("--registered-commit must be a full 40-hex SHA")
+ head = git("rev-parse", "HEAD")
+ if head != registered_commit:
+ raise ValueError("HEAD is not the registered commit")
+ if git("status", "--porcelain"):
+ raise ValueError("working tree must be clean for a registered run")
+ if output.exists() or output.with_suffix(".env.json").exists():
+ raise ValueError("registered artifact or sidecar already exists")
+ _check_specification(exercise, specification, config)
+ return {"head": head}
+
+
+def _parent(exercise: str) -> tuple[dict[str, Any], dict[str, Any], str]:
+ """Read only the committed model artifact and its binding sidecar."""
+ path = ROOT / "runs" / f"replication_urban2010_{exercise}_v1.json"
+ digest = _sha256(path)
+ if digest != PARENT_SHA256[exercise]:
+ raise ValueError("parent artifact SHA-256 differs from its frozen pin")
+ sidecar_path = path.with_suffix(".env.json")
+ sidecar_digest = _sha256(sidecar_path)
+ if sidecar_digest != PARENT_ENVIRONMENT_SHA256[exercise]:
+ raise ValueError(
+ "parent environment SHA-256 differs from its frozen pin"
+ )
+ sidecar = json.loads(sidecar_path.read_text(encoding="utf-8"))
+ if (
+ sidecar.get("artifact_sha256") != digest
+ or sidecar.get("artifact") != path.name
+ ):
+ raise ValueError("parent environment sidecar is not bound to artifact")
+ return (
+ json.loads(path.read_text(encoding="utf-8")),
+ sidecar["environment"],
+ sidecar_digest,
+ )
+
+
+def check_reproduction_environment(
+ current: dict[str, Any], expected: dict[str, Any], revision: str
+) -> None:
+ """Require the four recorded versions and the frozen parameter checkout."""
+ if current.get("python") != expected.get("python"):
+ raise ValueError("Python version differs from parent environment")
+ for package in ("numpy", "pandas", "scipy"):
+ wanted = expected.get("packages", {}).get(package)
+ if not wanted or current.get("packages", {}).get(package) != wanted:
+ raise ValueError(
+ f"{package} version differs from parent environment"
+ )
+ if revision != "a03e82e503":
+ raise ValueError("policyengine-us parameters require a03e82e503")
+
+
+def build_registered_inputs(
+ exercise: str,
+ config: Any,
+ *,
+ parent: dict[str, Any],
+ expected_environment: dict[str, Any],
+) -> tuple[Any, dict[str, Any], dict[int, Any], Any]:
+ """Compose existing cohort readers exactly, retaining marriage inputs.
+
+ No PSID reader is called outside psid2010 cohort code. Version guards run
+ before cohort loading. The raw marriage histories remain unused until the
+ common reproduction gate has passed; their later episode loader is lazy.
+ """
+ legacy = _script_module(
+ "run_track_a_registered"
+ if exercise == "cola"
+ else "run_fra68_registered"
+ )
+ track_config = config if exercise == "cola" else config.track_a_config()
+ base_params = legacy.load_ssa_parameters()
+ environment = _script_module("run_track_a_registered")._environment(
+ ssa_parameters_revision=base_params.pe_us_revision
+ )
+ check_reproduction_environment(
+ environment, expected_environment, base_params.pe_us_revision
+ )
+ claiming_pmf = legacy.load_claiming_pmf()
+ raw_inputs = {}
+ cohorts = {}
+ for wave in config.anchor_waves:
+ raw = legacy.psid2010.load_psid2010_inputs(anchor_wave=wave)
+ a3 = legacy.psid2010.build_psid2010_cohort(
+ raw, legacy.psid2010.Psid2010CohortSpec(anchor_wave=wave)
+ )
+ raw_inputs[wave] = raw
+ cohorts[wave] = legacy.prepare_track_a_cohort(
+ a3, data_provenance="registered_real", config=track_config
+ )
+ params = legacy.tr2008_ssa_parameters(
+ base_params, alternative=config.tr2008_alternative
+ )
+ realized = legacy.load_cola_history()
+ baseline = legacy.tr2008_baseline_cola(
+ realized,
+ first_year=config.tr2008_first_rate_year,
+ last_year=config.reference_year,
+ alternative=config.tr2008_alternative,
+ )
+ rate_check = _script_module("track_a_dry_run")._spec_rate_check(baseline)
+ di_rates = legacy.load_di_entitlement_rates(config.di_spec)
+ first_projection_year = min(track_config.start_years.values()) + 1
+ mortality = legacy.load_tr2008_mortality(
+ range(first_projection_year, config.reference_year + 1),
+ alternative=config.tr2008_alternative,
+ base_year=config.mortality_base_year,
+ )
+ provenance = {
+ "ssa_parameters": params.pe_us_revision,
+ "statutory_capture": {
+ "path": str(legacy.CAPTURE_PATH.relative_to(ROOT)),
+ "sha256": legacy.CAPTURE_SHA256,
+ },
+ "tr2008_file_sha256": dict(legacy.tr2008.FILE_SHA256),
+ "di_rates": {
+ key: value
+ for key, value in di_rates.provenance.items()
+ if key.endswith("sha256")
+ },
+ "cola_history_sha256": realized.provenance["sha256"],
+ "claiming_reference_sha256": _sha256(
+ legacy.psid2010.CLAIMING_REFERENCE_PATH
+ ),
+ "a1_specification_sha256": _sha256(legacy.A1_SPECIFICATION_PATH),
+ "mortality": dict(mortality.provenance),
+ "psid_inputs": {
+ str(wave): dict(raw.provenance) for wave, raw in raw_inputs.items()
+ },
+ }
+ if exercise == "fra68":
+ provenance["e1_specification_sha256"] = _sha256(
+ legacy.E1_SPECIFICATION_PATH
+ )
+ if provenance != parent.get("inputs_provenance"):
+ raise ValueError("registered inputs differ from parent provenance")
+ primary, *others = (cohorts[wave] for wave in config.anchor_waves)
+ inputs = legacy.TrackAInputs(
+ cohort=primary,
+ params=params,
+ baseline=baseline,
+ di_rates=di_rates,
+ population_mortality=mortality,
+ claiming_pmf=claiming_pmf,
+ provenance=provenance,
+ additional_cohorts=tuple(others),
+ )
+ return inputs, environment, raw_inputs, rate_check
+
+
+def write_new_pair(
+ output: Path, artifact: dict[str, Any], environment: dict[str, Any]
+) -> None:
+ """Serialize finite JSON first, then exclusively create the new pair.
+
+ A sidecar race or write failure removes only files this invocation created.
+ Existing artifact/sidecar bytes are never touched. The caller invokes this
+ only after reproduction and all group computations have succeeded.
+ """
+ artifact_text = json.dumps(artifact, indent=2, allow_nan=False) + "\n"
+ digest = hashlib.sha256(artifact_text.encode("utf-8")).hexdigest()
+ sidecar = output.with_suffix(".env.json")
+ sidecar_text = (
+ json.dumps(
+ {
+ "artifact": output.name,
+ "artifact_sha256": digest,
+ "environment": environment,
+ },
+ indent=2,
+ allow_nan=False,
+ )
+ + "\n"
+ )
+ created = []
+ try:
+ with output.open("x", encoding="utf-8") as artifact_handle:
+ created.append(output)
+ with sidecar.open("x", encoding="utf-8") as environment_handle:
+ created.append(sidecar)
+ artifact_handle.write(artifact_text)
+ environment_handle.write(sidecar_text)
+ except BaseException:
+ for path in reversed(created):
+ path.unlink(missing_ok=True)
+ raise
+
+
+def main(argv: list[str] | None = None) -> int:
+ parser = argparse.ArgumentParser(description=__doc__.splitlines()[0])
+ parser.add_argument("--exercise", choices=("cola", "fra68"), required=True)
+ parser.add_argument("--registration-pointer", required=True)
+ parser.add_argument("--registered-commit", required=True)
+ parser.add_argument("--output", type=Path)
+ args = parser.parse_args(argv)
+ output = args.output or (
+ ROOT
+ / "runs"
+ / f"replication_urban2010_{args.exercise}_groups_posthoc_v1.json"
+ )
+ started = datetime.datetime.now(datetime.timezone.utc).isoformat()
+ state = preflight(
+ exercise=args.exercise,
+ registration_pointer=args.registration_pointer,
+ registered_commit=args.registered_commit,
+ output=output,
+ )
+ posthoc_specification = (
+ ROOT / "docs/design/nasi_projection_groups_posthoc.md"
+ )
+ posthoc_specification_sha256 = _sha256(posthoc_specification)
+ if posthoc_specification_sha256 != POSTHOC_SPECIFICATION_SHA256:
+ raise ValueError("post hoc methodology specification SHA-256 differs")
+ parent, expected_environment, parent_environment_sha256 = _parent(
+ args.exercise
+ )
+ if args.registration_pointer == parent["registration_pointer"]:
+ raise ValueError(
+ "group outputs require a new issue #42 registration pointer"
+ )
+ if args.exercise == "cola":
+ from populace_dynamics.cola_track_a import TrackAConfig
+ from populace_dynamics.group_breakdowns.cola import reproduce_cola
+
+ config = TrackAConfig()
+ reproduce = reproduce_cola
+ else:
+ from populace_dynamics.fra68_track import FRA68Config
+ from populace_dynamics.group_breakdowns.fra68 import reproduce_fra68
+
+ config = FRA68Config()
+ reproduce = reproduce_fra68
+ from populace_dynamics.cohorts import psid2010
+ from populace_dynamics.group_breakdowns.common import (
+ LifetimeOptions,
+ run_group_breakdown,
+ )
+
+ inputs, environment, raw_inputs, rate_check = build_registered_inputs(
+ args.exercise,
+ config,
+ parent=parent,
+ expected_environment=expected_environment,
+ )
+ replay = reproduce(
+ inputs,
+ config=config,
+ registration_pointer=parent["registration_pointer"],
+ check_committed_parameters=True,
+ progress=lambda message: print(message, file=sys.stderr),
+ )
+
+ def marriage_episodes(wave: int) -> tuple[Any, tuple[int, ...]]:
+ history = raw_inputs[wave].marriage_history
+ return (
+ psid2010._episodes_with_separation(history),
+ tuple(int(value) for value in history.person_id.unique()),
+ )
+
+ result = run_group_breakdown(
+ replay,
+ parent,
+ inputs=inputs,
+ config=config,
+ parent_sha256=PARENT_SHA256[args.exercise],
+ registration_pointer=args.registration_pointer,
+ lifetime_options=LifetimeOptions(
+ marriage_episode_loader=marriage_episodes
+ ),
+ )
+ artifact = {
+ **result,
+ "header": f"{POSTHOC_LABEL}; report-only; exercise {args.exercise}",
+ "labels": list(
+ dict.fromkeys([*parent["labels"], POSTHOC_LABEL, "report-only"])
+ ),
+ "publishes_regardless": True,
+ "comparator_seal_opened_before_commit": False,
+ "data_provenance": "registered_real",
+ "registration_pointer": args.registration_pointer,
+ "parent_registration_pointer": parent["registration_pointer"],
+ "parent_environment_sha256": parent_environment_sha256,
+ "posthoc_specification": {
+ "path": str(posthoc_specification.relative_to(ROOT)),
+ "sha256": posthoc_specification_sha256,
+ },
+ "checks": {"registered_rate_path": rate_check},
+ "run": {
+ "started": started,
+ "finished": datetime.datetime.now(
+ datetime.timezone.utc
+ ).isoformat(),
+ "registration_pointer": args.registration_pointer,
+ "registered_commit": args.registered_commit,
+ "git_head": state["head"],
+ "git_clean": True,
+ "command": " ".join(
+ [
+ "python",
+ "scripts/run_projection_groups_registered.py",
+ *(sys.argv[1:] if argv is None else argv),
+ ]
+ ),
+ },
+ }
+ output.parent.mkdir(parents=True, exist_ok=True)
+ write_new_pair(output, artifact, environment)
+ print(output)
+ return 0
+
+
+if __name__ == "__main__":
+ raise SystemExit(main())
diff --git a/src/populace_dynamics/group_breakdowns/__init__.py b/src/populace_dynamics/group_breakdowns/__init__.py
new file mode 100644
index 00000000..34fda92b
--- /dev/null
+++ b/src/populace_dynamics/group_breakdowns/__init__.py
@@ -0,0 +1 @@
+"""Registered, report-only post hoc breakdowns of frozen projections."""
diff --git a/src/populace_dynamics/group_breakdowns/cola.py b/src/populace_dynamics/group_breakdowns/cola.py
new file mode 100644
index 00000000..1fdf0e32
--- /dev/null
+++ b/src/populace_dynamics/group_breakdowns/cola.py
@@ -0,0 +1,248 @@
+"""Frozen exercise-1 replay for registered post hoc group breakdowns.
+
+The orchestration follows ``cola_track_a/runner.py:953-1158`` and the
+registered A1 specification, ``docs/design/urban2010_cola_comparison.md``.
+All guards, projections, benefit calculations, diagnostics and five-age
+cell tabulations call the frozen modules read-only. The additional return
+values retain the final person state and benefit rows already produced by
+the registered runner; no new mechanism or measure is implemented here.
+
+Invariants: one projection per population and draw; every row reads its
+population's shared path; the aggregate result equals the frozen runner;
+no group attribute or group cell is computed before the caller's exact
+reproduction check. Invented differential and property tests establish
+these invariants without reading PSID or comparator outcomes.
+"""
+
+from __future__ import annotations
+
+from collections import Counter
+from collections.abc import Callable, Mapping
+from typing import Any
+
+import pandas as pd
+
+from populace_dynamics.cola_track_a import runner as legacy
+from populace_dynamics.group_breakdowns.common import ProjectionReplay
+
+__all__ = ["reproduce_cola"]
+
+
+def reproduce_cola(
+ inputs: legacy.TrackAInputs,
+ *,
+ config: legacy.TrackAConfig | None = None,
+ registration_pointer: str | None = None,
+ progress: Callable[[str], None] | None = None,
+ check_committed_parameters: bool = False,
+ specification: Mapping[str, Any] | None = None,
+) -> ProjectionReplay:
+ """Replay R0-R6 with frozen guards and retain person rows for G1/G2/G3.
+
+ The result equals ``cola_track_a.runner.run_track_a`` on identical
+ inputs. Populations share their projection across registered rows;
+ scenario benefits share simulated paths. The caller must compare the
+ reproduced result to its committed parent before computing any group
+ attribute or cell. This function neither loads group attributes nor
+ writes an artifact.
+ """
+
+ config = config or legacy.TrackAConfig()
+ config.check_runnable()
+ if specification is None:
+ specification = legacy.a1_parameter_block()
+ recorded_specification = legacy.specification_record(specification)
+ cohorts = legacy._population_cohorts(inputs, config)
+ data_provenance = next(iter(cohorts.values())).data_provenance
+ labels = tuple(next(iter(cohorts.values())).labels)
+ if data_provenance == legacy.REGISTERED_REAL and not registration_pointer:
+ raise ValueError(
+ "no real-data statistic before the issue #42 registration "
+ "comment exists; pass its pointer to run a registered_real cohort"
+ )
+ real = data_provenance == legacy.REGISTERED_REAL
+ departures = legacy.rulings_departures(config)
+ if real and departures:
+ raise ValueError(
+ f"the configuration departs from Max's rulings on {departures} "
+ "(2026-09-23, d074/d075); a registered run follows every ruling"
+ )
+ floor_minima = legacy._check_floor(inputs, config)
+ consistency = legacy._parameter_consistency(
+ inputs,
+ config,
+ cohorts,
+ compare_committed_values=real or check_committed_parameters,
+ )
+ if real and not consistency["consistent"]:
+ failed = sorted(
+ name
+ for name, check in consistency["checks"].items()
+ if not check["consistent"]
+ )
+ raise ValueError(
+ "the inputs differ from the parameters the configuration "
+ f"reports for this run: {failed}; a registered run refuses "
+ "a mismatch"
+ )
+ legacy._check_source_provenance(cohorts, data_provenance)
+ legacy._check_output_labels(labels, data_provenance)
+ schedule = legacy.claiming_schedule(
+ inputs.claiming_pmf, max_table_year=config.claim_table_max_year
+ )
+ fra_schedule = legacy.fra_schedule_from_parameters(inputs.params)
+ rows_by_row: dict[str, list[dict]] = {row: [] for row in config.rows}
+ counters_by_row: dict[str, Counter] = {
+ row: Counter() for row in config.rows
+ }
+ draws: dict[str, Any] = {}
+ states_by_wave: dict[int, list[pd.DataFrame]] = {
+ wave: [] for wave in cohorts
+ }
+ for wave, cohort in cohorts.items():
+ results, draws[str(wave)] = legacy._project_population(
+ cohort,
+ inputs,
+ config,
+ schedule=schedule,
+ fra_schedule=fra_schedule,
+ progress=progress,
+ )
+ context = legacy.BenefitContext(
+ cohort=cohort,
+ params=inputs.params,
+ baseline=inputs.baseline,
+ config=config,
+ )
+ pia_cache: dict = {}
+ for draw, result in results.items():
+ states_by_wave[wave].append(
+ result.slices[-1].copy().assign(draw=draw)
+ )
+ lookups = legacy.StateLookups(result, config.reference_year)
+ for row_id in config.rows_for_wave(wave):
+ rows, counters = legacy.reference_benefit_rows(
+ result,
+ draw=draw,
+ row=legacy.REGISTERED_ROWS[row_id],
+ context=context,
+ pia_cache=pia_cache,
+ lookups=lookups,
+ )
+ for row in rows:
+ row["age_reference"] = config.reference_year - int(
+ row["birth_year"]
+ )
+ rows_by_row[row_id].extend(rows)
+ counters_by_row[row_id].update(counters)
+ tabulations: dict[str, Any] = {}
+ for row_id in config.rows:
+ row = legacy.REGISTERED_ROWS[row_id]
+ if progress is not None:
+ progress(f"{row_id}: tabulating")
+ tabulation_config = legacy.ColaAgeProfileConfig(
+ reference_year=config.reference_year,
+ components=row.components,
+ benefit_period=row.tabulation_benefit_period,
+ headline_statistic=row.headline_statistic,
+ draw_indices=config.draw_indices,
+ floor_seeds=config.floor_seeds,
+ )
+ population = row.population
+ upstream = {
+ "row": row_id,
+ "field_changed": row.field_changed,
+ "first_reduced_determination_year": (
+ row.first_reduced_determination_year
+ ),
+ "exposure_clock": row.exposure_clock.value,
+ "benefit_period": row.benefit_period.value,
+ "benefit_scale": "annual_12_times_monthly",
+ "rate_path": f"TR2008 {config.tr2008_alternative}",
+ "behavior": "fixed_paths_shared_draws",
+ "population_wave": population.anchor_wave,
+ "population_weight": population.weight_variable,
+ "population_start_year": population.start_year,
+ "population_periods": population.periods(config.reference_year),
+ "family_unit_id": population.family_unit_variable,
+ }
+ try:
+ tabulation = legacy.tabulate_cola_age_profile(
+ pd.DataFrame(rows_by_row[row_id]),
+ data_provenance=data_provenance,
+ config=tabulation_config,
+ registration_pointer=registration_pointer,
+ labels=labels,
+ upstream_conventions=upstream,
+ specification=specification,
+ )
+ except legacy.ColaTabulationError as error:
+ tabulation = None
+ status = f"refused: {type(error).__name__}: {error}"
+ else:
+ undefined = tabulation["undefined_cells"]
+ status = (
+ "tabulated"
+ if not undefined
+ else "tabulated with undefined cells: "
+ + ", ".join(
+ f"{cell['group']} {cell['statistic']} "
+ f"({cell['n_defined_draws']} of "
+ f"{len(config.draw_indices)} draws defined)"
+ for cell in undefined
+ )
+ )
+ tabulations[row_id] = {
+ "status": status,
+ "row": row.as_dict(),
+ "benefit_counters": dict(sorted(counters_by_row[row_id].items())),
+ "reduced_increases_definition": legacy._REDUCED_INCREASE_DEFINITION,
+ "reduced_increases_by_age_group": legacy._reduced_increase_summary(
+ rows_by_row[row_id], row.components
+ ),
+ "component_shares_by_age_group": legacy._component_shares(
+ rows_by_row[row_id], row.components, config.draw_indices
+ ),
+ "tabulation": tabulation,
+ }
+ artifact = {
+ "schema_version": legacy.SCHEMA_VERSION,
+ "data_provenance": data_provenance,
+ "registration_pointer": registration_pointer,
+ "labels": list(labels),
+ "config": config.as_dict(),
+ "specification": recorded_specification,
+ "max_rulings": legacy.max_rulings(config),
+ "builder_defaults": legacy.builder_defaults(config),
+ "rows_not_built": dict(legacy.ROWS_NOT_BUILT),
+ "reduced_rate_minimum_by_row": floor_minima,
+ "parameter_consistency": consistency,
+ "cohorts": {
+ str(wave): {
+ **dict(cohort.diagnostics),
+ "source_provenance": dict(cohort.source_provenance),
+ }
+ for wave, cohort in cohorts.items()
+ },
+ "scheduled_entrants": 0,
+ "population_mortality": legacy._mortality_record(
+ inputs.population_mortality
+ ),
+ # The oracle's statutory parameter revision; the committed-value
+ # checks (parameter_consistency, when run) bind the values.
+ "ssa_parameters_revision": inputs.params.pe_us_revision,
+ "draws": draws,
+ "rows": tabulations,
+ "inputs_provenance": dict(inputs.provenance),
+ }
+ return ProjectionReplay(
+ exercise="cola",
+ result=artifact,
+ benefit_rows={
+ row: pd.DataFrame(values) for row, values in rows_by_row.items()
+ },
+ states={
+ wave: pd.concat(values, ignore_index=True)
+ for wave, values in states_by_wave.items()
+ },
+ )
diff --git a/src/populace_dynamics/group_breakdowns/common.py b/src/populace_dynamics/group_breakdowns/common.py
new file mode 100644
index 00000000..8e093b53
--- /dev/null
+++ b/src/populace_dynamics/group_breakdowns/common.py
@@ -0,0 +1,616 @@
+"""Reproduction gate and G1/G2/G3 composition for exercises 1 and 3.
+
+Sources: the frozen Track A and FRA68 runners, and the category definitions
+and citations in ``estimates.group_breakdown.MINT8_SCHEME``. No PSID reader
+is called here except through G1's cohort loader. No measure or statistic
+is reimplemented: G2 supplies measures and G3 supplies assignments/cells.
+
+Invariants: every frozen result field matches before any attribute loader
+or lifetime calculation runs; joins preserve unique (draw, person_id)
+keys; lifetime categories are fixed on the opening cohort within birth
+decade; the MINT variant contains current-law recipients aged 60+ only.
+Missing income, marriage history or sourced interest rates are unavailable,
+never invented zeroes. The caller writes only after this function returns.
+"""
+
+from __future__ import annotations
+
+import hashlib
+import json
+import math
+from collections.abc import Callable, Mapping
+from dataclasses import dataclass, field, replace
+from typing import Any
+
+import pandas as pd
+
+from populace_dynamics.cohorts import group_attributes as g1
+from populace_dynamics.cola_track_a.config import REGISTERED_ROWS
+from populace_dynamics.estimates import group_breakdown as g3
+from populace_dynamics.estimates import lifetime_measures as g2
+from populace_dynamics.estimates.cola_age_profile import ColaAgeProfileConfig
+
+POST_HOC_LABELS = (
+ "registered, one-shot, post hoc, not blind",
+ "report-only",
+)
+INVENTED_HEADER = "INVENTED DATA - NOT A COMPARISON"
+INCOME_REASON = (
+ "The closed projection has no household income or official poverty "
+ "measure; poverty status and household income quintiles are not computed."
+)
+
+
+class ReproductionMismatch(ValueError):
+ """Frozen parent reproduction failed; no group cells may be computed."""
+
+
+@dataclass(frozen=True)
+class ProjectionReplay:
+ """Frozen result plus in-memory rows and final states, before grouping."""
+
+ exercise: str
+ result: Mapping[str, Any]
+ benefit_rows: Mapping[str, pd.DataFrame]
+ states: Mapping[int, pd.DataFrame]
+ retention_sha256: str = field(init=False)
+
+ def __post_init__(self) -> None:
+ object.__setattr__(self, "retention_sha256", _retention_digest(self))
+
+
+@dataclass(frozen=True)
+class LifetimeOptions:
+ """Sourced G2 inputs; marriage history is resolved only after the gate.
+
+ The loader returns (episodes, history person ids) for an anchor wave.
+ Absent history is explicitly unavailable, rather than never-married.
+ """
+
+ rates: Any = None
+ interest: Any = None
+ marriage_episode_loader: (
+ Callable[[int], tuple[pd.DataFrame, tuple[int, ...]]] | None
+ ) = None
+
+
+def _canonical(value: Any) -> bytes:
+ return json.dumps(
+ value, sort_keys=True, separators=(",", ":"), allow_nan=False
+ ).encode()
+
+
+def _retention_digest(replay: ProjectionReplay) -> str:
+ """Seal retained content, including nested component dicts and keys."""
+ digest = hashlib.sha256(replay.exercise.encode())
+ for name, frames in (
+ ("benefits", replay.benefit_rows),
+ ("states", replay.states),
+ ):
+ for key in sorted(frames, key=str):
+ frame = frames[key]
+ digest.update(_canonical((name, key, list(frame.columns))))
+ digest.update(_canonical([str(dtype) for dtype in frame.dtypes]))
+ digest.update(
+ frame.to_csv(index=False, lineterminator="\n").encode()
+ )
+ return digest.hexdigest()
+
+
+def _first_difference(actual: Any, expected: Any, path: str) -> str | None:
+ if isinstance(actual, Mapping) and isinstance(expected, Mapping):
+ if set(actual) != set(expected):
+ return path + ".keys"
+ for key in sorted(actual):
+ difference = _first_difference(
+ actual[key], expected[key], f"{path}.{key}"
+ )
+ if difference:
+ return difference
+ return None
+ if isinstance(actual, (list, tuple)) and isinstance(
+ expected, (list, tuple)
+ ):
+ if len(actual) != len(expected):
+ return path + ".length"
+ for index, (left, right) in enumerate(
+ zip(actual, expected, strict=True)
+ ):
+ difference = _first_difference(left, right, f"{path}[{index}]")
+ if difference:
+ return difference
+ return None
+ return None if _canonical(actual) == _canonical(expected) else path
+
+
+def verify_reproduction(
+ replay: ProjectionReplay, parent: Mapping[str, Any]
+) -> dict[str, Any]:
+ """Require exact JSON values for every recomputed frozen result field.
+
+ Only the parent's entry-script wrapper (run times, environment, etc.)
+ lies outside the frozen runner's returned result. No floating tolerance
+ is used; differing summation requires a new reviewed convention.
+ Failure identifies a path, without printing any outcome value.
+ """
+
+ if not replay.result.get("rows") or not replay.result.get("draws"):
+ raise ReproductionMismatch("replay lacks frozen rows or draws")
+ absent = set(replay.result) - set(parent)
+ if absent:
+ raise ReproductionMismatch(f"parent lacks fields {sorted(absent)}")
+ expected = {key: parent[key] for key in replay.result}
+ difference = _first_difference(replay.result, expected, "result")
+ if difference:
+ raise ReproductionMismatch(
+ f"frozen reproduction mismatch at {difference}"
+ )
+ row_ids = set(replay.result["rows"])
+ if set(replay.benefit_rows) != row_ids:
+ raise ReproductionMismatch(
+ "retained rows differ from frozen row manifest"
+ )
+ if _retention_digest(replay) != replay.retention_sha256:
+ raise ReproductionMismatch(
+ "retained projection frames changed after replay"
+ )
+ return {
+ "identical": True,
+ "comparison": "exact canonical JSON values; no floating tolerance",
+ "absolute_tolerance": 0.0,
+ "relative_tolerance": 0.0,
+ "verified_result_fields": sorted(replay.result),
+ "verified_rows": sorted(row_ids),
+ "verified_payload_sha256": hashlib.sha256(
+ _canonical(expected)
+ ).hexdigest(),
+ "ordering": "all frozen fields verified before G1/G2/G3 grouping",
+ "retention_sha256": replay.retention_sha256,
+ }
+
+
+def benefit_type(components: Mapping[str, Mapping[str, float]]) -> str | None:
+ """Map baseline components to G3's four MINT benefit types.
+
+ A positive aged/disabled widow component takes the widow row, including
+ own-worker dual entitlement; otherwise a positive spouse excess takes
+ the spousal row. Worker-only means exactly one worker kind. A baseline
+ nonrecipient is unclassified. Concurrent widow/spouse or two worker
+ kinds are refused, since the frozen calculator should not emit them.
+ """
+
+ allowed = {
+ "retired_worker",
+ "disabled_worker",
+ "spouse",
+ "aged_widow",
+ "disabled_widow",
+ }
+ if set(components) - allowed:
+ raise ValueError("unknown benefit component")
+ if any(
+ isinstance(amounts["base"], bool)
+ or not math.isfinite(amounts["base"])
+ or amounts["base"] < 0
+ for amounts in components.values()
+ ):
+ raise ValueError(
+ "baseline components must be finite nonnegative amounts"
+ )
+ active = {
+ name for name, amounts in components.items() if amounts["base"] > 0
+ }
+ widow = bool(active & {"aged_widow", "disabled_widow"})
+ if widow and "spouse" in active:
+ raise ValueError(
+ "concurrent widow and spouse benefits have no mapping"
+ )
+ if {"retired_worker", "disabled_worker"} <= active:
+ raise ValueError("concurrent worker benefit kinds have no mapping")
+ if widow:
+ return "widower"
+ if "spouse" in active:
+ return "spousal"
+ if "retired_worker" in active:
+ return "retired_worker_only"
+ if "disabled_worker" in active:
+ return "disabled_worker_only"
+ return None
+
+
+def _label_codes(series: pd.Series, dimension: g3.Dimension) -> pd.Series:
+ codes = {category.label: category.key for category in dimension.categories}
+ unknown = set(series.dropna()) - set(codes)
+ if unknown:
+ raise ValueError(f"G1 labels outside G3 {dimension.key} scheme")
+ return series.map(codes)
+
+
+def _fixed_lifetime_dimensions() -> tuple[g3.Dimension, ...]:
+ return tuple(
+ replace(
+ dimension,
+ kind=g3.CATEGORICAL,
+ categories=tuple(
+ replace(category, codes=(category.key,), rank=None)
+ for category in dimension.categories
+ ),
+ notes=(*dimension.notes, "G3 categories fixed on opening cohort"),
+ )
+ for dimension in g3.LIFETIME_DIMENSIONS
+ )
+
+
+def _lifetime_side(
+ cohort: Any, params: Any, options: LifetimeOptions
+) -> tuple[pd.DataFrame, dict[str, Any]]:
+ persons = cohort.persons[["person_id", "birth_year", "weight"]].copy()
+ careers = pd.DataFrame(
+ [
+ (pid, year, earnings)
+ for pid, history in cohort.careers.items()
+ for year, earnings in history.items()
+ if year <= cohort.start_year
+ ],
+ columns=["person_id", "year", "earnings"],
+ )
+ rates = options.rates
+ if rates is None:
+ rates = g2.load_oasdi_tax_rates(
+ basis=g2.TaxRateBasis.EMPLOYEE_EMPLOYER_PAID
+ )
+ interest = options.interest
+ if interest is None:
+ interest = g2.load_trust_fund_interest_rates()
+ aime = g2.initial_aime_at_62(
+ careers,
+ persons,
+ params,
+ analysis_year=cohort.start_year,
+ convention=g2.AIME_CONVENTIONS["mint8_initial_aime"],
+ )
+ own = g2.lifetime_payroll_tax_pv_at_62(
+ careers,
+ persons,
+ params,
+ shared=False,
+ rates=rates,
+ interest=interest,
+ missing_rate=g2.MissingRatePolicy.NOT_COMPUTED,
+ )
+ measures = {
+ "initial_aime_quintile": (aime, "aime"),
+ "lifetime_payroll_tax_quintile": (own, "pv_at_62"),
+ }
+ provenance: dict[str, Any] = {}
+ absent_history_ids: set[int] = set()
+ shared_key = "lifetime_payroll_tax_quintile_shared"
+ if options.marriage_episode_loader is None:
+ persons[shared_key] = float("nan")
+ provenance[shared_key] = {
+ "status": "not computed",
+ "reason": "marriage history not supplied",
+ }
+ else:
+ episodes, history_ids = options.marriage_episode_loader(
+ cohort.anchor_wave
+ )
+ shared = g2.lifetime_payroll_tax_pv_at_62(
+ careers,
+ persons,
+ params,
+ shared=True,
+ rates=rates,
+ interest=interest,
+ marriage_episodes=episodes,
+ marriage_history_person_ids=history_ids,
+ missing_spouse=g2.MissingSpousePolicy.NOT_COMPUTED,
+ missing_rate=g2.MissingRatePolicy.NOT_COMPUTED,
+ )
+ # Do not mutate G2's measure frame/provenance. Its explicit history
+ # coverage flag governs which values this adapter may classify.
+ absent_history_ids = set(persons["person_id"]) - set(history_ids)
+ measures[shared_key] = (shared, "pv_at_62")
+ for key, (measure, column) in measures.items():
+ side = measure.frame[["person_id", column]].rename(
+ columns={column: key}
+ )
+ if key == shared_key:
+ side.loc[side["person_id"].isin(absent_history_ids), key] = float(
+ "nan"
+ )
+ persons = persons.merge(side, on="person_id", validate="one_to_one")
+ provenance[key] = {
+ "measure": measure.measure,
+ "provenance": measure.provenance,
+ "status_counts": measure.frame["status"].value_counts().to_dict(),
+ "unavailable_reasons": measure.frame["reason"]
+ .dropna()
+ .value_counts()
+ .to_dict(),
+ }
+ if key == shared_key:
+ provenance[key]["adapter_unavailable_marriage_history"] = len(
+ absent_history_ids
+ )
+ persons["birth_cohort"] = g2.ten_year_birth_cohort(persons["birth_year"])
+ scheme = g3.derive_scheme(
+ g3.MINT8_COHORT_SCHEME,
+ scheme_id="projection_opening_lifetime",
+ title="Opening-cohort lifetime categories",
+ drop=("sex", "race_ethnicity", "country_of_birth", "education"),
+ )
+ assignment = g3.assign_groups(
+ persons,
+ scheme,
+ {d.key: d.key for d in g3.LIFETIME_DIMENSIONS},
+ key_columns=("person_id",),
+ quintile_partition=("birth_cohort",),
+ )
+ for dimension in g3.LIFETIME_DIMENSIONS:
+ codes = assignment.long.loc[
+ assignment.long["dimension"] == dimension.key, "category"
+ ].reset_index(drop=True)
+ persons[dimension.key] = codes.where(codes != g3.UNCLASSIFIED, None)
+ provenance["quintiles"] = assignment.as_dict()
+ provenance["history_window"] = (
+ f"careers through opening year {cohort.start_year}"
+ )
+ return (
+ persons.drop(columns=["birth_year", "weight", "birth_cohort"]),
+ provenance,
+ )
+
+
+def run_group_breakdown(
+ replay: ProjectionReplay,
+ parent: Mapping[str, Any],
+ *,
+ inputs: Any,
+ config: Any,
+ attribute_loader: Callable[..., g1.GroupAttributes] | None = None,
+ lifetime_options: LifetimeOptions | None = None,
+ parent_sha256: str | None = None,
+ registration_pointer: str | None = None,
+) -> dict[str, Any]:
+ """Verify the entire replay, then join side frames and call G3.
+
+ Returns a JSON-safe report with no person rows. The registered entry
+ point handles exclusive writing; this function performs no writes.
+ """
+
+ identity = verify_reproduction(replay, parent)
+ if _first_difference(
+ config.as_dict(), replay.result.get("config"), "config"
+ ):
+ raise ReproductionMismatch(
+ "group configuration differs from verified replay"
+ )
+ cohorts = {
+ c.anchor_wave: c for c in (inputs.cohort, *inputs.additional_cohorts)
+ }
+ if any(c.data_provenance == g3.REGISTERED_REAL for c in cohorts.values()):
+ if not isinstance(
+ registration_pointer, str
+ ) or not g3.REGISTRATION_POINTER.fullmatch(registration_pointer):
+ raise ValueError(
+ "real-data grouping requires a new issue #42 pointer"
+ )
+ if registration_pointer == parent.get("registration_pointer"):
+ raise ValueError(
+ "parent registration cannot authorize post hoc grouping"
+ )
+ loader = (
+ g1.load_group_attributes
+ if attribute_loader is None
+ else attribute_loader
+ )
+ options = (
+ LifetimeOptions() if lifetime_options is None else lifetime_options
+ )
+ scheme = g3.derive_scheme(
+ g3.MINT8_SCHEME,
+ scheme_id="projection_mint8_with_opening_lifetime",
+ title="MINT8 annual rows with opening-cohort lifetime dimensions",
+ append=_fixed_lifetime_dimensions(),
+ notes=(
+ "Lifetime quintiles use G3's weighted percentile rule within "
+ "ten-year birth cohorts of the entire opening cohort, fixed "
+ "across draws, scenarios and population variants.",
+ INCOME_REASON,
+ ),
+ )
+ attributes: dict[int, pd.DataFrame] = {}
+ attribute_provenance = {}
+ lifetime_provenance = {}
+ for wave, cohort in cohorts.items():
+ ids = tuple(int(pid) for pid in cohort.persons["person_id"])
+ loaded = loader(ids, anchor_waves=(wave,))
+ if loaded.frame["person_id"].duplicated().any() or set(
+ loaded.frame["person_id"]
+ ) != set(ids):
+ raise ValueError(
+ "G1 side frame must retain each opening person exactly once"
+ )
+ recorded = loaded.provenance.get("content_sha256")
+ if recorded is not None and recorded != g1.content_sha256(
+ loaded.frame
+ ):
+ raise ValueError("G1 side frame changed after sealing")
+ side = loaded.frame[
+ [
+ "person_id",
+ "race_ethnicity_mint8",
+ "education_years",
+ "country_of_birth_mint8",
+ ]
+ ].copy()
+ side["race_ethnicity"] = _label_codes(
+ side.pop("race_ethnicity_mint8"), g3.RACE_ETHNICITY_DIMENSION
+ )
+ side["country_of_birth"] = _label_codes(
+ side.pop("country_of_birth_mint8"), g3.COUNTRY_OF_BIRTH_DIMENSION
+ )
+ lifetime, provenance = _lifetime_side(cohort, inputs.params, options)
+ attributes[wave] = side.merge(
+ lifetime, on="person_id", validate="one_to_one"
+ )
+ attribute_provenance[str(wave)] = loaded.provenance
+ lifetime_provenance[str(wave)] = provenance
+ if replay.exercise == "cola":
+ row_config = {key: REGISTERED_ROWS[key] for key in config.rows}
+ elif replay.exercise == "fra68":
+ row_config = {key: config.row(key) for key in config.rows}
+ else:
+ raise ValueError("unknown projection exercise")
+ output_rows = {}
+ for row_id, benefits in replay.benefit_rows.items():
+ row = row_config[row_id]
+ wave = row.anchor_wave
+ state = replay.states[wave][
+ ["draw", "person_id", "sex", "marital_status"]
+ ]
+ frame = benefits.merge(
+ state,
+ on=["draw", "person_id"],
+ how="left",
+ validate="one_to_one",
+ indicator=True,
+ )
+ if not frame["_merge"].eq("both").all():
+ raise ValueError("beneficiary absent from final projected state")
+ frame = frame.drop(columns="_merge").merge(
+ attributes[wave],
+ on="person_id",
+ how="left",
+ validate="many_to_one",
+ indicator=True,
+ )
+ if not frame["_merge"].eq("both").all():
+ raise ValueError("beneficiary absent from opening attribute frame")
+ frame = frame.drop(columns="_merge")
+ frame["age"] = config.reference_year - frame["birth_year"]
+ frame["benefit_type"] = frame["benefit_components"].map(benefit_type)
+ frame["poverty_status"] = None
+ frame["household_income_quintile"] = float("nan")
+ columns = {
+ d.key: d.key for d in scheme.dimensions if d.kind != g3.TOTAL
+ }
+ columns["education"] = "education_years"
+ cohort = cohorts[wave]
+ if replay.exercise == "cola":
+ tab_config = ColaAgeProfileConfig(
+ reference_year=config.reference_year,
+ components=row.components,
+ benefit_period=row.tabulation_benefit_period,
+ headline_statistic=row.headline_statistic,
+ draw_indices=config.draw_indices,
+ floor_seeds=config.floor_seeds,
+ )
+ else:
+ from populace_dynamics.fra68_track.runner import _tabulation_config
+
+ tab_config = _tabulation_config(row, config)
+ frozen_row = replay.result["rows"][row_id]
+ frozen_tabulation = frozen_row.get("tabulation") or {}
+ labels = (
+ tuple(
+ frozen_row.get(
+ "labels", frozen_tabulation.get("labels", cohort.labels)
+ )
+ )
+ + POST_HOC_LABELS
+ )
+ variants = {
+ "test_population": frame,
+ "mint_population": frame.loc[
+ (frame["age"] >= 60)
+ & frame["beneficiary_base"]
+ & (frame["benefit_base"] > 0)
+ ].copy(),
+ }
+ output_variants = {}
+ for name, selected in variants.items():
+ if set(selected["draw"]) != set(tab_config.draw_indices):
+ output_variants[name] = {
+ "status": "not computed",
+ "reason": "population absent in at least one registered draw",
+ "labels": list(labels),
+ }
+ continue
+ assignment = g3.assign_groups(
+ selected,
+ scheme,
+ columns,
+ key_columns=("draw", "person_id"),
+ unclassified_codes={
+ "marital_status": (
+ "unknown",
+ "no_marriage_history",
+ "separated",
+ )
+ },
+ )
+ result = g3.tabulate_projection_breakdown(
+ selected,
+ assignment,
+ config=tab_config,
+ data_provenance=cohort.data_provenance,
+ registration_pointer=registration_pointer,
+ labels=labels,
+ post_hoc_labels=POST_HOC_LABELS,
+ statistic_id=frozen_tabulation.get(
+ "statistic_id", g3.PROJECTION_STATISTIC_ID
+ ),
+ upstream={
+ "parent_sha256": parent_sha256,
+ "row": row_id,
+ "verified_payload_sha256": identity[
+ "verified_payload_sha256"
+ ],
+ "retention_sha256": replay.retention_sha256,
+ "reproduction_identical": True,
+ "absolute_tolerance": 0.0,
+ "relative_tolerance": 0.0,
+ },
+ ).as_dict()
+ result["population_variant"] = name
+ result["not_computed"] = {
+ "poverty_status": INCOME_REASON,
+ "household_income_quintile": INCOME_REASON,
+ "household_income_statistics": INCOME_REASON,
+ "official_poverty_statistics": INCOME_REASON,
+ }
+ output_variants[name] = result
+ output_rows[row_id] = {
+ "anchor_wave": wave,
+ "row": frozen_row["row"],
+ "labels": list(labels),
+ "variants": output_variants,
+ }
+ invented = all(c.data_provenance == g3.INVENTED for c in cohorts.values())
+ return {
+ "header": (
+ INVENTED_HEADER
+ if invented
+ else "Registered projection group breakdown"
+ ),
+ "exercise": replay.exercise,
+ "labels": ([INVENTED_HEADER] if invented else [])
+ + list(inputs.cohort.labels)
+ + list(POST_HOC_LABELS),
+ "report_only": True,
+ "parent_sha256": parent_sha256,
+ "reproduction": identity,
+ "attribute_provenance": attribute_provenance,
+ "lifetime_provenance": lifetime_provenance,
+ "conventions": {
+ "benefit_type": "baseline components; widow then spouse (dual entitlement included), otherwise worker only",
+ "marital_status": "2030 projected state; opening divorced/never married carried forward; inherited opening convention counts separated as married; literal separated and unknown codes unclassified",
+ "mint_population": "current-law beneficiaries aged 60 or older in 2030",
+ "test_population": "alive in 2030 with a positive benefit in either scenario; original scenario membership retained",
+ "artifact_choice": "one artifact per exercise, preserving separate parent identities and one-shot registrations",
+ "lifetime": "opening-cohort careers only; no projected earnings or unsourced interest rates",
+ },
+ "rows": output_rows,
+ }
diff --git a/src/populace_dynamics/group_breakdowns/fra68.py b/src/populace_dynamics/group_breakdowns/fra68.py
new file mode 100644
index 00000000..fea95bf5
--- /dev/null
+++ b/src/populace_dynamics/group_breakdowns/fra68.py
@@ -0,0 +1,392 @@
+"""Frozen exercise-3 replay for registered post hoc group breakdowns.
+
+The orchestration follows ``fra68_track/runner.py:841-1175``. Every
+projection, benefit, diagnostic and tabulation calculation is composed
+from the frozen modules. There is no new implementation of any measure.
+The retained state and benefit rows are inputs to G1/G2/G3 only after the
+common module's reproduction check succeeds.
+"""
+
+from __future__ import annotations
+
+from collections import Counter
+from collections.abc import Callable, Mapping
+from pathlib import Path
+from typing import Any
+
+import pandas as pd
+
+from populace_dynamics.cola_track_a import benefits as track_benefits
+from populace_dynamics.cola_track_a import runner as track_runner
+from populace_dynamics.cola_track_a.adapters import claiming_schedule
+from populace_dynamics.cola_track_a.benefits import (
+ BenefitContext,
+ StateLookups,
+)
+from populace_dynamics.cola_track_a.config import (
+ builder_defaults as track_a_builder_defaults,
+)
+from populace_dynamics.cola_track_a.config import (
+ max_rulings as track_a_max_rulings,
+)
+from populace_dynamics.cola_track_a.runner import TrackAInputs
+from populace_dynamics.engine.di_entitlement import (
+ fra_schedule_from_parameters,
+)
+from populace_dynamics.estimates.cola_age_profile import (
+ REGISTERED_REAL,
+ ColaTabulationError,
+ tabulate_cola_age_profile,
+)
+from populace_dynamics.fra68_track import runner as legacy
+from populace_dynamics.fra68_track.benefits import (
+ PersonScenario,
+ union_benefit_rows,
+)
+from populace_dynamics.fra68_track.config import (
+ E1_RULINGS,
+ SPECIFICATION_ID,
+ STATISTIC_ID,
+ TRACK_A_ROW_BY_WAVE,
+ FRA68Config,
+ builder_defaults,
+ max_rulings,
+ row_labels,
+)
+from populace_dynamics.fra68_track.reform import SCHEDULES
+from populace_dynamics.group_breakdowns.common import ProjectionReplay
+
+__all__ = ["reproduce_fra68"]
+
+
+def reproduce_fra68(
+ inputs: TrackAInputs,
+ *,
+ config: FRA68Config | None = None,
+ registration_pointer: str | None = None,
+ progress: Callable[[str], None] | None = None,
+ check_committed_parameters: bool = False,
+ specification: Mapping[str, Any] | None = None,
+ exercise_1_artifact: Path = legacy.EXERCISE_1_ARTIFACT_PATH,
+) -> ProjectionReplay:
+ """Replay frozen F0-F8 cells, retaining rows before attribute loading.
+
+ Reuses the registered projection, scenario calculator, union builder,
+ diagnostics and A7 tabulator without changing their implementations.
+ ``result`` reproduces :func:`fra68_track.runner.run_fra68`; callers
+ must verify every committed cell before deriving group attributes or
+ calling the G3 group tabulator. This function neither loads attributes
+ nor writes artifacts.
+ """
+
+ config = config or FRA68Config()
+ config.check_runnable()
+ track_config = config.track_a_config()
+ cohorts = track_runner._population_cohorts(inputs, track_config)
+ data_provenance = next(iter(cohorts.values())).data_provenance
+ labels = tuple(next(iter(cohorts.values())).labels)
+ real = data_provenance == REGISTERED_REAL
+ if real and not registration_pointer:
+ raise ValueError(
+ "no real-data statistic before the issue #42 registration "
+ "comment exists; pass its pointer to run a registered_real cohort"
+ )
+ block = (
+ legacy.e1_parameter_block() if specification is None else specification
+ )
+ if real:
+ legacy.check_specification_for_registered_run(block, config)
+ spec_check = legacy.specification_code_check(block, config)
+ consistency = track_runner._parameter_consistency(
+ inputs,
+ track_config,
+ cohorts,
+ compare_committed_values=real or check_committed_parameters,
+ )
+ if real and not consistency["consistent"]:
+ failed = sorted(
+ name
+ for name, check in consistency["checks"].items()
+ if not check["consistent"]
+ )
+ raise ValueError(
+ "the inputs differ from the parameters the configuration "
+ f"reports for this run: {failed}; a registered run refuses "
+ "a mismatch"
+ )
+ track_runner._check_source_provenance(cohorts, data_provenance)
+ track_runner._check_output_labels(labels, data_provenance)
+ baseline_params = inputs.params
+ base_scenario, reform_by_schedule, schedule_record = legacy._scenarios(
+ config, baseline_params
+ )
+ birth_month = int(config.di_spec.assumed_birth_month)
+ schedule = claiming_schedule(
+ inputs.claiming_pmf, max_table_year=config.claim_table_max_year
+ )
+ fra_schedule = fra_schedule_from_parameters(baseline_params)
+ rows_by_row: dict[str, list[dict]] = {row: [] for row in config.rows}
+ counters_by_row: dict[str, Counter] = {
+ row: Counter() for row in config.rows
+ }
+ states: dict[int, list[pd.DataFrame]] = {}
+ draws: dict[str, Any] = {}
+ di_window: dict[str, Any] = {}
+ for wave, cohort in cohorts.items():
+ results, draws[str(wave)] = track_runner._project_population(
+ cohort,
+ inputs,
+ track_config,
+ schedule=schedule,
+ fra_schedule=fra_schedule,
+ progress=progress,
+ )
+ states[wave] = []
+ track_row = legacy.TRACK_A_ROWS[TRACK_A_ROW_BY_WAVE[wave]]
+ wave_rows = config.rows_for_wave(wave)
+ reform_keys = {
+ config.row(row_id).reform_key: config.row(row_id)
+ for row_id in wave_rows
+ }
+ pia_cache: dict = {}
+ di_window[str(wave)] = {sid: {} for sid in config.schedule_ids}
+ for draw, result in results.items():
+ if progress is not None:
+ progress(f"anchor wave {wave}, draw {draw}: benefits")
+ lookups = StateLookups(result, config.reference_year)
+ final = lookups.final.reset_index(drop=True).copy(deep=True)
+ final["draw"] = int(draw)
+ states[wave].append(final)
+ shared = {
+ "cohort": cohort,
+ "inputs": inputs,
+ "track_config": track_config,
+ "track_row": track_row,
+ "result": result,
+ "lookups": lookups,
+ "pia_cache": pia_cache,
+ "assumed_birth_month": birth_month,
+ }
+ base_people, base_counters = legacy._compute_scenario(
+ base_scenario, **shared
+ )
+ reform_people: dict[tuple, dict[int, PersonScenario]] = {}
+ reform_counters: dict[tuple, Counter] = {}
+ for key, sample_row in reform_keys.items():
+ reform_people[key], reform_counters[key] = (
+ legacy._compute_scenario(
+ legacy._reform_scenario(
+ sample_row,
+ config,
+ baseline_params,
+ reform_by_schedule,
+ ),
+ **shared,
+ )
+ )
+ context = BenefitContext(
+ cohort=cohort,
+ params=baseline_params,
+ baseline=inputs.baseline,
+ config=track_config,
+ )
+ for row_id in wave_rows:
+ row = config.row(row_id)
+ rows, counters = union_benefit_rows(
+ base_people,
+ reform_people[row.reform_key],
+ draw=draw,
+ context=context,
+ lookups=lookups,
+ baseline_params=baseline_params,
+ reform_params=reform_by_schedule[row.schedule_id],
+ )
+ for item in rows:
+ item["age_reference"] = config.reference_year - int(
+ item["birth_year"]
+ )
+ rows_by_row[row_id].extend(rows)
+ counters_by_row[row_id].update(counters)
+ counters_by_row[row_id].update(
+ {f"baseline_{k}": v for k, v in base_counters.items()}
+ )
+ counters_by_row[row_id].update(
+ {
+ f"reform_{k}": v
+ for k, v in reform_counters[row.reform_key].items()
+ }
+ )
+ for sid in config.schedule_ids:
+ di_window[str(wave)][sid][str(draw)] = legacy._di_window(
+ result,
+ di_rates=inputs.di_rates,
+ baseline=baseline_params,
+ reform=reform_by_schedule[sid],
+ assumed_birth_month=birth_month,
+ )
+ # Membership first (E1 sections 7 and 12): every row's differences are
+ # sorted from A7's own recipient flags, and a C0 difference no named
+ # mechanism explains refuses the run before any row is tabulated.
+ membership = {
+ row_id: legacy._membership_differences(
+ rows_by_row[row_id],
+ config.row(row_id),
+ config,
+ baseline=baseline_params,
+ reform=reform_by_schedule[config.row(row_id).schedule_id],
+ assumed_birth_month=birth_month,
+ )
+ for row_id in config.rows
+ }
+ legacy._refuse_unexplained_c0_differences(membership)
+ tabulations: dict[str, Any] = {}
+ for row_id in config.rows:
+ row = config.row(row_id)
+ if progress is not None:
+ progress(f"{row_id}: tabulating")
+ output_labels = row_labels(row, labels)
+ tabulation_config = legacy._tabulation_config(row, config)
+ population = row.population
+ upstream = {
+ "specification": SPECIFICATION_ID,
+ "row": row_id,
+ "field_changed": row.field_changed,
+ "schedule": row.schedule_id,
+ "schedule_sha256": SCHEDULES[row.schedule_id].sha256(),
+ "survivor_retirement_age": row.survivor_retirement_age.value,
+ "claiming_response": row.claiming_response.value,
+ "benefit_period": "calendar_2030_payments",
+ "benefit_scale": "annual_12_times_monthly",
+ "rate_path": (
+ f"TR2008 {config.tr2008_alternative}, both scenarios"
+ ),
+ "behavior": row.behavior,
+ "population_wave": population.anchor_wave,
+ "population_weight": population.weight_variable,
+ "population_start_year": population.start_year,
+ "population_periods": population.periods(config.reference_year),
+ "family_unit_id": population.family_unit_variable,
+ }
+ try:
+ tabulation = tabulate_cola_age_profile(
+ pd.DataFrame(rows_by_row[row_id]),
+ data_provenance=data_provenance,
+ config=tabulation_config,
+ registration_pointer=registration_pointer,
+ labels=output_labels,
+ upstream_conventions=upstream,
+ statistic_id=STATISTIC_ID,
+ specification=block,
+ pending_rulings=E1_RULINGS,
+ )
+ except ColaTabulationError as error:
+ tabulation = None
+ status = f"refused: {type(error).__name__}: {error}"
+ else:
+ # A7's own count must equal the classification's, which read
+ # A7's flags: a difference would mean the two disagree.
+ summary = tabulation["input_summary"]
+ record = membership[row_id]["record"]
+ if not record.get("classified") or (
+ summary["n_rows_membership_differs"] != record["n_rows_differ"]
+ ):
+ raise ValueError(
+ f"{row_id}: A7 counts "
+ f"{summary['n_rows_membership_differs']} rows whose "
+ "membership differs, but the membership classification "
+ f"recorded {record.get('n_rows_differ')}"
+ )
+ undefined = tabulation["undefined_cells"]
+ status = (
+ "tabulated"
+ if not undefined
+ else "tabulated with undefined cells: "
+ + ", ".join(
+ f"{cell['group']} {cell['statistic']} "
+ f"({cell['n_defined_draws']} of "
+ f"{len(config.draw_indices)} draws defined)"
+ for cell in undefined
+ )
+ )
+ tabulations[row_id] = {
+ "status": status,
+ "row": row.as_dict(),
+ "labels": list(output_labels),
+ "membership_differences": membership[row_id]["record"],
+ "benefit_counters": dict(sorted(counters_by_row[row_id].items())),
+ "diagnostics": legacy._row_diagnostics(
+ rows_by_row[row_id], row, config.draw_indices
+ ),
+ "component_shares_by_age_group": track_runner._component_shares(
+ rows_by_row[row_id], row.components, config.draw_indices
+ ),
+ "tabulation": tabulation,
+ }
+ result = {
+ "schema_version": legacy.SCHEMA_VERSION,
+ "specification": SPECIFICATION_ID,
+ "data_provenance": data_provenance,
+ "registration_pointer": registration_pointer,
+ "labels": list(labels),
+ "config": config.as_dict(),
+ "max_rulings": max_rulings(config),
+ "builder_defaults": builder_defaults(config),
+ "track_a_conventions": {
+ "note": (
+ "the shared projection and the benefit-level and auxiliary "
+ "rules are Track A's; Max ruled these for exercise 1 (d074, "
+ "d075) and ruled on 2026-09-24 that exercise 3 runs exactly "
+ "like Track A (d188 item (a); d196 items (2) and (3) carry "
+ "over the COLA horizon and the opening-stock basis)"
+ ),
+ "config": track_config.as_dict(),
+ "benefit_computation_years": (
+ track_benefits.TRACK_A_COMPUTATION_YEARS.value
+ ),
+ "max_rulings_exercise_1": track_a_max_rulings(track_config),
+ "builder_defaults": track_a_builder_defaults(track_config),
+ },
+ "specification_check": spec_check,
+ "age_factor_parameters": schedule_record,
+ "cola_paths": legacy._cola_record(inputs.baseline),
+ "parameter_consistency": consistency,
+ "cohorts": {
+ str(wave): {
+ **dict(cohort.diagnostics),
+ "source_provenance": dict(cohort.source_provenance),
+ }
+ for wave, cohort in cohorts.items()
+ },
+ "scheduled_entrants": 0,
+ "population_mortality": track_runner._mortality_record(
+ inputs.population_mortality
+ ),
+ "ssa_parameters_revision": baseline_params.pe_us_revision,
+ "draws": draws,
+ "projection_identity_with_exercise_1": legacy.projection_identity_record(
+ draws,
+ data_provenance=data_provenance,
+ artifact_path=exercise_1_artifact,
+ ),
+ "di_window_diagnostic": {
+ "definition": legacy._DI_WINDOW_DEFINITION,
+ "incidence_at_start_ages": legacy._incidence_at(
+ inputs.di_rates, (65, 66, 67)
+ ),
+ "by_wave_schedule_draw_year": di_window,
+ },
+ "rows": tabulations,
+ "inputs_provenance": dict(inputs.provenance),
+ }
+
+ return ProjectionReplay(
+ exercise="fra68",
+ result=result,
+ benefit_rows={
+ row_id: pd.DataFrame(rows_by_row[row_id]) for row_id in config.rows
+ },
+ states={
+ wave: pd.concat(frames, ignore_index=True)
+ for wave, frames in states.items()
+ },
+ )
diff --git a/tests/group_breakdowns/test_cola.py b/tests/group_breakdowns/test_cola.py
new file mode 100644
index 00000000..b55aa0c5
--- /dev/null
+++ b/tests/group_breakdowns/test_cola.py
@@ -0,0 +1,179 @@
+"""Exercise-1 replay invariants on INVENTED data, with no PSID reads.
+
+Every person, career, benefit, weight, rate and probability comes from the
+existing invented Track A fixture. Differential checks compare only two
+implementations run on those invented inputs, never comparator values.
+"""
+
+from __future__ import annotations
+
+import json
+from collections import Counter
+
+import pandas as pd
+import pytest
+from hypothesis import given, settings
+from hypothesis import strategies as st
+
+from populace_dynamics.cohorts import psid2010
+from populace_dynamics.cola_track_a import invented, prepare_track_a_cohort
+from populace_dynamics.cola_track_a import runner as legacy
+from populace_dynamics.cola_track_a.config import TrackAConfig
+from populace_dynamics.group_breakdowns.cola import reproduce_cola
+from tests.cola_track_a.test_assembly import _inputs
+
+
+def _invented_inputs():
+ """The inherited INVENTED fixture, including both registered waves."""
+ cohort = prepare_track_a_cohort(
+ psid2010.build_psid2010_cohort(
+ invented.invented_psid2010_inputs(seed=7)
+ ),
+ data_provenance="invented",
+ config=TrackAConfig(draw_indices=(0, 1)),
+ )
+ return _inputs(cohort)
+
+
+def _normal(result):
+ return json.loads(json.dumps(result, default=str, allow_nan=False))
+
+
+@pytest.fixture(scope="module")
+def pair():
+ """All seven registered rows, both waves, and two INVENTED draws."""
+ inputs = _invented_inputs()
+ config = TrackAConfig(draw_indices=(0, 1))
+ project = legacy._project_population
+ benefit_rows = legacy.reference_benefit_rows
+ populations = []
+ paths = {}
+ retained_slices = []
+ initial = {
+ wave: cohort.initial_slice.copy(deep=True)
+ for wave, cohort in inputs.cohorts_by_wave().items()
+ }
+
+ def collect_population(*args, **kwargs):
+ output = project(*args, **kwargs)
+ populations.append(args[0].anchor_wave)
+ retained_slices.extend(
+ result.slices[-1] for result in output[0].values()
+ )
+ return output
+
+ def collect_rows(result, **kwargs):
+ key = (kwargs["context"].cohort.anchor_wave, kwargs["draw"])
+ paths.setdefault(key, []).append(id(result))
+ return benefit_rows(result, **kwargs)
+
+ with pytest.MonkeyPatch.context() as patch:
+ patch.setattr(legacy, "_project_population", collect_population)
+ patch.setattr(legacy, "reference_benefit_rows", collect_rows)
+ replay = reproduce_cola(inputs, config=config)
+ reference = legacy.run_track_a(inputs, config=config)
+ return (
+ replay,
+ reference,
+ config,
+ inputs,
+ {
+ "populations": populations,
+ "paths": paths,
+ "retained_slices": retained_slices,
+ "initial": initial,
+ },
+ )
+
+
+def test_complete_artifact_matches_frozen_runner(pair):
+ replay, reference, *_ = pair
+ # Entire payload: five age cells, per-draw diagnostics, counters,
+ # component shares, reduced-increase summaries, floors and provenance.
+ assert _normal(replay.result) == _normal(reference)
+
+
+def test_one_population_projection_shared_by_registered_rows(pair):
+ _, _, config, _, observed = pair
+ assert Counter(observed["populations"]) == {2011: 1, 2009: 1}
+ for (wave, _), paths in observed["paths"].items():
+ assert len(paths) == len(config.rows_for_wave(wave))
+ assert len(set(paths)) == 1
+
+
+def test_every_benefit_row_has_unique_reference_year_state(pair):
+ replay, _, config, _, _ = pair
+ assert replay.exercise == "cola"
+ assert tuple(replay.benefit_rows) == config.rows
+ assert set(replay.states) == {2011, 2009}
+ for wave, frame in replay.states.items():
+ assert not frame.duplicated(["draw", "person_id"]).any()
+ assert set(frame["draw"]) == set(config.draw_indices)
+ assert set(frame["year"]) == {config.reference_year}
+ assert set(frame["sex"]) <= {"female", "male"}
+ assert "marital_status" in frame
+ for row_id in config.rows_for_wave(wave):
+ rows = replay.benefit_rows[row_id]
+ assert not rows.duplicated(["draw", "person_id"]).any()
+ joined = rows.merge(
+ frame[["draw", "person_id"]],
+ on=["draw", "person_id"],
+ how="left",
+ indicator=True,
+ validate="one_to_one",
+ )
+ assert joined["_merge"].eq("both").all()
+ assert rows["benefit_base"].gt(0).all()
+ assert rows["benefit_reform"].gt(0).all()
+ assert rows["beneficiary_base"].all()
+ assert rows["beneficiary_reform"].all()
+ assert rows["age_reference"].equals(
+ config.reference_year - rows["birth_year"]
+ )
+
+
+def test_components_preserve_scenario_amounts(pair):
+ replay, *_ = pair
+ for rows in replay.benefit_rows.values():
+ for row in rows.itertuples(index=False):
+ for scenario in ("base", "reform"):
+ amount = sum(
+ component[scenario]
+ for component in row.benefit_components.values()
+ )
+ assert amount == pytest.approx(
+ getattr(row, f"benefit_{scenario}"), rel=1e-12
+ )
+
+
+def test_replay_retention_preserves_projection_and_opening_frames(pair):
+ replay, _, _, inputs, observed = pair
+ for wave, cohort in inputs.cohorts_by_wave().items():
+ pd.testing.assert_frame_equal(
+ cohort.initial_slice, observed["initial"][wave]
+ )
+ for frame in observed["retained_slices"]:
+ assert "draw" not in frame
+ frame = observed["retained_slices"][0]
+ before = frame.copy(deep=True)
+ saved = replay.states[2011]["sex"].copy()
+ try:
+ replay.states[2011]["sex"] = "INVENTED mutation"
+ pd.testing.assert_frame_equal(frame, before)
+ finally:
+ replay.states[2011]["sex"] = saved
+
+
+@settings(max_examples=3, deadline=None)
+@given(
+ row_id=st.sampled_from(("R0", "R1", "R2", "R3", "R4", "R5", "R6")),
+ draw=st.integers(min_value=0, max_value=5),
+)
+def test_replay_matches_frozen_semantics_for_draw_and_row(row_id, draw):
+ """Row selection and draw selection leave the frozen semantics intact."""
+ config = TrackAConfig(rows=(row_id,), draw_indices=(draw,))
+ inputs = _invented_inputs()
+ replay = reproduce_cola(inputs, config=config)
+ reference = legacy.run_track_a(inputs, config=config)
+ assert _normal(replay.result) == _normal(reference)
+ assert set(replay.benefit_rows[row_id]["draw"]) == {draw}
diff --git a/tests/group_breakdowns/test_common.py b/tests/group_breakdowns/test_common.py
new file mode 100644
index 00000000..7d430d61
--- /dev/null
+++ b/tests/group_breakdowns/test_common.py
@@ -0,0 +1,402 @@
+"""Reproduction ordering, joins and population invariants on INVENTED data.
+
+Synthetic parent cells here are INVENTED, never committed model outcomes.
+The side loader uses G1's builder and the lifetime inputs are invented G2
+rates. Property tests cover exact gating and fixed-category invariance.
+"""
+
+from __future__ import annotations
+
+import copy
+import importlib.util
+import json
+from dataclasses import replace
+from pathlib import Path
+from types import SimpleNamespace
+
+import pandas as pd
+import pytest
+from hypothesis import given, settings
+from hypothesis import strategies as st
+
+from populace_dynamics.cola_track_a.config import REGISTERED_ROWS, TrackAConfig
+from populace_dynamics.estimates import group_breakdown as g3
+from populace_dynamics.group_breakdowns import common
+from populace_dynamics.track_a_v2.invented import invented_parameters
+
+_SCRIPT = (
+ Path(__file__).resolve().parents[2]
+ / "scripts"
+ / "projection_groups_dry_run.py"
+)
+_SPEC = importlib.util.spec_from_file_location("_invented_groups", _SCRIPT)
+dry = importlib.util.module_from_spec(_SPEC)
+_SPEC.loader.exec_module(dry)
+
+
+def _case():
+ """Four INVENTED people, two draws; one person is under age60."""
+ persons = pd.DataFrame(
+ {
+ "person_id": [1, 2, 3, 4],
+ "birth_year": [1940, 1950, 1965, 1980],
+ "weight": [1.0, 2.0, 3.0, 4.0],
+ }
+ )
+ cohort = SimpleNamespace(
+ persons=persons,
+ careers={
+ pid: {year: 10000.0 * pid for year in range(2000, 2011)}
+ for pid in (1, 2, 3, 4)
+ },
+ start_year=2010,
+ anchor_wave=2011,
+ data_provenance="invented",
+ labels=(
+ common.INVENTED_HEADER,
+ "PSID-seeded closed cohort",
+ "Python oracle (not Axiom)",
+ "fixed-path mechanical incidence",
+ ),
+ )
+ config = TrackAConfig(rows=("R0",), draw_indices=(0, 1))
+ rows = []
+ states = []
+ for draw in config.draw_indices:
+ for pid, birth, weight in persons.itertuples(index=False, name=None):
+ amount = 10000.0 * pid
+ kind = "disabled_worker" if pid == 4 else "retired_worker"
+ rows.append(
+ {
+ "draw": draw,
+ "person_id": pid,
+ "family_unit_id": pid,
+ "birth_year": birth,
+ "age_reference": 2030 - birth,
+ "weight": weight,
+ "beneficiary_base": True,
+ "beneficiary_reform": True,
+ "benefit_base": amount,
+ "benefit_reform": amount * 0.9,
+ "benefit_components": {
+ kind: {"base": amount, "reform": amount * 0.9}
+ },
+ }
+ )
+ states.append(
+ {
+ "draw": draw,
+ "person_id": pid,
+ "sex": "female" if pid % 2 else "male",
+ "marital_status": (
+ "widowed",
+ "divorced",
+ "never_married",
+ "married",
+ )[pid - 1],
+ }
+ )
+ result = {
+ "labels": list(cohort.labels),
+ "config": config.as_dict(),
+ "rows": {
+ "R0": {
+ "row": REGISTERED_ROWS["R0"].as_dict(),
+ "tabulation": {"groups": [{"cell": -10.0}]},
+ }
+ },
+ "draws": {"2011": {"0": {"n": 4}, "1": {"n": 4}}},
+ }
+ replay = common.ProjectionReplay(
+ "cola",
+ result,
+ {"R0": pd.DataFrame(rows)},
+ {2011: pd.DataFrame(states)},
+ )
+ inputs = SimpleNamespace(
+ cohort=cohort, additional_cohorts=(), params=invented_parameters()
+ )
+ return replay, copy.deepcopy(result), inputs, config
+
+
+def _options():
+ # No marriage history is supplied in this fixture; this is unavailable,
+ # rather than an assumption that every INVENTED person is unmarried.
+ options = dry.invented_lifetime_options({})
+ return replace(options, marriage_episode_loader=None)
+
+
+def _run(replay, parent, inputs, config, **kwargs):
+ return common.run_group_breakdown(
+ replay,
+ parent,
+ inputs=inputs,
+ config=config,
+ attribute_loader=dry.invented_attribute_loader,
+ lifetime_options=_options(),
+ **kwargs,
+ )
+
+
+@pytest.fixture(scope="module")
+def report():
+ return _run(*_case())
+
+
+@pytest.mark.parametrize("field", ("cell", "draw", "labels", "row", "missing"))
+def test_every_mismatch_refuses_before_any_loader_and_writes_nothing(
+ tmp_path, monkeypatch, field
+):
+ replay, parent, inputs, config = _case()
+ if field == "cell":
+ parent["rows"]["R0"]["tabulation"]["groups"][0]["cell"] = -9.0
+ elif field == "draw":
+ parent["draws"]["2011"]["1"]["n"] = 5
+ elif field == "labels":
+ parent["labels"].append("INVENTED mutation")
+ elif field == "row":
+ parent["rows"]["R0"]["row"]["anchor_wave"] = 2009
+ else:
+ del parent["draws"]
+
+ def forbidden(*args, **kwargs):
+ pytest.fail(
+ "group attribute/measure loader called before reproduction"
+ )
+
+ monkeypatch.setattr(common.g1, "load_group_attributes", forbidden)
+ monkeypatch.setattr(common.g2, "initial_aime_at_62", forbidden)
+ monkeypatch.setattr(common.g3, "assign_groups", forbidden)
+ with pytest.raises(common.ReproductionMismatch):
+ common.run_group_breakdown(
+ replay, parent, inputs=inputs, config=config
+ )
+ assert not list(tmp_path.iterdir())
+
+
+@settings(max_examples=25, deadline=None)
+@given(
+ st.floats(
+ min_value=-1e6, max_value=1e6, allow_nan=False, allow_infinity=False
+ )
+)
+def test_exact_float_gate_has_no_tolerance(value):
+ replay, parent, _, _ = _case()
+ replay.result["rows"]["R0"]["tabulation"]["groups"][0]["cell"] = value
+ parent["rows"]["R0"]["tabulation"]["groups"][0]["cell"] = value
+ assert common.verify_reproduction(replay, parent)["identical"]
+ parent["rows"]["R0"]["tabulation"]["groups"][0]["cell"] = value + 1.0
+ with pytest.raises(common.ReproductionMismatch, match="cell"):
+ common.verify_reproduction(replay, parent)
+
+
+def test_outputs_use_g3_labels_statistics_and_population_variants(report):
+ variants = report["rows"]["R0"]["variants"]
+ assert variants["test_population"]["input_summary"]["n_rows"] == 8
+ assert variants["mint_population"]["input_summary"]["n_rows"] == 6
+ for variant in variants.values():
+ assert set(g3.MINT_BENEFIT_STATISTICS) <= set(variant["statistics"])
+ assert g3.RATIO_OF_SCENARIO_MEANS in variant["statistics"]
+ dimensions = {d["key"]: d for d in variant["dimensions"]}
+ for dimension in g3.MINT8_SCHEME.dimensions:
+ assert [
+ cell["label"] for cell in dimensions[dimension.key]["cells"]
+ ] == list(dimension.labels)
+ assert common.POST_HOC_LABELS[0] in variant["labels"]
+ assert "report-only" in variant["labels"]
+ assert "poverty_status" in variant["not_computed"]
+ assert (
+ "no household income" in variant["not_computed"]["poverty_status"]
+ )
+ for dimension in variant["dimensions"]:
+ assert dimension["ssa_subgroup_suppressed"]
+ json.dumps(report, allow_nan=False)
+
+
+def test_projected_marital_states_and_lifetime_unavailability(report):
+ assignment = report["rows"]["R0"]["variants"]["test_population"][
+ "assignment"
+ ]
+ dimensions = {d["key"]: d for d in assignment["dimensions"]}
+ assert dimensions["marital_status"]["n_by_category"] == {
+ "married": 2,
+ "divorced": 2,
+ "widowed": 2,
+ "never_married": 2,
+ }
+ assert (
+ dimensions["lifetime_payroll_tax_quintile_shared"]["n_unclassified"]
+ == 8
+ )
+ assert (
+ report["lifetime_provenance"]["2011"][
+ "lifetime_payroll_tax_quintile_shared"
+ ]["reason"]
+ == "marriage history not supplied"
+ )
+ assert report["reproduction"]["absolute_tolerance"] == 0
+
+
+def test_g3_total_differential_after_adapter_join(report):
+ replay, _, inputs, config = _case()
+ total_scheme = g3.derive_scheme(
+ g3.MINT8_SCHEME,
+ scheme_id="invented_total",
+ title="INVENTED total",
+ drop=tuple(
+ d.key for d in g3.MINT8_SCHEME.dimensions if d.key != g3.TOTAL_KEY
+ ),
+ )
+ rows = replay.benefit_rows["R0"]
+ assignment = g3.assign_groups(
+ rows, total_scheme, {}, key_columns=("draw", "person_id")
+ )
+ direct = g3.tabulate_projection_breakdown(
+ rows,
+ assignment,
+ data_provenance="invented",
+ config=common.ColaAgeProfileConfig(draw_indices=config.draw_indices),
+ labels=inputs.cohort.labels,
+ post_hoc_labels=common.POST_HOC_LABELS,
+ ).as_dict()
+ actual = report["rows"]["R0"]["variants"]["test_population"]["dimensions"][
+ 0
+ ]
+ assert actual == direct["dimensions"][0]
+
+
+@settings(max_examples=12, deadline=None)
+@given(st.integers(min_value=1, max_value=100), st.permutations((0, 1, 2, 3)))
+def test_fixed_lifetime_categories_preserve_weight_scaling_and_person_order(
+ scale, order
+):
+ _, _, inputs, _ = _case()
+ cohort = inputs.cohort
+ original, _ = common._lifetime_side(cohort, inputs.params, _options())
+ changed = copy.copy(cohort)
+ changed.persons = cohort.persons.iloc[list(order)].copy()
+ changed.persons["weight"] *= scale
+ actual, _ = common._lifetime_side(changed, inputs.params, _options())
+ pd.testing.assert_frame_equal(
+ original.sort_values("person_id").reset_index(drop=True),
+ actual.sort_values("person_id").reset_index(drop=True),
+ )
+
+
+@pytest.mark.parametrize(
+ "mapping,expected",
+ (
+ ({"retired_worker": 10}, "retired_worker_only"),
+ ({"disabled_worker": 10}, "disabled_worker_only"),
+ ({"retired_worker": 10, "spouse": 2}, "spousal"),
+ ({"retired_worker": 10, "aged_widow": 2}, "widower"),
+ ({"disabled_widow": 10}, "widower"),
+ ({"retired_worker": 0}, None),
+ ),
+)
+def test_current_law_benefit_mapping_includes_dual_entitlement(
+ mapping, expected
+):
+ assert (
+ common.benefit_type(
+ {k: {"base": v, "reform": 999} for k, v in mapping.items()}
+ )
+ == expected
+ )
+
+
+def test_inconsistent_components_refuse():
+ with pytest.raises(ValueError, match="concurrent"):
+ common.benefit_type(
+ {"retired_worker": {"base": 1}, "disabled_worker": {"base": 1}}
+ )
+ with pytest.raises(ValueError, match="unknown"):
+ common.benefit_type({"other": {"base": 1}})
+
+
+@pytest.mark.parametrize("amount", (-1, float("nan"), float("inf"), True))
+def test_invalid_baseline_components_refuse(amount):
+ with pytest.raises(ValueError, match="finite nonnegative"):
+ common.benefit_type({"retired_worker": {"base": amount}})
+
+
+@pytest.mark.parametrize("frame_kind", ("benefits", "states"))
+def test_retained_frames_are_sealed_before_attribute_loading(
+ frame_kind, monkeypatch
+):
+ replay, parent, inputs, config = _case()
+ if frame_kind == "benefits":
+ replay.benefit_rows["R0"].loc[0, "benefit_base"] = 1.0
+ else:
+ replay.states[2011].loc[0, "sex"] = "male"
+
+ def forbidden(*args, **kwargs):
+ pytest.fail("attribute loader called on tampered replay")
+
+ monkeypatch.setattr(common.g1, "load_group_attributes", forbidden)
+ with pytest.raises(common.ReproductionMismatch, match="frames changed"):
+ common.run_group_breakdown(
+ replay, parent, inputs=inputs, config=config
+ )
+
+
+def test_duplicate_or_missing_attribute_people_refuse():
+ replay, parent, inputs, config = _case()
+
+ def missing(person_ids, **kwargs):
+ return dry.invented_attribute_loader(person_ids[:-1], **kwargs)
+
+ with pytest.raises(ValueError, match="retain each opening person"):
+ common.run_group_breakdown(
+ replay,
+ parent,
+ inputs=inputs,
+ config=config,
+ attribute_loader=missing,
+ lifetime_options=_options(),
+ )
+
+
+@pytest.mark.parametrize("pointer", (None, "invalid", "parent"))
+def test_real_metadata_requires_new_pointer_before_group_loaders(
+ pointer, monkeypatch
+):
+ # Only INVENTED person rows exist; the real-data metadata is a refusal
+ # fixture and cannot reach an attribute, measure or outcome function.
+ replay, parent, inputs, config = _case()
+ inputs.cohort.data_provenance = "registered_real"
+ old = "https://github.com/PolicyEngine/microcosm-dynamics/issues/42#issuecomment-1"
+ parent["registration_pointer"] = old
+
+ def forbidden(*args, **kwargs):
+ pytest.fail("group loader reached without a new registration")
+
+ monkeypatch.setattr(common.g1, "load_group_attributes", forbidden)
+ with pytest.raises(ValueError, match="registration|pointer"):
+ common.run_group_breakdown(
+ replay,
+ parent,
+ inputs=inputs,
+ config=config,
+ registration_pointer=old if pointer == "parent" else pointer,
+ )
+
+
+@pytest.mark.parametrize(
+ "changes", ({"draw_indices": (0,)}, {"floor_seeds": (0, 1)})
+)
+def test_group_configuration_cannot_depart_from_verified_replay(
+ changes, monkeypatch
+):
+ replay, parent, inputs, config = _case()
+
+ def forbidden(*args, **kwargs):
+ pytest.fail(
+ "attribute loader reached with changed group configuration"
+ )
+
+ monkeypatch.setattr(common.g1, "load_group_attributes", forbidden)
+ with pytest.raises(common.ReproductionMismatch, match="configuration"):
+ common.run_group_breakdown(
+ replay, parent, inputs=inputs, config=replace(config, **changes)
+ )
diff --git a/tests/group_breakdowns/test_fra68.py b/tests/group_breakdowns/test_fra68.py
new file mode 100644
index 00000000..3cc6a54f
--- /dev/null
+++ b/tests/group_breakdowns/test_fra68.py
@@ -0,0 +1,121 @@
+"""Exercise-3 replay invariants and differential checks on INVENTED data.
+
+The existing FRA fixture supplies invented persons, earnings, weights,
+probabilities and rates. No real-data reader or outcome pipeline runs.
+"""
+
+from __future__ import annotations
+
+import json
+
+import pandas as pd
+import pytest
+from hypothesis import given, settings
+from hypothesis import strategies as st
+
+from populace_dynamics.cola_track_a import runner as track_runner
+from populace_dynamics.fra68_track import FRA68Config, run_fra68
+from populace_dynamics.group_breakdowns.fra68 import reproduce_fra68
+from tests.fra68_track.test_runner import _inputs
+
+
+def _normal(result):
+ return json.loads(json.dumps(result, default=str, allow_nan=False))
+
+
+@pytest.fixture(scope="module")
+def pair():
+ """All nine registered rows and both waves, with INVENTED data."""
+ inputs = _inputs()
+ config = FRA68Config(draw_indices=(0, 1))
+ return (
+ reproduce_fra68(inputs, config=config),
+ run_fra68(inputs, config=config),
+ config,
+ )
+
+
+def test_complete_artifact_matches_frozen_runner(pair):
+ replay, reference, _ = pair
+ # Entire payload: five age cells, per-draw diagnostics, counters,
+ # membership mechanisms, schedules, floors and provenance.
+ assert _normal(replay.result) == _normal(reference)
+
+
+def test_all_registered_rows_and_waves_are_retained(pair):
+ replay, _, config = pair
+ assert replay.exercise == "fra68"
+ assert tuple(replay.benefit_rows) == config.rows
+ assert set(replay.states) == {2011, 2009}
+ for wave, frame in replay.states.items():
+ assert not frame.duplicated(["draw", "person_id"]).any()
+ assert set(frame["draw"]) == set(config.draw_indices)
+ assert set(frame["year"]) == {config.reference_year}
+ assert set(frame["sex"]) <= {"female", "male"}
+ assert "marital_status" in frame
+ for row_id in config.rows_for_wave(wave):
+ rows = replay.benefit_rows[row_id]
+ assert not rows.duplicated(["draw", "person_id"]).any()
+ joined = rows.merge(
+ frame[["draw", "person_id"]],
+ on=["draw", "person_id"],
+ how="left",
+ indicator=True,
+ validate="one_to_one",
+ )
+ assert joined["_merge"].eq("both").all()
+ assert (
+ rows["benefit_base"].gt(0) | rows["benefit_reform"].gt(0)
+ ).all()
+ assert rows["age_reference"].equals(
+ config.reference_year - rows["birth_year"]
+ )
+
+
+def test_components_preserve_current_law_amounts(pair):
+ replay, _, _ = pair
+ for rows in replay.benefit_rows.values():
+ for row in rows.itertuples(index=False):
+ for scenario in ("base", "reform"):
+ amount = sum(
+ component[scenario]
+ for component in row.benefit_components.values()
+ )
+ assert amount == pytest.approx(
+ getattr(row, f"benefit_{scenario}"), rel=1e-12
+ )
+
+
+def test_projection_collection_does_not_mutate_frozen_slices(monkeypatch):
+ original = track_runner._project_population
+ observed = []
+
+ def collect(*args, **kwargs):
+ output = original(*args, **kwargs)
+ observed.extend(result.slices[-1] for result in output[0].values())
+ return output
+
+ monkeypatch.setattr(track_runner, "_project_population", collect)
+ replay = reproduce_fra68(
+ _inputs(), config=FRA68Config(rows=("F0",), draw_indices=(0,))
+ )
+ assert len(observed) == 1
+ assert "draw" not in observed[0]
+ before = observed[0].copy(deep=True)
+ replay.states[2011]["sex"] = "INVENTED mutation"
+ pd.testing.assert_frame_equal(observed[0], before)
+
+
+@settings(max_examples=6, deadline=None)
+@given(
+ row_id=st.sampled_from(("F0", "F3", "F4", "F7", "F8")),
+ draw=st.integers(min_value=0, max_value=5),
+)
+def test_replay_matches_frozen_semantics_for_draw_and_row(row_id, draw):
+ """Changing the retained row/draw cannot change frozen semantics."""
+ config = FRA68Config(rows=(row_id,), draw_indices=(draw,))
+ inputs = _inputs()
+ replay = reproduce_fra68(inputs, config=config)
+ reference = run_fra68(inputs, config=config)
+ assert _normal(replay.result) == _normal(reference)
+ assert set(replay.benefit_rows[row_id]["draw"]) == {draw}
diff --git a/tests/group_breakdowns/test_registration.py b/tests/group_breakdowns/test_registration.py
new file mode 100644
index 00000000..6f3d4090
--- /dev/null
+++ b/tests/group_breakdowns/test_registration.py
@@ -0,0 +1,421 @@
+"""INVENTED registration and exclusive-writer checks; no real-data run."""
+
+from __future__ import annotations
+
+import copy
+import hashlib
+import importlib.util
+import json
+from pathlib import Path
+from types import SimpleNamespace
+
+import pandas as pd
+import pytest
+from hypothesis import given
+from hypothesis import strategies as st
+
+ROOT = Path(__file__).resolve().parents[2]
+SPEC = importlib.util.spec_from_file_location(
+ "_projection_group_registration",
+ ROOT / "scripts" / "run_projection_groups_registered.py",
+)
+assert SPEC is not None and SPEC.loader is not None
+runner = importlib.util.module_from_spec(SPEC)
+SPEC.loader.exec_module(runner)
+
+POINTER = (
+ "https://github.com/PolicyEngine/microcosm-dynamics/issues/42"
+ "#issuecomment-123456789"
+)
+COMMIT = "a" * 40
+
+
+def invented_git(*args):
+ """INVENTED clean registered HEAD, without consulting any repository."""
+ return COMMIT if args == ("rev-parse", "HEAD") else ""
+
+
+@pytest.fixture
+def specification_spy(monkeypatch):
+ calls = []
+ monkeypatch.setattr(
+ runner,
+ "_check_specification",
+ lambda *args: calls.append(args),
+ )
+ return calls
+
+
+@pytest.mark.parametrize("exercise", ["cola", "fra68"])
+def test_valid_preflight_creates_nothing(
+ tmp_path, specification_spy, exercise
+):
+ output = tmp_path / "invented.json"
+ state = runner.preflight(
+ exercise=exercise,
+ registration_pointer=POINTER,
+ registered_commit=COMMIT,
+ output=output,
+ git=invented_git,
+ )
+ assert state == {"head": COMMIT}
+ assert len(specification_spy) == 1
+ assert list(tmp_path.iterdir()) == []
+
+
+@pytest.mark.parametrize(
+ "pointer",
+ [
+ "",
+ POINTER.replace("issues/42", "issues/41"),
+ POINTER.replace("#issuecomment-", "/comments/"),
+ POINTER + "/",
+ POINTER + "\n",
+ POINTER.replace("https://", "http://"),
+ ],
+)
+def test_bad_pointer_refuses_before_git_or_loaders(tmp_path, pointer):
+ def forbidden(*args):
+ pytest.fail(
+ "git or specification loader called before pointer refusal"
+ )
+
+ with pytest.raises(ValueError, match="issue #42"):
+ runner.preflight(
+ exercise="cola",
+ registration_pointer=pointer,
+ registered_commit=COMMIT,
+ output=tmp_path / "invented.json",
+ git=forbidden,
+ )
+ assert list(tmp_path.iterdir()) == []
+
+
+@given(
+ comment_id=st.integers(min_value=0, max_value=10**30),
+ corruption=st.sampled_from(["prefix", "suffix", "wrong_issue"]),
+)
+def test_pointer_near_matches_never_authorize(comment_id, corruption):
+ pointer = POINTER.rsplit("-", 1)[0] + f"-{comment_id}"
+ if corruption == "prefix":
+ pointer = " " + pointer
+ elif corruption == "suffix":
+ pointer += " "
+ else:
+ pointer = pointer.replace("issues/42", "issues/420")
+
+ def forbidden(*args):
+ pytest.fail("registration near-match reached repository access")
+
+ with pytest.raises(ValueError, match="issue #42"):
+ runner.preflight(
+ exercise="cola",
+ registration_pointer=pointer,
+ registered_commit=COMMIT,
+ output=Path("INVENTED_UNWRITTEN_OUTPUT.json"),
+ git=forbidden,
+ )
+
+
+@pytest.mark.parametrize("commit", ["a" * 39, "a" * 41, "A" * 40, "g" * 40])
+def test_full_lowercase_sha_required(tmp_path, specification_spy, commit):
+ with pytest.raises(ValueError, match="40-hex"):
+ runner.preflight(
+ exercise="cola",
+ registration_pointer=POINTER,
+ registered_commit=commit,
+ output=tmp_path / "invented.json",
+ git=invented_git,
+ )
+ assert specification_spy == []
+
+
+@pytest.mark.parametrize(
+ "git_result,message",
+ [("b" * 40, "HEAD"), (" M invented.txt", "clean")],
+)
+def test_head_and_clean_tree_are_required(
+ tmp_path, specification_spy, git_result, message
+):
+ def git(*args):
+ if args == ("rev-parse", "HEAD"):
+ return git_result if message == "HEAD" else COMMIT
+ return git_result
+
+ with pytest.raises(ValueError, match=message):
+ runner.preflight(
+ exercise="cola",
+ registration_pointer=POINTER,
+ registered_commit=COMMIT,
+ output=tmp_path / "invented.json",
+ git=git,
+ )
+ assert specification_spy == []
+ assert list(tmp_path.iterdir()) == []
+
+
+@pytest.mark.parametrize("sidecar", [False, True])
+def test_preflight_preserves_existing_pair(
+ tmp_path, specification_spy, sidecar
+):
+ output = tmp_path / "invented.json"
+ existing = output.with_suffix(".env.json") if sidecar else output
+ existing.write_text("INVENTED EXISTING BYTES", encoding="utf-8")
+ with pytest.raises(ValueError, match="already exists"):
+ runner.preflight(
+ exercise="cola",
+ registration_pointer=POINTER,
+ registered_commit=COMMIT,
+ output=output,
+ git=invented_git,
+ )
+ assert existing.read_text() == "INVENTED EXISTING BYTES"
+ assert specification_spy == []
+
+
+def test_specification_refusal_creates_nothing(tmp_path, monkeypatch):
+ def unratified(*args):
+ raise ValueError("INVENTED unratified specification")
+
+ monkeypatch.setattr(runner, "_check_specification", unratified)
+ with pytest.raises(ValueError, match="unratified"):
+ runner.preflight(
+ exercise="cola",
+ registration_pointer=POINTER,
+ registered_commit=COMMIT,
+ output=tmp_path / "invented.json",
+ git=invented_git,
+ )
+ assert list(tmp_path.iterdir()) == []
+
+
+def test_new_pair_has_binding_sha256(tmp_path):
+ output = tmp_path / "invented.json"
+ runner.write_new_pair(
+ output,
+ {"header": "INVENTED DATA - NOT A COMPARISON", "value": 1.0},
+ {"python": "INVENTED"},
+ )
+ sidecar = json.loads(output.with_suffix(".env.json").read_text())
+ assert sidecar["artifact"] == output.name
+ assert (
+ sidecar["artifact_sha256"]
+ == hashlib.sha256(output.read_bytes()).hexdigest()
+ )
+ assert output.read_bytes().endswith(b"\n")
+
+
+@pytest.mark.parametrize("sidecar", [False, True])
+def test_exclusive_writer_preserves_racing_existing_file(tmp_path, sidecar):
+ output = tmp_path / "invented.json"
+ existing = output.with_suffix(".env.json") if sidecar else output
+ existing.write_text("INVENTED RACING BYTES", encoding="utf-8")
+ with pytest.raises(FileExistsError):
+ runner.write_new_pair(output, {"value": 1.0}, {})
+ assert existing.read_text() == "INVENTED RACING BYTES"
+ assert set(tmp_path.iterdir()) == {existing}
+
+
+@pytest.mark.parametrize("nonfinite", [float("nan"), float("inf")])
+@pytest.mark.parametrize("environment", [False, True])
+def test_nonfinite_pair_refuses_before_any_write(
+ tmp_path, nonfinite, environment
+):
+ output = tmp_path / "invented.json"
+ bad = {"value": nonfinite}
+ with pytest.raises(ValueError, match="JSON compliant"):
+ runner.write_new_pair(
+ output,
+ {} if environment else bad,
+ bad if environment else {},
+ )
+ assert list(tmp_path.iterdir()) == []
+
+
+def test_environment_requires_exact_core_versions():
+ invented = {
+ "python": "INVENTED 1.0",
+ "packages": {"numpy": "1", "pandas": "2", "scipy": "3"},
+ }
+ runner.check_reproduction_environment(invented, invented, "a03e82e503")
+ for package in ("numpy", "pandas", "scipy"):
+ changed = {
+ **invented,
+ "packages": {**invented["packages"], package: "INVENTED CHANGE"},
+ }
+ with pytest.raises(ValueError, match=package):
+ runner.check_reproduction_environment(
+ changed, invented, "a03e82e503"
+ )
+ with pytest.raises(ValueError, match="Python"):
+ runner.check_reproduction_environment(
+ {**invented, "python": "INVENTED CHANGE"}, invented, "a03e82e503"
+ )
+ with pytest.raises(ValueError, match="a03e82e503"):
+ runner.check_reproduction_environment(invented, invented, "INVENTED")
+
+
+def test_environment_mismatch_refuses_before_cohort_reader(monkeypatch):
+ expected = {
+ "python": "INVENTED 1.0",
+ "packages": {"numpy": "1", "pandas": "2", "scipy": "3"},
+ }
+
+ def forbidden_reader(*args, **kwargs):
+ pytest.fail("cohort reader called before environment refusal")
+
+ legacy = SimpleNamespace(
+ load_ssa_parameters=lambda: SimpleNamespace(
+ pe_us_revision="a03e82e503"
+ ),
+ _environment=lambda **kwargs: {
+ **expected,
+ "python": "INVENTED CHANGE",
+ },
+ psid2010=SimpleNamespace(load_psid2010_inputs=forbidden_reader),
+ )
+ monkeypatch.setattr(runner, "_script_module", lambda name: legacy)
+ with pytest.raises(ValueError, match="Python"):
+ runner.build_registered_inputs(
+ "cola", object(), parent={}, expected_environment=expected
+ )
+
+
+def test_missing_registration_cli_never_reaches_builder(monkeypatch):
+ monkeypatch.setattr(
+ runner,
+ "build_registered_inputs",
+ lambda *args, **kwargs: pytest.fail("real-data builder called"),
+ )
+ with pytest.raises(SystemExit) as error:
+ runner.main(["--exercise", "cola"])
+ assert error.value.code == 2
+
+
+def test_methodology_pin_mismatch_refuses_before_parent_or_builder(
+ tmp_path, monkeypatch
+):
+ monkeypatch.setattr(runner, "preflight", lambda **kwargs: {"head": COMMIT})
+ monkeypatch.setattr(runner, "_sha256", lambda path: "INVENTED MISMATCH")
+
+ def forbidden(*args, **kwargs):
+ pytest.fail("parent or builder reached before methodology pin refusal")
+
+ monkeypatch.setattr(runner, "_parent", forbidden)
+ monkeypatch.setattr(runner, "build_registered_inputs", forbidden)
+ with pytest.raises(ValueError, match="methodology specification SHA-256"):
+ runner.main(
+ [
+ "--exercise",
+ "cola",
+ "--registration-pointer",
+ POINTER,
+ "--registered-commit",
+ COMMIT,
+ "--output",
+ str(tmp_path / "invented.json"),
+ ]
+ )
+ assert list(tmp_path.iterdir()) == []
+
+
+def test_parent_pointer_cannot_authorize_new_groups(tmp_path, monkeypatch):
+ monkeypatch.setattr(runner, "preflight", lambda **kwargs: {"head": COMMIT})
+ monkeypatch.setattr(
+ runner, "_sha256", lambda path: runner.POSTHOC_SPECIFICATION_SHA256
+ )
+ monkeypatch.setattr(
+ runner,
+ "_parent",
+ lambda exercise: ({"registration_pointer": POINTER}, {}, "SHA"),
+ )
+ monkeypatch.setattr(
+ runner,
+ "build_registered_inputs",
+ lambda *args, **kwargs: pytest.fail(
+ "old pointer reached cohort builder"
+ ),
+ )
+ with pytest.raises(ValueError, match="new issue #42"):
+ runner.main(
+ [
+ "--exercise",
+ "cola",
+ "--registration-pointer",
+ POINTER,
+ "--registered-commit",
+ COMMIT,
+ "--output",
+ str(tmp_path / "invented.json"),
+ ]
+ )
+ assert list(tmp_path.iterdir()) == []
+
+
+def test_mismatched_committed_cell_main_loads_no_attributes_or_writes(
+ tmp_path, monkeypatch
+):
+ """INVENTED parent mismatch exercises the actual common refusal gate."""
+ from populace_dynamics.group_breakdowns import cola, common
+
+ result = {
+ "rows": {
+ "R0": {
+ "tabulation": {"groups": [{"cell": {"percent_change": 1.0}}]}
+ }
+ },
+ "draws": {"2011": {"0": {"count": 1}}},
+ "registration_pointer": POINTER,
+ }
+ parent = copy.deepcopy(result)
+ parent["rows"]["R0"]["tabulation"]["groups"][0]["cell"][
+ "percent_change"
+ ] = 2.0
+ replay = common.ProjectionReplay(
+ exercise="cola",
+ result=result,
+ benefit_rows={"R0": pd.DataFrame({"person_id": [1]})},
+ states={},
+ )
+ loader_calls = []
+
+ def forbidden_attributes(*args, **kwargs):
+ loader_calls.append((args, kwargs))
+ pytest.fail("group attributes loaded before committed-cell refusal")
+
+ monkeypatch.setattr(
+ common.g1, "load_group_attributes", forbidden_attributes
+ )
+ monkeypatch.setattr(runner, "preflight", lambda **kwargs: {"head": COMMIT})
+ monkeypatch.setattr(
+ runner,
+ "_sha256",
+ lambda path: runner.POSTHOC_SPECIFICATION_SHA256,
+ )
+ monkeypatch.setattr(
+ runner, "_parent", lambda exercise: (parent, {}, "SHA")
+ )
+ monkeypatch.setattr(
+ runner,
+ "build_registered_inputs",
+ lambda *args, **kwargs: (object(), {}, {}, None),
+ )
+ monkeypatch.setattr(cola, "reproduce_cola", lambda *args, **kwargs: replay)
+ output = tmp_path / "invented_groups.json"
+ with pytest.raises(common.ReproductionMismatch, match="percent_change"):
+ runner.main(
+ [
+ "--exercise",
+ "cola",
+ "--registration-pointer",
+ POINTER + "0",
+ "--registered-commit",
+ COMMIT,
+ "--output",
+ str(output),
+ ]
+ )
+ assert loader_calls == []
+ assert not output.exists()
+ assert not output.with_suffix(".env.json").exists()
+ assert list(tmp_path.iterdir()) == []
From 72f72eb9bd7a8a25626438107feced23b8c5f5a3 Mon Sep 17 00:00:00 2001
From: Max Ghenis
Date: Sat, 3 Oct 2026 07:32:13 -0400
Subject: [PATCH 06/10] Add MINT group breakdowns for exercise 2 (uniform 13
percent cut)
group_breakdowns/uniform_cut.py re-executes the frozen Track U
specification for every registered row (build_age67_cohort ->
income_rows -> adjusted_incomes -> tabulation_rows ->
tabulate_uniform_cut with the design frame), checks the inputs digest
and every committed cell of runs/replication_boomers2004_uniform_cut_v1
with zero tolerance, and refuses before loading any attribute on a
mismatch. It then adds MINT groups (race and ethnicity, education,
lifetime earnings quintiles own and shared, poverty status, household
income quintile) and the report's own rows, with the adjusted-poverty
statistic, design standard errors and floors, plus MINT's poverty-table
statistics. Every observation must be born 1936-45; the 1946-55 cohort
(the blind U2 target) never enters.
scripts/run_track_u_groups_registered.py refuses without an issue #42
pointer, a clean HEAD equal to the registered commit and an absent
output. Invented data only; the large invented result is kept outside
the repository.
Co-Authored-By: Claude Opus 5.5
---
.../track_u_groups_invented.env.json | 9 +
scripts/make_nasi_repro_venv.sh | 169 +++++
scripts/run_track_u_groups_registered.py | 171 +++++
scripts/track_u_groups_invented.py | 165 ++++
.../group_breakdowns/__init__.py | 1 +
.../group_breakdowns/common.py | 308 ++++++++
.../group_breakdowns/uniform_cut.py | 716 ++++++++++++++++++
tests/group_breakdowns/test_common.py | 333 ++++++++
.../test_registration_script.py | 400 ++++++++++
tests/group_breakdowns/test_uniform_cut.py | 245 ++++++
.../test_uniform_cut_properties.py | 481 ++++++++++++
11 files changed, 2998 insertions(+)
create mode 100644 docs/analysis/nasi_group_breakdowns_invented_20261001/exercise_2_uniform_cut/track_u_groups_invented.env.json
create mode 100755 scripts/make_nasi_repro_venv.sh
create mode 100644 scripts/run_track_u_groups_registered.py
create mode 100644 scripts/track_u_groups_invented.py
create mode 100644 src/populace_dynamics/group_breakdowns/__init__.py
create mode 100644 src/populace_dynamics/group_breakdowns/common.py
create mode 100644 src/populace_dynamics/group_breakdowns/uniform_cut.py
create mode 100644 tests/group_breakdowns/test_common.py
create mode 100644 tests/group_breakdowns/test_registration_script.py
create mode 100644 tests/group_breakdowns/test_uniform_cut.py
create mode 100644 tests/group_breakdowns/test_uniform_cut_properties.py
diff --git a/docs/analysis/nasi_group_breakdowns_invented_20261001/exercise_2_uniform_cut/track_u_groups_invented.env.json b/docs/analysis/nasi_group_breakdowns_invented_20261001/exercise_2_uniform_cut/track_u_groups_invented.env.json
new file mode 100644
index 00000000..585c7c86
--- /dev/null
+++ b/docs/analysis/nasi_group_breakdowns_invented_20261001/exercise_2_uniform_cut/track_u_groups_invented.env.json
@@ -0,0 +1,9 @@
+{
+ "artifact": "track_u_groups_invented.json",
+ "artifact_sha256": "ee115ace11f52fe53223819bcb112b7a86d954cde7a7b059cc6507751aa1a6ce",
+ "environment": {
+ "python": "3.14.4",
+ "pandas": "3.0.3",
+ "kind": "INVENTED dry run; not recorded real-run environment"
+ }
+}
diff --git a/scripts/make_nasi_repro_venv.sh b/scripts/make_nasi_repro_venv.sh
new file mode 100755
index 00000000..8e808501
--- /dev/null
+++ b/scripts/make_nasi_repro_venv.sh
@@ -0,0 +1,169 @@
+#!/usr/bin/env bash
+# Build the recorded Track U environment without reading PSID or running it.
+# Source: runs/replication_boomers2004_uniform_cut_v1.env.json.
+# The sidecar does not record SciPy: a supplemental version is opt-in only.
+set -euo pipefail
+
+script_dir=$(cd -- "$(dirname -- "${BASH_SOURCE[0]}")" && pwd)
+repo_dir=$(cd -- "$script_dir/.." && pwd)
+venv_dir=${NASI_REPRO_VENV_DIR:-/Users/maxghenis/PolicyEngine/microcosm-dynamics/.venv-nasi-repro}
+sidecar=${NASI_REPRO_ENV_SIDECAR:-$repo_dir/runs/replication_boomers2004_uniform_cut_v1.env.json}
+bootstrap_python=${NASI_REPRO_BOOTSTRAP_PYTHON:-python3}
+scipy_version=${NASI_REPRO_SCIPY_VERSION:-}
+
+usage() {
+ cat <<'EOF'
+Usage: make_nasi_repro_venv.sh [--venv-dir PATH] [--scipy-version VERSION]
+
+Creates a NEW environment with the Python, NumPy, pandas, and project
+versions recorded in the Track U sidecar, then verifies each exact version.
+The project is installed editable with --no-deps; unrecorded dependencies
+are not guessed. No cohort or outcome pipeline is run.
+
+SciPy is absent from the recorded sidecar. --scipy-version supplies an
+explicit supplemental pin; this does not establish its original version.
+The build manifest labels the missing SciPy provenance and does not claim
+that the complete original environment has been reproduced.
+
+Environment overrides: NASI_REPRO_VENV_DIR, NASI_REPRO_SCIPY_VERSION,
+NASI_REPRO_BOOTSTRAP_PYTHON, NASI_REPRO_ENV_SIDECAR. Standard uv cache and
+Python installation directory overrides are also honored.
+EOF
+}
+
+while (($#)); do
+ case "$1" in
+ --venv-dir|--scipy-version)
+ if (($# < 2)) || [[ -z "$2" ]]; then
+ printf 'Missing value for %s\n' "$1" >&2
+ exit 2
+ fi
+ if [[ "$1" == --venv-dir ]]; then
+ venv_dir=$2
+ else
+ scipy_version=$2
+ fi
+ shift 2
+ ;;
+ --help|-h)
+ usage
+ exit 0
+ ;;
+ *)
+ printf 'Unknown argument: %s\n' "$1" >&2
+ usage >&2
+ exit 2
+ ;;
+ esac
+done
+
+if [[ -e "$venv_dir" || -L "$venv_dir" ]]; then
+ printf 'Refusing existing environment path: %s\n' "$venv_dir" >&2
+ exit 1
+fi
+command -v uv >/dev/null || { printf 'uv is required\n' >&2; exit 1; }
+
+# Validate first; printing only version pins prevents shell interpretation
+# of sidecar strings. The optional SciPy version must be a numeric release.
+pins=$("$bootstrap_python" - "$sidecar" "$scipy_version" <<'PY'
+import json
+import hashlib
+import re
+import sys
+from pathlib import Path
+
+source = Path(sys.argv[1])
+source_bytes = source.read_bytes()
+expected_sha256 = (
+ "4041a19ada1cb4c05628cf31bf10cbfdb364bdd2d5d62455a3121143913e0ed9"
+)
+if hashlib.sha256(source_bytes).hexdigest() != expected_sha256:
+ raise SystemExit("Refusing changed Track U environment sidecar")
+sidecar = json.loads(source_bytes)
+environment = sidecar["environment"]
+packages = environment["packages"]
+versions = [
+ environment["python"],
+ packages["numpy"],
+ packages["pandas"],
+ packages["policyengine-social-security-model"],
+]
+for version in versions + ([sys.argv[2]] if sys.argv[2] else []):
+ if not isinstance(version, str) or not re.fullmatch(
+ r"\d+\.\d+\.\d+", version
+ ):
+ raise SystemExit("Refusing malformed or non-exact version pin")
+recorded_scipy = packages.get("scipy", "")
+if recorded_scipy:
+ if not re.fullmatch(r"\d+\.\d+\.\d+", recorded_scipy):
+ raise SystemExit("Refusing malformed recorded SciPy pin")
+ if sys.argv[2] and sys.argv[2] != recorded_scipy:
+ raise SystemExit("Supplemental SciPy pin conflicts with sidecar")
+print("\n".join(versions + [recorded_scipy or sys.argv[2]]))
+PY
+)
+python_version=$(printf '%s\n' "$pins" | sed -n '1p')
+numpy_version=$(printf '%s\n' "$pins" | sed -n '2p')
+pandas_version=$(printf '%s\n' "$pins" | sed -n '3p')
+project_version=$(printf '%s\n' "$pins" | sed -n '4p')
+scipy_version=$(printf '%s\n' "$pins" | sed -n '5p')
+
+uv venv --python "$python_version" "$venv_dir"
+requirements=("numpy==$numpy_version" "pandas==$pandas_version")
+if [[ -n "$scipy_version" ]]; then
+ requirements+=("scipy==$scipy_version")
+fi
+uv pip install --python "$venv_dir/bin/python" "${requirements[@]}"
+uv pip install --python "$venv_dir/bin/python" --no-deps --editable "$repo_dir"
+
+"$venv_dir/bin/python" - "$sidecar" "$scipy_version" "$project_version" <<'PY'
+import hashlib
+import importlib.metadata
+import json
+import platform
+import sys
+from pathlib import Path
+
+source = Path(sys.argv[1]).resolve()
+source_bytes = source.read_bytes()
+sidecar = json.loads(source_bytes)
+recorded = sidecar["environment"]
+expected = {
+ "numpy": recorded["packages"]["numpy"],
+ "pandas": recorded["packages"]["pandas"],
+ "policyengine-social-security-model": sys.argv[3],
+}
+if sys.argv[2]:
+ expected["scipy"] = sys.argv[2]
+actual = {
+ name: importlib.metadata.version(name) for name in sorted(expected)
+}
+if platform.python_version() != recorded["python"] or actual != expected:
+ raise SystemExit("Refusing environment whose versions differ from pins")
+recorded_scipy = recorded["packages"].get("scipy")
+manifest = {
+ "purpose": "environment build verification only; no pipeline run",
+ "source_sidecar": str(source),
+ "source_sidecar_sha256": hashlib.sha256(source_bytes).hexdigest(),
+ "python": platform.python_version(),
+ "packages": actual,
+ "editable_project": True,
+ "recorded_versions_verified": True,
+ "complete_original_environment_reproduced": False,
+ "scipy_pin_provenance": (
+ "recorded sidecar" if recorded_scipy else
+ "explicit supplemental pin; original version unrecorded"
+ if sys.argv[2] else "not installed; original version unrecorded"
+ ),
+ "limitation": (
+ "Only sidecar-recorded versions and an optional supplemental "
+ "SciPy pin are installed; other project dependencies are not "
+ "reconstructed from unrecorded versions."
+ ),
+}
+destination = Path(sys.prefix) / "nasi_repro_build.json"
+with destination.open("x", encoding="utf-8") as handle:
+ json.dump(manifest, handle, indent=2, allow_nan=False)
+ handle.write("\n")
+print(json.dumps(manifest, indent=2, allow_nan=False))
+PY
diff --git a/scripts/run_track_u_groups_registered.py b/scripts/run_track_u_groups_registered.py
new file mode 100644
index 00000000..5d494a5e
--- /dev/null
+++ b/scripts/run_track_u_groups_registered.py
@@ -0,0 +1,171 @@
+#!/usr/bin/env python3
+# ruff: noqa: E402
+"""Registered, one-shot, post hoc, not blind Track U groups (report-only).
+
+Do not run before a new issue #42 registration. The frozen parent is
+reproduced before any group attribute is read or new cell calculated.
+See uniform_cut.py for the registered post hoc conventions.
+"""
+
+from __future__ import annotations
+
+import argparse
+import importlib.util
+import json
+import subprocess
+import sys
+from dataclasses import asdict
+from pathlib import Path
+
+ROOT = Path(__file__).resolve().parents[1]
+sys.path.insert(0, str(ROOT / "src"))
+
+from populace_dynamics.cohorts import age67
+from populace_dynamics.estimates import adjusted_poverty as ap
+from populace_dynamics.group_breakdowns import (
+ common,
+)
+from populace_dynamics.group_breakdowns import (
+ uniform_cut as groups,
+)
+from populace_dynamics.ss import params as ss_params
+from populace_dynamics.uniform_cut_track_u import rows, runner
+
+SIDECAR_SHA256 = (
+ "4041a19ada1cb4c05628cf31bf10cbfdb364bdd2d5d62455a3121143913e0ed9"
+)
+
+
+def check_environment(actual: dict, expected: dict) -> None:
+ """Match every recorded version; record the absent SciPy pin honestly."""
+
+ common.assert_exact_cells(
+ expected["python"], actual["python"], path="Python version"
+ )
+ for name, version in expected["packages"].items():
+ common.assert_exact_cells(
+ version, actual["packages"].get(name), path=f"package {name}"
+ )
+ common.assert_exact_cells(
+ expected["project"], actual["project"], path="project"
+ )
+
+
+def check_parameter_checkout() -> str:
+ """Require the existing checkout's full HEAD and clean tracked files."""
+
+ checkout = ss_params._resolve_pe_us(None)
+
+ def git(*args):
+ return subprocess.run(
+ ["git", "--no-optional-locks", "-C", str(checkout), *args],
+ check=True,
+ capture_output=True,
+ text=True,
+ ).stdout.strip()
+
+ head = git("rev-parse", "HEAD")
+ if not head.startswith("a03e82e503") or len(head) != 40:
+ raise common.GroupBreakdownRefusal(
+ "SSA checkout must be revision a03e82e503"
+ )
+ if git("status", "--porcelain", "--untracked-files=no"):
+ raise common.GroupBreakdownRefusal(
+ "SSA parameter checkout must have clean tracked files"
+ )
+ return head
+
+
+def main(argv: list[str] | None = None) -> int:
+ parser = argparse.ArgumentParser(description=__doc__.splitlines()[0])
+ parser.add_argument("--registration-pointer", required=True)
+ parser.add_argument("--registered-commit", required=True)
+ parser.add_argument(
+ "--output",
+ type=Path,
+ default=ROOT
+ / "runs/replication_boomers2004_uniform_cut_groups_posthoc_v1.json",
+ )
+ args = parser.parse_args(argv)
+ state = common.preflight(
+ registration_pointer=args.registration_pointer,
+ registered_commit=args.registered_commit,
+ output=args.output,
+ root=ROOT,
+ )
+ committed = groups.load_committed_artifact()
+ if args.registration_pointer == committed["registration_pointer"]:
+ raise common.GroupBreakdownRefusal(
+ "a new issue #42 registration is required"
+ )
+ sidecar = groups.BASE_ARTIFACT.with_suffix(".env.json")
+ common.assert_sha256(sidecar, SIDECAR_SHA256)
+ expected_environment = json.loads(sidecar.read_text())["environment"]
+ block = rows.specification_block()
+ spec = importlib.util.spec_from_file_location(
+ "_track_u_parent", ROOT / "scripts/run_track_u_registered.py"
+ )
+ if spec is None or spec.loader is None:
+ raise common.GroupBreakdownRefusal(
+ "cannot load frozen Track U registration guards"
+ )
+ parent = importlib.util.module_from_spec(spec)
+ spec.loader.exec_module(parent)
+ parent.check_specification_ratified(block)
+ common.assert_exact_cells(
+ committed["checks"]["specification_rows"],
+ rows.check_rows_against_block(block),
+ path="specification rows",
+ )
+ common.assert_exact_cells(
+ committed["checks"]["max_rulings"],
+ rows.check_rulings_against_block(block),
+ path="Max rulings",
+ )
+ common.assert_sha256(
+ rows.SPECIFICATION_PATH, committed["specification"]["sha256"]
+ )
+ # The existing parameter checkout is selected by the documented env
+ # variable. No checkout is created, changed, or fetched by this script.
+ full_parameter_head = check_parameter_checkout()
+ ssa = ss_params.load_ssa_parameters()
+ environment = common.environment(
+ ssa_parameters_revision=full_parameter_head
+ )
+ check_environment(environment, expected_environment)
+ parameters = runner.committed_parameters(ap.load_poverty_thresholds())
+ inputs = age67.load_age67_inputs()
+ parent.check_headline(committed["headline"]["row"], inputs)
+ artifact = groups.build_uniform_cut_groups(
+ inputs,
+ parameters,
+ committed,
+ data_provenance=ap.REGISTERED_REAL,
+ registration_pointer=args.registration_pointer,
+ lifetime_inputs=groups.LifetimeInputs(ssa),
+ )
+ run = asdict(state)
+ run["output_path"] = str(state.output_path.relative_to(ROOT))
+ run["sidecar_path"] = str(state.sidecar_path.relative_to(ROOT))
+ artifact["run"] = run
+ artifact["parent_artifact"] = {
+ "path": str(groups.BASE_ARTIFACT.relative_to(ROOT)),
+ "sha256": groups.BASE_ARTIFACT_SHA256,
+ }
+ artifact["environment_reproduction"] = {
+ "recorded_versions_matched": True,
+ "scipy_version_recorded_by_parent": False,
+ "note": (
+ "parent sidecar has no SciPy pin; current resolver versions "
+ "retained in sidecar"
+ ),
+ }
+ common.write_artifact_pair(
+ output=state.output_path, artifact=artifact, environment=environment
+ )
+ print(state.output_path)
+ return 0
+
+
+if __name__ == "__main__":
+ raise SystemExit(main())
diff --git a/scripts/track_u_groups_invented.py b/scripts/track_u_groups_invented.py
new file mode 100644
index 00000000..622fe515
--- /dev/null
+++ b/scripts/track_u_groups_invented.py
@@ -0,0 +1,165 @@
+#!/usr/bin/env python3
+# ruff: noqa: E402
+"""INVENTED DATA - NOT A COMPARISON: G4b complete adapter dry run."""
+
+from __future__ import annotations
+
+import argparse
+import platform
+import sys
+from pathlib import Path
+
+import pandas as pd
+
+ROOT = Path(__file__).resolve().parents[1]
+sys.path.insert(0, str(ROOT / "src"))
+
+from populace_dynamics.cohorts import (
+ age67,
+)
+from populace_dynamics.cohorts import (
+ group_attributes as ga,
+)
+from populace_dynamics.data import group_attributes_psid as gap
+from populace_dynamics.estimates import adjusted_poverty as ap
+from populace_dynamics.estimates import lifetime_measures as lm
+from populace_dynamics.group_breakdowns import (
+ common,
+)
+from populace_dynamics.group_breakdowns import (
+ uniform_cut as groups,
+)
+from populace_dynamics.ss.params import SSAParameters
+from populace_dynamics.uniform_cut_track_u import (
+ invented,
+ runner,
+)
+
+
+def invented_attribute_loader(inputs: age67.Age67Inputs):
+ """Build G1 side frames from explicit invented race/education reports."""
+
+ universe = (
+ inputs.death_records["person_id"].drop_duplicates().astype("int64")
+ )
+ report_columns = ["person_id", "wave", "role", *gap.REPORT_CODE_COLUMNS]
+ reports, education = [], []
+ for index, pid in enumerate(universe):
+ row = dict.fromkeys(report_columns, pd.NA)
+ row.update(
+ person_id=int(pid),
+ wave=2013,
+ role="head",
+ hispanic_code=int(index % 4 == 0),
+ family_education_code=12,
+ birth_state_code=1 if index % 2 else 0,
+ year_came_code=0 if index % 2 else 1960,
+ )
+ for mention in range(1, len(gap.FAMILY_ITEMS[2013].race["head"]) + 1):
+ row[f"race_code_{mention}"] = 1 + index % 3 if mention == 1 else 0
+ reports.append(row)
+ education.append(
+ {
+ "person_id": int(pid),
+ "wave": 2005,
+ "education_code": (10, 12, 14, 16, 17)[index % 5],
+ "role": pd.NA,
+ "family_education_code": pd.NA,
+ }
+ )
+ reports = pd.DataFrame(reports, columns=report_columns)
+ education = pd.DataFrame(education)
+ for frame in (reports, education):
+ for name in frame:
+ frame[name] = frame[name].astype(
+ "string" if name == "role" else "Int64"
+ )
+ side_inputs = ga.GroupAttributeInputs(
+ reports=reports,
+ education=education,
+ universe=universe,
+ provenance={
+ "kind": "INVENTED",
+ "fixture": "G4b deterministic categories, not PSID",
+ },
+ )
+
+ def loader(person_ids, *, anchor_waves):
+ return ga.build_group_attributes(
+ side_inputs, person_ids, anchor_waves=anchor_waves
+ )
+
+ return loader
+
+
+def invented_lifetime_inputs() -> groups.LifetimeInputs:
+ """INVENTED wage indices, contribution bases, tax and interest rates."""
+
+ params = SSAParameters(
+ nawi={
+ year: 1000.0 * 1.02 ** (year - 1951) for year in range(1951, 2024)
+ },
+ wage_base={1937: 50000.0},
+ pia_factors=(0.9, 0.32, 0.15),
+ fra_months_by_birth_year=[(1900, 792)],
+ early_monthly_rates=(5 / 900, 5 / 1200),
+ early_first_bracket_months=36,
+ pe_us_revision="INVENTED",
+ delayed_credit_by_birth_year=[(1900, 0.08)],
+ )
+ rates = lm.OASDITaxRates(
+ combined_percent_by_year={1937: 10.0},
+ open_ended_from=1938,
+ open_ended_combined_percent=10.0,
+ basis=lm.TaxRateBasis.EMPLOYEE_EMPLOYER_PAID,
+ provenance={"source": "INVENTED"},
+ )
+ interest = lm.TrustFundInterestRates(
+ percent_by_year={year: 3.0 for year in range(1937, 2024)},
+ series="INVENTED",
+ provenance={"source": "INVENTED"},
+ )
+ return groups.LifetimeInputs(params, rates, interest)
+
+
+def main(argv: list[str] | None = None) -> int:
+ parser = argparse.ArgumentParser(description=__doc__)
+ parser.add_argument(
+ "--output",
+ type=Path,
+ default=ROOT
+ / "docs/analysis/nasi_group_breakdowns_invented_20261001"
+ / "exercise_2_uniform_cut/track_u_groups_invented.json",
+ )
+ args = parser.parse_args(argv)
+ inputs = invented.invented_age67_inputs(supplement_waves_staged=True)
+ parameters = runner.committed_parameters(
+ invented.invented_poverty_thresholds()
+ )
+ committed = runner.run_track_u(
+ inputs, parameters, data_provenance=ap.INVENTED
+ )
+ artifact = groups.build_uniform_cut_groups(
+ inputs,
+ parameters,
+ committed,
+ data_provenance=ap.INVENTED,
+ attribute_loader=invented_attribute_loader(inputs),
+ lifetime_inputs=invented_lifetime_inputs(),
+ )
+ args.output.parent.mkdir(parents=True, exist_ok=True)
+ common.write_artifact_pair(
+ output=args.output,
+ artifact=artifact,
+ environment={
+ "python": platform.python_version(),
+ "pandas": pd.__version__,
+ "kind": "INVENTED dry run; not recorded real-run environment",
+ },
+ )
+ print(args.output)
+ return 0
+
+
+if __name__ == "__main__":
+ raise SystemExit(main())
diff --git a/src/populace_dynamics/group_breakdowns/__init__.py b/src/populace_dynamics/group_breakdowns/__init__.py
new file mode 100644
index 00000000..ad438c34
--- /dev/null
+++ b/src/populace_dynamics/group_breakdowns/__init__.py
@@ -0,0 +1 @@
+"""Registered, report-only adapters for the frozen blind-test pipelines."""
diff --git a/src/populace_dynamics/group_breakdowns/common.py b/src/populace_dynamics/group_breakdowns/common.py
new file mode 100644
index 00000000..7525989d
--- /dev/null
+++ b/src/populace_dynamics/group_breakdowns/common.py
@@ -0,0 +1,308 @@
+"""Refusal and publication guards for registered post hoc breakdowns.
+
+The registration and environment conventions compose the frozen Track A
+entry point, ``scripts/run_track_a_registered.py`` (preflight and
+``_environment``), without changing it. A reproduced parent artifact must
+match before its adapter loads attributes or computes subgroup cells.
+
+Invariants: comparisons have zero tolerance, including binary64 signed
+zero; preflight permits only a new paired artifact under this checkout's
+``runs`` directory; publication never overwrites either member of a pair,
+and removes only files it created if publication fails. No helper reads
+microdata or computes a measure.
+"""
+
+from __future__ import annotations
+
+import hashlib
+import importlib.util
+import json
+import math
+import os
+import re
+import struct
+import subprocess
+from collections.abc import Callable, Mapping
+from dataclasses import dataclass
+from pathlib import Path
+from typing import Any
+
+POSTHOC_LABELS = (
+ "registered, one-shot, post hoc, not blind",
+ "report-only",
+)
+REGISTRATION_POINTER = re.compile(
+ r"https://github\.com/PolicyEngine/microcosm-dynamics/issues/42"
+ r"#issuecomment-[0-9]+"
+)
+_ARTIFACT_NAME = re.compile(
+ r"[A-Za-z0-9][A-Za-z0-9_.-]*_groups_posthoc_v1\.json"
+)
+_ROOT = Path(__file__).resolve().parents[3]
+
+
+class GroupBreakdownRefusal(ValueError):
+ """A required reproduction or registration invariant was not met."""
+
+
+@dataclass(frozen=True)
+class ReproductionCheck:
+ """Recorded evidence of complete, exact parent-cell reproduction."""
+
+ compared_leaf_cells: int
+ comparison: str = "bit-exact binary64 floats and exact JSON scalars"
+ absolute_tolerance: float = 0.0
+ relative_tolerance: float = 0.0
+
+
+@dataclass(frozen=True)
+class RegisteredRun:
+ """The checked checkout and new output pair bound by a registration."""
+
+ registration_pointer: str
+ registered_commit: str
+ git_head: str
+ git_clean: bool
+ output_path: Path
+ sidecar_path: Path
+
+
+def assert_exact_cells(
+ committed: Any,
+ recomputed: Any,
+ *,
+ path: str = "cells",
+) -> ReproductionCheck:
+ """Refuse any changed, missing, extra, or nonfinite JSON cell.
+
+ Dictionary order is immaterial, and tuples have their JSON-array
+ semantics. Scalar types remain exact: ``True``, ``1`` and ``1.0`` are
+ distinct. Finite Python floats are compared by their binary64 bits,
+ so even a change from ``0.0`` to ``-0.0`` refuses. No summation
+ tolerance is introduced, and refusal messages expose paths only.
+ """
+
+ def refuse(location: str, reason: str = "mismatch") -> None:
+ raise GroupBreakdownRefusal(
+ f"committed cell {reason} at {location}; no group cells allowed"
+ )
+
+ def compare(expected: Any, actual: Any, location: str) -> int:
+ if isinstance(expected, Mapping) or isinstance(actual, Mapping):
+ if not isinstance(expected, Mapping) or not isinstance(
+ actual, Mapping
+ ):
+ refuse(location)
+ if any(not isinstance(key, str) for key in expected):
+ refuse(location, "has a non-JSON key")
+ if any(not isinstance(key, str) for key in actual):
+ refuse(location, "has a non-JSON key")
+ if expected.keys() != actual.keys():
+ refuse(location, "keys mismatch")
+ return sum(
+ compare(expected[key], actual[key], f"{location}.{key}")
+ for key in expected
+ )
+ arrays = (list, tuple)
+ if isinstance(expected, arrays) or isinstance(actual, arrays):
+ if not isinstance(expected, arrays) or not isinstance(
+ actual, arrays
+ ):
+ refuse(location)
+ if len(expected) != len(actual):
+ refuse(location, "array length mismatch")
+ return sum(
+ compare(left, right, f"{location}[{index}]")
+ for index, (left, right) in enumerate(
+ zip(expected, actual, strict=True)
+ )
+ )
+ if type(expected) is not type(actual):
+ refuse(location, "scalar type mismatch")
+ if isinstance(expected, float):
+ if not math.isfinite(expected) or not math.isfinite(actual):
+ refuse(location, "is nonfinite")
+ if struct.pack("!d", expected) != struct.pack("!d", actual):
+ refuse(location)
+ elif expected is None or isinstance(expected, str | bool | int):
+ if expected != actual:
+ refuse(location)
+ else:
+ refuse(location, "has a non-JSON value")
+ return 1
+
+ return ReproductionCheck(compare(committed, recomputed, path))
+
+
+def assert_sha256(path: Path, expected_sha256: str) -> str:
+ """Refuse to admit pinned non-microdata bytes with a different SHA-256."""
+
+ if not re.fullmatch(r"[0-9a-f]{64}", expected_sha256):
+ raise GroupBreakdownRefusal("expected SHA-256 must be full 64-hex")
+ actual = hashlib.sha256(Path(path).read_bytes()).hexdigest()
+ if actual != expected_sha256:
+ raise GroupBreakdownRefusal(f"SHA-256 mismatch for {path}")
+ return actual
+
+
+def _git(root: Path, *args: str) -> str:
+ return subprocess.run(
+ ["git", "-C", str(root), *args],
+ check=True,
+ capture_output=True,
+ text=True,
+ ).stdout.strip()
+
+
+def _exists(path: Path) -> bool:
+ """Include dangling symlinks among paths that must not be replaced."""
+
+ return path.exists() or path.is_symlink()
+
+
+def preflight(
+ *,
+ registration_pointer: str,
+ registered_commit: str,
+ output: Path,
+ root: Path,
+ git: Callable[..., str] | None = None,
+) -> RegisteredRun:
+ """Check registration, clean HEAD, and new paired paths before reads.
+
+ This verifies the pointer's form locally; it does not retrieve the
+ comment or imply that a new real-data run has been authorized. The
+ orchestrator must obtain the issue #42 registration before execution.
+ ``git`` is injectable solely to exercise guards on invented data.
+ """
+
+ if not REGISTRATION_POINTER.fullmatch(registration_pointer):
+ raise GroupBreakdownRefusal(
+ "registration pointer must be an issue #42 comment URL"
+ )
+ if not re.fullmatch(r"[0-9a-f]{40}", registered_commit):
+ raise GroupBreakdownRefusal(
+ "--registered-commit must be a full 40-hex SHA"
+ )
+ root = Path(root).resolve()
+ if git is None:
+
+ def git(*args: str) -> str:
+ return _git(root, *args)
+
+ head = git("rev-parse", "HEAD")
+ if head != registered_commit:
+ raise GroupBreakdownRefusal(
+ "HEAD is not the registered commit; no run allowed"
+ )
+ if git("status", "--porcelain", "--untracked-files=all") != "":
+ raise GroupBreakdownRefusal(
+ "working tree must be clean for a registered run"
+ )
+ runs = root / "runs"
+ if runs.is_symlink():
+ raise GroupBreakdownRefusal("runs directory must not be a symlink")
+ candidate = Path(output)
+ if not candidate.is_absolute():
+ candidate = root / candidate
+ if candidate.is_symlink():
+ raise GroupBreakdownRefusal("output must not be a symlink")
+ candidate = candidate.resolve()
+ if candidate.parent != runs.resolve() or not _ARTIFACT_NAME.fullmatch(
+ candidate.name
+ ):
+ raise GroupBreakdownRefusal(
+ "output must be runs/_groups_posthoc_v1.json"
+ )
+ sidecar = candidate.with_suffix(".env.json")
+ if _exists(candidate) or _exists(sidecar):
+ raise GroupBreakdownRefusal(
+ "artifact or environment sidecar already exists: one-shot"
+ )
+ return RegisteredRun(
+ registration_pointer,
+ registered_commit,
+ head,
+ True,
+ candidate,
+ sidecar,
+ )
+
+
+def environment(
+ *,
+ ssa_parameters_revision: str,
+ root: Path | None = None,
+) -> dict[str, Any]:
+ """Reuse Track A's installed-distribution resolver unchanged.
+
+ Loading the entry-point module through ``importlib`` composes its
+ resolver exactly as the frozen FRA-68 and Track M runners do. The
+ registered ``main`` function is never invoked.
+ """
+
+ script = (Path(root) if root is not None else _ROOT) / "scripts"
+ script = script / "run_track_a_registered.py"
+ spec = importlib.util.spec_from_file_location(
+ "_group_breakdowns_track_a_registered", script
+ )
+ if spec is None or spec.loader is None:
+ raise GroupBreakdownRefusal("cannot load Track A environment resolver")
+ module = importlib.util.module_from_spec(spec)
+ spec.loader.exec_module(module)
+ return module._environment(ssa_parameters_revision=ssa_parameters_revision)
+
+
+def write_artifact_pair(
+ *,
+ output: Path,
+ artifact: Mapping[str, Any],
+ environment: Mapping[str, Any],
+) -> tuple[Path, Path]:
+ """Exclusively create artifact and SHA-bound sidecar, or roll back.
+
+ Serialization happens before either file is opened. On any failure,
+ rollback removes only files this invocation created, and only while
+ their inode still matches. An existing artifact or sidecar is never
+ changed. A host interruption can leave an incomplete pair, which
+ preflight refuses; publication is not a transactional filesystem.
+ """
+
+ output = Path(output)
+ sidecar = output.with_suffix(".env.json")
+ artifact_text = json.dumps(artifact, indent=2, allow_nan=False) + "\n"
+ artifact_sha256 = hashlib.sha256(artifact_text.encode("utf-8")).hexdigest()
+ sidecar_text = (
+ json.dumps(
+ {
+ "artifact": output.name,
+ "artifact_sha256": artifact_sha256,
+ "environment": dict(environment),
+ },
+ indent=2,
+ allow_nan=False,
+ )
+ + "\n"
+ )
+ created: list[tuple[Path, os.stat_result]] = []
+ try:
+ for path, content in (
+ (output, artifact_text),
+ (sidecar, sidecar_text),
+ ):
+ with path.open("x", encoding="utf-8") as handle:
+ created.append((path, os.fstat(handle.fileno())))
+ handle.write(content)
+ except BaseException:
+ for path, stat in reversed(created):
+ try:
+ current = path.lstat()
+ if (current.st_dev, current.st_ino) == (
+ stat.st_dev,
+ stat.st_ino,
+ ):
+ path.unlink()
+ except FileNotFoundError:
+ pass
+ raise
+ return output, sidecar
diff --git a/src/populace_dynamics/group_breakdowns/uniform_cut.py b/src/populace_dynamics/group_breakdowns/uniform_cut.py
new file mode 100644
index 00000000..130d7339
--- /dev/null
+++ b/src/populace_dynamics/group_breakdowns/uniform_cut.py
@@ -0,0 +1,716 @@
+"""Exercise 2 post hoc groups, composing frozen Track U and G1--G3.
+
+Source locators: uniform_cut_track_u/runner.py:_compute_row and
+run_track_u; estimates/group_breakdown.py:MINT8_SCHEME and
+tabulate_poverty_breakdown; estimates/lifetime_measures.py. No PSID
+reader is called here except through cohorts.age67/group_attributes.
+
+Invariants: every registered row and F17 cell is reproduced bit-exactly
+before any group attribute loader or lifetime calculation runs. All
+observations are born in 1936--1945. Joins preserve observation identity,
+order and weights. Group SEs use the full design; floors split the full
+family-linked sample. Missing measures remain unavailable, never zero.
+"""
+
+from __future__ import annotations
+
+import json
+from collections.abc import Callable, Mapping
+from copy import deepcopy
+from dataclasses import dataclass, field
+from pathlib import Path
+from typing import Any
+
+import numpy as np
+import pandas as pd
+
+from populace_dynamics.cohorts import age67, psid2010
+from populace_dynamics.cohorts import group_attributes as ga
+from populace_dynamics.estimates import adjusted_poverty as ap
+from populace_dynamics.estimates import group_breakdown as gb
+from populace_dynamics.estimates import lifetime_measures as lm
+from populace_dynamics.estimates import uniform_cut_tabulation as ut
+from populace_dynamics.group_breakdowns.common import (
+ POSTHOC_LABELS,
+ GroupBreakdownRefusal,
+ assert_exact_cells,
+ assert_sha256,
+)
+from populace_dynamics.ss.params import SSAParameters
+from populace_dynamics.uniform_cut_track_u import diagnostics, runner
+from populace_dynamics.uniform_cut_track_u.rows import REGISTERED_ROWS
+
+ROOT = Path(__file__).resolve().parents[3]
+BASE_ARTIFACT = ROOT / "runs/replication_boomers2004_uniform_cut_v1.json"
+BASE_ARTIFACT_SHA256 = (
+ "48fb6cb108b09e19ed10c32586bd1b0d9521ec7746f693ba0241963ac24a4de3"
+)
+INPUT_FRAMES_SHA256 = (
+ "43c41f444128ecf72ce9f8c5e64a7cda61a6fccaa2a58ff60ae7ce91356c4271"
+)
+SCHEMA_VERSION = "populace_dynamics.uniform_cut_groups_posthoc.v1"
+INVENTED_HEADER = "INVENTED DATA - NOT A COMPARISON"
+MEASURES = (
+ "initial_aime_quintile",
+ "lifetime_payroll_tax_quintile",
+ "lifetime_payroll_tax_quintile_shared",
+ "lifetime_earnings_own",
+ "lifetime_earnings_shared",
+)
+
+
+@dataclass(frozen=True)
+class ReplayedRow:
+ """Frozen-pipeline outputs retained only in memory for group joins."""
+
+ result: Mapping[str, Any]
+ members: pd.DataFrame
+ adjusted: pd.DataFrame
+ design: pd.DataFrame
+
+
+@dataclass(frozen=True)
+class Replay:
+ """A complete reproduction; contains no newly loaded attributes."""
+
+ evidence: Mapping[str, Any]
+ rows: Mapping[str, ReplayedRow]
+
+
+@dataclass(frozen=True)
+class LifetimeInputs:
+ """G2 inputs, not a new measure implementation.
+
+ Careers are observed PSID head/spouse earnings, without filling gaps.
+ Taxes use G2's sourced combined OASDI taxes-paid series and trust-fund
+ interest series. Missing spouse careers/rates remain not computed.
+ """
+
+ params: SSAParameters
+ rates: Any = None
+ interest: Any = None
+ report_conventions: lm.ReportEarningsConventions = field(
+ default_factory=lambda: lm.ReportEarningsConventions(
+ missing_spouse=lm.MissingSpousePolicy.NOT_COMPUTED
+ )
+ )
+
+
+def load_committed_artifact(path: Path = BASE_ARTIFACT) -> dict[str, Any]:
+ """Read only the pinned model artifact; never comparator evidence."""
+
+ assert_sha256(path, BASE_ARTIFACT_SHA256)
+ artifact = json.loads(path.read_text())
+ if artifact["cohort_provenance"]["input_frames_sha256"] != (
+ INPUT_FRAMES_SHA256
+ ):
+ raise GroupBreakdownRefusal("committed input digest differs")
+ return artifact
+
+
+def check_observations(members: pd.DataFrame) -> None:
+ """Refuse held-out births, repeated observations and invalid ages."""
+
+ if members.empty or members["observation_id"].duplicated().any():
+ raise GroupBreakdownRefusal("empty or duplicate observations")
+ births = members["birth_year"]
+ ages = members["member_age"]
+ if members[["observation_id", "person_id"]].isna().any().any():
+ raise GroupBreakdownRefusal("missing observation or person identity")
+ if (
+ births.isna().any()
+ or not births.between(1936, 1945).all()
+ or not (births % 1 == 0).all()
+ ):
+ raise GroupBreakdownRefusal("only 1936--1945 observations allowed")
+ if not ages.isin((66, 67, 68)).all():
+ raise GroupBreakdownRefusal("Track U observations must be 66--68")
+
+
+def reexecute_track_u(
+ inputs: age67.Age67Inputs,
+ params: runner.TrackUParameters,
+ *,
+ data_provenance: str,
+ original_pointer: str | None,
+) -> Replay:
+ """Run each frozen row exactly once, retaining its cohort outputs.
+
+ Calls the frozen _compute_row directly (runner.py:479--559) so the
+ income, tabulation, pending decisions and full design stay identical.
+ The original pointer is used solely to reproduce old metadata.
+ """
+
+ checks = runner._check_inputs(
+ inputs, data_provenance, original_pointer, False
+ )
+ runner._check_parameters(params, data_provenance)
+ rows = {}
+ headline = runner.headline_row(inputs)
+ primary = None
+ for row_id, row in REGISTERED_ROWS.items():
+ if not row.built or runner.blocked_waves(row, inputs):
+ raise GroupBreakdownRefusal(f"registered row {row_id} blocked")
+ result, members, adjusted, cohort = runner._compute_row(
+ row,
+ inputs,
+ params,
+ data_provenance=data_provenance,
+ registration_pointer=original_pointer,
+ allow_blocked=False,
+ tabulation_config=None,
+ )
+ check_observations(members)
+ rows[row_id] = ReplayedRow(
+ result,
+ members,
+ adjusted,
+ runner.design_frame(inputs, row.age67_spec()),
+ )
+ if row_id == headline:
+ primary = (members, adjusted, cohort)
+ if primary is None:
+ raise GroupBreakdownRefusal("headline not computed")
+ members, adjusted, cohort = primary
+ if data_provenance == ap.INVENTED:
+ checks["invented_cohort"] = runner.invented.check_invented_cohort(
+ cohort
+ )
+ evidence = {
+ "cohort_provenance": {
+ key: value
+ for key, value in cohort.provenance.items()
+ if key != "psid_files_sha256"
+ },
+ "psid_files_sha256": dict(
+ cohort.provenance.get("psid_files_sha256") or {}
+ ),
+ "input_checks": checks,
+ "rows": {key: value.result for key, value in rows.items()},
+ "headline": {
+ "row": headline,
+ "rule": runner.HEADLINE_RULE,
+ "wealth_refused_waves": sorted(inputs.wealth_refusals),
+ },
+ "parameters": params.provenance(),
+ "f17_diagnostics": {
+ "row": headline,
+ "official_concept_poverty_rate": runner.official_concept_rates(
+ members, adjusted
+ ),
+ "components": diagnostics.component_diagnostics(cohort, inputs),
+ },
+ }
+ return Replay(evidence, rows)
+
+
+def _join(left: pd.DataFrame, right: pd.DataFrame, key: str) -> pd.DataFrame:
+ """An exact left join: no lost people, repeats or changed row order."""
+
+ if right[key].duplicated().any():
+ raise GroupBreakdownRefusal(f"duplicate side-frame {key}")
+ if not set(left[key]).issubset(set(right[key])):
+ raise GroupBreakdownRefusal(f"side frame missing {key}")
+ overlap = (set(left) & set(right)) - {key}
+ if overlap:
+ raise GroupBreakdownRefusal(f"ambiguous side-frame columns {overlap}")
+ result = left.merge(
+ right, on=key, how="left", validate="many_to_one", sort=False
+ )
+ if not left[key].tolist() == result[key].tolist():
+ raise GroupBreakdownRefusal("join changed observation order")
+ result.attrs = deepcopy(left.attrs)
+ return result
+
+
+def _attribute_frame(
+ members: pd.DataFrame, loader: Callable[..., ga.GroupAttributes]
+) -> tuple[pd.DataFrame, list[dict[str, Any]]]:
+ """Resolve education at each observation wave through G1."""
+
+ frames, provenance = [], []
+ for wave, part in members.groupby("wave", sort=True):
+ attributes = loader(
+ part["person_id"].drop_duplicates().tolist(),
+ anchor_waves=(int(wave),),
+ )
+ if attributes.provenance.get("content_sha256") != ga.content_sha256(
+ attributes.frame
+ ):
+ raise GroupBreakdownRefusal("group attribute seal differs")
+ side = attributes.frame[
+ [
+ "person_id",
+ "race_ethnicity_mint8",
+ "country_of_birth_mint8",
+ "education_years",
+ ]
+ ].copy()
+ side["race_ethnicity"] = side.pop("race_ethnicity_mint8").replace(
+ {
+ c.label: c.key
+ for c in gb.MINT8_SCHEME.dimension("race_ethnicity").categories
+ }
+ )
+ side["country_of_birth_group"] = side.pop(
+ "country_of_birth_mint8"
+ ).replace(
+ {
+ c.label: c.key
+ for c in gb.MINT8_SCHEME.dimension(
+ "country_of_birth"
+ ).categories
+ }
+ )
+ frames.append(
+ _join(part[["observation_id", "person_id"]], side, "person_id")
+ )
+ provenance.append(dict(attributes.provenance))
+ combined = pd.concat(frames, ignore_index=True).drop(columns="person_id")
+ return _join(members, combined, "observation_id"), provenance
+
+
+def lifetime_frame(
+ members: pd.DataFrame,
+ inputs: age67.Age67Inputs,
+ supplied: LifetimeInputs | None,
+) -> tuple[pd.DataFrame, dict[str, Any]]:
+ """Compute all five lifetime measures only with G2, by income year."""
+
+ side = members[["observation_id"]].copy()
+ if supplied is None:
+ for name in MEASURES:
+ side[name] = np.nan
+ return side, {
+ "status": "not computed",
+ "reason": "SSA lifetime inputs not supplied",
+ }
+ careers = inputs.observed_earnings[
+ ["person_id", "period", "earnings"]
+ ].rename(columns={"period": "year"})
+ rates = supplied.rates or lm.load_oasdi_tax_rates(
+ basis=lm.TaxRateBasis.EMPLOYEE_EMPLOYER_PAID
+ )
+ interest = supplied.interest or lm.load_trust_fund_interest_rates()
+ history_ids = set(inputs.marriage_history["person_id"])
+ episodes = psid2010._episodes_with_separation(inputs.marriage_history)
+ frames, records = [], []
+ for year, part in members.groupby("income_year", sort=True):
+ persons = part[["person_id", "birth_year"]].drop_duplicates()
+ careers_to_date = careers[careers["year"] <= int(year)]
+ results = {
+ MEASURES[0]: lm.initial_aime_at_62(
+ careers_to_date,
+ persons,
+ supplied.params,
+ analysis_year=int(year),
+ convention=lm.AIME_CONVENTIONS["mint8_initial_aime"],
+ )
+ }
+ for shared in (False, True):
+ suffix = "_shared" if shared else ""
+ results[f"lifetime_payroll_tax_quintile{suffix}"] = (
+ lm.lifetime_payroll_tax_pv_at_62(
+ careers_to_date,
+ persons,
+ supplied.params,
+ shared=shared,
+ rates=rates,
+ interest=interest,
+ marriage_episodes=episodes if shared else None,
+ missing_spouse=lm.MissingSpousePolicy.NOT_COMPUTED,
+ missing_rate=lm.MissingRatePolicy.NOT_COMPUTED,
+ marriage_history_person_ids=history_ids,
+ )
+ )
+ results[f"lifetime_earnings_{'shared' if shared else 'own'}"] = (
+ lm.report_average_indexed_earnings_22_62(
+ careers_to_date,
+ persons,
+ supplied.params,
+ shared=shared,
+ marriage_episodes=episodes if shared else None,
+ conventions=supplied.report_conventions,
+ marriage_history_person_ids=history_ids,
+ )
+ )
+ joined = part[["observation_id", "person_id"]].copy()
+ for name, measure in results.items():
+ value = (
+ "aime"
+ if name == MEASURES[0]
+ else (
+ "pv_at_62"
+ if "payroll" in name
+ else "average_indexed_earnings"
+ )
+ )
+ frame = measure.frame.copy()
+ unavailable = frame["status"] != lm.COMPUTED
+ exclusions = {}
+ if name.endswith("shared"):
+ absent_history = frame["person_id"].map(
+ lambda pid: pid not in history_ids
+ )
+ unavailable |= absent_history
+ exclusions["marriage_history_absent"] = int(
+ absent_history.sum()
+ )
+ for count in (
+ "years_marital_unknown",
+ "married_years_spouse_year_absent",
+ "married_years_own_year_absent",
+ "married_years_spouse_record_disagrees",
+ "years_multiple_marriages_in_force",
+ ):
+ unavailable |= frame[count] > 0
+ exclusions[count] = int((frame[count] > 0).sum())
+ frame.loc[unavailable, value] = np.nan
+ joined = _join(
+ joined,
+ frame[["person_id", value]].rename(columns={value: name}),
+ "person_id",
+ )
+ records.append(
+ {
+ "income_year": int(year),
+ "dimension": name,
+ "provenance": dict(measure.provenance),
+ "n_unavailable": int(unavailable.sum()),
+ "adapter_exclusion_counts": exclusions,
+ }
+ )
+ frames.append(joined.drop(columns="person_id"))
+ return _join(
+ side, pd.concat(frames, ignore_index=True), "observation_id"
+ ), {
+ "records": records,
+ "coverage": (
+ "observed head/spouse earnings only; missing ages unfilled; "
+ "careers through each observation income year"
+ ),
+ "shared_history": (
+ "absent marriage history is unavailable, not never married; "
+ "unknown marital years, missing partner years and ambiguous "
+ "marriage links are unavailable"
+ ),
+ }
+
+
+def _scheme(row_id: str) -> gb.CategoryScheme:
+ append = list(gb.LIFETIME_DIMENSIONS)
+ for kind in ("own", "shared"):
+ record = ut.NOT_COMPUTED_REPORT_ROWS[f"lifetime_earnings_{kind}"]
+ append.append(
+ gb.Dimension(
+ key=f"lifetime_earnings_{kind}",
+ label=record["section"],
+ kind=gb.QUINTILE,
+ categories=tuple(
+ gb.Category(f"quintile_{i}", label, rank=i)
+ for i, label in enumerate(record["rows"], 1)
+ ),
+ source=(
+ "estimates/uniform_cut_tabulation.py "
+ "NOT_COMPUTED_REPORT_ROWS"
+ ),
+ notes=(
+ "post hoc convention: 1st is lowest; cut over this "
+ "row's observation sample",
+ ),
+ )
+ )
+ replace = {}
+ if row_id == "U1":
+ replace["age"] = gb.Dimension(
+ key="age",
+ label="Age",
+ kind=gb.BAND,
+ categories=tuple(
+ gb.Category(f"age_{age}", str(age), lower=age, upper=age)
+ for age in (66, 67, 68)
+ ),
+ source=(
+ "uniform_cut_track_u/rows.py U1; "
+ "cohorts/age67.py observation_plan"
+ ),
+ notes=("U1 observes 66/67/68; U0 is exact age 67",),
+ )
+ return gb.derive_scheme(
+ gb.MINT8_SCHEME,
+ scheme_id=f"track_u_{row_id.lower().replace('-', '_')}_posthoc",
+ title=(
+ "Exercise 2 MINT8 groups with cohort and Report lifetime measures"
+ ),
+ append=append,
+ replace=replace,
+ population="Track U registered row, births 1936--1945 only",
+ notes=(
+ "household income quintiles use the analysis unit's baseline "
+ "income, weighted across the observation sample; lifetime "
+ "MINT quintiles within birth decades",
+ ),
+ )
+
+
+def tabulate_row_groups(
+ row_id: str,
+ replayed: ReplayedRow,
+ *,
+ inputs: age67.Age67Inputs,
+ attribute_loader: Callable[..., ga.GroupAttributes],
+ lifetime_inputs: LifetimeInputs | None,
+ data_provenance: str,
+ registration_pointer: str | None,
+) -> dict[str, Any]:
+ """Join verified cohort outputs and delegate every new cell to G3."""
+
+ members = replayed.members
+ check_observations(members)
+ frame, provenance = _attribute_frame(members, attribute_loader)
+ measures, lifetime_provenance = lifetime_frame(
+ members, inputs, lifetime_inputs
+ )
+ frame = _join(frame, measures, "observation_id")
+ joined = ut.tabulation_rows(members, replayed.adjusted)
+ frame = _join(
+ joined,
+ frame.drop(
+ columns=[c for c in frame if c in joined and c != "observation_id"]
+ ),
+ "observation_id",
+ )
+ adjusted_side = replayed.adjusted[
+ [
+ "observation_id",
+ "baseline_income",
+ "money_income",
+ "threshold",
+ "cut",
+ "ssi_offset",
+ "ssi_new",
+ ]
+ ]
+ frame = _join(frame, adjusted_side, "observation_id")
+ frame["birth_cohort"] = lm.ten_year_birth_cohort(frame["birth_year"])
+ frame["benefit_type_group"] = (
+ members["benefit_type"].to_numpy()
+ if "benefit_type" in members
+ else None
+ )
+ scheme = _scheme(row_id)
+ columns = {
+ "sex": "sex",
+ "race_ethnicity": "race_ethnicity",
+ "country_of_birth": "country_of_birth_group",
+ "age": "member_age",
+ "marital_status": "marital_status_4",
+ "education": "education_years",
+ "poverty_status": "poverty_status_group",
+ "household_income_quintile": "group_income",
+ "benefit_type": "benefit_type_group",
+ **{name: name for name in MEASURES},
+ }
+ result = {}
+ seeds = replayed.result["tabulation"]["config"]["floor_seeds"]
+ for concept in ("adjusted", "reported_money_official_style"):
+ rows = frame.copy()
+ if concept == "adjusted":
+ rows["group_income"] = rows["baseline_income"]
+ else:
+ if "total_family_income" not in rows:
+ result[concept] = {
+ "status": "not computed",
+ "reason": (
+ "cohort output lacks reported family money income"
+ ),
+ "labels": [*ap.OUTPUT_LABELS, *POSTHOC_LABELS],
+ }
+ continue
+ rows["group_income"] = rows["total_family_income"]
+ rows["poor_baseline"] = (
+ rows["total_family_income"] < rows["threshold"]
+ )
+ rows["poor_reform"] = (
+ rows["total_family_income"]
+ - rows["cut"]
+ + rows["ssi_offset"]
+ + rows["ssi_new"]
+ ) < rows["threshold"]
+ rows["poverty_status_group"] = np.where(
+ rows["poor_baseline"], "in_poverty", "above_poverty"
+ )
+ assignment = gb.assign_groups(
+ rows,
+ scheme,
+ columns,
+ key_columns=("observation_id",),
+ quintile_partitions={
+ name: ("birth_cohort",) for name in MEASURES[:3]
+ },
+ unclassified_codes={"marital_status": ("unclassified",)},
+ )
+ result[concept] = gb.tabulate_poverty_breakdown(
+ rows,
+ assignment,
+ design=replayed.design,
+ data_provenance=data_provenance,
+ id_column="observation_id",
+ floor_seeds=seeds,
+ registration_pointer=registration_pointer,
+ labels=ap.OUTPUT_LABELS,
+ post_hoc_labels=POSTHOC_LABELS,
+ statistic_id=ut.STATISTIC_ID,
+ upstream={
+ "row": row_id,
+ "concept": concept,
+ "official_style_caveat": (
+ "uses reported TOTAL FAMILY INCOME and the frozen "
+ "row's 65+ threshold, cut and SSI response; "
+ "head/wife sensitivity units can differ from the "
+ "family; not the Census householder-age convention"
+ ),
+ },
+ ).as_dict()
+ return {
+ "labels": [*ap.OUTPUT_LABELS, *POSTHOC_LABELS],
+ "groups": result,
+ "attributes_provenance": provenance,
+ "lifetime_provenance": lifetime_provenance,
+ "frozen_row": dict(replayed.result.get("row", {})),
+ "age67_spec": dict(replayed.result.get("age67_spec", {})),
+ "income_spec": dict(replayed.result.get("income_spec", {})),
+ "not_computed": (
+ {}
+ if "benefit_type" in members
+ else {
+ "benefit_type": (
+ "Track U cohort outputs have no Social Security type "
+ "flags; unavailable rows retained as unclassified"
+ )
+ }
+ ),
+ "age_note": (
+ "U1 age dimension is 66/67/68; "
+ "U0 and remaining rows are exact age 67"
+ ),
+ }
+
+
+def build_uniform_cut_groups(
+ inputs: age67.Age67Inputs,
+ params: runner.TrackUParameters,
+ committed: Mapping[str, Any],
+ *,
+ data_provenance: str,
+ registration_pointer: str | None = None,
+ attribute_loader: Callable[..., ga.GroupAttributes] | None = None,
+ lifetime_inputs: LifetimeInputs | None = None,
+ reexecute: Callable[..., Replay] = reexecute_track_u,
+) -> dict[str, Any]:
+ """Refuse before groups on any mismatch; this function writes nothing."""
+
+ actual_digest = age67.input_frames_sha256(inputs)
+ expected_digest = committed["cohort_provenance"]["input_frames_sha256"]
+ if actual_digest != expected_digest or (
+ data_provenance == ap.REGISTERED_REAL
+ and expected_digest != INPUT_FRAMES_SHA256
+ ):
+ raise GroupBreakdownRefusal("Track U input digest mismatch")
+ if data_provenance != committed["data_provenance"]:
+ raise GroupBreakdownRefusal("committed data provenance differs")
+ if data_provenance == ap.REGISTERED_REAL:
+ if (
+ not registration_pointer
+ or not runner.REGISTRATION_POINTER.fullmatch(registration_pointer)
+ ):
+ raise GroupBreakdownRefusal(
+ "groups require a new issue #42 pointer"
+ )
+ if registration_pointer == committed.get("registration_pointer"):
+ raise GroupBreakdownRefusal(
+ "groups require a new registration, not the original pointer"
+ )
+ assert_exact_cells(
+ load_committed_artifact(),
+ committed,
+ path="pinned parent artifact",
+ )
+ assert_exact_cells(
+ committed["psid_files_sha256"],
+ inputs.provenance["psid_files_sha256"],
+ path="PSID file hashes",
+ )
+ replay = reexecute(
+ inputs,
+ params,
+ data_provenance=data_provenance,
+ original_pointer=committed.get("registration_pointer"),
+ )
+ expected_rows = set(REGISTERED_ROWS)
+ if (
+ set(replay.rows) != expected_rows
+ or set(committed["rows"]) != expected_rows
+ ):
+ raise GroupBreakdownRefusal("every registered row is required")
+ for row in replay.rows.values():
+ check_observations(row.members)
+ check = assert_exact_cells(
+ {key: committed[key] for key in replay.evidence},
+ replay.evidence,
+ path="frozen Track U reproduction",
+ )
+ # All reproduction checks finish before this first attribute read.
+ if attribute_loader is None:
+ side_inputs = ga.load_group_attribute_inputs()
+
+ def attribute_loader(person_ids, *, anchor_waves):
+ return ga.build_group_attributes(
+ side_inputs, person_ids, anchor_waves=anchor_waves
+ )
+
+ results = {
+ row_id: tabulate_row_groups(
+ row_id,
+ row,
+ inputs=inputs,
+ attribute_loader=attribute_loader,
+ lifetime_inputs=lifetime_inputs,
+ data_provenance=data_provenance,
+ registration_pointer=registration_pointer,
+ )
+ for row_id, row in replay.rows.items()
+ }
+ return {
+ "header": (
+ INVENTED_HEADER
+ if data_provenance == ap.INVENTED
+ else "REGISTERED POST HOC GROUP BREAKDOWN - Track U exercise 2"
+ ),
+ "schema_version": SCHEMA_VERSION,
+ "statistic_id": ut.STATISTIC_ID,
+ "data_provenance": data_provenance,
+ "registration_pointer": registration_pointer,
+ "labels": [
+ *(
+ [INVENTED_HEADER, gb.INVENTED_DATA_LABEL]
+ if data_provenance == ap.INVENTED
+ else []
+ ),
+ *ap.OUTPUT_LABELS,
+ *POSTHOC_LABELS,
+ ],
+ "reproduction": {
+ "comparison": check.comparison,
+ "absolute_tolerance": check.absolute_tolerance,
+ "relative_tolerance": check.relative_tolerance,
+ "compared_leaf_cells": check.compared_leaf_cells,
+ "input_frames_sha256": actual_digest,
+ "original_registration_pointer": committed.get(
+ "registration_pointer"
+ ),
+ "all_registered_rows_checked_before_groups": True,
+ },
+ "named_deltas": list(runner.NAMED_DELTAS),
+ "rows": results,
+ }
diff --git a/tests/group_breakdowns/test_common.py b/tests/group_breakdowns/test_common.py
new file mode 100644
index 00000000..48ae43f4
--- /dev/null
+++ b/tests/group_breakdowns/test_common.py
@@ -0,0 +1,333 @@
+"""INVENTED DATA tests of parent identity and one-shot registration guards.
+
+All JSON documents and checkout states below are invented. No committed
+outcome artifact or microdata is opened. The static tier classifier assigns
+artifact because the registration tests name temporary runs JSON paths.
+"""
+
+import hashlib
+import json
+import math
+from pathlib import Path
+
+import pytest
+from hypothesis import given
+from hypothesis import strategies as st
+
+from populace_dynamics.group_breakdowns.common import (
+ POSTHOC_LABELS,
+ GroupBreakdownRefusal,
+ assert_exact_cells,
+ assert_sha256,
+ environment,
+ preflight,
+ write_artifact_pair,
+)
+
+INVENTED_POINTER = (
+ "https://github.com/PolicyEngine/microcosm-dynamics/issues/42"
+ "#issuecomment-123456789"
+)
+INVENTED_COMMIT = "a" * 40
+_SCALARS = st.one_of(
+ st.none(),
+ st.booleans(),
+ st.integers(),
+ st.floats(allow_nan=False, allow_infinity=False),
+ st.text(),
+)
+_JSON = st.recursive(
+ _SCALARS,
+ lambda children: st.one_of(
+ st.lists(children, max_size=4),
+ st.dictionaries(st.text(), children, max_size=4),
+ ),
+ max_leaves=12,
+)
+
+
+def _invented_git(*args):
+ return INVENTED_COMMIT if args == ("rev-parse", "HEAD") else ""
+
+
+def _preflight(tmp_path, **overrides):
+ arguments = {
+ "registration_pointer": INVENTED_POINTER,
+ "registered_commit": INVENTED_COMMIT,
+ "output": tmp_path / "runs" / "invented_groups_posthoc_v1.json",
+ "root": tmp_path,
+ "git": _invented_git,
+ }
+ arguments.update(overrides)
+ return preflight(**arguments)
+
+
+@given(_JSON)
+def test_exact_identity_after_json_round_trip(invented):
+ """Every finite invented JSON tree reproduces all of its leaves."""
+
+ restored = json.loads(json.dumps(invented, allow_nan=False))
+ checked = assert_exact_cells(invented, restored)
+ assert checked.compared_leaf_cells >= 0
+ assert checked.absolute_tolerance == checked.relative_tolerance == 0.0
+
+
+@given(_JSON, _JSON)
+def test_identity_differential_against_canonical_json(left, right):
+ """The independent JSON representation has the same exact semantics."""
+
+ same = json.dumps(left, sort_keys=True, allow_nan=False) == json.dumps(
+ right, sort_keys=True, allow_nan=False
+ )
+ if same:
+ assert_exact_cells(left, right)
+ else:
+ with pytest.raises(GroupBreakdownRefusal):
+ assert_exact_cells(left, right)
+
+
+@given(st.floats(allow_nan=False, allow_infinity=False))
+def test_one_binary64_step_never_gets_a_summation_tolerance(value):
+ changed = math.nextafter(value, math.inf)
+ with pytest.raises(GroupBreakdownRefusal):
+ assert_exact_cells(
+ {"invented_rate": value}, {"invented_rate": changed}
+ )
+
+
+@pytest.mark.parametrize(
+ ("committed", "recomputed"),
+ [
+ (0.0, -0.0),
+ (True, 1),
+ (1, 1.0),
+ ({"x": None}, {}),
+ ({}, {"x": None}),
+ ([1], [1, 2]),
+ (float("nan"), float("nan")),
+ (float("inf"), float("inf")),
+ ({1: "x"}, {1: "x"}),
+ (object(), object()),
+ ],
+)
+def test_structural_and_nonfinite_mismatches_refuse(committed, recomputed):
+ with pytest.raises(GroupBreakdownRefusal):
+ assert_exact_cells(committed, recomputed)
+
+
+def test_mapping_order_and_json_array_semantics():
+ check = assert_exact_cells(
+ {"a": (1, 2), "b": "x"}, {"b": "x", "a": [1, 2]}
+ )
+ assert check.compared_leaf_cells == 3
+
+
+def test_mismatch_reports_path_without_cell_values():
+ with pytest.raises(GroupBreakdownRefusal) as caught:
+ assert_exact_cells({"invented": [123]}, {"invented": [456]})
+ assert "cells.invented[0]" in str(caught.value)
+ assert "123" not in str(caught.value)
+ assert "456" not in str(caught.value)
+
+
+def test_labels_state_the_post_hoc_report_only_status():
+ assert POSTHOC_LABELS == (
+ "registered, one-shot, post hoc, not blind",
+ "report-only",
+ )
+
+
+def test_preflight_returns_bound_new_pair(tmp_path):
+ result = _preflight(tmp_path)
+ assert result.git_head == result.registered_commit == INVENTED_COMMIT
+ assert result.git_clean
+ assert result.output_path.name == "invented_groups_posthoc_v1.json"
+ assert result.sidecar_path.name == "invented_groups_posthoc_v1.env.json"
+ assert not result.output_path.exists()
+ assert not result.sidecar_path.exists()
+
+
+@pytest.mark.parametrize(
+ "pointer",
+ [
+ "",
+ INVENTED_POINTER.replace("/42#", "/41#"),
+ INVENTED_POINTER.replace("https://", "http://"),
+ INVENTED_POINTER.replace("microcosm-dynamics", "other-project"),
+ INVENTED_POINTER + "?other=1",
+ INVENTED_POINTER + "\n",
+ INVENTED_POINTER.replace("#issuecomment-", "#discussioncomment-"),
+ ],
+)
+def test_registration_pointer_refuses_before_git(tmp_path, pointer):
+ def must_not_call_git(*args):
+ pytest.fail("invalid pointer must refuse before reading git state")
+
+ with pytest.raises(GroupBreakdownRefusal, match="issue #42"):
+ _preflight(
+ tmp_path, registration_pointer=pointer, git=must_not_call_git
+ )
+
+
+@pytest.mark.parametrize("commit", ["", "a" * 39, "A" * 40, "g" * 40])
+def test_registered_commit_must_be_full_lowercase_sha(tmp_path, commit):
+ with pytest.raises(GroupBreakdownRefusal, match="40-hex"):
+ _preflight(tmp_path, registered_commit=commit)
+
+
+def test_registered_commit_must_equal_head(tmp_path):
+ with pytest.raises(GroupBreakdownRefusal, match="HEAD"):
+ _preflight(tmp_path, registered_commit="b" * 40)
+
+
+@given(st.text(min_size=1).filter(lambda value: bool(value.strip())))
+def test_any_dirty_status_refuses_before_writing(status):
+ """A tracked or untracked status has no permitted silent exception."""
+
+ def dirty_git(*args):
+ return INVENTED_COMMIT if args == ("rev-parse", "HEAD") else status
+
+ with pytest.raises(GroupBreakdownRefusal, match="clean"):
+ preflight(
+ registration_pointer=INVENTED_POINTER,
+ registered_commit=INVENTED_COMMIT,
+ output=Path("invented_groups_posthoc_v1.json"),
+ root=Path("INVENTED-NONEXISTENT-CHECKOUT"),
+ git=dirty_git,
+ )
+
+
+@pytest.mark.parametrize(
+ "relative",
+ [
+ "invented_groups_posthoc_v1.json",
+ "runs/invented.json",
+ "runs/subdirectory/invented_groups_posthoc_v1.json",
+ "runs/../invented_groups_posthoc_v1.json",
+ "runs/.invented_groups_posthoc_v1.json",
+ ],
+)
+def test_output_must_be_a_named_artifact_in_checkout_runs(tmp_path, relative):
+ with pytest.raises(GroupBreakdownRefusal, match="output must be"):
+ _preflight(tmp_path, output=Path(relative))
+
+
+@pytest.mark.parametrize("sidecar", [False, True])
+def test_either_existing_member_blocks_registration(tmp_path, sidecar):
+ path = tmp_path / "runs" / "invented_groups_posthoc_v1.json"
+ path.parent.mkdir()
+ if sidecar:
+ path = path.with_suffix(".env.json")
+ path.write_text("INVENTED existing bytes", encoding="utf-8")
+ with pytest.raises(GroupBreakdownRefusal, match="one-shot"):
+ _preflight(tmp_path)
+ assert path.read_text(encoding="utf-8") == "INVENTED existing bytes"
+
+
+def test_symlinked_runs_cannot_escape_checkout(tmp_path):
+ target = tmp_path / "elsewhere"
+ target.mkdir()
+ (tmp_path / "runs").symlink_to(target, target_is_directory=True)
+ with pytest.raises(GroupBreakdownRefusal, match="symlink"):
+ _preflight(tmp_path)
+
+
+def test_dangling_sidecar_symlink_blocks_registration(tmp_path):
+ output = tmp_path / "runs" / "invented_groups_posthoc_v1.json"
+ output.parent.mkdir()
+ output.with_suffix(".env.json").symlink_to(tmp_path / "absent")
+ with pytest.raises(GroupBreakdownRefusal, match="one-shot"):
+ _preflight(tmp_path)
+
+
+def test_write_pair_binds_the_exact_artifact_bytes(tmp_path):
+ output = tmp_path / "invented.json"
+ written = write_artifact_pair(
+ output=output,
+ artifact={"header": "INVENTED DATA - NOT A COMPARISON", "x": 0.0},
+ environment={"python": "INVENTED VERSION"},
+ )
+ assert written == (output, output.with_suffix(".env.json"))
+ sidecar = json.loads(written[1].read_text(encoding="utf-8"))
+ assert sidecar["artifact"] == output.name
+ assert (
+ sidecar["artifact_sha256"]
+ == hashlib.sha256(output.read_bytes()).hexdigest()
+ )
+ assert sidecar["environment"] == {"python": "INVENTED VERSION"}
+
+
+@pytest.mark.parametrize("sidecar", [False, True])
+def test_write_collision_preserves_existing_member_and_removes_new(
+ tmp_path, sidecar
+):
+ output = tmp_path / "invented.json"
+ existing = output.with_suffix(".env.json") if sidecar else output
+ existing.write_text("INVENTED original bytes", encoding="utf-8")
+ with pytest.raises(FileExistsError):
+ write_artifact_pair(output=output, artifact={"x": 1}, environment={})
+ assert existing.read_text(encoding="utf-8") == "INVENTED original bytes"
+ absent = output if sidecar else output.with_suffix(".env.json")
+ assert not absent.exists()
+
+
+@pytest.mark.parametrize("in_environment", [False, True])
+def test_nonfinite_json_refuses_before_either_file_is_created(
+ tmp_path, in_environment
+):
+ output = tmp_path / "invented.json"
+ artifact = {} if in_environment else {"x": float("nan")}
+ env = {"x": float("inf")} if in_environment else {}
+ with pytest.raises(ValueError):
+ write_artifact_pair(output=output, artifact=artifact, environment=env)
+ assert not output.exists()
+ assert not output.with_suffix(".env.json").exists()
+
+
+def test_failed_sidecar_open_rolls_back_new_artifact(tmp_path, monkeypatch):
+ output = tmp_path / "invented.json"
+ original_open = Path.open
+
+ def failing_open(path, *args, **kwargs):
+ if path == output.with_suffix(".env.json"):
+ raise OSError("INVENTED storage failure")
+ return original_open(path, *args, **kwargs)
+
+ monkeypatch.setattr(Path, "open", failing_open)
+ with pytest.raises(OSError, match="INVENTED storage failure"):
+ write_artifact_pair(output=output, artifact={"x": 1}, environment={})
+ assert not output.exists()
+ assert not output.with_suffix(".env.json").exists()
+
+
+def test_environment_delegates_to_track_a_resolver(tmp_path):
+ script = tmp_path / "scripts" / "run_track_a_registered.py"
+ script.parent.mkdir()
+ script.write_text(
+ "def _environment(*, ssa_parameters_revision):\n"
+ " return {'delegated_revision': ssa_parameters_revision}\n"
+ "def main():\n"
+ " raise AssertionError('registered main must not run')\n",
+ encoding="utf-8",
+ )
+ assert environment(ssa_parameters_revision="INVENTED", root=tmp_path) == {
+ "delegated_revision": "INVENTED"
+ }
+
+
+def test_sha_pin_refuses_changed_invented_artifact(tmp_path):
+ path = tmp_path / "invented.json"
+ path.write_bytes(b"INVENTED file")
+ digest = hashlib.sha256(path.read_bytes()).hexdigest()
+ assert assert_sha256(path, digest) == digest
+ path.write_bytes(b"INVENTED changed file")
+ with pytest.raises(GroupBreakdownRefusal, match="SHA-256 mismatch"):
+ assert_sha256(path, digest)
+
+
+@pytest.mark.parametrize("digest", ["", "a" * 63, "A" * 64, "g" * 64])
+def test_sha_pin_requires_full_lowercase_sha_before_file_read(
+ tmp_path, digest
+):
+ with pytest.raises(GroupBreakdownRefusal, match="64-hex"):
+ assert_sha256(tmp_path / "absent", digest)
diff --git a/tests/group_breakdowns/test_registration_script.py b/tests/group_breakdowns/test_registration_script.py
new file mode 100644
index 00000000..dd72715d
--- /dev/null
+++ b/tests/group_breakdowns/test_registration_script.py
@@ -0,0 +1,400 @@
+"""INVENTED registration-state tests; no checkout or microdata is opened.
+
+Subprocess calls, parameter loaders, parent evidence and outcome writers
+are replaced. Only the entry-point source and invented temporary files
+are read. This module belongs to tier unit.
+"""
+
+from __future__ import annotations
+
+import copy
+import importlib.util
+import json
+from pathlib import Path
+from types import SimpleNamespace
+
+import pytest
+
+from populace_dynamics.group_breakdowns.common import (
+ GroupBreakdownRefusal,
+ RegisteredRun,
+)
+
+ROOT = Path(__file__).resolve().parents[2]
+INVENTED_COMMIT = "b" * 40
+# Required production prefix with an explicitly invented full-SHA suffix.
+INVENTED_PARAMETER_HEAD = "a03e82e503" + "0" * 30
+INVENTED_OLD_POINTER = (
+ "https://github.com/PolicyEngine/microcosm-dynamics/issues/42"
+ "#issuecomment-1"
+)
+INVENTED_NEW_POINTER = INVENTED_OLD_POINTER[:-1] + "2"
+INVENTED_ENVIRONMENT = {
+ "python": "INVENTED Python version",
+ "packages": {
+ "numpy": "INVENTED NumPy version",
+ "pandas": "INVENTED pandas version",
+ "policyengine-social-security-model": "INVENTED project version",
+ },
+ "project": {"name": "INVENTED project", "version": "INVENTED version"},
+}
+
+
+def _entry():
+ source = ROOT / "scripts" / "run_track_u_groups_registered.py"
+ spec = importlib.util.spec_from_file_location(
+ "_invented_track_u_group_registration", source
+ )
+ module = importlib.util.module_from_spec(spec)
+ assert spec.loader is not None
+ spec.loader.exec_module(module)
+ return module
+
+
+def _fake_checkout(entry, monkeypatch, *, head, status):
+ calls = []
+ checkout = Path("INVENTED-NONEXISTENT-PARAMETER-CHECKOUT")
+ monkeypatch.setattr(entry.ss_params, "_resolve_pe_us", lambda _: checkout)
+
+ def run(command, **kwargs):
+ assert command[:4] == [
+ "git",
+ "--no-optional-locks",
+ "-C",
+ str(checkout),
+ ]
+ operation = tuple(command[4:])
+ calls.append(operation)
+ if operation == ("rev-parse", "HEAD"):
+ return SimpleNamespace(stdout=head + "\n")
+ if operation == ("status", "--porcelain", "--untracked-files=no"):
+ return SimpleNamespace(stdout=status)
+ raise AssertionError(f"unexpected git operation: {operation}")
+
+ monkeypatch.setattr(entry.subprocess, "run", run)
+ return calls
+
+
+def test_clean_full_parameter_head_is_used_independent_of_abbreviation(
+ monkeypatch,
+):
+ entry = _entry()
+ calls = _fake_checkout(
+ entry, monkeypatch, head=INVENTED_PARAMETER_HEAD, status=""
+ )
+ assert entry.check_parameter_checkout() == INVENTED_PARAMETER_HEAD
+ assert calls == [
+ ("rev-parse", "HEAD"),
+ ("status", "--porcelain", "--untracked-files=no"),
+ ]
+
+
+@pytest.mark.parametrize(
+ "head",
+ [
+ "a03e82e5",
+ "a03e82e503",
+ "a03e82e503" + "0" * 29,
+ INVENTED_PARAMETER_HEAD + "0",
+ "c" * 40,
+ ],
+)
+def test_short_or_wrong_parameter_head_refuses_before_status(
+ monkeypatch, head
+):
+ entry = _entry()
+ calls = _fake_checkout(entry, monkeypatch, head=head, status="")
+ with pytest.raises(GroupBreakdownRefusal, match="revision a03e82e503"):
+ entry.check_parameter_checkout()
+ assert calls == [("rev-parse", "HEAD")]
+
+
+@pytest.mark.parametrize(
+ "status",
+ [
+ " M parameters/INVENTED.yaml\n",
+ "M parameters/INVENTED.yaml\n",
+ " D parameters/INVENTED.yaml\n",
+ ],
+)
+def test_dirty_parameter_checkout_refuses_changed_tracked_parameters(
+ monkeypatch, status
+):
+ entry = _entry()
+ _fake_checkout(
+ entry, monkeypatch, head=INVENTED_PARAMETER_HEAD, status=status
+ )
+ with pytest.raises(GroupBreakdownRefusal, match="clean tracked files"):
+ entry.check_parameter_checkout()
+
+
+def test_recorded_environment_matches_without_guessing_unrecorded_packages():
+ entry = _entry()
+ actual = copy.deepcopy(INVENTED_ENVIRONMENT)
+ actual["packages"]["scipy"] = "INVENTED unrecorded supplemental version"
+ entry.check_environment(actual, INVENTED_ENVIRONMENT)
+
+
+@pytest.mark.parametrize(
+ ("change", "match"),
+ [
+ ("python", "Python version"),
+ ("numpy", "package numpy"),
+ ("pandas_absent", "package pandas"),
+ ("project_version", "project.version"),
+ ("project_name", "project.name"),
+ ],
+)
+def test_environment_version_and_project_mismatch_refuse(change, match):
+ entry = _entry()
+ actual = copy.deepcopy(INVENTED_ENVIRONMENT)
+ if change == "python":
+ actual["python"] = "INVENTED changed Python"
+ elif change == "numpy":
+ actual["packages"]["numpy"] = "INVENTED changed NumPy"
+ elif change == "pandas_absent":
+ del actual["packages"]["pandas"]
+ elif change == "project_version":
+ actual["project"]["version"] = "INVENTED changed project version"
+ else:
+ actual["project"]["name"] = "INVENTED changed project name"
+ with pytest.raises(GroupBreakdownRefusal, match=match):
+ entry.check_environment(actual, INVENTED_ENVIRONMENT)
+
+
+def test_spent_original_pointer_refuses_before_inputs_or_environment(
+ monkeypatch,
+):
+ entry = _entry()
+ calls = []
+ monkeypatch.setattr(entry.common, "preflight", lambda **kwargs: object())
+ monkeypatch.setattr(
+ entry.groups,
+ "load_committed_artifact",
+ lambda: {"registration_pointer": INVENTED_OLD_POINTER},
+ )
+
+ def forbidden(*args, **kwargs):
+ calls.append("read or write after spent pointer")
+ raise AssertionError("spent registration must refuse first")
+
+ monkeypatch.setattr(entry.age67, "load_age67_inputs", forbidden)
+ monkeypatch.setattr(entry.common, "environment", forbidden)
+ monkeypatch.setattr(entry.common, "assert_sha256", forbidden)
+ monkeypatch.setattr(entry.common, "write_artifact_pair", forbidden)
+ monkeypatch.setattr(entry, "check_parameter_checkout", forbidden)
+ with pytest.raises(GroupBreakdownRefusal, match="new issue #42"):
+ entry.main(
+ [
+ "--registration-pointer",
+ INVENTED_OLD_POINTER,
+ "--registered-commit",
+ INVENTED_COMMIT,
+ ]
+ )
+ assert calls == []
+
+
+@pytest.mark.parametrize(
+ "arguments",
+ [
+ [],
+ ["--registration-pointer", INVENTED_NEW_POINTER],
+ ["--registered-commit", INVENTED_COMMIT],
+ ],
+)
+def test_missing_required_cli_refuses_before_preflight(
+ monkeypatch, capsys, arguments
+):
+ entry = _entry()
+
+ def forbidden(**kwargs):
+ raise AssertionError("missing CLI must not reach preflight")
+
+ monkeypatch.setattr(entry.common, "preflight", forbidden)
+ with pytest.raises(SystemExit) as caught:
+ entry.main(arguments)
+ assert caught.value.code == 2
+ assert "required" in capsys.readouterr().err
+
+
+def _invented_main_state(entry, tmp_path, monkeypatch, actual_environment):
+ """A fake parent and pipeline for testing entry-point guard ordering."""
+
+ monkeypatch.setattr(entry, "ROOT", tmp_path)
+ output = tmp_path / "invented_groups_posthoc_v1.json"
+ state = RegisteredRun(
+ registration_pointer=INVENTED_NEW_POINTER,
+ registered_commit=INVENTED_COMMIT,
+ git_head=INVENTED_COMMIT,
+ git_clean=True,
+ output_path=output,
+ sidecar_path=output.with_suffix(".env.json"),
+ )
+ monkeypatch.setattr(entry.common, "preflight", lambda **kwargs: state)
+ parent_path = tmp_path / "invented_parent.json"
+ monkeypatch.setattr(entry.groups, "BASE_ARTIFACT", parent_path)
+ parent_path.with_suffix(".env.json").write_text(
+ json.dumps({"environment": INVENTED_ENVIRONMENT}), encoding="utf-8"
+ )
+ checks = {"specification_rows": {"INVENTED": True}, "max_rulings": {}}
+ committed = {
+ "registration_pointer": INVENTED_OLD_POINTER,
+ "checks": checks,
+ "specification": {"sha256": "f" * 64},
+ "headline": {"row": "INVENTED"},
+ }
+ monkeypatch.setattr(
+ entry.groups, "load_committed_artifact", lambda: committed
+ )
+ monkeypatch.setattr(
+ entry.common, "assert_sha256", lambda path, expected: expected
+ )
+ monkeypatch.setattr(entry.rows, "specification_block", lambda: {})
+ monkeypatch.setattr(
+ entry.rows,
+ "check_rows_against_block",
+ lambda block: checks["specification_rows"],
+ )
+ monkeypatch.setattr(
+ entry.rows,
+ "check_rulings_against_block",
+ lambda block: checks["max_rulings"],
+ )
+ parent_source = tmp_path / "scripts" / "run_track_u_registered.py"
+ parent_source.parent.mkdir()
+ parent_source.write_text(
+ "def check_specification_ratified(block):\n"
+ " pass\n"
+ "def check_headline(row, inputs):\n"
+ " pass\n",
+ encoding="utf-8",
+ )
+ monkeypatch.setattr(
+ entry, "check_parameter_checkout", lambda: INVENTED_PARAMETER_HEAD
+ )
+ monkeypatch.setattr(
+ entry.ss_params,
+ "load_ssa_parameters",
+ lambda: SimpleNamespace(pe_us_revision="a03e82e5"),
+ )
+ resolver_calls = []
+
+ def resolver(*, ssa_parameters_revision):
+ resolver_calls.append(ssa_parameters_revision)
+ return actual_environment
+
+ monkeypatch.setattr(entry.common, "environment", resolver)
+ return state, resolver_calls
+
+
+def test_environment_refusal_precedes_psid_loader(tmp_path, monkeypatch):
+ entry = _entry()
+ actual = copy.deepcopy(INVENTED_ENVIRONMENT)
+ actual["packages"]["numpy"] = "INVENTED changed NumPy"
+ _, resolver_calls = _invented_main_state(
+ entry, tmp_path, monkeypatch, actual
+ )
+ calls = []
+
+ def forbidden(*args, **kwargs):
+ calls.append("cohort or outcome computation")
+ raise AssertionError("environment mismatch must refuse first")
+
+ monkeypatch.setattr(entry.age67, "load_age67_inputs", forbidden)
+ monkeypatch.setattr(entry.ap, "load_poverty_thresholds", forbidden)
+ monkeypatch.setattr(entry.groups, "build_uniform_cut_groups", forbidden)
+ monkeypatch.setattr(entry.common, "write_artifact_pair", forbidden)
+ with pytest.raises(GroupBreakdownRefusal, match="package numpy"):
+ entry.main(
+ [
+ "--registration-pointer",
+ INVENTED_NEW_POINTER,
+ "--registered-commit",
+ INVENTED_COMMIT,
+ ]
+ )
+ assert resolver_calls == [INVENTED_PARAMETER_HEAD]
+ assert calls == []
+ assert not (tmp_path / "invented_groups_posthoc_v1.json").exists()
+
+
+def test_dirty_parameter_refusal_precedes_parameter_and_psid_loaders(
+ tmp_path, monkeypatch
+):
+ entry = _entry()
+ _invented_main_state(
+ entry, tmp_path, monkeypatch, copy.deepcopy(INVENTED_ENVIRONMENT)
+ )
+ calls = []
+
+ def dirty_checkout():
+ raise GroupBreakdownRefusal(
+ "SSA parameter checkout must have clean tracked files"
+ )
+
+ def forbidden(*args, **kwargs):
+ calls.append("parameter or cohort loader")
+ raise AssertionError("dirty parameter checkout must refuse first")
+
+ monkeypatch.setattr(entry, "check_parameter_checkout", dirty_checkout)
+ monkeypatch.setattr(entry.ss_params, "load_ssa_parameters", forbidden)
+ monkeypatch.setattr(entry.age67, "load_age67_inputs", forbidden)
+ monkeypatch.setattr(entry.common, "environment", forbidden)
+ monkeypatch.setattr(entry.common, "write_artifact_pair", forbidden)
+ with pytest.raises(GroupBreakdownRefusal, match="clean tracked files"):
+ entry.main(
+ [
+ "--registration-pointer",
+ INVENTED_NEW_POINTER,
+ "--registered-commit",
+ INVENTED_COMMIT,
+ ]
+ )
+ assert calls == []
+ assert not (tmp_path / "invented_groups_posthoc_v1.json").exists()
+
+
+def test_full_parameter_head_reaches_environment_provenance(
+ tmp_path, monkeypatch
+):
+ entry = _entry()
+ state, resolver_calls = _invented_main_state(
+ entry, tmp_path, monkeypatch, copy.deepcopy(INVENTED_ENVIRONMENT)
+ )
+ monkeypatch.setattr(entry.ap, "load_poverty_thresholds", lambda: object())
+ monkeypatch.setattr(
+ entry.runner, "committed_parameters", lambda _: object()
+ )
+ monkeypatch.setattr(entry.age67, "load_age67_inputs", lambda: object())
+ monkeypatch.setattr(
+ entry.groups,
+ "build_uniform_cut_groups",
+ lambda *args, **kwargs: {"header": "INVENTED DATA - NOT A COMPARISON"},
+ )
+ written = []
+ monkeypatch.setattr(
+ entry.common,
+ "write_artifact_pair",
+ lambda **kwargs: written.append(kwargs),
+ )
+ assert (
+ entry.main(
+ [
+ "--registration-pointer",
+ INVENTED_NEW_POINTER,
+ "--registered-commit",
+ INVENTED_COMMIT,
+ ]
+ )
+ == 0
+ )
+ assert resolver_calls == [INVENTED_PARAMETER_HEAD]
+ assert len(written) == 1
+ assert written[0]["output"] == state.output_path
+ assert written[0]["environment"] == INVENTED_ENVIRONMENT
+ assert written[0]["artifact"]["run"]["registration_pointer"] == (
+ INVENTED_NEW_POINTER
+ )
+ assert not state.output_path.exists()
+ assert not state.sidecar_path.exists()
diff --git a/tests/group_breakdowns/test_uniform_cut.py b/tests/group_breakdowns/test_uniform_cut.py
new file mode 100644
index 00000000..6e80f572
--- /dev/null
+++ b/tests/group_breakdowns/test_uniform_cut.py
@@ -0,0 +1,245 @@
+"""G4b INVENTED-data tests, with pinned parameter/category captures.
+
+All observations, incomes and attribute reports are invented. Public
+builders transitively read committed "data" / "external" definitions;
+the static tier classifier therefore assigns this module artifact.
+No real-data entry point or microdata reader is executed.
+"""
+
+from __future__ import annotations
+
+import copy
+from dataclasses import replace
+
+import pandas as pd
+import pytest
+from pandas.testing import assert_frame_equal
+
+from populace_dynamics.cohorts import age67
+from populace_dynamics.estimates import adjusted_poverty as ap
+from populace_dynamics.estimates import group_breakdown as gb
+from populace_dynamics.group_breakdowns import uniform_cut as groups
+from populace_dynamics.group_breakdowns.common import (
+ GroupBreakdownRefusal,
+ assert_exact_cells,
+)
+from populace_dynamics.uniform_cut_track_u import invented, runner
+from scripts.track_u_groups_invented import (
+ invented_attribute_loader,
+ invented_lifetime_inputs,
+)
+
+
+@pytest.fixture(scope="module")
+def replay():
+ inputs = invented.invented_age67_inputs(supplement_waves_staged=True)
+ parameters = runner.committed_parameters(
+ invented.invented_poverty_thresholds()
+ )
+ committed = runner.run_track_u(
+ inputs, parameters, data_provenance=ap.INVENTED
+ )
+ retained = groups.reexecute_track_u(
+ inputs,
+ parameters,
+ data_provenance=ap.INVENTED,
+ original_pointer=None,
+ )
+ return inputs, parameters, committed, retained
+
+
+def test_every_frozen_row_and_f17_cell_differential(replay):
+ _, _, committed, retained = replay
+ check = assert_exact_cells(
+ {key: committed[key] for key in retained.evidence},
+ retained.evidence,
+ )
+ assert check.compared_leaf_cells > 1000
+ assert set(retained.rows) == set(groups.REGISTERED_ROWS)
+ assert set(retained.rows["U0"].members.member_age) == {67}
+ assert set(retained.rows["U1"].members.member_age) <= {66, 67, 68}
+
+
+@pytest.mark.parametrize("row_id", ["U0", "U10-F"])
+def test_mismatched_committed_cell_refuses_before_load_or_write(
+ replay,
+ row_id,
+ tmp_path,
+ monkeypatch,
+):
+ inputs, parameters, committed, retained = replay
+ mismatched = copy.deepcopy(committed)
+ mismatched["rows"][row_id]["tabulation"]["cells"][0]["delta"] += 1.0
+ calls = []
+
+ def forbidden_loader(*args, **kwargs):
+ calls.append((args, kwargs))
+ raise AssertionError("group attribute loader called before refusal")
+
+ monkeypatch.setattr(
+ groups.ga, "load_group_attribute_inputs", forbidden_loader
+ )
+ monkeypatch.setattr(groups, "lifetime_frame", forbidden_loader)
+ with pytest.raises(GroupBreakdownRefusal, match="committed cell"):
+ groups.build_uniform_cut_groups(
+ inputs,
+ parameters,
+ mismatched,
+ data_provenance=ap.INVENTED,
+ reexecute=lambda *args, **kwargs: retained,
+ )
+ assert calls == []
+ assert list(tmp_path.iterdir()) == []
+
+
+def test_input_digest_refuses_before_reexecution(replay):
+ inputs, parameters, committed, _ = replay
+ mismatched = copy.deepcopy(committed)
+ mismatched["cohort_provenance"]["input_frames_sha256"] = "0" * 64
+
+ def forbidden(*args, **kwargs):
+ raise AssertionError("reexecuted mismatched inputs")
+
+ with pytest.raises(GroupBreakdownRefusal, match="digest"):
+ groups.build_uniform_cut_groups(
+ inputs,
+ parameters,
+ mismatched,
+ data_provenance=ap.INVENTED,
+ reexecute=forbidden,
+ attribute_loader=forbidden,
+ )
+
+
+def test_missing_registered_row_refuses_before_attributes(replay):
+ inputs, parameters, committed, retained = replay
+ rows = dict(retained.rows)
+ rows.pop("U10-F")
+ with pytest.raises(GroupBreakdownRefusal, match="every registered"):
+ groups.build_uniform_cut_groups(
+ inputs,
+ parameters,
+ committed,
+ data_provenance=ap.INVENTED,
+ reexecute=lambda *args, **kwargs: replace(retained, rows=rows),
+ attribute_loader=lambda *args, **kwargs: pytest.fail("loaded"),
+ )
+
+
+def test_held_out_births_refuse_before_attributes(replay):
+ inputs, parameters, committed, retained = replay
+ rows = dict(retained.rows)
+ members = rows["U0"].members.copy()
+ members.loc[members.index[0], "birth_year"] = 1946
+ rows["U0"] = replace(rows["U0"], members=members)
+ with pytest.raises(GroupBreakdownRefusal, match="1936--1945"):
+ groups.build_uniform_cut_groups(
+ inputs,
+ parameters,
+ committed,
+ data_provenance=ap.INVENTED,
+ reexecute=lambda *args, **kwargs: replace(retained, rows=rows),
+ attribute_loader=lambda *args, **kwargs: pytest.fail("loaded"),
+ )
+
+
+@pytest.fixture(scope="module")
+def u0_breakdown(replay):
+ inputs, _, _, retained = replay
+ return groups.tabulate_row_groups(
+ "U0",
+ retained.rows["U0"],
+ inputs=inputs,
+ attribute_loader=invented_attribute_loader(inputs),
+ lifetime_inputs=invented_lifetime_inputs(),
+ data_provenance=ap.INVENTED,
+ registration_pointer=None,
+ )
+
+
+def _dimension(result, key):
+ return next(d for d in result["dimensions"] if d["key"] == key)
+
+
+def test_g3_total_agrees_with_frozen_rates_se_and_floor(replay, u0_breakdown):
+ original = replay[2]["rows"]["U0"]["tabulation"]["cells"][0]
+ result = u0_breakdown["groups"]["adjusted"]
+ total = _dimension(result, "total")["cells"][0]
+ statistics = {s["statistic"]: s for s in total["statistics"]}
+ for statistic, field in (
+ (gb.POVERTY_RATE_CURRENT_LAW, "baseline_rate"),
+ (gb.POVERTY_RATE_PROPOSAL, "reform_rate"),
+ (gb.POVERTY_RATE_CHANGE, "delta"),
+ ):
+ assert statistics[statistic]["value"] == original[field]
+ assert statistics[statistic]["uncertainty"]["design_se"] == (
+ original["design_se"][field]
+ )
+ assert statistics[statistic]["uncertainty"]["floor"] == (
+ original["floor"][field]
+ )
+
+
+def test_mint_order_and_unclassified_retained(u0_breakdown):
+ result = u0_breakdown["groups"]["adjusted"]
+ marital = _dimension(result, "marital_status")
+ assert [c["label"] for c in marital["cells"]][:4] == [
+ "Married",
+ "Divorced",
+ "Widowed",
+ "Never married",
+ ]
+ assert marital["unclassified"]["n_rows"] > 0
+ assert "benefit_type" in u0_breakdown["not_computed"]
+ assert result["post_hoc_labels"] == list(groups.POSTHOC_LABELS)
+ assert gb.INVENTED_DATA_LABEL in result["labels"]
+
+
+def test_missing_lifetime_is_unavailable_and_sources_are_unmutated(replay):
+ inputs, _, _, retained = replay
+ members = retained.rows["U0"].members
+ before = members.copy(deep=True)
+ side, provenance = groups.lifetime_frame(members, inputs, None)
+ assert side[list(groups.MEASURES)].isna().all().all()
+ assert provenance["status"] == "not computed"
+ assert_frame_equal(members, before)
+
+
+def test_side_frame_must_be_sealed(replay):
+ inputs, _, _, retained = replay
+ loader = invented_attribute_loader(inputs)
+
+ def tampered(person_ids, *, anchor_waves):
+ attributes = loader(person_ids, anchor_waves=anchor_waves)
+ attributes.frame.loc[0, "education_years"] = 1
+ return attributes
+
+ with pytest.raises(GroupBreakdownRefusal, match="seal"):
+ groups._attribute_frame(retained.rows["U0"].members, tampered)
+
+
+def test_declared_official_style_concept_is_separate(u0_breakdown):
+ adjusted = u0_breakdown["groups"]["adjusted"]
+ official = u0_breakdown["groups"]["reported_money_official_style"]
+ assert official["upstream"]["concept"] == "reported_money_official_style"
+ assert "householder" in official["upstream"]["official_style_caveat"]
+ assert official["statistics"] == adjusted["statistics"]
+
+
+def test_observation_plan_does_not_extend_to_u2():
+ births = {
+ birth
+ for row in groups.REGISTERED_ROWS.values()
+ for birth, _, _, _ in age67.observation_plan(row.age67_spec())
+ }
+ assert min(births) == 1936
+ assert max(births) == 1945
+
+
+def test_missing_side_person_refuses():
+ with pytest.raises(GroupBreakdownRefusal, match="missing"):
+ groups._join(
+ pd.DataFrame({"person_id": [1, 2]}),
+ pd.DataFrame({"person_id": [1], "attr": [0]}),
+ "person_id",
+ )
diff --git a/tests/group_breakdowns/test_uniform_cut_properties.py b/tests/group_breakdowns/test_uniform_cut_properties.py
new file mode 100644
index 00000000..1ea9575b
--- /dev/null
+++ b/tests/group_breakdowns/test_uniform_cut_properties.py
@@ -0,0 +1,481 @@
+"""INVENTED DATA - NOT A COMPARISON: exercise-2 adapter invariants.
+
+All frames and measure stubs here are invented. These tests call no reader,
+load no committed outcome and do not execute a real-data pipeline.
+"""
+
+from __future__ import annotations
+
+from collections.abc import Callable
+
+import numpy as np
+import pandas as pd
+import pytest
+from hypothesis import given, settings
+from hypothesis import strategies as st
+from pandas.testing import assert_frame_equal
+
+from populace_dynamics.cohorts import age67
+from populace_dynamics.cohorts import group_attributes as ga
+from populace_dynamics.estimates import group_breakdown as gb
+from populace_dynamics.estimates import lifetime_measures as lm
+from populace_dynamics.estimates import uniform_cut_tabulation as ut
+from populace_dynamics.group_breakdowns import uniform_cut as adapter
+from populace_dynamics.group_breakdowns.common import GroupBreakdownRefusal
+from populace_dynamics.ss.params import SSAParameters
+
+FAST = settings(max_examples=20, deadline=None)
+TABULATIONS = settings(max_examples=8, deadline=None)
+
+
+def _inputs(
+ careers: pd.DataFrame, person_ids: tuple[int, ...] = (1, 2)
+) -> age67.Age67Inputs:
+ """An INVENTED materialized cohort input, with known never-married IDs."""
+
+ history = pd.DataFrame(
+ {
+ "person_id": person_ids,
+ "is_marriage": False,
+ "marriage_order": pd.array(
+ [None] * len(person_ids), dtype="Int64"
+ ),
+ "start_year": pd.array([None] * len(person_ids), dtype="Int64"),
+ "start_month": pd.array([None] * len(person_ids), dtype="Int64"),
+ "end_year": pd.array([None] * len(person_ids), dtype="Int64"),
+ "separation_year": pd.array(
+ [None] * len(person_ids), dtype="Int64"
+ ),
+ "how_ended": "never_married",
+ "spouse_person_id": pd.array(
+ [None] * len(person_ids), dtype="Int64"
+ ),
+ "last_known_status": "never_married",
+ }
+ )
+ return age67.Age67Inputs(
+ anchors={},
+ design=pd.DataFrame(),
+ death_records=pd.DataFrame(),
+ marriage_history=history,
+ observed_earnings=careers,
+ family_income={},
+ family_wealth={},
+ wealth_refusals={},
+ provenance={"kind": "invented"},
+ )
+
+
+def _params() -> SSAParameters:
+ """INVENTED flat schedules; G2 calls are stubbed in the cutoff test."""
+
+ return SSAParameters(
+ nawi={year: 100.0 for year in range(1951, 2026)},
+ wage_base={1937: 10_000.0},
+ pia_factors=(0.9, 0.32, 0.15),
+ fra_months_by_birth_year=[(1900, 792)],
+ early_monthly_rates=(5 / 900, 5 / 1200),
+ early_first_bracket_months=36,
+ pe_us_revision="INVENTED",
+ )
+
+
+@FAST
+@given(st.lists(st.integers(1936, 1945), min_size=1, max_size=15))
+def test_valid_observation_births_stay_inside_registered_cohort(births):
+ members = pd.DataFrame(
+ {
+ "observation_id": range(len(births)),
+ "person_id": range(1, len(births) + 1),
+ "birth_year": births,
+ "member_age": 67,
+ }
+ )
+ original = members.copy(deep=True)
+ adapter.check_observations(members)
+ assert_frame_equal(members, original)
+
+
+@FAST
+@given(st.integers(1946, 2100))
+def test_every_held_out_birth_refuses(birth):
+ members = pd.DataFrame(
+ {
+ "observation_id": [1],
+ "person_id": [1],
+ "birth_year": [birth],
+ "member_age": [67],
+ }
+ )
+ with pytest.raises(GroupBreakdownRefusal, match="1936"):
+ adapter.check_observations(members)
+
+
+@FAST
+@given(
+ st.integers(1936, 1944),
+ st.sampled_from((0.125, 0.25, 0.5, 0.75)),
+)
+def test_fractional_birth_is_not_a_registered_birth_year(birth, fraction):
+ members = pd.DataFrame(
+ {
+ "observation_id": [1],
+ "person_id": [1],
+ "birth_year": [birth + fraction],
+ "member_age": [67],
+ }
+ )
+ with pytest.raises(GroupBreakdownRefusal):
+ adapter.check_observations(members)
+
+
+@FAST
+@given(st.data(), st.integers(1, 12))
+def test_many_to_one_join_preserves_keys_order_weights_and_inputs(data, n):
+ order = data.draw(st.permutations(tuple(range(n))))
+ left = pd.DataFrame(
+ {
+ "observation_id": order,
+ "person_id": [i // 2 + 1 for i in order],
+ "weight": [float(i + 1) for i in order],
+ },
+ index=[i + 100 for i in range(n)],
+ )
+ persons = sorted(set(left["person_id"]))
+ right = pd.DataFrame(
+ {"person_id": persons[::-1], "invented_group": persons[::-1]}
+ )
+ original_left, original_right = left.copy(), right.copy()
+ joined = adapter._join(left, right, "person_id")
+ assert joined["observation_id"].tolist() == left["observation_id"].tolist()
+ assert joined["weight"].tolist() == left["weight"].tolist()
+ assert joined["invented_group"].tolist() == left["person_id"].tolist()
+ assert len(joined) == len(left)
+ assert_frame_equal(left, original_left)
+ assert_frame_equal(right, original_right)
+
+
+@FAST
+@given(st.integers(1, 10_000))
+def test_duplicate_or_missing_side_key_refuses(person_id):
+ left = pd.DataFrame({"person_id": [person_id]})
+ duplicate = pd.DataFrame(
+ {"person_id": [person_id, person_id], "attribute": [1, 1]}
+ )
+ with pytest.raises(GroupBreakdownRefusal, match="duplicate"):
+ adapter._join(left, duplicate, "person_id")
+ missing = pd.DataFrame({"person_id": [person_id + 1], "attribute": [1]})
+ with pytest.raises(GroupBreakdownRefusal, match="missing"):
+ adapter._join(left, missing, "person_id")
+
+
+def test_join_preserves_and_isolates_the_cohort_source_guard():
+ left = pd.DataFrame({"person_id": [1]})
+ left.attrs = {"provenance_kind": "psid_files", "audit": {"sealed": True}}
+ right = pd.DataFrame({"person_id": [1], "attribute": [None]})
+ joined = adapter._join(left, right, "person_id")
+ assert joined.attrs == left.attrs
+ joined.attrs["audit"]["sealed"] = False
+ assert left.attrs["audit"]["sealed"] is True
+
+
+def _measure_stub(
+ value_column: str, calls: list[pd.DataFrame]
+) -> Callable[..., lm.MeasureResult]:
+ """INVENTED reduction solely to observe the adapter's G2 boundary."""
+
+ def compute(careers, persons, params, **kwargs):
+ del params
+ calls.append(careers.copy(deep=True))
+ totals = careers.groupby("person_id")["earnings"].sum()
+ frame = persons[["person_id"]].copy()
+ frame[value_column] = frame["person_id"].map(totals)
+ frame["status"] = lm.COMPUTED
+ if kwargs.get("shared"):
+ frame["marriage_history_absent"] = False
+ for column in (
+ "years_marital_unknown",
+ "married_years_spouse_unavailable",
+ "married_years_spouse_year_absent",
+ "married_years_own_year_absent",
+ "married_years_spouse_record_disagrees",
+ "years_multiple_marriages_in_force",
+ ):
+ frame[column] = 0
+ return lm.MeasureResult("INVENTED", frame, {"kind": "INVENTED"})
+
+ return compute
+
+
+@FAST
+@given(st.floats(min_value=0.0, max_value=1e9, allow_nan=False))
+def test_future_earnings_cannot_enter_any_of_the_five_g2_measures(future):
+ members = pd.DataFrame(
+ {
+ "observation_id": ["early", "later", "second"],
+ "person_id": [1, 1, 2],
+ "birth_year": [1937, 1937, 1938],
+ "income_year": [2003, 2005, 2005],
+ "member_age": [66, 68, 67],
+ }
+ )
+ adapter.check_observations(members)
+ careers = pd.DataFrame(
+ {
+ "person_id": [1, 1, 1, 2, 2],
+ "period": [2002, 2003, 2005, 2003, 2005],
+ "earnings": [10.0, 20.0, 30.0, 40.0, 50.0],
+ }
+ )
+ later = pd.concat(
+ [
+ careers,
+ pd.DataFrame(
+ {
+ "person_id": [1, 2],
+ "period": [2020, 2022],
+ "earnings": [future, future],
+ }
+ ),
+ ],
+ ignore_index=True,
+ )
+ calls = []
+ supplied = adapter.LifetimeInputs(
+ _params(), rates=object(), interest=object()
+ )
+ with pytest.MonkeyPatch.context() as patch:
+ for name, column in (
+ ("initial_aime_at_62", "aime"),
+ ("lifetime_payroll_tax_pv_at_62", "pv_at_62"),
+ (
+ "report_average_indexed_earnings_22_62",
+ "average_indexed_earnings",
+ ),
+ ):
+ patch.setattr(adapter.lm, name, _measure_stub(column, calls))
+ baseline, _ = adapter.lifetime_frame(
+ members, _inputs(careers), supplied
+ )
+ with_future, _ = adapter.lifetime_frame(
+ members, _inputs(later), supplied
+ )
+ assert_frame_equal(baseline, with_future)
+ assert set(adapter.MEASURES).issubset(baseline)
+ assert len(calls) == 20 # five measures, two years, two input histories
+ for position, called in enumerate(calls):
+ cutoff = 2003 if position % 10 < 5 else 2005
+ assert called["year"].le(cutoff).all()
+ assert "year" not in careers and "year" not in later
+
+
+@pytest.mark.parametrize(
+ "reason",
+ (
+ "marriage_history_absent",
+ "years_marital_unknown",
+ "married_years_spouse_year_absent",
+ "married_years_own_year_absent",
+ "married_years_spouse_record_disagrees",
+ "years_multiple_marriages_in_force",
+ ),
+)
+def test_incomplete_shared_inputs_remain_unavailable(reason, monkeypatch):
+ members = pd.DataFrame(
+ {
+ "observation_id": ["o1"],
+ "person_id": [1],
+ "birth_year": [1937],
+ "income_year": [2004],
+ }
+ )
+ careers = pd.DataFrame(
+ {"person_id": [1], "period": [2000], "earnings": [10.0]}
+ )
+ for name, column in (
+ ("initial_aime_at_62", "aime"),
+ ("lifetime_payroll_tax_pv_at_62", "pv_at_62"),
+ ("report_average_indexed_earnings_22_62", "average_indexed_earnings"),
+ ):
+ stub = _measure_stub(column, [])
+
+ def compute(*args, _stub=stub, **kwargs):
+ result = _stub(*args, **kwargs)
+ if kwargs.get("shared") and reason != "marriage_history_absent":
+ result.frame[reason] = 1
+ return result
+
+ monkeypatch.setattr(adapter.lm, name, compute)
+ inputs = _inputs(
+ careers, (2,) if reason == "marriage_history_absent" else (1,)
+ )
+ side, provenance = adapter.lifetime_frame(
+ members,
+ inputs,
+ adapter.LifetimeInputs(_params(), rates=object(), interest=object()),
+ )
+ for name in adapter.MEASURES:
+ if name.endswith("shared"):
+ assert pd.isna(side.loc[0, name])
+ record = next(
+ r for r in provenance["records"] if r["dimension"] == name
+ )
+ assert record["adapter_exclusion_counts"][reason] == 1
+ else:
+ assert side.loc[0, name] == 10.0
+
+
+def _rows(incomes, weights):
+ """INVENTED U1 observations and already-computed poverty flags."""
+
+ births = [1936, 1937, 1940, 1941, 1944, 1945]
+ ages = [66, 67, 68, 66, 67, 68]
+ members = pd.DataFrame(
+ {
+ "observation_id": [f"o{i}" for i in range(6)],
+ "person_id": range(1, 7),
+ "birth_year": births,
+ "member_age": ages,
+ "income_year": np.array(births) + ages,
+ "wave": np.array(births) + ages + 1,
+ "family_unit_id": range(1, 7),
+ "sex": ["female", "male"] * 3,
+ "marital_status_4": [
+ "unclassified",
+ "married",
+ "widowed",
+ "divorced",
+ "never_married",
+ "married",
+ ],
+ "weight": weights,
+ "stratum": 1,
+ "cluster": [1, 2, 1, 2, 1, 2],
+ }
+ )
+ money = np.asarray(incomes, dtype=float)
+ members["total_family_income"] = money
+ adjusted = pd.DataFrame(
+ {
+ "observation_id": members["observation_id"],
+ "baseline_income": money + 5.0,
+ "money_income": money,
+ "threshold": 100.0,
+ "cut": 6.5,
+ "ssi_offset": 0.0,
+ "ssi_new": 0.0,
+ "poor_baseline": money + 5.0 < 100.0,
+ "poor_reform": money + 5.0 - 6.5 < 100.0,
+ }
+ )
+ design = pd.DataFrame({"stratum": [1, 1, 1], "cluster": [1, 2, 3]})
+ result = {"tabulation": {"config": {"floor_seeds": [0, 1]}}}
+ return adapter.ReplayedRow(result, members, adjusted, design)
+
+
+def _attribute_loader(person_ids, *, anchor_waves):
+ """INVENTED G1-shaped labels, including explicitly missing attributes."""
+
+ frame = pd.DataFrame(
+ {
+ "person_id": person_ids,
+ "race_ethnicity_mint8": [
+ gb.MINT8_SCHEME.dimension("race_ethnicity")
+ .categories[(pid - 1) % 4]
+ .label
+ for pid in person_ids
+ ],
+ "country_of_birth_mint8": [
+ "United States" if pid % 2 else None for pid in person_ids
+ ],
+ "education_years": [12 if pid % 2 else None for pid in person_ids],
+ }
+ )
+ return ga.GroupAttributes(
+ frame,
+ {
+ "content_sha256": ga.content_sha256(frame),
+ "anchor_waves": list(anchor_waves),
+ "kind": "INVENTED",
+ },
+ )
+
+
+@TABULATIONS
+@given(
+ st.lists(st.integers(0, 200), min_size=6, max_size=6),
+ st.lists(st.integers(1, 10), min_size=6, max_size=6),
+)
+def test_adapter_totals_match_frozen_rates_design_se_and_floors(
+ incomes, weights
+):
+ replayed = _rows(incomes, weights)
+ original_members = replayed.members.copy(deep=True)
+ original_adjusted = replayed.adjusted.copy(deep=True)
+ result = adapter.tabulate_row_groups(
+ "U1",
+ replayed,
+ inputs=_inputs(pd.DataFrame(), tuple(range(1, 7))),
+ attribute_loader=_attribute_loader,
+ lifetime_inputs=None,
+ data_provenance="invented",
+ registration_pointer=None,
+ )
+ rows = ut.tabulation_rows(replayed.members, replayed.adjusted)
+ reference = ut.tabulate_uniform_cut(
+ rows,
+ data_provenance="invented",
+ design=replayed.design,
+ config=ut.TabulationConfig(floor_seeds=(0, 1)),
+ )["cells"][0]
+ grouped = result["groups"]["adjusted"]
+ dimensions = {d["key"]: d for d in grouped["dimensions"]}
+ total = dimensions["total"]["cells"][0]
+ statistics = {s["statistic"]: s for s in total["statistics"]}
+ for name, legacy in (
+ (gb.POVERTY_RATE_CURRENT_LAW, "baseline_rate"),
+ (gb.POVERTY_RATE_PROPOSAL, "reform_rate"),
+ (gb.POVERTY_RATE_CHANGE, "delta"),
+ ):
+ assert statistics[name]["value"] == reference[legacy]
+ assert statistics[name]["uncertainty"]["design_se"] == (
+ reference["design_se"][legacy]
+ )
+ assert statistics[name]["uncertainty"]["floor"] == (
+ reference["floor"][legacy]
+ )
+ assert dimensions["marital_status"]["unclassified"]["n_rows"] == 1
+ assert [c["label"] for c in dimensions["age"]["cells"]] == [
+ "66",
+ "67",
+ "68",
+ ]
+ for dimension in dimensions.values():
+ classified = sum(
+ c["counts"]["unweighted_n"] for c in dimension["cells"]
+ )
+ assert classified + dimension["unclassified"]["n_rows"] == 6
+ partitions = grouped["assignment"]["quintile_partitions"]
+ for name in adapter.MEASURES[:3]:
+ assert partitions[name] == ["birth_cohort"]
+ assert partitions["household_income_quintile"] == []
+ assert_frame_equal(replayed.members, original_members)
+ assert_frame_equal(replayed.adjusted, original_adjusted)
+
+
+@pytest.mark.parametrize("row_id", ("U0", "U1", "U3", "U8"))
+def test_scheme_preserves_mint_dimensions_and_states_u1_age_variant(row_id):
+ scheme = adapter._scheme(row_id)
+ assert scheme.composite
+ assert scheme.keys[-5:] == adapter.MEASURES
+ for dimension in gb.MINT8_SCHEME.dimensions:
+ if row_id == "U1" and dimension.key == "age":
+ ages = scheme.dimension("age")
+ assert ages.labels == ("66", "67", "68")
+ assert [(c.lower, c.upper) for c in ages.categories] == [
+ (66, 66),
+ (67, 67),
+ (68, 68),
+ ]
+ else:
+ assert scheme.dimension(dimension.key) == dimension
From 64e368acce951ce763e048a1d65f22f0059daafc Mon Sep 17 00:00:00 2001
From: Max Ghenis
Date: Sat, 3 Oct 2026 07:33:08 -0400
Subject: [PATCH 07/10] Add MINT group breakdowns for exercise 4 (minimum
benefit)
group_breakdowns/min_benefit.py re-executes Registration 17's
registered computation (cohort, careers, evaluate and the unchanged
tabulate_track_m for every registered row and the d430 sensitivity) and
refuses before loading any attribute unless every committed cell,
n column, floor, design SE and cohort structure of
runs/replication_urban2006_minimum_benefit_v1.json matches. It then
adds MINT groups (marital status in MINT order with unclassified
counted, MINT age bands from 2022 - birth year, sex, race and
ethnicity, education, initial AIME and lifetime payroll tax quintiles,
benefit type) to the share receiving the minimum under each option.
tests/group_breakdowns/test_min_benefit_reproduction_psid.py is a
host-only check that the committed (already published) cells reproduce
exactly; it computes no group cell. Invented data otherwise.
Co-Authored-By: Claude Opus 5.5
---
.../exercise_4_min_benefit/RESULTS.md | 90 +
scripts/make_nasi_repro_venv.sh | 183 ++
scripts/min_benefit_groups_dry_run.py | 340 +++
scripts/run_min_benefit_groups_registered.py | 457 ++++
.../group_breakdowns/__init__.py | 26 +
.../group_breakdowns/common.py | 636 ++++++
.../group_breakdowns/min_benefit.py | 1882 +++++++++++++++++
tests/group_breakdowns/__init__.py | 0
.../test_min_benefit_common.py | 509 +++++
.../test_min_benefit_groups.py | 1001 +++++++++
.../test_min_benefit_parent_artifact.py | 181 ++
.../test_min_benefit_registered_script.py | 567 +++++
.../test_min_benefit_reproduction_psid.py | 113 +
13 files changed, 5985 insertions(+)
create mode 100644 docs/analysis/nasi_group_breakdowns_invented_20261001/exercise_4_min_benefit/RESULTS.md
create mode 100755 scripts/make_nasi_repro_venv.sh
create mode 100644 scripts/min_benefit_groups_dry_run.py
create mode 100644 scripts/run_min_benefit_groups_registered.py
create mode 100644 src/populace_dynamics/group_breakdowns/__init__.py
create mode 100644 src/populace_dynamics/group_breakdowns/common.py
create mode 100644 src/populace_dynamics/group_breakdowns/min_benefit.py
create mode 100644 tests/group_breakdowns/__init__.py
create mode 100644 tests/group_breakdowns/test_min_benefit_common.py
create mode 100644 tests/group_breakdowns/test_min_benefit_groups.py
create mode 100644 tests/group_breakdowns/test_min_benefit_parent_artifact.py
create mode 100644 tests/group_breakdowns/test_min_benefit_registered_script.py
create mode 100644 tests/group_breakdowns/test_min_benefit_reproduction_psid.py
diff --git a/docs/analysis/nasi_group_breakdowns_invented_20261001/exercise_4_min_benefit/RESULTS.md b/docs/analysis/nasi_group_breakdowns_invented_20261001/exercise_4_min_benefit/RESULTS.md
new file mode 100644
index 00000000..e8b0a6a5
--- /dev/null
+++ b/docs/analysis/nasi_group_breakdowns_invented_20261001/exercise_4_min_benefit/RESULTS.md
@@ -0,0 +1,90 @@
+INVENTED DATA - NOT A COMPARISON
+
+# Exercise 4 (Track M) by MINT8 subgroups: INVENTED dry run
+
+INVENTED DATA - NOT A COMPARISON. Every number below comes from an INVENTED PSID-shaped cohort, INVENTED parameters and an INVENTED side frame. None is a PSID, SSA, Census or comparator value, and none is a result.
+
+## What ran
+
+- An INVENTED stand-in for the committed run (the registered computation, once).
+- Re-execution and exact reproduction: 32 checks, all identical (exact: both sides are encoded with json.dumps(allow_nan=False) and decoded, then every leaf must be equal in type and value, every float bit for bit (float.hex, so -0.0 differs from 0.0), every mapping with the same keys and every list with the same length and order).
+- Consistency of G3's Total/Female/Male cells with Track M's All/Women/Men cells: 96 checks, all identical.
+- Refusal shown: a parent with `$.rows.MS0.tabulation.cells[0].share_percent` moved one ulp was refused (side-frame loader called: False).
+- Location rule (for the real parent, which records the COLA history by an absolute path): the location fields ($.inputs.source.cola_history.path) are compared by their repository-relative path: an absolute path ending in /data/external/ssa_cola_history.json reads as data/external/ssa_cola_history.json on both sides, any other path is compared as written; the same record's sha256 and content_sha256 are compared exactly, and nothing else is normalized.
+
+## Group attributes (INVENTED)
+
+| Dimension | Classified | Unclassified | Reasons |
+|---|---:|---:|---|
+| Total | 120 | 0 | - |
+| Sex | 120 | 0 | - |
+| Race and ethnicity | 105 | 15 | code:attribute_unknown:dk_na_refused: 8, code:attribute_unknown:never_head_or_spouse: 7 |
+| Country of birth | 111 | 9 | code:attribute_unknown:dk_na_refused: 1, code:attribute_unknown:never_head_or_spouse: 7, code:unresolved:us_territory: 1 |
+| Age | 120 | 0 | - |
+| Marital status | 120 | 0 | - |
+| Highest education level | 116 | 4 | missing: 4 |
+| Current-law benefit type | 120 | 0 | - |
+| Current-law initial AIME quintile | 106 | 14 | missing: 14 |
+| Lifetime payroll tax quintile | 106 | 14 | missing: 14 |
+| Lifetime payroll tax quintile (shared) | 106 | 14 | missing: 14 |
+
+## MS0, share receiving a minimum by group (INVENTED, percent)
+
+| Group | Row | Option 2 | Option 3 | Option 4 | Option 5 | n |
+|---|---|---:|---:|---:|---:|---:|
+| Total | Total | 0.0 | 2.4 | 4.3 | 12.1 | 120 |
+| Sex | Female | 0.0 | 1.0 | 1.9 | 5.8 | 66 |
+| Sex | Male | 0.0 | 4.3 | 7.5 | 20.6 | 54 |
+| Race and ethnicity | Hispanic or Latino, any race | 0.0 | 11.7 | 11.7 | 29.0 | 17 |
+| Race and ethnicity | White, non-Hispanic | 0.0 | 1.5 | 1.5 | 8.0 | 47 |
+| Race and ethnicity | Black or African American, non-Hispanic | 0.0 | 0.0 | 3.2 | 13.3 | 19 |
+| Race and ethnicity | All other races, non-Hispanic | 0.0 | 0.0 | 4.4 | 9.6 | 22 |
+| Country of birth | United States | 0.0 | 2.9 | 4.5 | 11.0 | 101 |
+| Country of birth | Other countries | 0.0 | 0.0 | 6.9 | 38.6 | 10 |
+| Age | 60–69 | 0.0 | 9.7 | 9.7 | 29.3 | 30 |
+| Age | 70–79 | 0.0 | 0.0 | 6.3 | 6.3 | 28 |
+| Age | 80–89 | 0.0 | 0.0 | 1.2 | 7.7 | 51 |
+| Age | 90 or older | 0.0 | 0.0 | 0.0 | 0.0 | 11 |
+| Marital status | Married | 0.0 | 3.0 | 3.9 | 16.3 | 70 |
+| Marital status | Divorced | 0.0 | 2.6 | 8.6 | 8.6 | 28 |
+| Marital status | Widowed | 0.0 | 0.0 | 0.0 | 2.8 | 15 |
+| Marital status | Never married | 0.0 | 0.0 | 0.0 | 0.0 | 7 |
+| Highest education level | Graduate | 0.0 | 0.0 | 0.0 | 0.0 | 11 |
+| Highest education level | Bachelor | 0.0 | 0.0 | 8.9 | 27.1 | 11 |
+| Highest education level | Associate | 0.0 | 0.0 | 9.1 | 24.0 | 17 |
+| Highest education level | High school | 0.0 | 0.0 | 0.0 | 4.2 | 25 |
+| Highest education level | Less than high school | 0.0 | 5.5 | 5.5 | 13.2 | 52 |
+| Current-law benefit type | Retired worker only | 0.0 | 3.0 | 5.3 | 14.1 | 95 |
+| Current-law benefit type | Widow(er) (includes dually entitled) | 0.0 | 0.0 | 0.0 | 2.8 | 15 |
+| Current-law benefit type | Spousal (includes dually entitled) | 0.0 | 0.0 | 0.0 | 0.0 | 7 |
+| Current-law benefit type | Disabled worker only | 0.0 | 0.0 | 0.0 | 30.3 | 3 |
+| Current-law initial AIME quintile | Highest | 0.0 | 0.0 | 0.0 | 0.0 | 21 |
+| Current-law initial AIME quintile | Second highest | 0.0 | 0.0 | 0.0 | 0.0 | 20 |
+| Current-law initial AIME quintile | Middle | 0.0 | 0.0 | 0.0 | 0.0 | 21 |
+| Current-law initial AIME quintile | Second lowest | 0.0 | 0.0 | 5.9 | 28.9 | 20 |
+| Current-law initial AIME quintile | Lowest | 0.0 | 12.7 | 17.0 | 36.3 | 24 |
+| Lifetime payroll tax quintile | Highest | 0.0 | 0.0 | 0.0 | 0.0 | 21 |
+| Lifetime payroll tax quintile | Second highest | 0.0 | 0.0 | 0.0 | 0.0 | 20 |
+| Lifetime payroll tax quintile | Middle | 0.0 | 0.0 | 0.0 | 0.0 | 21 |
+| Lifetime payroll tax quintile | Second lowest | 0.0 | 0.0 | 0.0 | 16.4 | 19 |
+| Lifetime payroll tax quintile | Lowest | 0.0 | 12.0 | 21.3 | 46.2 | 25 |
+| Lifetime payroll tax quintile (shared) | Highest | 0.0 | 0.0 | 0.0 | 0.0 | 20 |
+| Lifetime payroll tax quintile (shared) | Second highest | 0.0 | 0.0 | 0.0 | 0.0 | 23 |
+| Lifetime payroll tax quintile (shared) | Middle | 0.0 | 0.0 | 0.0 | 12.2 | 20 |
+| Lifetime payroll tax quintile (shared) | Second lowest | 0.0 | 11.9 | 15.4 | 31.2 | 19 |
+| Lifetime payroll tax quintile (shared) | Lowest | 0.0 | 2.7 | 9.1 | 24.7 | 24 |
+
+## Not computed
+
+- `poverty_status`: not computed: MINT8's poverty status and household income quintile need each person's 2022 household (family) money income and, for poverty, the official threshold of the family's size and composition; Track M's inputs (min_benefit_track_m.structure.TrackMStructureInputs: the 2023 anchor, the family file's 2022 Social Security of the reference person and spouse, death records, marriage history, the labor-income panel and the design; and the receipt and next-wave labor-income frames of cohort.TrackMCohortInputs) carry neither, and this package reads no new PSID item
+- `household_income_quintile`: not computed: MINT8's poverty status and household income quintile need each person's 2022 household (family) money income and, for poverty, the official threshold of the family's size and composition; Track M's inputs (min_benefit_track_m.structure.TrackMStructureInputs: the 2023 anchor, the family file's 2022 Social Security of the reference person and spouse, death records, marriage history, the labor-income panel and the design; and the receipt and next-wave labor-income frames of cohort.TrackMCohortInputs) carry neither, and this package reads no new PSID item
+- `mint8_benefit_statistics`: not computed: MINT8's benefit statistics compare each person's benefit under the option with the same person's current-law benefit. Track M's evaluate() returns, per person, only the receipt flags receives_2..receives_5, the receipt basis and the exposure flags (min_benefit_track_m/evaluation.py evaluate); per worker record it returns the PIA under the row's PIA rule, before any option's cut or minimum (WorkerOutcome.pia), and each option's PIA (WorkerOutcome.outcomes[k].option_pia, rules.evaluate_worker), but no person-level benefit: the own benefit after the claim factor is not formed, and a spouse's or survivor's benefit is decided only as paid or not (rules.spouse_excess_paid and rules.survivor_excess_paid return booleans). The registered options also have no current-law column: option 1 is 'Reduced current law', a 12.45 percent uniform cut (policy.OPTIONS[1]). Computing these statistics would need a person-level benefit model this breakdown does not add
+- `mint8_poverty_statistics`: not computed: MINT8's poverty status and household income quintile need each person's 2022 household (family) money income and, for poverty, the official threshold of the family's size and composition; Track M's inputs (min_benefit_track_m.structure.TrackMStructureInputs: the 2023 anchor, the family file's 2022 Social Security of the reference person and spouse, death records, marriage history, the labor-income panel and the design; and the receipt and next-wave labor-income frames of cohort.TrackMCohortInputs) carry neither, and this package reads no new PSID item
+
+## Reproduce
+
+```
+python scripts/min_benefit_groups_dry_run.py
+```
+
+`result.json` holds the full document. INVENTED DATA - NOT A COMPARISON.
diff --git a/scripts/make_nasi_repro_venv.sh b/scripts/make_nasi_repro_venv.sh
new file mode 100755
index 00000000..295b1452
--- /dev/null
+++ b/scripts/make_nasi_repro_venv.sh
@@ -0,0 +1,183 @@
+#!/usr/bin/env bash
+# Build the environment a bit-for-bit reproduction of a committed blind-test
+# run needs (NASI follow-ups, 2026-10-01; package G4c).
+#
+# The group-breakdown adapters (src/populace_dynamics/group_breakdowns/)
+# re-execute a committed registered run and refuse unless every recomputed
+# cell equals the committed artifact exactly, floats bit for bit. Design
+# standard errors go through numpy, so the reproduction needs the Python,
+# numpy, pandas and scipy versions the committed run's .env.json sidecar
+# records. This script creates
+#
+# /.venv-nasi-repro
+#
+# with uv at exactly those versions (the default, GIL-enabled CPython build:
+# a free-threaded 3.14t interpreter takes other numpy, pandas and scipy
+# wheels, so uv is asked for "+gil" and the build is checked), installs this
+# repository editable
+# (its other dependencies resolved under those pins) plus pytest and
+# hypothesis, and then checks every pinned version against the sidecar.
+# It reads no PSID file and runs no pipeline.
+#
+# The versions below are exercise 4's (Track M) sidecar,
+# runs/replication_urban2006_minimum_benefit_v1.env.json;
+# tests/group_breakdowns/test_min_benefit_parent_artifact.py holds them
+# equal to it. A reproduction also needs, outside this venv:
+# * the staged PSID files the committed artifact records by SHA-256
+# (inputs.source.psid_files_sha256), under POPULACE_DYNAMICS_PSID_DIR
+# or ~/PolicyEngine/psid-data;
+# * policyengine-us at revision a03e82e503, read as a git checkout by
+# populace_dynamics.ss.params (not installed): set
+# POPULACE_DYNAMICS_PE_US_DIR to it. This script checks the checkout
+# named by PE_US_DIR (default below) when it exists and never creates
+# it.
+#
+# Usage:
+# scripts/make_nasi_repro_venv.sh [--force]
+#
+# --force rebuilds the venv even when it already matches. Environment
+# overrides: NASI_REPRO_VENV (venv path), PE_US_DIR (checkout to check).
+set -euo pipefail
+
+PYTHON_VERSION="3.14.7"
+NUMPY_VERSION="2.5.3"
+PANDAS_VERSION="3.0.6"
+SCIPY_VERSION="1.18.1"
+PE_US_REVISION="a03e82e503"
+
+SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
+REPO="$(cd "${SCRIPT_DIR}/.." && pwd)"
+# The main checkout: the first entry of `git worktree list` (a worktree's
+# own path otherwise). The venv and the policyengine-us checkout live
+# beside the main checkout's .venv, shared by every worktree.
+MAIN="$(git -C "${REPO}" worktree list --porcelain | awk 'NR==1 {print $2}')"
+VENV="${NASI_REPRO_VENV:-${MAIN}/.venv-nasi-repro}"
+PE_US_DIR="${PE_US_DIR:-${MAIN}/.claude/pe-us-a03e82e503}"
+SIDECAR="${REPO}/runs/replication_urban2006_minimum_benefit_v1.env.json"
+
+FORCE=0
+for arg in "$@"; do
+ case "${arg}" in
+ --force) FORCE=1 ;;
+ *)
+ echo "unknown argument ${arg} (usage: $0 [--force])" >&2
+ exit 2
+ ;;
+ esac
+done
+
+command -v uv >/dev/null 2>&1 || {
+ echo "uv is required (https://docs.astral.sh/uv/)" >&2
+ exit 1
+}
+
+# Print each pinned version the venv's interpreter sees, one per line.
+installed() {
+ "${VENV}/bin/python" - <<'PY'
+import platform
+import sysconfig
+from importlib.metadata import version
+
+print("python", platform.python_version())
+print("gil_disabled", int(sysconfig.get_config_var("Py_GIL_DISABLED") or 0))
+for name in ("numpy", "pandas", "scipy"):
+ print(name, version(name))
+PY
+}
+
+expected() {
+ printf 'python %s\ngil_disabled 0\nnumpy %s\npandas %s\nscipy %s\n' \
+ "${PYTHON_VERSION}" "${NUMPY_VERSION}" "${PANDAS_VERSION}" \
+ "${SCIPY_VERSION}"
+}
+
+if [[ -x "${VENV}/bin/python" && "${FORCE}" -eq 0 ]] \
+ && [[ "$(installed 2>/dev/null || true)" == "$(expected)" ]]; then
+ echo "${VENV} already holds the pinned versions; --force rebuilds it"
+else
+ uv venv --clear --python "${PYTHON_VERSION}+gil" "${VENV}"
+ CONSTRAINTS="$(mktemp)"
+ trap 'rm -f "${CONSTRAINTS}"' EXIT
+ printf 'numpy==%s\npandas==%s\nscipy==%s\n' \
+ "${NUMPY_VERSION}" "${PANDAS_VERSION}" "${SCIPY_VERSION}" \
+ >"${CONSTRAINTS}"
+ uv pip install --python "${VENV}/bin/python" \
+ --constraint "${CONSTRAINTS}" \
+ "numpy==${NUMPY_VERSION}" \
+ "pandas==${PANDAS_VERSION}" \
+ "scipy==${SCIPY_VERSION}" \
+ --editable "${REPO}" \
+ pytest hypothesis
+fi
+
+# The venv's versions must equal the pins, and the pins the sidecar's.
+if [[ "$(installed)" != "$(expected)" ]]; then
+ echo "the venv's versions differ from the pins:" >&2
+ installed >&2
+ exit 1
+fi
+"${VENV}/bin/python" - "${SIDECAR}" "${PYTHON_VERSION}" "${NUMPY_VERSION}" \
+ "${PANDAS_VERSION}" "${SCIPY_VERSION}" "${PE_US_REVISION}" <<'PY'
+import json
+import sys
+
+sidecar, python, numpy, pandas, scipy, revision = sys.argv[1:]
+environment = json.load(open(sidecar, encoding="utf-8"))["environment"]
+recorded = {
+ "python": environment["python"],
+ "numpy": environment["packages"]["numpy"],
+ "pandas": environment["packages"]["pandas"],
+ "scipy": environment["packages"]["scipy"],
+ "policyengine-us": environment["policyengine_us_parameters"]["revision"],
+}
+pinned = {
+ "python": python,
+ "numpy": numpy,
+ "pandas": pandas,
+ "scipy": scipy,
+ "policyengine-us": revision,
+}
+if recorded != pinned:
+ sys.exit(f"the pins {pinned} are not the sidecar's {recorded}")
+print("pins equal the sidecar:", json.dumps(recorded, sort_keys=True))
+import platform
+
+# Recorded, not compared (the registered scripts compare the same fields).
+print("platform: sidecar", environment["platform"], "| venv", platform.platform())
+PY
+
+# The editable install must import this repository's source.
+"${VENV}/bin/python" - "${REPO}" <<'PY'
+import sys
+from pathlib import Path
+
+import populace_dynamics
+
+source = Path(populace_dynamics.__file__).resolve()
+repo = Path(sys.argv[1]).resolve()
+print("populace_dynamics from", source)
+if not source.is_relative_to(repo):
+ print(
+ f"note: populace_dynamics resolves outside {repo}; the registered "
+ "scripts put this checkout's src first on sys.path"
+ )
+PY
+
+if [[ -d "${PE_US_DIR}/.git" || -f "${PE_US_DIR}/.git" ]]; then
+ head="$(git -C "${PE_US_DIR}" rev-parse HEAD)"
+ if [[ "${head}" != "${PE_US_REVISION}"* ]]; then
+ echo "${PE_US_DIR} is at ${head}, not ${PE_US_REVISION}" >&2
+ exit 1
+ fi
+ echo "policyengine-us checkout ${PE_US_DIR} is at ${head}"
+else
+ echo "note: no policyengine-us checkout at ${PE_US_DIR}; a" \
+ "reproduction needs one at ${PE_US_REVISION}"
+fi
+
+cat < str:
+ return subprocess.run(
+ ["git", "-C", str(ROOT), *args],
+ check=True,
+ capture_output=True,
+ text=True,
+ ).stdout.strip()
+
+
+def invented_parent(
+ frames: Any, parameters: Any, cola: Any
+) -> common.CommittedArtifact:
+ """The INVENTED stand-in for the committed run: the registered
+ computation once, through the registered runner's JSON encoding."""
+
+ run = mb.reexecute_track_m(
+ frames,
+ parameters,
+ cola_rates=cola,
+ data_provenance=INVENTED,
+ registration_pointer=None,
+ provenance_kind=INVENTED,
+ source=dict(frames.provenance),
+ )
+ return common.CommittedArtifact(
+ document=common.json_normalized(dict(run.result))
+ )
+
+
+def tampered(
+ parent: common.CommittedArtifact,
+) -> tuple[common.CommittedArtifact, str]:
+ """``parent`` with MS0's first defined share moved by one ulp."""
+
+ document = copy.deepcopy(dict(parent.document))
+ cells = document["rows"]["MS0"]["tabulation"]["cells"]
+ index = next(i for i, cell in enumerate(cells) if cell["defined"])
+ value = cells[index]["share_percent"]
+ cells[index]["share_percent"] = math.nextafter(value, math.inf)
+ return (
+ common.CommittedArtifact(document=document),
+ f"$.rows.MS0.tabulation.cells[{index}].share_percent",
+ )
+
+
+def refusal_check(
+ frames: Any, parameters: Any, cola: Any, parent: common.CommittedArtifact
+) -> dict[str, Any]:
+ """Run against a tampered parent: it must refuse before the loader."""
+
+ bad, path = tampered(parent)
+ calls: list[int] = []
+
+ def loader(person_ids):
+ calls.append(len(person_ids))
+ raise AssertionError("the side frame was requested")
+
+ try:
+ mb.run_group_breakdowns(
+ parameters=parameters,
+ parent=bad,
+ data_provenance=INVENTED,
+ registration_pointer=None,
+ load_cohort_inputs=lambda: frames,
+ load_side_frame=loader,
+ cola_rates=cola,
+ provenance_kind=INVENTED,
+ source=dict(frames.provenance),
+ )
+ except common.ReproductionMismatchError as error:
+ return {
+ "tampered_path": path,
+ "tampering": "one unit in the last place (math.nextafter)",
+ "refused": True,
+ "side_frame_loader_called": bool(calls),
+ "document_returned": False,
+ "error": str(error),
+ }
+ raise AssertionError("the tampered parent was not refused")
+
+
+def _cell_value(cell: dict[str, Any]) -> str:
+ statistic = cell["statistics"][0]
+ if not statistic["defined"]:
+ return "undefined"
+ return f"{statistic['value']:.1f}"
+
+
+def _markdown(document: dict[str, Any]) -> str:
+ result = document["result"]
+ ms0 = result["breakdowns"]["MS0"]
+ lines = [
+ HEADER,
+ "",
+ "# Exercise 4 (Track M) by MINT8 subgroups: INVENTED dry run",
+ "",
+ HEADER + ". Every number below comes from an INVENTED PSID-shaped "
+ "cohort, INVENTED parameters and an INVENTED side frame. None is "
+ "a PSID, SSA, Census or comparator value, and none is a result.",
+ "",
+ "## What ran",
+ "",
+ "- An INVENTED stand-in for the committed run (the registered "
+ "computation, once).",
+ f"- Re-execution and exact reproduction: "
+ f"{len(result['reproduction']['checks'])} checks, all identical "
+ f"({result['reproduction']['comparison']['rule']}).",
+ f"- Consistency of G3's Total/Female/Male cells with Track M's "
+ f"All/Women/Men cells: "
+ f"{result['consistency_with_registered_cells']['n_checks']} checks, "
+ "all identical.",
+ f"- Refusal shown: a parent with "
+ f"`{document['refusal']['tampered_path']}` moved one ulp was "
+ f"refused (side-frame loader called: "
+ f"{document['refusal']['side_frame_loader_called']}).",
+ f"- Location rule (for the real parent, which records the COLA "
+ f"history by an absolute path): "
+ f"{result['reproduction']['location_fields']['rule']}.",
+ "",
+ "## Group attributes (INVENTED)",
+ "",
+ "| Dimension | Classified | Unclassified | Reasons |",
+ "|---|---:|---:|---|",
+ ]
+ for summary in result["breakdown"]["assignment"]["dimensions"]:
+ reasons = ", ".join(
+ f"{reason}: {count}"
+ for reason, count in summary["unclassified_reasons"].items()
+ )
+ lines.append(
+ f"| {summary['label']} | {summary['n_classified']} | "
+ f"{summary['n_unclassified']} | {reasons or '-'} |"
+ )
+ lines += [
+ "",
+ "## MS0, share receiving a minimum by group (INVENTED, percent)",
+ "",
+ "| Group | Row | Option 2 | Option 3 | Option 4 | Option 5 | n |",
+ "|---|---|---:|---:|---:|---:|---:|",
+ ]
+ by_option = {
+ number: {d["key"]: d for d in ms0[number]["dimensions"]}
+ for number in ("2", "3", "4", "5")
+ }
+ for dimension in ms0["2"]["dimensions"]:
+ for position, cell in enumerate(dimension["cells"]):
+ values = [
+ _cell_value(by_option[n][dimension["key"]]["cells"][position])
+ for n in ("2", "3", "4", "5")
+ ]
+ n = cell["statistics"][0]["unweighted_n"]
+ lines.append(
+ f"| {dimension['label']} | {cell['label']} | "
+ + " | ".join(values)
+ + f" | {n} |"
+ )
+ lines += [
+ "",
+ "## Not computed",
+ "",
+ ]
+ for name, reason in result["not_computed"]["dimensions"].items():
+ lines.append(f"- `{name}`: {reason}")
+ for name, item in result["not_computed"]["statistics"].items():
+ lines.append(f"- `{name}`: {item['reason']}")
+ lines += [
+ "",
+ "## Reproduce",
+ "",
+ "```",
+ document["run"]["command"],
+ "```",
+ "",
+ f"`result.json` holds the full document. {HEADER}.",
+ "",
+ ]
+ return "\n".join(lines)
+
+
+def main(argv: list[str] | None = None) -> int:
+ parser = argparse.ArgumentParser(description=__doc__.splitlines()[0])
+ parser.add_argument("--output-dir", type=Path, default=DEFAULT_OUTPUT_DIR)
+ parser.add_argument("--seed", type=int, default=DEFAULT_SEED)
+ parser.add_argument(
+ "--family-units", type=int, default=DEFAULT_FAMILY_UNITS
+ )
+ parser.add_argument(
+ "--side-frame-seed", type=int, default=DEFAULT_SIDE_FRAME_SEED
+ )
+ args = parser.parse_args(argv)
+ started = datetime.datetime.now(datetime.timezone.utc).isoformat()
+ parameters, cola = invented.invented_parameters()
+ frames = invented_psid.invented_cohort_inputs(
+ seed=args.seed, n_family_units=args.family_units
+ )
+ parent = invented_parent(frames, parameters, cola)
+ side_frames = mb.invented_group_attribute_loader(
+ frames, seed=args.side_frame_seed
+ )
+ result = mb.run_group_breakdowns(
+ parameters=parameters,
+ parent=parent,
+ data_provenance=INVENTED,
+ registration_pointer=None,
+ load_cohort_inputs=lambda: frames,
+ load_side_frame=side_frames,
+ cola_rates=cola,
+ provenance_kind=INVENTED,
+ source=dict(frames.provenance),
+ )
+ document = {
+ "header": HEADER,
+ "description": (
+ "Exercise 4 (Track M) group breakdowns on an INVENTED "
+ "PSID-shaped cohort with INVENTED parameters and an INVENTED "
+ "side frame, through the real Track M, G1, G2 and G3 code. Not "
+ "PSID values, not a result, not a comparison."
+ ),
+ "inputs": {
+ "cohort": (
+ "min_benefit_track_m.invented_psid.invented_cohort_inputs"
+ ),
+ "seed": args.seed,
+ "family_units": args.family_units,
+ "parameters": (
+ "min_benefit_track_m.invented.invented_parameters (INVENTED)"
+ ),
+ "side_frame": (
+ "group_breakdowns.common.invented_group_attribute_inputs "
+ "through cohorts.group_attributes.build_group_attributes"
+ ),
+ "side_frame_seed": args.side_frame_seed,
+ },
+ "refusal": refusal_check(frames, parameters, cola, parent),
+ "result": result,
+ "run": {
+ "started": started,
+ "finished": datetime.datetime.now(
+ datetime.timezone.utc
+ ).isoformat(),
+ "git_head": _git("rev-parse", "HEAD"),
+ "git_clean": _git("status", "--porcelain") == "",
+ "python": platform.python_version(),
+ "command": " ".join(
+ [
+ "python",
+ "scripts/min_benefit_groups_dry_run.py",
+ *(sys.argv[1:] if argv is None else argv),
+ ]
+ ),
+ },
+ }
+ args.output_dir.mkdir(parents=True, exist_ok=True)
+ path = args.output_dir / "result.json"
+ path.write_text(
+ json.dumps(document, indent=1, allow_nan=False) + "\n",
+ encoding="utf-8",
+ )
+ (args.output_dir / "RESULTS.md").write_text(
+ _markdown(document), encoding="utf-8"
+ )
+ print(path, hashlib.sha256(path.read_bytes()).hexdigest())
+ return 0
+
+
+if __name__ == "__main__":
+ raise SystemExit(main())
diff --git a/scripts/run_min_benefit_groups_registered.py b/scripts/run_min_benefit_groups_registered.py
new file mode 100644
index 00000000..76a3773e
--- /dev/null
+++ b/scripts/run_min_benefit_groups_registered.py
@@ -0,0 +1,457 @@
+"""Exercise 4 (Track M) by MINT8 subgroups: the registered post hoc run.
+
+NASI follow-up package G4c. The only entry point that computes exercise
+4's group breakdowns on real data
+(:mod:`populace_dynamics.group_breakdowns.min_benefit`). It runs once,
+after its own issue #42 registration comment exists, at exactly the commit
+that comment registers, and writes a NEW artifact beside the committed
+exercise-4 run, which it never edits. It mirrors
+``scripts/run_track_m_registered.py``.
+
+**Preflight** (:func:`preflight`, before anything is loaded):
+
+* the registration pointer must be a comment on issue #42, and not
+ Registration 17's (the breakdown's cells are new outcomes, so they need
+ a new registration);
+* the working tree must be clean and ``HEAD`` must equal
+ ``--registered-commit``;
+* the output artifact and its sidecar must not exist yet (both are
+ created exclusively, so a file that appears during the run is not
+ overwritten either);
+* the parent artifact ``runs/replication_urban2006_minimum_benefit_v1.json``
+ must be the committed bytes (SHA-256 pinned) with its sidecar binding
+ them, Registration 17's pointer and commit, and the M1 specification it
+ recorded must be the file at this commit; the M1 block must still pass
+ Track M's registered-run gate.
+
+**Before any PSID file is read:** Track M's components and parameter pins
+(``run_track_m_registered.check_runnable`` and ``committed_parameters``,
+reused through importlib); the pins must equal the parent's
+``checks.parameter_pins``; the environment (Track A's ``_environment``
+resolver, as exercises 3 and 4 reuse it) must match the parent's sidecar
+in Python, numpy, pandas and scipy versions and the policyengine-us
+revision, because the reproduction is compared bit for bit
+(``scripts/make_nasi_repro_venv.sh`` builds that environment).
+
+**The computation** is :func:`populace_dynamics.group_breakdowns.
+min_benefit.run_group_breakdowns` with the staged-PSID loaders: Track M's
+``cohort.load_cohort_inputs`` and G1's ``group_attributes.
+load_group_attributes`` (anchor wave 2023). It re-executes Registration
+17's computation, refuses unless every recomputed block equals the parent
+exactly (before the group-attribute loader is called), then computes the
+cells and refuses unless G3's Total, Female and Male cells equal the
+registered All, Women and Men cells. A refusal writes nothing.
+
+The artifact carries Track M's labels, "registered, one-shot, post hoc,
+not blind" and "report-only", the covered-earnings disclosure (d280), the
+reproduction record, the SHA-256 of every module the run composes and the
+run block; it publishes regardless of outcome. Nothing here has been run
+on real data.
+
+Usage::
+
+ python scripts/run_min_benefit_groups_registered.py \\
+ --registration-pointer \\
+ --registered-commit \\
+ [--output runs/replication_urban2006_minimum_benefit_groups_posthoc_v1.json]
+
+Writes the artifact and a ``.env.json`` sidecar next to it.
+"""
+
+from __future__ import annotations
+
+import argparse
+import datetime
+import importlib.util
+import json
+import re
+import subprocess
+import sys
+import sysconfig
+from collections.abc import Callable, Mapping, Sequence
+from pathlib import Path
+from typing import Any
+
+ROOT = Path(__file__).resolve().parents[1]
+if str(ROOT / "src") not in sys.path:
+ sys.path.insert(0, str(ROOT / "src"))
+
+from populace_dynamics.group_breakdowns import common # noqa: E402
+from populace_dynamics.group_breakdowns import ( # noqa: E402
+ min_benefit as mb,
+)
+from populace_dynamics.min_benefit_track_m import ( # noqa: E402
+ COVERED_EARNINGS_DISCLOSURE,
+)
+from populace_dynamics.min_benefit_track_m.evaluation import ( # noqa: E402
+ PSID_FILES,
+ TrackMParameters,
+)
+from populace_dynamics.min_benefit_track_m.specification import ( # noqa: E402
+ M1_SPECIFICATION_PATH,
+ check_specification_for_registered_run,
+ m1_parameter_block,
+)
+from populace_dynamics.min_benefit_track_m.tabulation import ( # noqa: E402
+ REGISTERED_REAL,
+)
+
+REGISTERED_HEADER = (
+ "REGISTERED ONE-SHOT POST HOC RUN - exercise 4 (Track M: the share of "
+ "OASDI beneficiaries 62+ receiving a minimum benefit, options 2-5, "
+ "income year 2022) by MINT8 characteristic subgroups. "
+ + "; ".join(mb.LABELS)
+ + ". "
+ + COVERED_EARNINGS_DISCLOSURE
+ + ". Publishes regardless of outcome."
+)
+DEFAULT_OUTPUT = mb.OUTPUT_ARTIFACT_PATH
+#: The environment fields the reproduction needs to match the parent's.
+ENVIRONMENT_FIELDS: tuple[str, ...] = (
+ "python",
+ "packages.numpy",
+ "packages.pandas",
+ "packages.scipy",
+ "policyengine_us_parameters.revision",
+)
+#: The files this run composes, beyond the packages in
+#: :data:`COMPOSED_PACKAGES`; their SHA-256 go in the artifact. The
+#: registered commit on a clean tree fixes every byte; these digests make
+#: the composition auditable from the artifact alone.
+COMPOSED_SOURCES: tuple[str, ...] = (
+ "scripts/run_min_benefit_groups_registered.py",
+ "scripts/run_track_m_registered.py",
+ "scripts/run_track_a_registered.py",
+ "src/populace_dynamics/estimates/group_breakdown.py",
+ "src/populace_dynamics/estimates/lifetime_measures.py",
+ "src/populace_dynamics/estimates/uniform_cut_tabulation.py",
+ "src/populace_dynamics/estimates/cola_age_profile.py",
+ "src/populace_dynamics/estimates/career.py",
+ "src/populace_dynamics/estimates/parameters.py",
+ "src/populace_dynamics/cohorts/group_attributes.py",
+ "src/populace_dynamics/cohorts/psid2010.py",
+ "src/populace_dynamics/data/group_attributes_psid.py",
+ "src/populace_dynamics/data/social_security_receipt.py",
+ "src/populace_dynamics/data/prior_year_labor_income.py",
+ "data/external/group_category_schemes_v1.json",
+ "data/external/mint8_row_categories.json",
+ "data/external/mint8_lifetime_quintile_definitions.json",
+ "data/external/ssa_oasdi_tax_rates_2026.json",
+ "data/external/ssa_trust_fund_interest_rates_2026.json",
+ "docs/design/minimum_benefits_comparison.md",
+)
+#: Packages every module of which the run composes (all ``*.py``).
+COMPOSED_PACKAGES: tuple[str, ...] = (
+ "src/populace_dynamics/group_breakdowns",
+ "src/populace_dynamics/min_benefit_track_m",
+ "src/populace_dynamics/ss",
+)
+
+
+def _git(*args: str) -> str:
+ return subprocess.run(
+ ["git", "-C", str(ROOT), *args],
+ check=True,
+ capture_output=True,
+ text=True,
+ ).stdout.strip()
+
+
+def _script_module(name: str) -> Any:
+ path = ROOT / "scripts" / f"{name}.py"
+ spec = importlib.util.spec_from_file_location(f"_{name}", path)
+ module = importlib.util.module_from_spec(spec)
+ assert spec.loader is not None
+ spec.loader.exec_module(module)
+ return module
+
+
+def preflight(
+ *,
+ registration_pointer: str,
+ registered_commit: str,
+ output: Path,
+ git: Callable[..., str] = _git,
+ parent_path: Path = mb.PARENT_ARTIFACT_PATH,
+ parent_sha256: str = mb.PARENT_ARTIFACT_SHA256,
+ specification_path: Path = M1_SPECIFICATION_PATH,
+) -> dict[str, Any]:
+ """Refuse to run unless this is the registered one-shot state."""
+
+ if not isinstance(
+ registration_pointer, str
+ ) or not common.REGISTRATION_POINTER.fullmatch(registration_pointer):
+ raise ValueError(
+ "the registration pointer must be an issue #42 comment URL "
+ "(https://github.com/PolicyEngine/microcosm-dynamics/issues/42"
+ "#issuecomment-)"
+ )
+ if registration_pointer == mb.PARENT_REGISTRATION_POINTER:
+ raise ValueError(
+ "the breakdown's cells are new outcomes: they need their own "
+ "issue #42 registration, not Registration 17's"
+ )
+ if not re.fullmatch(r"[0-9a-f]{40}", registered_commit):
+ raise ValueError("--registered-commit must be a full 40-hex SHA")
+ head = git("rev-parse", "HEAD")
+ if head != registered_commit:
+ raise ValueError(
+ f"HEAD {head} is not the registered commit {registered_commit}"
+ )
+ if git("status", "--porcelain") != "":
+ raise ValueError("the working tree must be clean for a registered run")
+ output = Path(output)
+ if output.exists() or output.with_suffix(".env.json").exists():
+ raise ValueError(
+ f"{output} already exists: the registered run is one-shot"
+ )
+ parent = mb.load_parent_artifact(
+ parent_path, expected_sha256=parent_sha256
+ )
+ binding = mb.check_parent_binding(
+ parent,
+ specification_sha256=common.file_sha256(specification_path),
+ )
+ check_specification_for_registered_run(m1_parameter_block())
+ return {"head": head, "parent": parent, "binding": binding}
+
+
+def _field(record: Mapping[str, Any], dotted: str) -> Any:
+ value: Any = record
+ for key in dotted.split("."):
+ value = (value or {}).get(key) if isinstance(value, Mapping) else None
+ return value
+
+
+def check_environment(
+ environment: Mapping[str, Any], parent_sidecar: Mapping[str, Any]
+) -> dict[str, Any]:
+ """Refuse an environment other than the parent run's.
+
+ Compares :data:`ENVIRONMENT_FIELDS` with the parent sidecar's
+ ``environment``; the platform string is recorded, not compared.
+ """
+
+ recorded = parent_sidecar.get("environment") or {}
+ fields = {
+ name: {
+ "parent": _field(recorded, name),
+ "this_run": _field(environment, name),
+ }
+ for name in ENVIRONMENT_FIELDS
+ }
+ differ = [
+ name
+ for name, pair in fields.items()
+ if pair["parent"] is None or pair["parent"] != pair["this_run"]
+ ]
+ if differ:
+ raise ValueError(
+ f"the environment differs from the parent run's in {differ}; "
+ "the reproduction is compared bit for bit (build the recorded "
+ "environment with scripts/make_nasi_repro_venv.sh)"
+ )
+ return {
+ "fields": fields,
+ "matches_parent": True,
+ "platform_parent": recorded.get("platform"),
+ "platform_this_run": environment.get("platform"),
+ "gil_disabled_this_run": bool(
+ sysconfig.get_config_var("Py_GIL_DISABLED")
+ ),
+ "gil_note": (
+ "recorded, not compared: the parent's sidecar does not record "
+ "the interpreter build; scripts/make_nasi_repro_venv.sh builds "
+ "the default (GIL) CPython, and a build that changed any float "
+ "would show as a reproduction mismatch"
+ ),
+ }
+
+
+def check_parameter_pins(
+ pins: Mapping[str, str], parent: common.CommittedArtifact
+) -> dict[str, str]:
+ """The Track M parameter pins must equal the parent's recorded pins."""
+
+ recorded = (parent.document.get("checks") or {}).get("parameter_pins")
+ if dict(pins) != dict(recorded or {}):
+ raise ValueError(
+ "the Track M parameter files' SHA-256 differ from the parent "
+ "run's checks.parameter_pins"
+ )
+ return dict(pins)
+
+
+def composed_paths(
+ paths: Sequence[str] = COMPOSED_SOURCES,
+ packages: Sequence[str] = COMPOSED_PACKAGES,
+) -> tuple[str, ...]:
+ """Every composed file, repository-relative and sorted."""
+
+ found = set(paths)
+ for package in packages:
+ found.update(
+ str(path.relative_to(ROOT))
+ for path in (ROOT / package).glob("*.py")
+ )
+ return tuple(sorted(found))
+
+
+def source_sha256(paths: Sequence[str] | None = None) -> dict[str, str]:
+ """SHA-256 of every composed file (refuses a missing one)."""
+
+ return {
+ path: common.file_sha256(ROOT / path)
+ for path in (composed_paths() if paths is None else paths)
+ }
+
+
+def _default_cohort_loader() -> Any:
+ from populace_dynamics.min_benefit_track_m import cohort
+
+ return cohort.load_cohort_inputs()
+
+
+def _default_side_frame_loader(person_ids: Sequence[int]) -> Any:
+ from populace_dynamics.cohorts import group_attributes
+
+ return group_attributes.load_group_attributes(
+ person_ids, anchor_waves=mb.ANCHOR_WAVES
+ )
+
+
+def execute(
+ *,
+ registration_pointer: str,
+ registered_commit: str,
+ output: Path,
+ argv: Sequence[str],
+ git: Callable[..., str] = _git,
+ track_m_script: Any = None,
+ environment: Callable[..., dict[str, Any]] | None = None,
+ load_parameters: (
+ Callable[[Mapping[str, Any]], tuple[TrackMParameters, dict]] | None
+ ) = None,
+ load_cola: Callable[[], Any] | None = None,
+ load_cohort_inputs: Callable[[], Any] = _default_cohort_loader,
+ load_side_frame: Callable[
+ [Sequence[int]], Any
+ ] = _default_side_frame_loader,
+ parent_path: Path = mb.PARENT_ARTIFACT_PATH,
+ parent_sha256: str = mb.PARENT_ARTIFACT_SHA256,
+) -> dict[str, Any]:
+ """Preflight, the pre-PSID checks, the run, then the exclusive writes.
+
+ The keyword hooks default to the staged-PSID loaders, the real
+ resolvers and the committed parent; tests replace them with INVENTED
+ ones. Nothing is written unless :func:`~populace_dynamics.
+ group_breakdowns.min_benefit.run_group_breakdowns` returns.
+ """
+
+ started = datetime.datetime.now(datetime.timezone.utc).isoformat()
+ state = preflight(
+ registration_pointer=registration_pointer,
+ registered_commit=registered_commit,
+ output=output,
+ git=git,
+ parent_path=parent_path,
+ parent_sha256=parent_sha256,
+ )
+ parent: common.CommittedArtifact = state["parent"]
+ track_m = track_m_script or _script_module("run_track_m_registered")
+ block = m1_parameter_block()
+ track_m.check_runnable()
+ parameters, pins = (load_parameters or track_m.committed_parameters)(block)
+ pins = check_parameter_pins(pins, parent)
+ resolve = (
+ environment or _script_module("run_track_a_registered")._environment
+ )
+ env = resolve(ssa_parameters_revision=parameters.params.pe_us_revision)
+ env_check = check_environment(env, parent.sidecar or {})
+ if load_cola is None:
+ from populace_dynamics.estimates.parameters import load_cola_history
+
+ load_cola = load_cola_history
+ cola = load_cola()
+ document = mb.run_group_breakdowns(
+ parameters=parameters,
+ parent=parent,
+ data_provenance=REGISTERED_REAL,
+ registration_pointer=registration_pointer,
+ load_cohort_inputs=load_cohort_inputs,
+ load_side_frame=load_side_frame,
+ cola_rates=cola,
+ provenance_kind=PSID_FILES,
+ source={"cola_history": dict(cola.provenance)},
+ )
+ artifact = {
+ **document,
+ "header": REGISTERED_HEADER,
+ "reads_comparator_values": False,
+ "registration": {
+ "pointer": registration_pointer,
+ "binding_to_parent": state["binding"],
+ },
+ "checks": {
+ "parameter_pins": pins,
+ "environment": env_check,
+ },
+ "code_sha256": source_sha256(),
+ "run": {
+ "started": started,
+ "finished": datetime.datetime.now(
+ datetime.timezone.utc
+ ).isoformat(),
+ "registration_pointer": registration_pointer,
+ "registered_commit": registered_commit,
+ "git_head": state["head"],
+ "git_clean": True,
+ "command": " ".join(
+ [
+ "python",
+ "scripts/run_min_benefit_groups_registered.py",
+ *argv,
+ ]
+ ),
+ },
+ }
+ output = Path(output)
+ output.parent.mkdir(parents=True, exist_ok=True)
+ common.write_new(
+ output,
+ json.dumps(artifact, indent=2, sort_keys=False, allow_nan=False)
+ + "\n",
+ )
+ common.write_new(
+ output.with_suffix(".env.json"),
+ json.dumps(
+ {
+ "artifact": output.name,
+ "artifact_sha256": common.file_sha256(output),
+ "environment": env,
+ },
+ indent=2,
+ )
+ + "\n",
+ )
+ return artifact
+
+
+def main(argv: list[str] | None = None) -> int:
+ parser = argparse.ArgumentParser(description=__doc__.splitlines()[0])
+ parser.add_argument("--registration-pointer", required=True)
+ parser.add_argument("--registered-commit", required=True)
+ parser.add_argument("--output", type=Path, default=DEFAULT_OUTPUT)
+ args = parser.parse_args(argv)
+ execute(
+ registration_pointer=args.registration_pointer,
+ registered_commit=args.registered_commit,
+ output=args.output,
+ argv=list(sys.argv[1:] if argv is None else argv),
+ )
+ print(args.output)
+ return 0
+
+
+if __name__ == "__main__":
+ raise SystemExit(main())
diff --git a/src/populace_dynamics/group_breakdowns/__init__.py b/src/populace_dynamics/group_breakdowns/__init__.py
new file mode 100644
index 00000000..2e7de2a1
--- /dev/null
+++ b/src/populace_dynamics/group_breakdowns/__init__.py
@@ -0,0 +1,26 @@
+"""Post hoc group breakdowns of the finished DYNASIM3 blind tests (NASI).
+
+After the NASI meeting of 2026-10-01 Max asked for each finished blind
+test's results by SSA's MINT8 characteristic subgroups. Each module of
+this package is the adapter of one test: it re-executes that test's frozen
+registered computation exactly, composing the existing code read-only,
+refuses unless every committed cell it recomputes equals the committed
+artifact, and only then joins the person attributes (the cohort side frame
+of :mod:`populace_dynamics.cohorts.group_attributes` and the lifetime
+measures of :mod:`populace_dynamics.estimates.lifetime_measures`) and cuts
+the MINT-scheme cells with :mod:`populace_dynamics.estimates.
+group_breakdown`.
+
+* :mod:`.common`: the post hoc labels, the exact comparison with a
+ committed artifact, the mapping of the cohort side frame onto the MINT8
+ scheme's codes, and INVENTED side-frame inputs for dry runs.
+* :mod:`.min_benefit`: exercise 4 (Track M, the minimum-benefit share of
+ OASDI beneficiaries 62 and older, PSID 2023 wave, income year 2022).
+
+A real-data breakdown is new outcomes: it runs only through its registered
+entry script, under its own issue #42 registration, and every output
+carries the test's own labels plus "registered, one-shot, post hoc, not
+blind" and "report-only". Development uses INVENTED data.
+
+Submodules are imported explicitly; this initializer imports none of them.
+"""
diff --git a/src/populace_dynamics/group_breakdowns/common.py b/src/populace_dynamics/group_breakdowns/common.py
new file mode 100644
index 00000000..3b800c51
--- /dev/null
+++ b/src/populace_dynamics/group_breakdowns/common.py
@@ -0,0 +1,636 @@
+"""Shared pieces of the post hoc group-breakdown adapters (NASI follow-up).
+
+Every adapter in :mod:`populace_dynamics.group_breakdowns` follows one
+order, and this module holds what the order needs that is not specific to
+a test:
+
+1. **Exact reproduction before any group work.** The adapter re-executes
+ its test's registered computation and compares what it recomputes with
+ the committed artifact (:func:`compare_exact`): every JSON leaf must be
+ equal in type and value, and every float equal bit for bit
+ (``float.hex``), after both sides pass through the same JSON encoding
+ the registered runners use (``json.dumps(..., allow_nan=False)``). No
+ tolerance is applied (:data:`EXACT_COMPARISON`). A difference raises
+ :class:`ReproductionMismatchError` (:func:`require_identical`) before
+ any group attribute is read; the error and the check records name the
+ differing paths only, never a value, so a refusal on real data prints
+ no outcome.
+2. **Labels.** Every output carries the test's own labels plus
+ :data:`POST_HOC_LABELS` ("registered, one-shot, post hoc, not blind",
+ the d479 label of Track A v2, ``track_a_v2/manifest.py``; and
+ "report-only", the paper's term for cells published with no tolerance).
+3. **The cohort side frame on the MINT8 scheme.** The person attributes
+ of :mod:`populace_dynamics.cohorts.group_attributes` (G1) carry SSA's
+ verbatim MINT8 row labels; :func:`side_frame_codes` maps each to the
+ category of :mod:`populace_dynamics.estimates.group_breakdown` (G3)
+ whose label equals it, and names every row G1 does not place
+ (``attribute_unknown:`` or G1's ``unresolved:``) as
+ a declared unclassified code, so G3 counts it by reason and enters it
+ in Total only. Education enters G3 as G1's integer years, and
+ :func:`education_agreement` holds G3's band equal to G1's own MINT8
+ label for every person.
+4. **INVENTED inputs.** :func:`invented_group_attribute_inputs` draws
+ head/spouse reports and education rows in the shapes and code domains
+ G1's builder validates, so a dry run passes INVENTED data through G1's
+ real builder. Its provenance says INVENTED.
+
+This module reads no PSID file and computes no outcome.
+"""
+
+from __future__ import annotations
+
+import hashlib
+import json
+from collections.abc import Iterable, Mapping, Sequence
+from dataclasses import dataclass
+from pathlib import Path
+from typing import Any
+
+import numpy as np
+import pandas as pd
+
+from populace_dynamics.cohorts import group_attributes as g1
+from populace_dynamics.data import group_attributes_psid as gap
+from populace_dynamics.estimates import group_breakdown as g3
+
+__all__ = [
+ "CommittedArtifact",
+ "EXACT_COMPARISON",
+ "INVENTED_DATA_HEADER",
+ "MARITAL_STATUS_CODES",
+ "POST_HOC_LABEL",
+ "POST_HOC_LABELS",
+ "REGISTRATION_POINTER",
+ "REPORT_ONLY_LABEL",
+ "ReproductionCheck",
+ "ReproductionMismatchError",
+ "SchemeCodes",
+ "age_in_year",
+ "compare_exact",
+ "education_agreement",
+ "education_years",
+ "file_sha256",
+ "invented_group_attribute_inputs",
+ "json_normalized",
+ "leaf_differences",
+ "load_committed_artifact",
+ "marital_codes",
+ "require_identical",
+ "side_frame_codes",
+ "write_new",
+]
+
+#: Max's d479 label for a post hoc rerun (``track_a_v2/manifest.py``).
+POST_HOC_LABEL = "registered, one-shot, post hoc, not blind"
+#: The paper's term for cells published with no tolerance (paper.qmd).
+REPORT_ONLY_LABEL = "report-only"
+POST_HOC_LABELS: tuple[str, ...] = (POST_HOC_LABEL, REPORT_ONLY_LABEL)
+#: The header of every invented-data output: Track M's ``DRY_RUN_HEADER``
+#: and Track A v2's ``INVENTED_HEADER`` (a test holds the three equal).
+#: G3 also puts its own invented label first among its labels.
+INVENTED_DATA_HEADER = "INVENTED DATA - NOT A COMPARISON"
+#: An issue #42 comment, the registration issue (G3's pattern).
+REGISTRATION_POINTER = g3.REGISTRATION_POINTER
+
+#: How a recomputed result is compared with the committed artifact.
+EXACT_COMPARISON: dict[str, Any] = {
+ "rule": (
+ "exact: both sides are encoded with json.dumps(allow_nan=False) "
+ "and decoded, then every leaf must be equal in type and value, "
+ "every float bit for bit (float.hex, so -0.0 differs from 0.0), "
+ "every mapping with the same keys and every list with the same "
+ "length and order"
+ ),
+ "tolerance": None,
+ "tolerance_note": (
+ "none: the registered runners write floats with Python's "
+ "shortest round-trip repr, so a reproduction in the recorded "
+ "environment is bit-identical; any difference, including a "
+ "last-bit difference from floating-point summation in another "
+ "numpy build, refuses the run before any group attribute is read"
+ ),
+ "reports": (
+ "differing paths only, never a value (a refusal on real data "
+ "prints no outcome)"
+ ),
+}
+
+
+class ReproductionMismatchError(ValueError):
+ """A recomputed cell differs from the committed artifact: refuse."""
+
+
+# =========================================================================
+# Exact comparison
+# =========================================================================
+def json_normalized(value: Any) -> Any:
+ """``value`` through the registered runners' JSON encoding."""
+
+ return json.loads(json.dumps(value, allow_nan=False))
+
+
+def _same_leaf(left: Any, right: Any) -> bool:
+ if type(left) is not type(right):
+ return False
+ if isinstance(left, float):
+ return left.hex() == right.hex()
+ return left == right
+
+
+def leaf_differences(
+ committed: Any, recomputed: Any, path: str = "$"
+) -> tuple[list[str], int]:
+ """``(paths that differ, leaves compared)`` of two decoded documents.
+
+ Mappings must have the same keys (a key on one side only is a
+ difference at that key's path), lists the same length; leaves must be
+ equal in type and value, floats bit for bit. Paths name keys and list
+ positions only, never a value.
+ """
+
+ if isinstance(committed, dict) and isinstance(recomputed, dict):
+ differences: list[str] = []
+ compared = 0
+ for key in sorted(set(committed) | set(recomputed), key=str):
+ where = f"{path}.{key}"
+ if key not in committed or key not in recomputed:
+ differences.append(f"{where} (present on one side only)")
+ continue
+ found, n = leaf_differences(committed[key], recomputed[key], where)
+ differences.extend(found)
+ compared += n
+ return differences, compared
+ if isinstance(committed, list) and isinstance(recomputed, list):
+ if len(committed) != len(recomputed):
+ return [f"{path} (lengths differ)"], 0
+ differences = []
+ compared = 0
+ for index, (left, right) in enumerate(
+ zip(committed, recomputed, strict=True)
+ ):
+ found, n = leaf_differences(left, right, f"{path}[{index}]")
+ differences.extend(found)
+ compared += n
+ return differences, compared
+ if isinstance(committed, dict | list) or isinstance(
+ recomputed, dict | list
+ ):
+ return [f"{path} (structure differs)"], 0
+ return ([] if _same_leaf(committed, recomputed) else [path]), 1
+
+
+@dataclass(frozen=True)
+class ReproductionCheck:
+ """One comparison of a recomputed block with the committed one."""
+
+ name: str
+ n_leaves: int
+ differences: tuple[str, ...]
+
+ @property
+ def identical(self) -> bool:
+ return not self.differences
+
+ def as_dict(self, *, max_paths: int = 50) -> dict[str, Any]:
+ return {
+ "name": self.name,
+ "identical": self.identical,
+ "n_leaves_compared": self.n_leaves,
+ "n_differences": len(self.differences),
+ "differing_paths": list(self.differences[:max_paths]),
+ }
+
+
+def compare_exact(
+ name: str, committed: Any, recomputed: Any
+) -> ReproductionCheck:
+ """``recomputed`` against ``committed`` (:data:`EXACT_COMPARISON`)."""
+
+ found, n = leaf_differences(
+ json_normalized(committed), json_normalized(recomputed)
+ )
+ return ReproductionCheck(name, n, tuple(found))
+
+
+def require_identical(
+ checks: Iterable[ReproductionCheck], *, stage: str
+) -> None:
+ """Refuse unless every check is identical (paths only in the error)."""
+
+ failed = [check for check in checks if not check.identical]
+ if failed:
+ listed = "; ".join(
+ f"{check.name}: {len(check.differences)} differing paths, "
+ f"first {list(check.differences[:5])}"
+ for check in failed
+ )
+ raise ReproductionMismatchError(
+ f"{stage}: the recomputed result differs from the committed "
+ f"artifact ({listed}); refused before any group attribute or "
+ "group cell, and nothing is written"
+ )
+
+
+# =========================================================================
+# Files
+# =========================================================================
+def file_sha256(path: Path | str) -> str:
+ return hashlib.sha256(Path(path).read_bytes()).hexdigest()
+
+
+def write_new(path: Path, text: str) -> None:
+ """Create ``path`` exclusively: never overwrite (one shot)."""
+
+ with Path(path).open("x", encoding="utf-8") as handle:
+ handle.write(text)
+
+
+@dataclass(frozen=True)
+class CommittedArtifact:
+ """A committed run artifact, its SHA-256 and its sidecar.
+
+ ``path``, ``sha256`` and ``sidecar`` are ``None`` for an INVENTED
+ parent held in memory by a dry run.
+ """
+
+ document: Mapping[str, Any]
+ path: Path | None = None
+ sha256: str | None = None
+ sidecar: Mapping[str, Any] | None = None
+
+ def record(self, root: Path | None = None) -> dict[str, Any]:
+ where = None
+ if self.path is not None:
+ where = str(self.path)
+ if root is not None:
+ try:
+ where = str(Path(self.path).relative_to(root))
+ except ValueError:
+ pass
+ return {
+ "path": where,
+ "sha256": self.sha256,
+ "sidecar_artifact_sha256": (
+ None
+ if self.sidecar is None
+ else self.sidecar.get("artifact_sha256")
+ ),
+ "registration_pointer": self.document.get("registration_pointer"),
+ "registered_commit": (self.document.get("run") or {}).get(
+ "registered_commit"
+ ),
+ }
+
+
+def load_committed_artifact(
+ path: Path, *, expected_sha256: str
+) -> CommittedArtifact:
+ """Read a committed artifact, refusing other bytes or an unbound sidecar.
+
+ The file's SHA-256 must equal ``expected_sha256`` and its ``.env.json``
+ sidecar must name the file and bind the same SHA-256.
+ """
+
+ path = Path(path)
+ digest = file_sha256(path)
+ if digest != expected_sha256:
+ raise ValueError(
+ f"{path.name} has SHA-256 {digest}, not the committed "
+ f"{expected_sha256}"
+ )
+ sidecar_path = path.with_suffix(".env.json")
+ sidecar = json.loads(sidecar_path.read_text(encoding="utf-8"))
+ if sidecar.get("artifact") != path.name:
+ raise ValueError(f"{sidecar_path.name} does not name {path.name}")
+ if sidecar.get("artifact_sha256") != digest:
+ raise ValueError(
+ f"{sidecar_path.name} does not bind {path.name}'s bytes"
+ )
+ return CommittedArtifact(
+ document=json.loads(path.read_text(encoding="utf-8")),
+ path=path,
+ sha256=digest,
+ sidecar=sidecar,
+ )
+
+
+# =========================================================================
+# The cohort side frame (G1) on G3's MINT8 codes
+# =========================================================================
+#: Per side-frame dimension: G1's MINT8 label column, its status column
+#: and the attribute's own status column.
+_SIDE_FRAME_COLUMNS: dict[str, tuple[str, str, str]] = {
+ "race_ethnicity": (
+ "race_ethnicity_mint8",
+ "race_ethnicity_mint8_status",
+ "race_ethnicity_status",
+ ),
+ "country_of_birth": (
+ "country_of_birth_mint8",
+ "country_of_birth_mint8_status",
+ "country_of_birth_status",
+ ),
+}
+
+
+@dataclass(frozen=True)
+class SchemeCodes:
+ """Codes for one G3 dimension and the unclassified codes declared."""
+
+ codes: tuple[Any, ...]
+ unclassified_codes: tuple[str, ...]
+ rule: str
+
+
+def _text(value: Any) -> str | None:
+ if value is None or value is pd.NA:
+ return None
+ if isinstance(value, float) and np.isnan(value):
+ return None
+ return str(value)
+
+
+def side_frame_codes(
+ attributes: pd.DataFrame, dimension: g3.Dimension
+) -> SchemeCodes:
+ """G1's MINT8 labels as the codes of G3's ``dimension``.
+
+ A row G1 assigns takes the code of the G3 category whose printed
+ label equals G1's label (both are SSA's verbatim MINT8 row labels; a
+ label no category prints is refused). A row G1 leaves unresolved
+ takes G1's status (``unresolved:``, for example a U.S.
+ territory birth); a row whose attribute is unknown takes
+ ``attribute_unknown:`` (for example
+ ``never_head_or_spouse``: the PSID asks race and birthplace of heads
+ and spouses only). Those codes are declared unclassified.
+ """
+
+ if dimension.key not in _SIDE_FRAME_COLUMNS:
+ raise ValueError(
+ f"{dimension.key} is not a side-frame dimension "
+ f"({sorted(_SIDE_FRAME_COLUMNS)})"
+ )
+ label_column, status_column, attribute_column = _SIDE_FRAME_COLUMNS[
+ dimension.key
+ ]
+ by_label = {
+ category.label: category.codes[0] for category in dimension.categories
+ }
+ codes: list[Any] = []
+ declared: set[str] = set()
+ for label, status, attribute in zip(
+ attributes[label_column],
+ attributes[status_column],
+ attributes[attribute_column],
+ strict=True,
+ ):
+ label, status, attribute = (
+ _text(label),
+ _text(status),
+ _text(attribute),
+ )
+ if status == g1.ASSIGNED:
+ if label not in by_label:
+ raise ValueError(
+ f"{dimension.key}: G1 label {label!r} is not one of "
+ f"G3's labels {sorted(by_label)}"
+ )
+ codes.append(by_label[label])
+ continue
+ if label is not None:
+ raise ValueError(
+ f"{dimension.key}: a row G1 does not assign carries a label"
+ )
+ if status is not None and status.startswith("unresolved:"):
+ code = status
+ elif status == g1.ATTRIBUTE_UNKNOWN:
+ code = f"{g1.ATTRIBUTE_UNKNOWN}:{attribute or 'unknown'}"
+ else:
+ raise ValueError(
+ f"{dimension.key}: undocumented G1 status {status!r}"
+ )
+ codes.append(code)
+ declared.add(code)
+ return SchemeCodes(
+ codes=tuple(codes),
+ unclassified_codes=tuple(sorted(declared)),
+ rule=(
+ f"G1 column {label_column!r} (SSA's verbatim MINT8 label) to the "
+ f"G3 {dimension.key!r} category printing that label; a row G1 "
+ f"does not assign is unclassified with code "
+ f"'unresolved:' ({status_column}) or "
+ f"'attribute_unknown:<{attribute_column}>'"
+ ),
+ )
+
+
+def education_years(attributes: pd.DataFrame) -> tuple[int | None, ...]:
+ """G1's resolved years of schooling, ``None`` where unknown."""
+
+ out: list[int | None] = []
+ for value in attributes["education_years"]:
+ if value is None or value is pd.NA or pd.isna(value):
+ out.append(None)
+ else:
+ out.append(int(value))
+ return tuple(out)
+
+
+def education_agreement(
+ attributes: pd.DataFrame, dimension: g3.Dimension
+) -> dict[str, Any]:
+ """Hold G3's MINT8 education band equal to G1's MINT8 label.
+
+ Both encode SSA's definition (Graduate more than 16 years, Bachelor
+ 16, Associate 14-15, High school 12-13, Less than high school under
+ 12); this is the differential check between the two implementations.
+ Refuses a person they place differently.
+ """
+
+ agree = 0
+ for years, label in zip(
+ education_years(attributes),
+ attributes["education_mint8"],
+ strict=True,
+ ):
+ label = _text(label)
+ if years is None:
+ if label is not None:
+ raise ValueError("G1 labels an education it has no years for")
+ continue
+ hits = [c.label for c in dimension.categories if c.contains(years)]
+ g3_label = hits[0] if hits else None
+ if g3_label != label:
+ raise ValueError(
+ "G1's and G3's MINT8 education bands place a person "
+ "differently"
+ )
+ agree += 1
+ return {
+ "n_with_years": agree,
+ "g3_band_equals_g1_label": True,
+ "rule": (
+ "for every person with resolved years, the G3 band containing "
+ "the years prints G1's education_mint8 label"
+ ),
+ }
+
+
+#: MINT8's four marital statuses, as G3's codes.
+MARITAL_STATUS_CODES: tuple[str, ...] = (
+ "married",
+ "divorced",
+ "widowed",
+ "never_married",
+)
+
+
+def marital_codes(
+ statuses: Sequence[Any], *, unclassified: Sequence[str]
+) -> SchemeCodes:
+ """Marital statuses as G3 codes; ``unclassified`` statuses declared.
+
+ A status that is neither one of MINT8's four nor declared is refused
+ by G3 (``assign_groups``), never dropped.
+ """
+
+ declared = tuple(sorted(set(unclassified)))
+ held = set(declared) & set(MARITAL_STATUS_CODES)
+ if held:
+ raise ValueError(f"MINT8 statuses {sorted(held)} cannot be declared")
+ codes = tuple(_text(status) for status in statuses)
+ present = tuple(code for code in declared if code in set(codes))
+ return SchemeCodes(
+ codes=codes,
+ unclassified_codes=present,
+ rule=(
+ "the status in the analysis year, MINT8's four statuses as "
+ f"themselves; {list(declared)} unclassified"
+ ),
+ )
+
+
+def age_in_year(birth_years: Iterable[Any], year: int) -> tuple[int, ...]:
+ """Age in ``year`` as ``year - birth_year`` (whole years)."""
+
+ out = []
+ for birth in birth_years:
+ if isinstance(birth, bool | np.bool_):
+ raise ValueError("birth years must be integers")
+ out.append(int(year) - int(birth))
+ return tuple(out)
+
+
+# =========================================================================
+# INVENTED side-frame inputs
+# =========================================================================
+_REPORT_COLUMNS = ["person_id", "wave", "role", *gap.REPORT_CODE_COLUMNS]
+_EDUCATION_COLUMNS = [
+ "person_id",
+ "wave",
+ "education_code",
+ "role",
+ "family_education_code",
+]
+
+
+def _frame(rows: list[dict[str, Any]], columns: list[str]) -> pd.DataFrame:
+ frame = pd.DataFrame(rows, columns=columns)
+ for column in frame:
+ frame[column] = frame[column].astype(
+ "string" if column == "role" else "Int64"
+ )
+ return frame
+
+
+def invented_group_attribute_inputs(
+ persons: pd.DataFrame, *, wave: int, seed: int = 0
+) -> g1.GroupAttributeInputs:
+ """INVENTED inputs for G1's builder (INVENTED DATA - NOT A COMPARISON).
+
+ ``persons`` has ``person_id`` and ``role`` (``"head"``, ``"spouse"`` or
+ missing for another family-unit member). Each head and spouse gets one
+ report in ``wave`` with a Spanish-descent code (0 not Hispanic, 1
+ Hispanic, 9 DK/NA), one or two race mentions, a completed-education
+ code and a birthplace pair (a state with year-came 0, a territory, a
+ foreign country with a year of arrival, or DK/NA); every person gets an
+ individual education row, with code 99 (NA) for a few. Codes are
+ drawn inside the wave's documented domains, which G1's builder checks
+ against the committed codebook values. No PSID value enters.
+ """
+
+ items = gap.FAMILY_ITEMS[int(wave)]
+ if int(wave) not in gap.BIRTHPLACE_WAVES or not items.asks_hispanic_origin:
+ raise ValueError(
+ "invented side-frame inputs are drawn for a wave that asks "
+ "Spanish descent and birthplace (2013-2023)"
+ )
+ rng = np.random.default_rng(seed)
+ reports: list[dict[str, Any]] = []
+ education: list[dict[str, Any]] = []
+ for pid, role in zip(persons["person_id"], persons["role"], strict=True):
+ pid = int(pid)
+ role = _text(role)
+ years = int(rng.choice(np.arange(6, 18)))
+ code = 99 if rng.random() < 0.03 else years
+ family_code = None
+ if role is not None:
+ if role not in gap.ROLES:
+ raise ValueError(f"undocumented role {role!r}")
+ family_code = code
+ row: dict[str, Any] = dict.fromkeys(_REPORT_COLUMNS, pd.NA)
+ row.update(person_id=pid, wave=int(wave), role=role)
+ row["hispanic_code"] = int(
+ rng.choice([0, 1, 9], p=[0.8, 0.14, 0.06])
+ )
+ first = int(
+ rng.choice([1, 2, 4, 7, 9], p=[0.6, 0.2, 0.08, 0.1, 0.02])
+ )
+ second = int(rng.choice([1, 2, 3])) if rng.random() < 0.06 else 0
+ mentions = [first, second if first != 9 else 0, 0, 0]
+ for index in range(len(items.race[role])):
+ row[f"race_code_{index + 1}"] = mentions[index]
+ row["family_education_code"] = family_code
+ place = rng.random()
+ if place < 0.84:
+ state, came = int(rng.integers(1, 57)), 0
+ elif place < 0.87:
+ state, came = 0, 0
+ elif place < 0.98:
+ state, came = 0, int(rng.integers(1950, 2016))
+ else:
+ state, came = 99, 0
+ row["birth_state_code"] = state
+ row["year_came_code"] = came
+ reports.append(row)
+ education.append(
+ {
+ "person_id": pid,
+ "wave": int(wave),
+ "education_code": code,
+ "role": pd.NA if role is None else role,
+ "family_education_code": (
+ pd.NA if family_code is None else family_code
+ ),
+ }
+ )
+ universe = pd.Series(
+ np.sort(persons["person_id"].astype("int64").unique()),
+ name="person_id",
+ dtype="int64",
+ )
+ return g1.GroupAttributeInputs(
+ reports=_frame(reports, _REPORT_COLUMNS),
+ education=_frame(education, _EDUCATION_COLUMNS),
+ universe=universe,
+ provenance={
+ "kind": "INVENTED",
+ "label": INVENTED_DATA_HEADER,
+ "generator": (
+ "populace_dynamics.group_breakdowns.common."
+ "invented_group_attribute_inputs"
+ ),
+ "wave": int(wave),
+ "seed": int(seed),
+ },
+ )
diff --git a/src/populace_dynamics/group_breakdowns/min_benefit.py b/src/populace_dynamics/group_breakdowns/min_benefit.py
new file mode 100644
index 00000000..69bddf9e
--- /dev/null
+++ b/src/populace_dynamics/group_breakdowns/min_benefit.py
@@ -0,0 +1,1882 @@
+"""Exercise 4 (Track M) by MINT8 characteristic subgroups (NASI G4c).
+
+Exercise 4 of the DYNASIM3 scorecard measured, on the PSID's 2023 wave
+(income year 2022), the share of OASDI beneficiaries aged 62 and older who
+would receive a minimum benefit under Table 5's options 2-5, for All, Men
+and Women (Registration 17, issue #42 comment 5855753783, registered
+commit ``2e4e08be``; committed artifact
+``runs/replication_urban2006_minimum_benefit_v1.json``). This module
+breaks that statistic down by SSA's MINT8 characteristic subgroups. It is
+post hoc: the committed run's comparator values are public, so every
+output is labelled "registered, one-shot, post hoc, not blind" and
+"report-only" beside Track M's own labels, and none is scored.
+
+The order, which :func:`run_group_breakdowns` enforces
+------------------------------------------------------
+1. **Parameters before data.** The parameters' provenance
+ (``TrackMParameters.source()``: the oracle's policyengine-us revision,
+ the quarter-of-coverage capture and the Census thresholds) must equal
+ the committed artifact's ``parameters`` block.
+2. **The same PSID files.** The cohort inputs are read (Track M's
+ ``cohort.load_cohort_inputs``, which records every file's SHA-256) and
+ the files' SHA-256 must equal the committed ``inputs.source.
+ psid_files_sha256``.
+3. **Re-execution of the frozen registered computation**
+ (:func:`reexecute_track_m`), composing Track M read-only exactly as
+ ``scripts/run_track_m_registered.py``'s ``run_pipeline`` does: the M4
+ cohort under the scored reading and under cos d430's sensitivity
+ reading, the M5 careers, and ``pipeline.run_track_m`` with the
+ committed run's own registration pointer, plus each cohort's
+ ``cohort_structure``; and, for the person rows the groups need,
+ ``evaluation.evaluate(inputs, policy_for_row(row), params)`` and the
+ unchanged ``tabulation.tabulate_track_m`` for every registered row
+ MS0-MS6 and for d430's sensitivity.
+4. **Exact reproduction before any group work**
+ (:func:`reproduction_checks`). Every block of the re-executed pipeline
+ result (every key ``run_track_m`` returns, the rows' tabulations and
+ diagnostics, the sensitivity, both cohort structures, the inputs'
+ provenance; not ``header`` and ``specification``, which the registered
+ runner overwrote) and every per-row tabulation and diagnostics block
+ must equal the committed artifact exactly
+ (:data:`.common.EXACT_COMPARISON`: equal JSON leaves, floats bit for
+ bit, no tolerance). On any difference the run raises
+ :class:`.common.ReproductionMismatchError` before the group-attribute
+ loader is called, and nothing is written.
+5. **Person attributes** (:func:`person_attributes`), joined by
+ ``person_id``: sex, age, marital status and the 2022 benefit type from
+ Track M's own cohort; race and ethnicity, country of birth and
+ education from the cohort side frame
+ (:mod:`populace_dynamics.cohorts.group_attributes`, G1, anchor wave
+ 2023); the initial AIME at 62 and the lifetime payroll-tax present
+ values, own and shared, from
+ :mod:`populace_dynamics.estimates.lifetime_measures` (G2).
+6. **Cells** (:func:`tabulate_breakdowns`): G3's
+ ``tabulate_share_breakdown`` for every registered row and option 2-5
+ and for d430's sensitivity, with the design-based standard error and
+ the five-seed half-sample floor Track M registered.
+7. **Consistency with the registered cells**
+ (:func:`consistency_checks`): G3's Total, Female and Male cells must
+ equal ``tabulate_track_m``'s All, Women and Men cells exactly (share,
+ weighted and unweighted counts, design SE, floor); otherwise refuse.
+
+The scheme (:data:`SCHEME`)
+---------------------------
+G3's ``MINT8_ANNUAL_WITH_LIFETIME_SCHEME`` (MINT8's annual beneficiary
+rows, with MINT8's three cohort-table quintile measures appended as the
+lifetime-earnings dimension) without the two rows Track M cannot compute
+(:data:`NOT_COMPUTED_DIMENSIONS`). Its conventions here:
+
+* **Sex**: Track M's (ER32000 through the death records); ``unknown``
+ (code 9) is unclassified, as Track M counts it in All only.
+* **Age**: 2022 minus the birth year Track M resolved (the first-estimates
+ birth-year law). The universe is born 1960 or earlier, so every age is
+ 62 or more and MINT8's "60–69" row holds ages 62-69 here; "90 or older"
+ is MINT8's top row.
+* **Marital status**: Track M's ``marital_status_2022`` (the marriage
+ history's state at the end of 2022, separated counted as married by
+ ``psid2010.marital_state_at``'s default), MINT8's four as themselves;
+ ``unknown`` and ``no_marriage_history`` are unclassified.
+* **Race and ethnicity, country of birth, education**: G1's MINT8 scheme
+ (``data/external/group_category_schemes_v1.json``) through
+ :func:`.common.side_frame_codes` and G1's years of schooling, education
+ resolved as of the 2023 wave.
+* **Current-law benefit type** (:func:`benefit_type_2022`,
+ :data:`BENEFIT_TYPE_RULES`): Track M's own reading of the 2022 receipt
+ mapped to MINT8's four rows. A survivor or dependent mention among the
+ six G33A items ER35213-ER35218 the cohort reads makes the person a
+ widow(er) or a spouse (dually entitled included); otherwise the benefit
+ is the own worker benefit Track M pays (its scored reading counts an
+ unknown or "other" type as own receipt, cos d430), disabled or retired
+ by the worker type mentioned or, when that does not settle it, by the
+ basis of Track M's own worker record, and a disabled worker at or over
+ the full retirement age is a retired worker (MINT8). "Only" that
+ cannot be established is unclassified, with its reason.
+* **Lifetime earnings** (MINT8's cohort-table measures):
+
+ - *initial AIME quintile*: G2's ``initial_aime_at_62`` under the
+ exercise-4 convention (``AIME_CONVENTIONS["exercise_4_min_benefit"]``:
+ statutory computation years, earnings through the year of attaining
+ 61) over the MS0 evaluation's one history of the person's own worker
+ record, which re-expresses MS0's AIME (indexed to the year of
+ attaining 60, computed through the year before entitlement) at age 62;
+ for an old-age record entitled in the year of attaining 62 the two are
+ the same number, which the run checks. A person with no own record
+ (paid only a spouse's or survivor's benefit) has no initial AIME and
+ is unclassified;
+ - *lifetime payroll tax quintile*, own and shared: G2's
+ ``lifetime_payroll_tax_pv_at_62`` over each person's lifetime history
+ through 2022 built by Track M's own rule (``careers.
+ observed_histories`` of the earnings panel, ``prior_year_labor_income.
+ next_wave_histories``, then ``coverage.one_history`` with MS0's
+ policy), spouses included for the shared measure, G2's registered
+ builder defaults (combined OASDI rate on earnings capped at the
+ taxable maximum, the OASDI trust funds' effective interest rate,
+ separated counted as married, a spouse with no history counted own).
+
+ Quintiles are cut by G3's weighted-quintile rule with the 2023
+ cross-section weight within each 10-year birth cohort of the universe
+ (MINT8: "We calculate the AIME quintiles for each birth cohort"), so the
+ "1960–1969" cohort holds only the 1960 births here.
+
+Group attributes are person attributes: computed once (MS0's histories,
+the scored reading's cohort) and used for every row and for d430's
+sensitivity, whose universe and person order are the same.
+
+What is not computed (:data:`NOT_COMPUTED_DIMENSIONS`,
+:data:`NOT_COMPUTED_STATISTICS`): current-law poverty status and household
+income quintile (Track M's inputs carry no 2022 family money income), and
+MINT8's benefit-change statistics (``evaluate`` returns receipt flags,
+not person-level benefits). Each is recorded with its reason.
+
+Nothing here reads a PSID file itself, and nothing here writes: the
+cohort-input and group-attribute loaders are passed in (the registered
+entry script passes the staged-PSID loaders, the dry run INVENTED ones),
+and the caller writes the returned document. It never reads a comparator
+value.
+"""
+
+from __future__ import annotations
+
+import copy
+from collections import Counter
+from collections.abc import Callable, Mapping, Sequence
+from dataclasses import dataclass
+from pathlib import Path
+from typing import Any
+
+import pandas as pd
+
+from populace_dynamics.cohorts import group_attributes as g1
+from populace_dynamics.cohorts import psid2010
+from populace_dynamics.data import prior_year_labor_income as pyl
+from populace_dynamics.data import social_security_receipt as ssr
+from populace_dynamics.estimates import group_breakdown as g3
+from populace_dynamics.estimates import lifetime_measures as g2
+from populace_dynamics.estimates.parameters import COLA_HISTORY_PATH
+from populace_dynamics.group_breakdowns import common
+from populace_dynamics.min_benefit_track_m import (
+ COVERED_EARNINGS_DISCLOSURE,
+ OUTPUT_LABELS,
+ careers,
+ cohort,
+ coverage,
+ pipeline,
+ structure,
+ tabulation,
+)
+from populace_dynamics.min_benefit_track_m.evaluation import (
+ INVENTED,
+ PSID_FILES,
+ Evaluation,
+ TrackMInputs,
+ TrackMParameters,
+ evaluate,
+)
+from populace_dynamics.min_benefit_track_m.policy import (
+ OWN_RECEIPT_PRE62_UNKNOWN_OR_OTHER_UNOBSERVED,
+ OWN_RECEIPT_SENSITIVITY_ID,
+ REGISTERED_ROWS,
+ SENSITIVITIES,
+ TABLE6_OPTIONS,
+ policy_for_row,
+)
+from populace_dynamics.ss.params import SSAParameters
+
+__all__ = [
+ "AIME_CONVENTION",
+ "ANALYSIS_YEAR",
+ "COLA_HISTORY_RELATIVE_PATH",
+ "ANCHOR_WAVES",
+ "BENEFIT_TYPE_RULES",
+ "COLUMNS_MAP",
+ "GROUP_NAMED_DELTAS",
+ "LABELS",
+ "LOCATION_FIELDS",
+ "LOCATION_RULE",
+ "NOT_COMPUTED_DIMENSIONS",
+ "NOT_COMPUTED_STATISTICS",
+ "OUTPUT_ARTIFACT_PATH",
+ "PARENT_ARTIFACT_PATH",
+ "PARENT_ARTIFACT_SHA256",
+ "PARENT_REGISTERED_COMMIT",
+ "PARENT_REGISTRATION_POINTER",
+ "PARENT_SPECIFICATION_SHA256",
+ "PersonAttributes",
+ "SCHEMA_VERSION",
+ "SCHEME",
+ "SENSITIVITY_KEY",
+ "YOUNGEST_AGE",
+ "STATISTIC_ID",
+ "TrackMReexecution",
+ "benefit_type_2022",
+ "check_parent_binding",
+ "located",
+ "relative_location",
+ "consistency_checks",
+ "initial_aime_measure",
+ "invented_group_attribute_loader",
+ "lifetime_careers",
+ "load_parent_artifact",
+ "person_attributes",
+ "reexecute_track_m",
+ "reproduce_parent",
+ "reproduction_checks",
+ "run_group_breakdowns",
+ "tabulate_breakdowns",
+]
+
+SCHEMA_VERSION = "populace_dynamics.group_breakdowns.min_benefit.v1"
+_ROOT = Path(__file__).resolve().parents[3]
+#: The committed exercise-4 artifact (Registration 17) and its pins, as
+#: ``tests/test_replication_urban2006_minimum_benefit.py`` pins them.
+PARENT_ARTIFACT_PATH = (
+ _ROOT / "runs" / "replication_urban2006_minimum_benefit_v1.json"
+)
+PARENT_ARTIFACT_SHA256 = (
+ "b2c2806254618bc8c4cf6c20e652ec2a06ce8a7c05de01a82c151a5d7921cafc"
+)
+PARENT_REGISTRATION_POINTER = (
+ "https://github.com/PolicyEngine/microcosm-dynamics/issues/42"
+ "#issuecomment-5855753783"
+)
+PARENT_REGISTERED_COMMIT = "2e4e08beaad1614da089d76b2ad639d0a4884e2e"
+PARENT_SPECIFICATION_SHA256 = (
+ "2e55afc109bcd65d6cf7307b5f52ba91c4c9039891e80388b26214dcce29cc57"
+)
+#: The new artifact; the parent is never edited.
+OUTPUT_ARTIFACT_PATH = (
+ _ROOT
+ / "runs"
+ / "replication_urban2006_minimum_benefit_groups_posthoc_v1.json"
+)
+#: Track M's four labels and the two post hoc labels, on every output.
+LABELS: tuple[str, ...] = (*OUTPUT_LABELS, *common.POST_HOC_LABELS)
+ANALYSIS_YEAR = cohort.INCOME_YEAR
+#: The universe's anchor wave: G1 resolves education as of it.
+ANCHOR_WAVES: tuple[int, ...] = (cohort.WAVE,)
+#: The youngest age in the universe (born 1960 or earlier).
+YOUNGEST_AGE = ANALYSIS_YEAR - structure.LAST_BIRTH_YEAR
+#: G2's exercise-4 AIME convention (Track M's ``history_pia`` for an
+#: entitlement at 62), which must equal MINT8's initial-AIME convention.
+AIME_CONVENTION = g2.AIME_CONVENTIONS["exercise_4_min_benefit"]
+_MINT8_AIME_CONVENTION = g2.AIME_CONVENTIONS["mint8_initial_aime"]
+STATISTIC_ID = f"{tabulation.STATISTIC_ID}_by_mint8_group"
+SENSITIVITY_KEY = OWN_RECEIPT_SENSITIVITY_ID
+_SENSITIVITY = SENSITIVITIES[OWN_RECEIPT_SENSITIVITY_ID]
+#: The sensitivity's tabulation row id, as ``pipeline`` names it.
+SENSITIVITY_ROW_ID = f"{_SENSITIVITY['row']}:{OWN_RECEIPT_SENSITIVITY_ID}"
+#: The pipeline result's keys the registered runner replaced, so they are
+#: compared separately (the specification) or not at all (the header).
+_RUNNER_REPLACED_KEYS = ("header", "specification")
+#: The COLA history's repository-relative path.
+COLA_HISTORY_RELATIVE_PATH = str(
+ COLA_HISTORY_PATH.relative_to(COLA_HISTORY_PATH.parents[2])
+)
+#: The one field of the committed artifact that records where a file was
+#: read rather than what was read: the registered runner recorded the COLA
+#: history by the absolute path of the worktree it ran in
+#: (``estimates.parameters.load_cola_history`` writes ``path.resolve()``).
+#: The same record pins the file's bytes (``sha256``, ``content_sha256``),
+#: which are compared exactly.
+LOCATION_FIELDS: tuple[tuple[str, ...], ...] = (
+ ("inputs", "source", "cola_history", "path"),
+)
+LOCATION_RULE = (
+ "the location fields ("
+ + ", ".join("$." + ".".join(keys) for keys in LOCATION_FIELDS)
+ + ") are compared by their repository-relative path: an absolute "
+ "path ending in /"
+ + COLA_HISTORY_RELATIVE_PATH
+ + " reads as "
+ + COLA_HISTORY_RELATIVE_PATH
+ + " on both sides, any other path is "
+ "compared as written; the same record's sha256 and content_sha256 are "
+ "compared exactly, and nothing else is normalized"
+)
+
+# =========================================================================
+# The scheme
+# =========================================================================
+_POVERTY_REASON = (
+ "not computed: MINT8's poverty status and household income quintile "
+ "need each person's 2022 household (family) money income and, for "
+ "poverty, the official threshold of the family's size and composition; "
+ "Track M's inputs (min_benefit_track_m.structure.TrackMStructureInputs: "
+ "the 2023 anchor, the family file's 2022 Social Security of the "
+ "reference person and spouse, death records, marriage history, the "
+ "labor-income panel and the design; and the receipt and next-wave "
+ "labor-income frames of cohort.TrackMCohortInputs) carry neither, and "
+ "this package reads no new PSID item"
+)
+#: MINT8 dimensions the exercise-4 breakdown does not compute, with why.
+NOT_COMPUTED_DIMENSIONS: dict[str, str] = {
+ "poverty_status": _POVERTY_REASON,
+ "household_income_quintile": _POVERTY_REASON,
+}
+#: MINT8 statistics not computed, with why.
+NOT_COMPUTED_STATISTICS: dict[str, dict[str, Any]] = {
+ "mint8_benefit_statistics": {
+ "statistics": list(g3.MINT_BENEFIT_STATISTICS),
+ "reason": (
+ "not computed: MINT8's benefit statistics compare each "
+ "person's benefit under the option with the same person's "
+ "current-law benefit. Track M's evaluate() returns, per "
+ "person, only the receipt flags receives_2..receives_5, the "
+ "receipt basis and the exposure flags (min_benefit_track_m/"
+ "evaluation.py evaluate); per worker record it returns the PIA "
+ "under the row's PIA rule, before any option's cut or minimum "
+ "(WorkerOutcome.pia), and each option's PIA "
+ "(WorkerOutcome.outcomes[k].option_pia, rules.evaluate_worker), "
+ "but no person-level benefit: the own benefit after the claim "
+ "factor is not formed, and a spouse's or survivor's benefit is "
+ "decided only as paid or not (rules.spouse_excess_paid and "
+ "rules.survivor_excess_paid return booleans). The registered "
+ "options also have no current-law column: option 1 is "
+ "'Reduced current law', a 12.45 percent uniform cut "
+ "(policy.OPTIONS[1]). Computing these statistics would need a "
+ "person-level benefit model this breakdown does not add"
+ ),
+ },
+ "mint8_poverty_statistics": {
+ "statistics": list(g3.POVERTY_STATISTICS),
+ "reason": _POVERTY_REASON,
+ },
+}
+
+_AGE_NOTE = (
+ "age is 2022 minus the birth year Track M resolves (the first-estimates "
+ "birth-year law, cohort._universe); the universe is born 1960 or "
+ "earlier (structure.LAST_BIRTH_YEAR), so every age is 62 or more and "
+ "MINT8's '60–69' row holds ages 62-69 in this breakdown"
+)
+_QUINTILE_NOTE = (
+ "the three lifetime quintiles are cut within each 10-year birth cohort "
+ "of Track M's universe (lifetime_measures.ten_year_birth_cohort; "
+ "MINT8: 'We calculate the AIME quintiles for each birth cohort') with "
+ "the 2023 cross-section weight ER35265; the '1960–1969' cohort holds "
+ "only the 1960 births here"
+)
+_DROPPED_NOTE = (
+ "poverty status and household income quintile are dropped: Track M's "
+ "inputs carry no 2022 family money income (not_computed)"
+)
+SCHEME = g3.derive_scheme(
+ g3.MINT8_ANNUAL_WITH_LIFETIME_SCHEME,
+ scheme_id="mint8_beneficiary_annual_with_lifetime__exercise4_track_m",
+ title=(
+ "Exercise 4 (Track M) by MINT8's annual beneficiary rows and its "
+ "lifetime-earnings quintiles, without poverty and household income"
+ ),
+ drop=tuple(NOT_COMPUTED_DIMENSIONS),
+ population=(
+ "Track M's universe: persons of the PSID's 2023 wave in a "
+ "responding family unit (sequence 1-20), not aged 0-5, with a "
+ "positive cross-section weight ER35265, born 1960 or earlier, with "
+ "a person-level 2022 Social Security amount (ER35219) above zero "
+ "(M1 specification section 10)"
+ ),
+ notes=(_AGE_NOTE, _QUINTILE_NOTE, _DROPPED_NOTE),
+)
+#: The attribute column of every non-total dimension of :data:`SCHEME`.
+COLUMNS_MAP: dict[str, str] = {
+ "sex": "sex",
+ "race_ethnicity": "race_ethnicity",
+ "country_of_birth": "country_of_birth",
+ "age": "age",
+ "marital_status": "marital_status",
+ "education": "education_years",
+ "benefit_type": "benefit_type",
+ "initial_aime_quintile": "initial_aime_at_62",
+ "lifetime_payroll_tax_quintile": "lifetime_payroll_tax_pv_at_62",
+ "lifetime_payroll_tax_quintile_shared": (
+ "lifetime_payroll_tax_pv_at_62_shared"
+ ),
+}
+_QUINTILE_PARTITION = ("birth_cohort_10y",)
+_SEX_UNCLASSIFIED = ("unknown",)
+_MARITAL_UNCLASSIFIED = ("no_marriage_history", "unknown")
+
+#: Deltas this breakdown adds to Track M's named deltas.
+GROUP_NAMED_DELTAS: tuple[str, ...] = (
+ "a static 2022 PSID universe of beneficiaries 62 and older against "
+ "MINT8's projected beneficiaries 60 and older in 2030, 2050 and 2070",
+ "age 60-69 holds ages 62-69 (the universe is born 1960 or earlier)",
+ "benefit type from Track M's reading of the self-reported 2022 "
+ "receipt types (G33A, ER35213-ER35218) and its own worker records, not "
+ "SSA's administrative beneficiary type",
+ "persons without an own worker record (paid only a spouse's or "
+ "survivor's benefit) have no initial AIME and are unclassified in the "
+ "AIME quintile",
+ "race and ethnicity and country of birth are asked of reference "
+ "persons and spouses only: other family-unit members are unclassified",
+ "lifetime quintiles cut within 10-year birth cohorts of this universe "
+ "with the PSID cross-section weight, not over MINT8's population",
+ "lifetime histories are PSID labor income of the reference person and "
+ "spouse treated as covered earnings (d280), with Track M's next-wave "
+ "odd years and gap rule; years as another family member are "
+ "unobserved",
+ "lifetime payroll taxes apply the combined OASDI rate to every dollar "
+ "of labor income up to the taxable maximum; self-employment is not "
+ "separated (G2's registered builder defaults)",
+ "current-law poverty status and household income quintile not "
+ "computed (no 2022 family money income in Track M's inputs)",
+ "MINT8's benefit-change statistics not computed (no person-level "
+ "benefits in Track M's evaluation)",
+)
+
+
+# =========================================================================
+# The parent artifact
+# =========================================================================
+def load_parent_artifact(
+ path: Path = PARENT_ARTIFACT_PATH,
+ *,
+ expected_sha256: str = PARENT_ARTIFACT_SHA256,
+) -> common.CommittedArtifact:
+ """The committed exercise-4 artifact, its bytes pinned and its sidecar
+ bound (:func:`.common.load_committed_artifact`)."""
+
+ return common.load_committed_artifact(
+ path, expected_sha256=expected_sha256
+ )
+
+
+def check_parent_binding(
+ parent: common.CommittedArtifact,
+ *,
+ specification_sha256: str,
+) -> dict[str, Any]:
+ """Refuse a parent other than Registration 17's committed run.
+
+ Its registration pointer and registered commit must be Registration
+ 17's, and the M1 specification it recorded must be the file the
+ re-execution's guard reads now (``specification_sha256``, the current
+ file's SHA-256) and the committed one.
+ """
+
+ document = parent.document
+ run = document.get("run") or {}
+ recorded = document.get("specification") or {}
+ if document.get("registration_pointer") != PARENT_REGISTRATION_POINTER:
+ raise ValueError("the parent is not Registration 17's artifact")
+ if run.get("registered_commit") != PARENT_REGISTERED_COMMIT:
+ raise ValueError("the parent's registered commit is not 2e4e08be")
+ if recorded.get("sha256") != PARENT_SPECIFICATION_SHA256:
+ raise ValueError("the parent recorded another M1 specification")
+ if specification_sha256 != PARENT_SPECIFICATION_SHA256:
+ raise ValueError(
+ "the M1 specification changed since Registration 17 "
+ f"({specification_sha256}): the re-execution would not be the "
+ "registered computation"
+ )
+ return {
+ "registration_pointer": PARENT_REGISTRATION_POINTER,
+ "registered_commit": PARENT_REGISTERED_COMMIT,
+ "specification_sha256": PARENT_SPECIFICATION_SHA256,
+ "specification_version": recorded.get("version"),
+ }
+
+
+# =========================================================================
+# Re-execution
+# =========================================================================
+@dataclass(frozen=True, eq=False)
+class TrackMReexecution:
+ """The registered computation re-executed, with its person rows.
+
+ ``result`` is what ``run_pipeline`` returns (``pipeline.run_track_m``
+ plus both cohorts' ``cohort_structure``); ``evaluations`` and
+ ``tabulations`` are every registered row's ``evaluate`` and
+ ``tabulate_track_m``; ``sensitivity_*`` are d430's.
+ """
+
+ cohort_inputs: Any
+ cohort: cohort.TrackMCohort
+ cohort_sensitivity: cohort.TrackMCohort
+ records: TrackMInputs
+ records_sensitivity: TrackMInputs
+ result: Mapping[str, Any]
+ evaluations: Mapping[str, Evaluation]
+ tabulations: Mapping[str, Mapping[str, Any]]
+ sensitivity_evaluation: Evaluation
+ sensitivity_tabulation: Mapping[str, Any]
+ data_provenance: str
+ registration_pointer: str | None
+
+
+def reexecute_track_m(
+ cohort_inputs: Any,
+ parameters: TrackMParameters,
+ *,
+ cola_rates: Mapping[int, float],
+ data_provenance: str,
+ registration_pointer: str | None,
+ provenance_kind: str,
+ source: Mapping[str, Any],
+) -> TrackMReexecution:
+ """``run_pipeline``'s computation on ``cohort_inputs``, and the rows.
+
+ Mirrors ``scripts/run_track_m_registered.py``'s ``run_pipeline`` step
+ for step (both cohorts, ``careers.build_track_m_inputs`` with
+ ``provenance_kind`` and ``source``, ``pipeline.run_track_m`` with the
+ sensitivity, the two ``cohort_structure`` blocks) and then evaluates
+ and tabulates every registered row and d430's sensitivity with the
+ unchanged ``evaluate`` and ``tabulate_track_m``, as ``run_track_m``
+ does inside. The registered path passes ``registered_real``,
+ Registration 17's pointer, ``psid_files`` and
+ ``{"cola_history": cola.provenance}``.
+ """
+
+ built = cohort.build_cohort(cohort_inputs)
+ built_sensitivity = cohort.build_cohort(
+ cohort_inputs,
+ own_receipt_reading=OWN_RECEIPT_PRE62_UNKNOWN_OR_OTHER_UNOBSERVED,
+ )
+
+ def records_of(built_cohort: cohort.TrackMCohort) -> TrackMInputs:
+ return careers.build_track_m_inputs(
+ built_cohort,
+ earnings=cohort_inputs.earnings,
+ prior_year=cohort_inputs.prior_year_labor,
+ params=parameters.params,
+ cola_rates=cola_rates,
+ provenance_kind=provenance_kind,
+ source=dict(source),
+ )
+
+ records = records_of(built)
+ records_sensitivity = records_of(built_sensitivity)
+ result = pipeline.run_track_m(
+ records,
+ parameters,
+ data_provenance=data_provenance,
+ registration_pointer=registration_pointer,
+ own_receipt_sensitivity=records_sensitivity,
+ )
+ result["cohort_structure"] = cohort.cohort_structure(built)
+ result["sensitivities"][SENSITIVITY_KEY]["cohort_structure"] = (
+ cohort.cohort_structure(built_sensitivity)
+ )
+ evaluations: dict[str, Evaluation] = {}
+ tabulations: dict[str, Mapping[str, Any]] = {}
+ for row in REGISTERED_ROWS:
+ evaluation = evaluate(records, policy_for_row(row), parameters)
+ evaluations[row] = evaluation
+ tabulations[row] = tabulation.tabulate_track_m(
+ evaluation.rows,
+ row_id=row,
+ data_provenance=data_provenance,
+ design=records.design,
+ registration_pointer=registration_pointer,
+ labels=OUTPUT_LABELS,
+ )
+ sensitivity_evaluation = evaluate(
+ records_sensitivity, policy_for_row(_SENSITIVITY["row"]), parameters
+ )
+ sensitivity_tabulation = tabulation.tabulate_track_m(
+ sensitivity_evaluation.rows,
+ row_id=SENSITIVITY_ROW_ID,
+ data_provenance=data_provenance,
+ design=records_sensitivity.design,
+ registration_pointer=registration_pointer,
+ labels=OUTPUT_LABELS,
+ scored=False,
+ )
+ return TrackMReexecution(
+ cohort_inputs=cohort_inputs,
+ cohort=built,
+ cohort_sensitivity=built_sensitivity,
+ records=records,
+ records_sensitivity=records_sensitivity,
+ result=result,
+ evaluations=evaluations,
+ tabulations=tabulations,
+ sensitivity_evaluation=sensitivity_evaluation,
+ sensitivity_tabulation=sensitivity_tabulation,
+ data_provenance=data_provenance,
+ registration_pointer=registration_pointer,
+ )
+
+
+_ABSENT = object()
+
+
+def _check(
+ name: str, committed: Any, recomputed: Any
+) -> common.ReproductionCheck:
+ if committed is _ABSENT:
+ return common.ReproductionCheck(
+ name, 0, ("$ (absent from the committed artifact)",)
+ )
+ return common.compare_exact(name, committed, recomputed)
+
+
+def _dig(document: Mapping[str, Any], *keys: str) -> Any:
+ value: Any = document
+ for key in keys:
+ if not isinstance(value, Mapping) or key not in value:
+ return _ABSENT
+ value = value[key]
+ return value
+
+
+def parameter_check(
+ parameters: TrackMParameters, parent: Mapping[str, Any]
+) -> common.ReproductionCheck:
+ """The parameters' provenance against the committed ``parameters``."""
+
+ return _check(
+ "parameters", _dig(parent, "parameters"), parameters.source()
+ )
+
+
+def psid_files_check(
+ cohort_inputs: Any, parent: Mapping[str, Any]
+) -> common.ReproductionCheck:
+ """The PSID files read against the committed run's SHA-256 map."""
+
+ key = pipeline.PSID_FILES_SOURCE_KEY
+ committed = _dig(parent, "inputs", "source")
+ committed = _ABSENT if committed is _ABSENT else dict(committed).get(key)
+ return _check(
+ f"inputs.source.{key}",
+ committed,
+ dict(getattr(cohort_inputs, "provenance", {}) or {}).get(key),
+ )
+
+
+def _recorded(document: Mapping[str, Any], keys: Sequence[str]) -> Any:
+ """The leaf at ``keys``, or ``None`` where the document has none."""
+
+ value = _dig(document, *keys)
+ return None if value is _ABSENT else copy.deepcopy(value)
+
+
+def relative_location(path: Any) -> Any:
+ """``path`` as :data:`LOCATION_RULE` compares it."""
+
+ if isinstance(path, str) and path.startswith("/"):
+ if path.endswith("/" + COLA_HISTORY_RELATIVE_PATH):
+ return COLA_HISTORY_RELATIVE_PATH
+ return path
+
+
+def located(document: Any) -> Any:
+ """``document`` (a whole artifact) with its :data:`LOCATION_FIELDS`
+ read by :func:`relative_location`; every other leaf unchanged."""
+
+ if not isinstance(document, Mapping):
+ return document
+ out = copy.deepcopy(dict(document))
+ for keys in LOCATION_FIELDS:
+ node: Any = out
+ for key in keys[:-1]:
+ node = node.get(key) if isinstance(node, dict) else None
+ if isinstance(node, dict) and keys[-1] in node:
+ node[keys[-1]] = relative_location(node[keys[-1]])
+ return out
+
+
+def reproduction_checks(
+ reexecution: TrackMReexecution, parent: Mapping[str, Any]
+) -> tuple[common.ReproductionCheck, ...]:
+ """Every recomputed block against the committed artifact, exactly.
+
+ One check per key of the re-executed pipeline result (every key but
+ ``header`` and ``specification``, which the registered runner
+ replaced), and one per registered row's and the sensitivity's
+ tabulation and diagnostics from the rows the groups use. Both sides
+ pass through :func:`located` first (:data:`LOCATION_RULE`).
+ """
+
+ checks = []
+ committed = located(parent)
+ recomputed = located(dict(reexecution.result))
+ for key, value in recomputed.items():
+ if key in _RUNNER_REPLACED_KEYS:
+ continue
+ checks.append(_check(f"pipeline.{key}", _dig(committed, key), value))
+ for row in REGISTERED_ROWS:
+ evaluation = reexecution.evaluations[row]
+ checks.append(
+ _check(
+ f"group_rows.{row}.tabulation",
+ _dig(parent, "rows", row, "tabulation"),
+ reexecution.tabulations[row],
+ )
+ )
+ checks.append(
+ _check(
+ f"group_rows.{row}.diagnostics",
+ _dig(parent, "rows", row, "diagnostics"),
+ evaluation.diagnostics,
+ )
+ )
+ checks.append(
+ _check(
+ f"group_rows.{SENSITIVITY_KEY}.tabulation",
+ _dig(parent, "sensitivities", SENSITIVITY_KEY, "tabulation"),
+ reexecution.sensitivity_tabulation,
+ )
+ )
+ checks.append(
+ _check(
+ f"group_rows.{SENSITIVITY_KEY}.diagnostics",
+ _dig(parent, "sensitivities", SENSITIVITY_KEY, "diagnostics"),
+ reexecution.sensitivity_evaluation.diagnostics,
+ )
+ )
+ return tuple(checks)
+
+
+# =========================================================================
+# Person attributes
+# =========================================================================
+#: How Track M's 2022 receipt maps to MINT8's four benefit types, in the
+#: order :func:`benefit_type_2022` applies them.
+BENEFIT_TYPE_RULES: tuple[str, ...] = (
+ "inputs, all Track M's own: the six G33A 'whether Social Security "
+ "type' items of income year 2022 (ER35213 disability, ER35214 "
+ "retirement, ER35215 survivor, ER35216 dependent of disabled, ER35217 "
+ "dependent of retired, ER35218 other), decoded as Track M's cohort "
+ "decodes them (cohort._anchor_types: 1 mentioned, 5 not mentioned, any "
+ "other code unknown); Track M's paid_own_worker_benefit (the 2022 "
+ "receipt is own receipt by cohort.observation and the person has an "
+ "own worker record); and the basis of that record "
+ "(cohort.classify_own_record: 'disability' when the first own receipt "
+ "mentions disability or precedes the year of attaining 62, otherwise "
+ "'old_age')",
+ "1. survivor and a dependent type both mentioned: unclassified "
+ "(survivor_and_dependent_mentioned)",
+ "2. survivor mentioned: 'Widow(er) (includes dually entitled)', "
+ "whatever worker type is also mentioned (the item does not say whose "
+ "survivor; at 62 and older it is read as a widow(er)'s benefit)",
+ "3. dependent of disabled or of retired mentioned: 'Spousal (includes "
+ "dually entitled)', whatever worker type is also mentioned (the item "
+ "does not say whose dependent; at 62 and older it is read as a "
+ "spouse's benefit)",
+ "4. no auxiliary type mentioned but one of the three auxiliary items "
+ "(survivor, dependent of disabled, dependent of retired) unknown: "
+ "unclassified (auxiliary_item_unknown), since MINT8's 'only' cannot be "
+ "established",
+ "5. otherwise the benefit is read as the person's own worker benefit; "
+ "a person Track M does not pay their own worker benefit is "
+ "unclassified (no_own_worker_benefit)",
+ "6. the worker type: disability when disability is mentioned and "
+ "retirement is not; retirement when retirement is mentioned and "
+ "disability is not; when both or neither are mentioned (an 'other' or "
+ "unknown type, which Track M's scored reading counts as own receipt, "
+ "cos d430), the basis of Track M's own worker record",
+ "7. disability: 'Disabled worker only' when the person is under the "
+ "full retirement age in 2022 (12 times the whole-year age 2022 minus "
+ "the birth year, Track M's claim-age convention, below the oracle's "
+ "params.fra_months of the birth year); at or over it 'Retired worker "
+ "only' (MINT8: 'Disabled workers convert to retired workers at FRA')",
+ "8. retirement: 'Retired worker only'",
+)
+_UNCLASSIFIED_PREFIX = "unclassified:"
+_AUXILIARY_TYPES = frozenset(ssr.AUXILIARY_TYPES)
+_DEPENDENT_TYPES = frozenset({"dependent_of_disabled", "dependent_of_retired"})
+_OWN_RECORD_BASES = (cohort.BASIS_OLD_AGE, cohort.BASIS_DISABILITY)
+
+
+def benefit_type_2022(
+ types: Mapping[str, Any],
+ *,
+ birth_year: int,
+ paid_own_worker_benefit: bool,
+ own_record_basis: str | None,
+ params: SSAParameters,
+) -> str:
+ """One person's MINT8 benefit-type code (:data:`BENEFIT_TYPE_RULES`).
+
+ ``types`` maps each of the six SS types to True (mentioned), False
+ (not mentioned) or None (unknown), as ``cohort._anchor_types`` decodes
+ the 2022 items; ``paid_own_worker_benefit`` and ``own_record_basis``
+ (``None`` without an own record) are Track M's. Returns a G3
+ ``benefit_type`` code or ``unclassified:``.
+ """
+
+ unknown = {name for name in ssr.SS_TYPES if types.get(name) is None}
+ mentioned = {name for name in ssr.SS_TYPES if types.get(name) is True}
+ if paid_own_worker_benefit and own_record_basis not in _OWN_RECORD_BASES:
+ raise ValueError(
+ "a person paid their own worker benefit has an old-age or "
+ f"disability own record, not {own_record_basis!r}"
+ )
+ dependent = mentioned & _DEPENDENT_TYPES
+ if "survivor" in mentioned and dependent:
+ return f"{_UNCLASSIFIED_PREFIX}survivor_and_dependent_mentioned"
+ if "survivor" in mentioned:
+ return "widower"
+ if dependent:
+ return "spousal"
+ if unknown & _AUXILIARY_TYPES:
+ return f"{_UNCLASSIFIED_PREFIX}auxiliary_item_unknown"
+ if not paid_own_worker_benefit:
+ return f"{_UNCLASSIFIED_PREFIX}no_own_worker_benefit"
+ worker, _ = _own_worker_type(mentioned, own_record_basis)
+ if worker == cohort.BASIS_DISABILITY:
+ months = 12 * (ANALYSIS_YEAR - int(birth_year))
+ if months < params.fra_months(int(birth_year)):
+ return "disabled_worker_only"
+ return "retired_worker_only"
+
+
+def _own_worker_type(
+ mentioned: set[str] | frozenset[str], own_record_basis: str | None
+) -> tuple[str, str]:
+ """``(worker type, what decided it)`` (:data:`BENEFIT_TYPE_RULES` 6).
+
+ The type is ``"disability"`` or ``"old_age"``; it is decided by the
+ 2022 mentions when exactly one worker type is mentioned, and by Track
+ M's own-record basis otherwise.
+ """
+
+ retirement = "retirement" in mentioned
+ disability = "disability" in mentioned
+ if disability != retirement:
+ return (
+ cohort.BASIS_DISABILITY if disability else cohort.BASIS_OLD_AGE,
+ "2022_mentions",
+ )
+ return (
+ (
+ cohort.BASIS_DISABILITY
+ if own_record_basis == cohort.BASIS_DISABILITY
+ else cohort.BASIS_OLD_AGE
+ ),
+ "own_record_basis",
+ )
+
+
+def _benefit_types(
+ built: cohort.TrackMCohort, cohort_inputs: Any, params: SSAParameters
+) -> tuple[list[str], dict[str, Any]]:
+ """Every universe person's benefit type, in ``built.persons`` order.
+
+ Each person's 2022 observation is rebuilt from the anchor exactly as
+ ``cohort.build_cohort`` builds it, and must reproduce the cohort's
+ ``types_2022``, ``status_2022`` and ``paid_own_worker_benefit``
+ (refused otherwise).
+ """
+
+ anchor = cohort_inputs.anchor.set_index("person_id")
+ bases = dict(
+ zip(built.records["record_id"], built.records["basis"], strict=True)
+ )
+ codes = []
+ by_rule: Counter[str] = Counter()
+ for pid, birth, types_2022, status_2022, paid, own in zip(
+ built.persons["person_id"],
+ built.persons["birth_year"],
+ built.persons["types_2022"],
+ built.persons["status_2022"],
+ built.persons["paid_own_worker_benefit"],
+ built.persons["own_record_id"],
+ strict=True,
+ ):
+ types = cohort._anchor_types(anchor.loc[int(pid)])
+ observed = cohort.observation(
+ ANALYSIS_YEAR, receipt=True, types=types, source="individual"
+ )
+ own_id = None if own is None or pd.isna(own) else str(own)
+ if (
+ ",".join(sorted(observed.types)) != str(types_2022)
+ or observed.status != str(status_2022)
+ or bool(paid) != (observed.status == cohort.OWN and bool(own_id))
+ ):
+ raise ValueError(
+ f"person {pid}: the 2022 receipt rebuilt from the anchor is "
+ "not the cohort's"
+ )
+ basis = None if own_id is None else str(bases[own_id])
+ code = benefit_type_2022(
+ types,
+ birth_year=int(birth),
+ paid_own_worker_benefit=bool(paid),
+ own_record_basis=basis,
+ params=params,
+ )
+ if code in ("retired_worker_only", "disabled_worker_only"):
+ mentioned = {name for name, value in types.items() if value}
+ worker, decided_by = _own_worker_type(mentioned, basis)
+ by_rule[f"worker_type_from_{decided_by}"] += 1
+ if (
+ code == "retired_worker_only"
+ and worker == cohort.BASIS_DISABILITY
+ ):
+ by_rule["disability_at_or_over_fra_read_as_retired"] += 1
+ codes.append(code)
+ counts = Counter(codes)
+ return codes, {
+ "rules": list(BENEFIT_TYPE_RULES),
+ "n_by_code": dict(sorted(counts.items())),
+ "n_by_rule": dict(sorted(by_rule.items())),
+ }
+
+
+def _careers_rows(
+ pid: int, history: coverage.OneHistory
+) -> list[dict[str, Any]]:
+ observed = set(history.observed_years)
+ next_wave = set(history.next_wave_years)
+ rows = []
+ for year, value in sorted(history.values.items()):
+ rows.append(
+ {
+ "person_id": int(pid),
+ "year": int(year),
+ "earnings": float(value),
+ "provenance": (
+ "observed"
+ if year in observed
+ else "next_wave" if year in next_wave else "imputed"
+ ),
+ }
+ )
+ return rows
+
+
+_CAREER_COLUMNS = ["person_id", "year", "earnings", "provenance"]
+
+
+def initial_aime_measure(
+ reexecution: TrackMReexecution, params: SSAParameters
+) -> tuple[g2.MeasureResult, dict[str, Any]]:
+ """G2's initial AIME at 62 over the MS0 own records' one histories.
+
+ Returns the measure and its checks: the persons without an own record,
+ and the exercise-4 identity (an old-age record entitled in the year of
+ attaining 62 has the same AIME at 62 as MS0's ``history_pia.aime``;
+ refused otherwise).
+ """
+
+ mint8 = (
+ _MINT8_AIME_CONVENTION.computation_years,
+ _MINT8_AIME_CONVENTION.last_earnings_age,
+ )
+ if (
+ AIME_CONVENTION.computation_years,
+ AIME_CONVENTION.last_earnings_age,
+ ) != mint8:
+ raise ValueError(
+ "G2's exercise-4 AIME convention is no longer MINT8's initial "
+ "AIME convention (statutory computation years, earnings through "
+ "the year of attaining 61)"
+ )
+ persons = reexecution.cohort.persons
+ ms0 = reexecution.evaluations["MS0"]
+ rows: list[dict[str, Any]] = []
+ own_ids: dict[int, str] = {}
+ for pid, own in zip(
+ persons["person_id"], persons["own_record_id"], strict=True
+ ):
+ if own is None or pd.isna(own):
+ continue
+ own_ids[int(pid)] = str(own)
+ rows.extend(_careers_rows(int(pid), ms0.workers[str(own)].history))
+ measure = g2.initial_aime_at_62(
+ pd.DataFrame(rows, columns=_CAREER_COLUMNS),
+ persons[["person_id", "birth_year"]],
+ params,
+ analysis_year=ANALYSIS_YEAR,
+ convention=AIME_CONVENTION,
+ )
+ by_person = measure.frame.set_index("person_id")
+ compared = 0
+ for pid, record_id in own_ids.items():
+ outcome = ms0.workers[record_id]
+ years = outcome.years
+ if (
+ years.basis != "old_age"
+ or years.window_year != years.birth_year + 62
+ or by_person.loc[pid, "status"] != g2.COMPUTED
+ ):
+ continue
+ aime = by_person.loc[pid, "aime"]
+ if outcome.history_pia.aime is None or not (
+ float(aime) == float(outcome.history_pia.aime)
+ ):
+ raise ValueError(
+ f"person {pid}: G2's AIME at 62 is not MS0's AIME for an "
+ "old-age entitlement at 62"
+ )
+ compared += 1
+ no_rows = sum(
+ 1
+ for pid, record_id in own_ids.items()
+ if by_person.loc[pid, "status"] != g2.COMPUTED
+ )
+ return measure, {
+ "persons_without_own_record": int(len(persons) - len(own_ids)),
+ "own_records_not_computed": int(no_rows),
+ "own_records_not_computed_note": (
+ "G2 leaves an own record whose history has no year through the "
+ "year of attaining 61 not computed (earnings unobserved, not "
+ "zero), where MS0's oracle reads the empty history as zeros; "
+ "such persons are unclassified in the AIME quintile"
+ ),
+ "history": (
+ "the MS0 evaluation's one history of the person's own worker "
+ "record (WorkerOutcome.history: observed panel years, "
+ "next-wave odd years, gap-rule years), through the record's "
+ "section 4a last year"
+ ),
+ "convention": {
+ "name": AIME_CONVENTION.name,
+ "equals_mint8_initial_aime_convention": True,
+ "rule": (
+ "G2's exercise_4_min_benefit and mint8_initial_aime "
+ "conventions both read statutory computation years and "
+ "earnings through the year of attaining 61 (checked)"
+ ),
+ },
+ "identity_with_ms0_aime": {
+ "rule": (
+ "an old-age own record entitled in the year of attaining 62 "
+ "whose AIME G2 computes has G2's AIME at 62 equal to MS0's "
+ "history_pia.aime"
+ ),
+ "n_compared": compared,
+ "all_equal": True,
+ },
+ "own_records_by_basis": dict(
+ sorted(
+ Counter(
+ ms0.workers[record].years.basis
+ for record in own_ids.values()
+ ).items()
+ )
+ ),
+ }
+
+
+def lifetime_careers(
+ cohort_inputs: Any, person_ids: Sequence[int]
+) -> pd.DataFrame:
+ """Each person's lifetime history through 2022 by Track M's own rule.
+
+ ``careers.observed_histories`` (the labor-income panel, years through
+ 2022) and ``prior_year_labor_income.next_wave_histories`` (observed
+ next-wave odd years), joined by ``coverage.one_history`` with
+ ``last_year=2022`` and MS0's policy (the gap rule), exactly the
+ sources and rule a worker record's history reads, but not cut at the
+ record's last year: G2's payroll taxes are lifetime taxes.
+ """
+
+ ids = {int(pid) for pid in person_ids}
+ panel = careers.observed_histories(cohort_inputs.earnings, ids)
+ prior = cohort_inputs.prior_year_labor
+ next_wave = pyl.next_wave_histories(prior[prior["person_id"].isin(ids)])
+ policy = policy_for_row("MS0")
+ rows: list[dict[str, Any]] = []
+ for pid in sorted(ids):
+ history = coverage.one_history(
+ panel.get(pid, {}),
+ next_wave.get(pid, {}),
+ last_year=ANALYSIS_YEAR,
+ policy=policy,
+ )
+ rows.extend(_careers_rows(pid, history))
+ return pd.DataFrame(rows, columns=_CAREER_COLUMNS)
+
+
+def _marriage_inputs(
+ cohort_inputs: Any, universe: set[int]
+) -> tuple[pd.DataFrame, set[int], set[int]]:
+ """Marriage episodes of the universe and of every spouse they had."""
+
+ history = cohort_inputs.structure_inputs.marriage_history
+ mine = history[history["person_id"].isin(universe)]
+ # The episodes Track M's cohort reads (cohort._marital_states).
+ own_episodes = psid2010._episodes_with_separation(mine)
+ spouses = {
+ int(spouse) for spouse in own_episodes["spouse_person_id"].dropna()
+ }
+ everyone = universe | spouses
+ episodes = psid2010._episodes_with_separation(
+ history[history["person_id"].isin(everyone)]
+ )
+ with_history = {int(pid) for pid in mine["person_id"]}
+ return episodes, spouses, with_history
+
+
+@dataclass(frozen=True, eq=False)
+class PersonAttributes:
+ """The group attributes of every universe person, in row order.
+
+ ``frame`` has ``person_id`` (the evaluation rows' string id),
+ ``weight``, one column per :data:`COLUMNS_MAP` value and
+ ``birth_cohort_10y``; ``unclassified_codes`` the declared codes per
+ categorical dimension; ``provenance`` every source and rule.
+ """
+
+ frame: pd.DataFrame
+ unclassified_codes: Mapping[str, tuple[str, ...]]
+ provenance: Mapping[str, Any]
+
+
+def _files_agree(
+ side_frame: Mapping[str, Any], cohort_inputs: Any
+) -> dict[str, Any]:
+ """Files both readers record must be the same bytes (refused if not)."""
+
+ side = dict(
+ (side_frame.get("inputs") or {}).get("psid_files_sha256") or {}
+ )
+ track_m = dict(
+ (getattr(cohort_inputs, "provenance", {}) or {}).get(
+ pipeline.PSID_FILES_SOURCE_KEY
+ )
+ or {}
+ )
+ shared = sorted(set(side) & set(track_m))
+ differ = [name for name in shared if side[name] != track_m[name]]
+ if differ:
+ raise ValueError(
+ f"the side frame and Track M read different bytes of {differ}"
+ )
+ return {"files_read_by_both": shared, "all_equal": True}
+
+
+def person_attributes(
+ reexecution: TrackMReexecution,
+ parameters: TrackMParameters,
+ side_frame: g1.GroupAttributes,
+ *,
+ tax_rates: Any,
+ interest_rates: Any,
+) -> PersonAttributes:
+ """Every group attribute of every universe person (module docstring).
+
+ ``side_frame`` is G1's frame for the universe's person ids (int).
+ Refuses a side frame that does not hold exactly the universe, persons
+ out of the evaluation rows' order, files the side frame and Track M
+ read with different bytes, a G1/G3 education disagreement and a
+ broken AIME identity.
+ """
+
+ built = reexecution.cohort
+ persons = built.persons
+ rows = reexecution.evaluations["MS0"].rows
+ ids = [int(pid) for pid in persons["person_id"]]
+ if [str(pid) for pid in ids] != rows["person_id"].tolist():
+ raise ValueError("the cohort's persons are not the rows' persons")
+ frame = side_frame.frame.set_index("person_id")
+ if sorted(int(pid) for pid in frame.index) != sorted(ids):
+ raise ValueError(
+ "the side frame must hold exactly the universe's persons"
+ )
+ frame = frame.loc[ids].reset_index()
+ files_record = _files_agree(
+ side_frame.provenance, reexecution.cohort_inputs
+ )
+ params = parameters.params
+ dimensions = {d.key: d for d in SCHEME.dimensions}
+ race = common.side_frame_codes(frame, dimensions["race_ethnicity"])
+ country = common.side_frame_codes(frame, dimensions["country_of_birth"])
+ education_check = common.education_agreement(
+ frame, dimensions["education"]
+ )
+ marital = common.marital_codes(
+ persons["marital_status_2022"].tolist(),
+ unclassified=_MARITAL_UNCLASSIFIED,
+ )
+ benefit, benefit_record = _benefit_types(
+ built, reexecution.cohort_inputs, params
+ )
+ universe = set(ids)
+ aime, aime_checks = initial_aime_measure(reexecution, params)
+ episodes, spouses, with_history = _marriage_inputs(
+ reexecution.cohort_inputs, universe
+ )
+ lifetime = lifetime_careers(
+ reexecution.cohort_inputs, sorted(universe | spouses)
+ )
+ person_frame = persons[["person_id", "birth_year"]]
+ own_tax = g2.lifetime_payroll_tax_pv_at_62(
+ lifetime,
+ person_frame,
+ params,
+ shared=False,
+ rates=tax_rates,
+ interest=interest_rates,
+ )
+ shared_tax = g2.lifetime_payroll_tax_pv_at_62(
+ lifetime,
+ person_frame,
+ params,
+ shared=True,
+ rates=tax_rates,
+ interest=interest_rates,
+ marriage_episodes=episodes,
+ separated_is_married=True,
+ marriage_history_person_ids=with_history,
+ )
+
+ ages = common.age_in_year(persons["birth_year"], ANALYSIS_YEAR)
+ if min(ages) < YOUNGEST_AGE:
+ raise ValueError(
+ f"an age under {YOUNGEST_AGE}: the universe is born "
+ f"{structure.LAST_BIRTH_YEAR} or earlier"
+ )
+
+ def measure(result: g2.MeasureResult, column: str) -> list[float]:
+ values = result.frame.set_index("person_id")[column]
+ return [float(values.loc[pid]) for pid in ids]
+
+ out = pd.DataFrame(
+ {
+ "person_id": rows["person_id"].to_numpy(),
+ "weight": rows["weight"].to_numpy(dtype=float),
+ "sex": rows["sex"].tolist(),
+ "race_ethnicity": list(race.codes),
+ "country_of_birth": list(country.codes),
+ "age": list(ages),
+ "marital_status": list(marital.codes),
+ "education_years": pd.array(
+ common.education_years(frame), dtype="Int64"
+ ),
+ "benefit_type": benefit,
+ "initial_aime_at_62": measure(aime, "aime"),
+ "lifetime_payroll_tax_pv_at_62": measure(own_tax, "pv_at_62"),
+ "lifetime_payroll_tax_pv_at_62_shared": measure(
+ shared_tax, "pv_at_62"
+ ),
+ "birth_cohort_10y": g2.ten_year_birth_cohort(
+ persons["birth_year"].reset_index(drop=True)
+ ).to_numpy(),
+ }
+ )
+ present_sex = tuple(
+ code for code in _SEX_UNCLASSIFIED if code in set(out["sex"])
+ )
+ unclassified = {
+ "sex": present_sex,
+ "race_ethnicity": race.unclassified_codes,
+ "country_of_birth": country.unclassified_codes,
+ "marital_status": marital.unclassified_codes,
+ "benefit_type": tuple(
+ sorted(
+ {code for code in benefit if code.startswith("unclassified:")}
+ )
+ ),
+ }
+ provenance = {
+ "person_key": (
+ "person_id: Track M's evaluation rows' id (str of the cohort's "
+ "integer person id); the side frame and the lifetime measures "
+ "are joined on the integer id"
+ ),
+ "fixed_across_rows": (
+ "attributes are computed once, from the scored reading's cohort "
+ "and MS0's histories, and used for every registered row and for "
+ "d430's sensitivity (same universe, same person order)"
+ ),
+ "sex": (
+ "Track M's sex (ER32000 through the death records); 'unknown' "
+ "(code 9) unclassified, as Track M counts it in All only"
+ ),
+ "age": _AGE_NOTE,
+ "marital_status": {
+ "rule": marital.rule,
+ "source": (
+ "Track M's marital_status_2022 (cohort._marital_states: "
+ "psid2010.marital_state_at at the end of 2022, separated "
+ "counted as married by default)"
+ ),
+ },
+ "race_ethnicity": {"rule": race.rule},
+ "country_of_birth": {"rule": country.rule},
+ "education": {
+ "rule": (
+ "G1's education_years (the most recent reported years of "
+ "schooling at or before the 2023 wave) in G3's MINT8 bands"
+ ),
+ "agreement_with_g1_label": education_check,
+ },
+ "benefit_type": benefit_record,
+ "side_frame": copy.deepcopy(dict(side_frame.provenance)),
+ "side_frame_files": files_record,
+ "initial_aime_at_62": {
+ **aime_checks,
+ "measure_provenance": aime.provenance,
+ },
+ "lifetime_payroll_tax_pv_at_62": {
+ "history": (
+ "lifetime_careers: Track M's observed panel and next-wave "
+ "odd years through 2022 with MS0's gap rule, for the "
+ "universe and every spouse in their marriage histories"
+ ),
+ "n_persons_with_history": int(lifetime["person_id"].nunique()),
+ "measure_provenance": own_tax.provenance,
+ },
+ "lifetime_payroll_tax_pv_at_62_shared": {
+ "measure_provenance": shared_tax.provenance,
+ },
+ "quintile_partition": {
+ "columns": list(_QUINTILE_PARTITION),
+ "rule": _QUINTILE_NOTE,
+ "persons_by_cohort": {
+ str(key): int(value)
+ for key, value in sorted(
+ Counter(out["birth_cohort_10y"]).items()
+ )
+ },
+ },
+ }
+ return PersonAttributes(
+ frame=out,
+ unclassified_codes=unclassified,
+ provenance=provenance,
+ )
+
+
+def assign(attributes: PersonAttributes) -> g3.GroupAssignment:
+ """G3's assignment of every person to :data:`SCHEME`'s categories."""
+
+ quintiles = {
+ dimension.key: _QUINTILE_PARTITION
+ for dimension in SCHEME.dimensions
+ if dimension.kind == g3.QUINTILE
+ }
+ return g3.assign_groups(
+ attributes.frame,
+ SCHEME,
+ COLUMNS_MAP,
+ key_columns=("person_id",),
+ weight_column="weight",
+ quintile_partitions=quintiles,
+ unclassified_codes={
+ key: codes
+ for key, codes in attributes.unclassified_codes.items()
+ if codes
+ },
+ )
+
+
+def quintile_rule_sensitivity(
+ attributes: PersonAttributes, assignment: g3.GroupAssignment
+) -> dict[str, Any]:
+ """Persons whose quintile G2's midpoint rule would place differently.
+
+ The cells use G3's rule (thresholds at the weighted 20/40/60/80th
+ percentiles; a value equal to a threshold takes the lower quintile).
+ G2's ``weighted_quintiles`` places a value by the midpoint of its
+ cumulative weight span. The two differ only for persons whose weight
+ straddles a boundary; this counts them per dimension. Diagnostic
+ only: no cell uses G2's rule.
+ """
+
+ out = {}
+ frame = attributes.frame
+ for dimension in SCHEME.dimensions:
+ if dimension.kind != g3.QUINTILE:
+ continue
+ values = pd.Series(
+ frame[COLUMNS_MAP[dimension.key]].to_numpy(dtype=float)
+ )
+ g2_labels = g2.weighted_quintiles(
+ values,
+ pd.Series(frame["weight"].to_numpy(dtype=float)),
+ by=pd.Series(frame["birth_cohort_10y"].to_numpy()),
+ ).astype(str)
+ long = assignment.long[assignment.long["dimension"] == dimension.key]
+ g3_labels = long.sort_values("row_id")["label"].to_numpy()
+ present = values.notna().to_numpy()
+ differ = int(
+ (g2_labels.to_numpy()[present] != g3_labels[present]).sum()
+ )
+ out[dimension.key] = {
+ "n_with_measure": int(present.sum()),
+ "n_placed_differently_by_g2_midpoint_rule": differ,
+ }
+ return {
+ "rule": (
+ "cells use G3's quintile rule; G2's midpoint rule "
+ "(lifetime_measures.weighted_quintiles) is computed only to "
+ "count the persons the two rules place differently"
+ ),
+ "by_dimension": out,
+ }
+
+
+# =========================================================================
+# Cells
+# =========================================================================
+def _breakdown_keys() -> list[tuple[str, str, int]]:
+ keys = [
+ (row, row, number)
+ for row in REGISTERED_ROWS
+ for number in TABLE6_OPTIONS
+ ]
+ keys.extend(
+ (SENSITIVITY_KEY, SENSITIVITY_ROW_ID, number)
+ for number in TABLE6_OPTIONS
+ )
+ return keys
+
+
+def tabulate_breakdowns(
+ reexecution: TrackMReexecution,
+ assignment: g3.GroupAssignment,
+ *,
+ data_provenance: str,
+ registration_pointer: str | None,
+ parent_sha256: str | None,
+) -> dict[tuple[str, int], g3.GroupBreakdownResult]:
+ """G3's share breakdown for every registered row, option and d430.
+
+ ``tabulate_share_breakdown`` over the row's evaluation rows with
+ ``receives_`` as the indicator, the design frame of the records and
+ G3's defaults (Track M's floor seeds 0-4, family units linked through
+ persons). A real-data breakdown carries the new registration's
+ pointer, Track M's labels and the post hoc labels.
+ """
+
+ out = {}
+ for key, row_id, number in _breakdown_keys():
+ if key == SENSITIVITY_KEY:
+ rows = reexecution.sensitivity_evaluation.rows
+ design = reexecution.records_sensitivity.design
+ else:
+ rows = reexecution.evaluations[key].rows
+ design = reexecution.records.design
+ out[(key, number)] = g3.tabulate_share_breakdown(
+ rows,
+ assignment,
+ indicator_column=f"receives_{number}",
+ design=design,
+ data_provenance=data_provenance,
+ registration_pointer=registration_pointer,
+ labels=OUTPUT_LABELS,
+ post_hoc_labels=common.POST_HOC_LABELS,
+ statistic_id=STATISTIC_ID,
+ upstream={
+ "row": row_id,
+ "option": int(number),
+ "scored": False,
+ "parent_artifact_sha256": parent_sha256,
+ },
+ )
+ return out
+
+
+_SEX_CELLS = (
+ ("total", "total", "all"),
+ ("sex", "female", "women"),
+ ("sex", "male", "men"),
+)
+
+
+def _track_m_cell(table: Mapping[str, Any], number: int, row: str) -> dict:
+ for cell in table["cells"]:
+ if cell["option"] == number and cell["row"] == row:
+ return cell
+ raise KeyError((number, row))
+
+
+def _comparable_track_m(cell: Mapping[str, Any]) -> dict[str, Any]:
+ out = {"defined": cell["defined"], "floor": cell["floor"]}
+ if cell["defined"]:
+ out.update(
+ share_percent=cell["share_percent"],
+ weighted_n=cell["weighted_n"],
+ unweighted_n=cell["unweighted_n"],
+ weighted_receiving=cell["weighted_receiving"],
+ unweighted_receiving=cell["unweighted_receiving"],
+ design_se=cell["design_se"],
+ )
+ return out
+
+
+def _comparable_g3(cell: g3.GroupCell) -> dict[str, Any]:
+ statistic = cell.statistic(g3.SHARE)
+ out = {
+ "defined": statistic.defined,
+ "floor": statistic.uncertainty["floor"],
+ }
+ if statistic.defined:
+ out.update(
+ share_percent=statistic.value,
+ weighted_n=statistic.weighted_n,
+ unweighted_n=statistic.unweighted_n,
+ weighted_receiving=cell.counts["weighted_indicator"],
+ unweighted_receiving=cell.counts["numerators"]["n_indicator"],
+ design_se=statistic.uncertainty["design_se"],
+ )
+ return out
+
+
+def consistency_checks(
+ reexecution: TrackMReexecution,
+ breakdowns: Mapping[tuple[str, int], g3.GroupBreakdownResult],
+) -> tuple[common.ReproductionCheck, ...]:
+ """G3's Total/Female/Male cells against Track M's All/Women/Men.
+
+ For every breakdown: the share, the weighted and unweighted counts of
+ the cell and of those receiving, the design-based standard error and
+ the floor must equal ``tabulate_track_m``'s cell exactly (which in
+ turn equals the committed artifact's).
+ """
+
+ checks = []
+ for (key, number), result in breakdowns.items():
+ table = (
+ reexecution.sensitivity_tabulation
+ if key == SENSITIVITY_KEY
+ else reexecution.tabulations[key]
+ )
+ for dimension, category, row in _SEX_CELLS:
+ checks.append(
+ common.compare_exact(
+ f"{key}.option_{number}.{dimension}.{category}",
+ _comparable_track_m(_track_m_cell(table, number, row)),
+ _comparable_g3(result.cell(dimension, category)),
+ )
+ )
+ return tuple(checks)
+
+
+# =========================================================================
+# The run
+# =========================================================================
+_SHARED_RESULT_KEYS = (
+ "schema_version",
+ "kind",
+ "statistic_id",
+ "data_provenance",
+ "registration_pointer",
+ "labels",
+ "post_hoc_labels",
+ "scheme",
+ "statistics",
+ "statistic_definitions",
+ "assignment",
+)
+
+
+def _factor(
+ breakdowns: Mapping[tuple[str, int], g3.GroupBreakdownResult],
+) -> tuple[dict[str, Any], dict[str, Any]]:
+ """Blocks shared by every breakdown once, and each breakdown's cells.
+
+ Refuses (rather than drops) a shared block that differs between
+ breakdowns, so the factoring is lossless.
+ """
+
+ shared: dict[str, Any] | None = None
+ conventions: dict[str, Any] | None = None
+ cells: dict[str, Any] = {}
+ for (key, number), result in breakdowns.items():
+ document = result.as_dict()
+ mine = {name: document[name] for name in _SHARED_RESULT_KEYS}
+ convention = dict(document["conventions"])
+ indicator = convention.pop("indicator_column")
+ if shared is None:
+ shared, conventions = mine, convention
+ else:
+ differing = [
+ check.name
+ for check in (
+ common.compare_exact("shared blocks", shared, mine),
+ common.compare_exact(
+ "conventions", conventions, convention
+ ),
+ )
+ if not check.identical
+ ]
+ if differing:
+ raise ValueError(
+ f"breakdown {key} option {number}: {differing} differ "
+ "from the first breakdown's; they cannot be factored"
+ )
+ cells.setdefault(key, {})[str(number)] = {
+ "option": int(number),
+ "indicator_column": indicator,
+ "scored": False,
+ "upstream": document["upstream"],
+ "input_summary": document["input_summary"],
+ "dimensions": document["dimensions"],
+ }
+ assert shared is not None and conventions is not None
+ return {**shared, "conventions": conventions}, cells
+
+
+def invented_group_attribute_loader(
+ cohort_inputs: Any, *, seed: int = 0
+) -> Callable[[Sequence[int]], g1.GroupAttributes]:
+ """A loader of INVENTED side frames for an INVENTED Track M cohort.
+
+ Roles come from the invented 2023 anchor's relationship codes (10
+ reference person: head; 20 and 22: spouse; others none); the inputs
+ are :func:`.common.invented_group_attribute_inputs`, passed through
+ G1's real builder.
+ """
+
+ anchor = cohort_inputs.anchor
+ roles = {
+ int(pid): {10: "head", 20: "spouse", 22: "spouse"}.get(int(code))
+ for pid, code in zip(
+ anchor["person_id"], anchor["relationship"], strict=True
+ )
+ }
+
+ def load(person_ids: Sequence[int]) -> g1.GroupAttributes:
+ ids = sorted(int(pid) for pid in person_ids)
+ persons = pd.DataFrame(
+ {
+ "person_id": ids,
+ "role": pd.array(
+ [roles.get(pid) for pid in ids], dtype="string"
+ ),
+ }
+ )
+ inputs = common.invented_group_attribute_inputs(
+ persons, wave=ANCHOR_WAVES[0], seed=seed
+ )
+ return g1.build_group_attributes(
+ inputs, ids, anchor_waves=ANCHOR_WAVES
+ )
+
+ return load
+
+
+def _check_run_provenance(
+ *,
+ data_provenance: str,
+ registration_pointer: str | None,
+ provenance_kind: str,
+ parent: common.CommittedArtifact,
+) -> None:
+ """Refuse a breakdown run whose provenance does not fit together.
+
+ ``registered_real`` needs the new registration's pointer (an issue #42
+ comment other than the parent run's) and ``psid_files`` records;
+ ``invented`` needs INVENTED records.
+ """
+
+ if data_provenance == tabulation.REGISTERED_REAL:
+ if not isinstance(
+ registration_pointer, str
+ ) or not common.REGISTRATION_POINTER.fullmatch(registration_pointer):
+ raise ValueError(
+ "a real-data breakdown needs its own issue #42 registration "
+ "pointer"
+ )
+ if registration_pointer == parent.document.get("registration_pointer"):
+ raise ValueError(
+ "the breakdown's new outcomes need a new registration: the "
+ "pointer must not be the parent run's"
+ )
+ if provenance_kind != PSID_FILES:
+ raise ValueError("a real-data breakdown reads psid_files records")
+ elif data_provenance == INVENTED:
+ if provenance_kind != INVENTED:
+ raise ValueError("an invented breakdown reads invented records")
+ else:
+ raise ValueError(f"unknown data_provenance {data_provenance!r}")
+
+
+def reproduce_parent(
+ *,
+ parameters: TrackMParameters,
+ parent: common.CommittedArtifact,
+ data_provenance: str,
+ load_cohort_inputs: Callable[[], Any],
+ cola_rates: Mapping[int, float],
+ provenance_kind: str,
+ source: Mapping[str, Any],
+) -> tuple[TrackMReexecution, tuple[common.ReproductionCheck, ...]]:
+ """Steps 1-4 of the module docstring: re-execute and prove it exact.
+
+ The parameters' provenance is compared before the cohort inputs are
+ loaded, the PSID files' SHA-256 before anything is computed, and every
+ re-executed block (:func:`reproduction_checks`) before the caller does
+ anything else. The re-execution runs under the parent's own
+ registration pointer, so its output is comparable with the committed
+ artifact byte for byte. Raises
+ :class:`.common.ReproductionMismatchError` on any difference; computes
+ no group attribute and no group cell, and writes nothing.
+ """
+
+ document = parent.document
+ early = [parameter_check(parameters, document)]
+ common.require_identical(early, stage="parameters")
+ cohort_inputs = load_cohort_inputs()
+ files = psid_files_check(cohort_inputs, document)
+ common.require_identical([files], stage="PSID files")
+ reexecution = reexecute_track_m(
+ cohort_inputs,
+ parameters,
+ cola_rates=cola_rates,
+ data_provenance=data_provenance,
+ registration_pointer=document.get("registration_pointer"),
+ provenance_kind=provenance_kind,
+ source=source,
+ )
+ checks = (*early, files, *reproduction_checks(reexecution, document))
+ common.require_identical(checks, stage="reproduction")
+ return reexecution, checks
+
+
+def run_group_breakdowns(
+ *,
+ parameters: TrackMParameters,
+ parent: common.CommittedArtifact,
+ data_provenance: str,
+ registration_pointer: str | None,
+ load_cohort_inputs: Callable[[], Any],
+ load_side_frame: Callable[[Sequence[int]], g1.GroupAttributes],
+ cola_rates: Mapping[int, float],
+ provenance_kind: str,
+ source: Mapping[str, Any],
+ tax_rates: Any = None,
+ interest_rates: Any = None,
+) -> dict[str, Any]:
+ """Re-execute, prove exact reproduction, then the group cells.
+
+ The order of the module docstring: :func:`reproduce_parent` first, and
+ ``load_side_frame`` is called only after every reproduction check is
+ identical. ``registered_real`` needs the new registration's pointer
+ (an issue #42 comment other than Registration 17's) and
+ ``psid_files`` records; ``invented`` needs INVENTED records. Returns
+ the document; writes nothing.
+ """
+
+ _check_run_provenance(
+ data_provenance=data_provenance,
+ registration_pointer=registration_pointer,
+ provenance_kind=provenance_kind,
+ parent=parent,
+ )
+ document = parent.document
+ reexecution, checks = reproduce_parent(
+ parameters=parameters,
+ parent=parent,
+ data_provenance=data_provenance,
+ load_cohort_inputs=load_cohort_inputs,
+ cola_rates=cola_rates,
+ provenance_kind=provenance_kind,
+ source=source,
+ )
+
+ # -- only now: group attributes and group cells ------------------------
+ universe = [int(pid) for pid in reexecution.cohort.persons["person_id"]]
+ side_frame = load_side_frame(universe)
+ attributes = person_attributes(
+ reexecution,
+ parameters,
+ side_frame,
+ tax_rates=(
+ g2.load_oasdi_tax_rates() if tax_rates is None else tax_rates
+ ),
+ interest_rates=(
+ g2.load_trust_fund_interest_rates()
+ if interest_rates is None
+ else interest_rates
+ ),
+ )
+ assignment = assign(attributes)
+ breakdowns = tabulate_breakdowns(
+ reexecution,
+ assignment,
+ data_provenance=data_provenance,
+ registration_pointer=registration_pointer,
+ parent_sha256=parent.sha256,
+ )
+ consistency = consistency_checks(reexecution, breakdowns)
+ common.require_identical(consistency, stage="consistency")
+ shared, cells = _factor(breakdowns)
+ labels = list(shared["labels"]) + list(common.POST_HOC_LABELS)
+ return {
+ "header": (
+ common.INVENTED_DATA_HEADER
+ if data_provenance == INVENTED
+ else None
+ ),
+ "schema_version": SCHEMA_VERSION,
+ "description": (
+ "Exercise 4 (Track M: the share of OASDI beneficiaries 62 and "
+ "older receiving a minimum benefit, options 2-5, PSID 2023 "
+ "wave, income year 2022) broken down by MINT8's characteristic "
+ "subgroups, every registered row MS0-MS6 and d430's "
+ "sensitivity; post hoc, not blind, report-only, unscored"
+ ),
+ "publishes_regardless": True,
+ "blind": False,
+ "scored": False,
+ "data_provenance": data_provenance,
+ "registration_pointer": registration_pointer,
+ "labels": labels,
+ "disclosure": COVERED_EARNINGS_DISCLOSURE,
+ "parent": {
+ **parent.record(_ROOT),
+ "specification": copy.deepcopy(document.get("specification")),
+ "role": (
+ "the committed run this breakdown re-executes and must "
+ "reproduce exactly; never edited"
+ ),
+ },
+ "reproduction": {
+ "comparison": dict(common.EXACT_COMPARISON),
+ "reexecuted_under_registration_pointer": document.get(
+ "registration_pointer"
+ ),
+ "reexecution_note": (
+ "the registered computation is re-executed under its own "
+ "registration's pointer so that its output is comparable "
+ "byte for byte with the committed artifact; every group "
+ "cell carries this breakdown's pointer"
+ ),
+ "location_fields": {
+ "fields": ["$." + ".".join(keys) for keys in LOCATION_FIELDS],
+ "rule": LOCATION_RULE,
+ "committed": [
+ _recorded(document, keys) for keys in LOCATION_FIELDS
+ ],
+ "this_run": [
+ _recorded(reexecution.result, keys)
+ for keys in LOCATION_FIELDS
+ ],
+ },
+ "identical": True,
+ "checks": [check.as_dict() for check in checks],
+ },
+ "consistency_with_registered_cells": {
+ "rule": (
+ "G3's Total, Female and Male cells equal tabulate_track_m's "
+ "All, Women and Men cells exactly (share, weighted and "
+ "unweighted counts of the cell and of those receiving, "
+ "design SE, floor), for every row, option and d430"
+ ),
+ "identical": True,
+ "n_checks": len(consistency),
+ "checks": [check.as_dict(max_paths=5) for check in consistency],
+ },
+ "statistic": {
+ **tabulation.statistic_block(),
+ "id": STATISTIC_ID,
+ "formula": (
+ "100 * sum_i w_i A_k,i / sum_i w_i over the persons of the "
+ "group (tabulation._share's formula; G3 "
+ "tabulate_share_breakdown)"
+ ),
+ "uncertainty": tabulation.uncertainty_block(),
+ },
+ "group_attributes": attributes.provenance,
+ "quintile_rule_sensitivity": quintile_rule_sensitivity(
+ attributes, assignment
+ ),
+ "not_computed": {
+ "dimensions": dict(NOT_COMPUTED_DIMENSIONS),
+ "statistics": copy.deepcopy(NOT_COMPUTED_STATISTICS),
+ },
+ "named_deltas": [
+ *document.get("named_deltas", pipeline.NAMED_DELTAS),
+ *GROUP_NAMED_DELTAS,
+ ],
+ "breakdown": shared,
+ "breakdowns": cells,
+ }
diff --git a/tests/group_breakdowns/__init__.py b/tests/group_breakdowns/__init__.py
new file mode 100644
index 00000000..e69de29b
diff --git a/tests/group_breakdowns/test_min_benefit_common.py b/tests/group_breakdowns/test_min_benefit_common.py
new file mode 100644
index 00000000..8a48fe89
--- /dev/null
+++ b/tests/group_breakdowns/test_min_benefit_common.py
@@ -0,0 +1,509 @@
+"""The shared pieces of the group-breakdown adapters (NASI G4c).
+
+Unit tier: everything here is INVENTED or synthetic, and nothing reads a
+PSID file or a committed artifact. The exact comparison
+(:func:`common.compare_exact`) is load-bearing: it is what stops a
+breakdown before any group attribute is read, so its invariants are
+property-tested:
+
+* a document compared with itself has no difference, and every leaf is
+ compared;
+* moving any one leaf (a float by one unit in the last place, an integer
+ by one, a string, a sign of zero) is found, at exactly that leaf's path;
+* the refusal names paths only, never a value.
+"""
+
+from __future__ import annotations
+
+import copy
+import json
+import math
+from typing import Any
+
+import numpy as np
+import pandas as pd
+import pytest
+from hypothesis import given, settings
+from hypothesis import strategies as st
+
+from populace_dynamics.cohorts import group_attributes as g1
+from populace_dynamics.data import group_attributes_psid as gap
+from populace_dynamics.estimates import group_breakdown as g3
+from populace_dynamics.group_breakdowns import common
+from populace_dynamics.min_benefit_track_m import DRY_RUN_HEADER
+from populace_dynamics.track_a_v2 import INVENTED_HEADER
+
+# -------------------------------------------------------------------------
+# Strategies: JSON documents as the registered runners write them
+# -------------------------------------------------------------------------
+_LEAVES = st.one_of(
+ st.none(),
+ st.booleans(),
+ st.integers(min_value=-(10**12), max_value=10**12),
+ st.floats(allow_nan=False, allow_infinity=False),
+ st.text(max_size=8),
+)
+_DOCUMENTS = st.recursive(
+ _LEAVES,
+ lambda children: st.one_of(
+ st.lists(children, max_size=4),
+ st.dictionaries(st.text(min_size=1, max_size=6), children, max_size=4),
+ ),
+ max_leaves=25,
+)
+
+
+def _leaf_paths(value: Any, path: str = "$") -> list[tuple[str, Any]]:
+ if isinstance(value, dict):
+ out = []
+ for key in sorted(value, key=str):
+ out.extend(_leaf_paths(value[key], f"{path}.{key}"))
+ return out
+ if isinstance(value, list):
+ out = []
+ for index, item in enumerate(value):
+ out.extend(_leaf_paths(item, f"{path}[{index}]"))
+ return out
+ return [(path, value)]
+
+
+def _set(document: Any, path: str, new: Any) -> Any:
+ """``document`` with the leaf at ``path`` replaced (paths as above)."""
+
+ out = copy.deepcopy(document)
+ for candidate, _ in _leaf_paths(out):
+ if candidate != path:
+ continue
+ holder: Any = None
+ key: Any = None
+ node: Any = out
+ tokens = _tokens(path)
+ for token in tokens:
+ holder, key = node, token
+ node = node[token]
+ if holder is None:
+ return new
+ holder[key] = new
+ return out
+ raise KeyError(path)
+
+
+def _tokens(path: str) -> list[Any]:
+ tokens: list[Any] = []
+ rest = path[1:]
+ while rest:
+ if rest.startswith("["):
+ end = rest.index("]")
+ tokens.append(int(rest[1:end]))
+ rest = rest[end + 1 :]
+ else:
+ rest = rest[1:]
+ cut = min(
+ [i for i in (rest.find("."), rest.find("[")) if i >= 0]
+ or [len(rest)]
+ )
+ tokens.append(rest[:cut])
+ rest = rest[cut:]
+ return tokens
+
+
+def _moved(value: Any) -> Any:
+ """A leaf one step away from ``value``, of the same JSON type."""
+
+ if value is None:
+ return 0
+ if isinstance(value, bool):
+ return not value
+ if isinstance(value, int):
+ return value + 1
+ if isinstance(value, float):
+ if value == 0.0:
+ return -value if math.copysign(1.0, value) > 0 else 0.0
+ up = math.nextafter(value, math.inf)
+ return up if math.isfinite(up) else math.nextafter(value, -math.inf)
+ return value + "x"
+
+
+def _simple_keys(document: Any) -> bool:
+ """Keys the path notation can address unambiguously."""
+
+ if isinstance(document, dict):
+ return all(
+ not any(ch in key for ch in ".[]") and _simple_keys(value)
+ for key, value in document.items()
+ )
+ if isinstance(document, list):
+ return all(_simple_keys(item) for item in document)
+ return True
+
+
+# -------------------------------------------------------------------------
+# The exact comparison
+# -------------------------------------------------------------------------
+@settings(max_examples=100, deadline=None)
+@given(_DOCUMENTS)
+def test_property_a_document_equals_itself_leaf_by_leaf(document):
+ normalized = common.json_normalized(document)
+ check = common.compare_exact("self", document, copy.deepcopy(document))
+ assert check.identical
+ assert check.differences == ()
+ assert check.n_leaves == len(_leaf_paths(normalized))
+
+
+@settings(max_examples=150, deadline=None)
+@given(_DOCUMENTS, st.data())
+def test_property_any_one_moved_leaf_is_found_at_its_path(document, data):
+ document = common.json_normalized(document)
+ if not _simple_keys(document):
+ return
+ leaves = _leaf_paths(document)
+ if not leaves:
+ return
+ path, value = data.draw(st.sampled_from(leaves))
+ moved = _set(document, path, _moved(value))
+ check = common.compare_exact("moved", document, moved)
+ assert check.differences == (path,)
+ with pytest.raises(common.ReproductionMismatchError) as error:
+ common.require_identical([check], stage="test")
+ message = str(error.value)
+ assert repr(path) in message
+ assert "nothing is written" in message
+ shown = (repr(value), repr(_moved(value)))
+ if isinstance(value, float) and not any(text in path for text in shown):
+ # Paths only: the refusal never prints the committed value.
+ assert not any(text in message for text in shown)
+
+
+def test_floats_compare_bit_for_bit():
+ assert not common.compare_exact("z", 0.0, -0.0).identical
+ nudged = math.nextafter(0.1 + 0.2, math.inf)
+ assert not common.compare_exact("u", 0.1 + 0.2, nudged).identical
+ assert common.compare_exact(
+ "same", 0.30000000000000004, 0.1 + 0.2
+ ).identical
+
+
+@pytest.mark.parametrize(
+ ("committed", "recomputed", "expected"),
+ [
+ (1, 1.0, ["$"]),
+ (True, 1, ["$"]),
+ (None, 0, ["$"]),
+ ({"a": 1}, {"a": 1, "b": 2}, ["$.b (present on one side only)"]),
+ ({"a": [1, 2]}, {"a": [1]}, ["$.a (lengths differ)"]),
+ ({"a": {"b": 1}}, {"a": 1}, ["$.a (structure differs)"]),
+ ({"a": [1, 2.5]}, {"a": [1, 2.5]}, []),
+ ],
+)
+def test_type_key_and_length_differences(committed, recomputed, expected):
+ check = common.compare_exact("x", committed, recomputed)
+ assert list(check.differences) == expected
+
+
+def test_comparison_goes_through_the_runners_json_encoding():
+ """Tuples and lists encode alike, so they compare equal; NaN refuses."""
+
+ assert common.compare_exact("t", {"a": (1, 2)}, {"a": [1, 2]}).identical
+ with pytest.raises(ValueError):
+ common.compare_exact("nan", float("nan"), float("nan"))
+
+
+def test_require_identical_passes_identical_checks():
+ common.require_identical(
+ [common.compare_exact("a", {"x": 1.5}, {"x": 1.5})], stage="test"
+ )
+
+
+def test_check_record_caps_paths():
+ check = common.ReproductionCheck(
+ "many", 10, tuple(f"$[{i}]" for i in range(80))
+ )
+ record = check.as_dict()
+ assert record["n_differences"] == 80
+ assert len(record["differing_paths"]) == 50
+ assert record["identical"] is False
+
+
+def test_exact_comparison_declares_no_tolerance():
+ assert common.EXACT_COMPARISON["tolerance"] is None
+
+
+# -------------------------------------------------------------------------
+# Labels
+# -------------------------------------------------------------------------
+def test_the_invented_headers_agree():
+ assert common.INVENTED_DATA_HEADER == DRY_RUN_HEADER == INVENTED_HEADER
+
+
+def test_post_hoc_labels():
+ assert common.POST_HOC_LABELS == (
+ "registered, one-shot, post hoc, not blind",
+ "report-only",
+ )
+
+
+# -------------------------------------------------------------------------
+# Files
+# -------------------------------------------------------------------------
+def test_write_new_never_overwrites(tmp_path):
+ path = tmp_path / "artifact.json"
+ common.write_new(path, "one")
+ with pytest.raises(FileExistsError):
+ common.write_new(path, "two")
+ assert path.read_text() == "one"
+
+
+def _artifact(tmp_path, document=None, sidecar=None):
+ path = tmp_path / "parent_v1.json"
+ path.write_text(json.dumps(document or {"cells": [1.5]}))
+ digest = common.file_sha256(path)
+ record = {"artifact": path.name, "artifact_sha256": digest}
+ if sidecar is not None:
+ record.update(sidecar)
+ path.with_suffix(".env.json").write_text(json.dumps(record))
+ return path, digest
+
+
+def test_a_committed_artifact_loads_with_its_pin_and_sidecar(tmp_path):
+ path, digest = _artifact(tmp_path)
+ parent = common.load_committed_artifact(path, expected_sha256=digest)
+ assert parent.document == {"cells": [1.5]}
+ assert parent.sha256 == digest
+ assert parent.record(tmp_path)["path"] == path.name
+ assert parent.record()["sidecar_artifact_sha256"] == digest
+
+
+def test_other_bytes_are_refused(tmp_path):
+ path, _ = _artifact(tmp_path)
+ with pytest.raises(ValueError, match="not the committed"):
+ common.load_committed_artifact(path, expected_sha256="0" * 64)
+
+
+@pytest.mark.parametrize(
+ ("sidecar", "match"),
+ [
+ ({"artifact": "another.json"}, "does not name"),
+ ({"artifact_sha256": "0" * 64}, "does not bind"),
+ ],
+)
+def test_an_unbound_sidecar_is_refused(tmp_path, sidecar, match):
+ path, digest = _artifact(tmp_path, sidecar=sidecar)
+ with pytest.raises(ValueError, match=match):
+ common.load_committed_artifact(path, expected_sha256=digest)
+
+
+def test_an_invented_parent_records_no_path():
+ record = common.CommittedArtifact(document={"x": 1}).record()
+ assert record["path"] is None
+ assert record["sha256"] is None
+
+
+# -------------------------------------------------------------------------
+# The cohort side frame (G1) on G3's codes
+# -------------------------------------------------------------------------
+_RACE = {d.key: d for d in g3.MINT8_SCHEME.dimensions}["race_ethnicity"]
+_COUNTRY = {d.key: d for d in g3.MINT8_SCHEME.dimensions}["country_of_birth"]
+_EDUCATION = {d.key: d for d in g3.MINT8_SCHEME.dimensions}["education"]
+
+
+def _side(rows):
+ return pd.DataFrame(
+ rows,
+ columns=[
+ "race_ethnicity_mint8",
+ "race_ethnicity_mint8_status",
+ "race_ethnicity_status",
+ ],
+ )
+
+
+def test_side_frame_labels_become_g3_codes():
+ frame = _side(
+ [
+ ("White, non-Hispanic", g1.ASSIGNED, g1.KNOWN),
+ ("Hispanic or Latino, any race", g1.ASSIGNED, g1.KNOWN),
+ (pd.NA, g1.ATTRIBUTE_UNKNOWN, g1.NEVER_HEAD_OR_SPOUSE),
+ (pd.NA, "unresolved:multiple_races", g1.KNOWN),
+ ]
+ )
+ codes = common.side_frame_codes(frame, _RACE)
+ assert codes.codes == (
+ "white_non_hispanic",
+ "hispanic_any_race",
+ "attribute_unknown:never_head_or_spouse",
+ "unresolved:multiple_races",
+ )
+ assert codes.unclassified_codes == (
+ "attribute_unknown:never_head_or_spouse",
+ "unresolved:multiple_races",
+ )
+
+
+@pytest.mark.parametrize(
+ ("row", "match"),
+ [
+ (("Martian", g1.ASSIGNED, g1.KNOWN), "is not one of"),
+ (("White, non-Hispanic", g1.ATTRIBUTE_UNKNOWN, g1.KNOWN), "carries"),
+ ((pd.NA, "guessed", g1.KNOWN), "undocumented"),
+ ],
+)
+def test_side_frame_codes_refuse_what_g1_does_not_say(row, match):
+ with pytest.raises(ValueError, match=match):
+ common.side_frame_codes(_side([row]), _RACE)
+
+
+def test_only_side_frame_dimensions_map():
+ with pytest.raises(ValueError, match="not a side-frame dimension"):
+ common.side_frame_codes(_side([]), _EDUCATION)
+
+
+def test_every_g1_mint8_label_is_a_g3_label():
+ """G1's scheme file and G3's MINT8 scheme print the same labels."""
+
+ schemes = g1.load_schemes()["schemes"]["mint8"]["dimensions"]
+ for key, dimension in (
+ ("race_ethnicity", _RACE),
+ ("country_of_birth", _COUNTRY),
+ ):
+ g3_labels = {category.label for category in dimension.categories}
+ assert set(schemes[key]["categories"].values()) <= g3_labels
+
+
+#: G1's years-of-schooling domain: 0 ("completed no grades", from the
+#: family-file recode) through the individual file's top code.
+_YEARS_DOMAIN = (0, gap.EDUCATION_YEARS_RANGE[1])
+_YEARS = st.integers(min_value=_YEARS_DOMAIN[0], max_value=_YEARS_DOMAIN[1])
+
+
+@settings(max_examples=60, deadline=None)
+@given(st.lists(_YEARS, max_size=40))
+def test_property_g1_and_g3_place_every_years_value_alike(years):
+ """Differential: G1's MINT8 education label and G3's band agree on
+ every value of G1's documented domain."""
+
+ spec = g1.load_schemes()["schemes"]["mint8"]["dimensions"]["education"]
+ frame = pd.DataFrame(
+ {"education_years": pd.array(years + [pd.NA], dtype="Int64")}
+ )
+ labels, status = g1.apply_category_scheme(frame, "education", spec)
+ frame["education_mint8"] = labels
+ record = common.education_agreement(frame, _EDUCATION)
+ assert record["n_with_years"] == len(years)
+ assert (status.iloc[:-1] == g1.ASSIGNED).all()
+
+
+def test_g1_and_g3_agree_on_the_whole_domain():
+ low, high = _YEARS_DOMAIN
+ spec = g1.load_schemes()["schemes"]["mint8"]["dimensions"]["education"]
+ frame = pd.DataFrame(
+ {
+ "education_years": pd.array(
+ list(range(low, high + 1)), dtype="Int64"
+ )
+ }
+ )
+ frame["education_mint8"], _ = g1.apply_category_scheme(
+ frame, "education", spec
+ )
+ record = common.education_agreement(frame, _EDUCATION)
+ assert record["n_with_years"] == high - low + 1
+
+
+def test_g1_refuses_years_outside_its_domain():
+ spec = g1.load_schemes()["schemes"]["mint8"]["dimensions"]["education"]
+ frame = pd.DataFrame(
+ {"education_years": pd.array([_YEARS_DOMAIN[1] + 1], dtype="Int64")}
+ )
+ with pytest.raises(ValueError, match="undocumented"):
+ g1.apply_category_scheme(frame, "education", spec)
+
+
+def test_an_education_disagreement_is_refused():
+ frame = pd.DataFrame(
+ {
+ "education_years": pd.array([12], dtype="Int64"),
+ "education_mint8": ["Bachelor"],
+ }
+ )
+ with pytest.raises(ValueError, match="differently"):
+ common.education_agreement(frame, _EDUCATION)
+
+
+def test_marital_codes_declare_only_present_non_mint_statuses():
+ codes = common.marital_codes(
+ ["married", "unknown", "widowed"],
+ unclassified=("unknown", "no_marriage_history"),
+ )
+ assert codes.codes == ("married", "unknown", "widowed")
+ assert codes.unclassified_codes == ("unknown",)
+ with pytest.raises(ValueError, match="cannot be declared"):
+ common.marital_codes(["married"], unclassified=("married",))
+
+
+@given(
+ st.lists(st.integers(min_value=1880, max_value=2030), max_size=20),
+ st.integers(min_value=1950, max_value=2100),
+)
+def test_property_age_in_year(births, year):
+ ages = common.age_in_year(births, year)
+ assert [age + birth for age, birth in zip(ages, births, strict=True)] == [
+ year
+ ] * len(births)
+
+
+def test_age_refuses_booleans():
+ with pytest.raises(ValueError, match="integers"):
+ common.age_in_year([True], 2022)
+ with pytest.raises(ValueError, match="integers"):
+ common.age_in_year([np.bool_(False)], 2022)
+
+
+# -------------------------------------------------------------------------
+# INVENTED side-frame inputs
+# -------------------------------------------------------------------------
+def _invented_persons(n: int = 30) -> pd.DataFrame:
+ roles = ["head", "spouse", None]
+ return pd.DataFrame(
+ {
+ "person_id": [1001 + i for i in range(n)],
+ "role": pd.array([roles[i % 3] for i in range(n)], dtype="string"),
+ }
+ )
+
+
+def test_invented_inputs_pass_g1s_real_builder():
+ persons = _invented_persons()
+ inputs = common.invented_group_attribute_inputs(persons, wave=2023)
+ assert inputs.provenance["kind"] == "INVENTED"
+ assert inputs.provenance["label"] == common.INVENTED_DATA_HEADER
+ built = g1.build_group_attributes(
+ inputs, persons["person_id"].tolist(), anchor_waves=(2023,)
+ )
+ assert built.frame["person_id"].tolist() == sorted(persons["person_id"])
+ # The PSID asks race and birthplace of heads and spouses only.
+ others = built.frame.set_index("person_id").loc[
+ persons.loc[persons["role"].isna(), "person_id"]
+ ]
+ assert (others["race_ethnicity_status"] == g1.NEVER_HEAD_OR_SPOUSE).all()
+ record = common.education_agreement(built.frame, _EDUCATION)
+ assert record["g3_band_equals_g1_label"] is True
+
+
+def test_invented_inputs_are_deterministic():
+ persons = _invented_persons()
+ first = common.invented_group_attribute_inputs(persons, wave=2023, seed=5)
+ second = common.invented_group_attribute_inputs(persons, wave=2023, seed=5)
+ pd.testing.assert_frame_equal(first.reports, second.reports)
+ pd.testing.assert_frame_equal(first.education, second.education)
+
+
+def test_invented_inputs_refuse_a_wave_without_birthplace():
+ with pytest.raises(ValueError, match="2013-2023"):
+ common.invented_group_attribute_inputs(_invented_persons(), wave=1999)
+
+
+def test_invented_inputs_refuse_an_undocumented_role():
+ persons = _invented_persons(3)
+ persons["role"] = pd.array(["head", "boarder", None], dtype="string")
+ with pytest.raises(ValueError, match="undocumented role"):
+ common.invented_group_attribute_inputs(persons, wave=2023)
diff --git a/tests/group_breakdowns/test_min_benefit_groups.py b/tests/group_breakdowns/test_min_benefit_groups.py
new file mode 100644
index 00000000..2bb9a070
--- /dev/null
+++ b/tests/group_breakdowns/test_min_benefit_groups.py
@@ -0,0 +1,1001 @@
+"""Exercise 4 (Track M) by MINT8 subgroups, on INVENTED data (NASI G4c).
+
+INVENTED DATA - NOT A COMPARISON. Unit tier: the cohort is
+``min_benefit_track_m.invented_psid``'s PSID-shaped frames, the
+parameters ``min_benefit_track_m.invented.invented_parameters`` (no
+policyengine-us checkout is read) and the side frame
+``group_breakdowns.common.invented_group_attribute_inputs`` through G1's
+real builder. The "committed parent" is the registered computation run
+once on those frames, through the registered runner's JSON encoding.
+
+Invariants stated and tested here:
+
+1. **Order.** Nothing about groups is computed unless the re-executed
+ registered computation equals the parent exactly: a parent with any one
+ committed cell moved (a share, a design SE, a floor, a cohort-structure
+ count, a d430 cell, the parameters, the PSID file hashes) is refused,
+ and the side-frame loader is never called.
+2. **Partition.** In every breakdown and every dimension, the categories
+ and the unclassified rows partition Total: unweighted counts exactly,
+ weighted counts and weighted numerators to float summation.
+3. **Differential.** Every group cell's share equals the direct formula
+ ``100 * fsum(w A) / fsum(w)`` over the group's rows (bit for bit), and
+ G3's Total, Female and Male cells equal Track M's All, Women and Men
+ cells exactly (the run's own consistency checks, recomputed here).
+4. **Bounds.** Shares lie in [0, 100]; standard errors and floors are
+ nonnegative; every age is at least 62.
+5. **Benefit type** (:func:`min_benefit.benefit_type_2022`): property-
+ tested over every combination of the six type items, receipt and
+ record basis.
+"""
+
+from __future__ import annotations
+
+import copy
+import dataclasses
+import itertools
+import json
+import math
+from collections.abc import Callable
+from types import SimpleNamespace
+from typing import Any
+
+import numpy as np
+import pandas as pd
+import pytest
+from hypothesis import given, settings
+from hypothesis import strategies as st
+
+from populace_dynamics.data import social_security_receipt as ssr
+from populace_dynamics.estimates import group_breakdown as g3
+from populace_dynamics.estimates import lifetime_measures as g2
+from populace_dynamics.group_breakdowns import common
+from populace_dynamics.group_breakdowns import min_benefit as mb
+from populace_dynamics.min_benefit_track_m import (
+ OUTPUT_LABELS,
+ cohort,
+ invented,
+ invented_psid,
+ tabulation,
+)
+from populace_dynamics.min_benefit_track_m.evaluation import (
+ INVENTED,
+ PSID_FILES,
+)
+from populace_dynamics.min_benefit_track_m.policy import (
+ REGISTERED_ROWS,
+ TABLE6_OPTIONS,
+)
+
+SEED = 11
+FAMILY_UNITS = 40
+SIDE_FRAME_SEED = 3
+POINTER = (
+ "https://github.com/PolicyEngine/microcosm-dynamics/issues/42"
+ "#issuecomment-7"
+)
+_BENEFIT_CODES = {
+ "retired_worker_only",
+ "widower",
+ "spousal",
+ "disabled_worker_only",
+}
+
+
+def _frames():
+ # Receipt of unknown or "other" type before 62 for a third of persons,
+ # so d430's sensitivity reading differs from the scored one.
+ return invented_psid.invented_cohort_inputs(
+ seed=SEED,
+ n_family_units=FAMILY_UNITS,
+ unknown_or_other_before_62=0.3,
+ )
+
+
+def _reexecute(frames, params, cola):
+ return mb.reexecute_track_m(
+ frames,
+ params,
+ cola_rates=cola,
+ data_provenance=INVENTED,
+ registration_pointer=None,
+ provenance_kind=INVENTED,
+ source=dict(frames.provenance),
+ )
+
+
+def _run(world, parent=None, *, load_side_frame=None, **overrides):
+ kwargs = dict(
+ parameters=world.params,
+ parent=parent or world.parent,
+ data_provenance=INVENTED,
+ registration_pointer=None,
+ load_cohort_inputs=lambda: world.frames,
+ load_side_frame=load_side_frame or world.spy([]),
+ cola_rates=world.cola,
+ provenance_kind=INVENTED,
+ source=dict(world.frames.provenance),
+ )
+ kwargs.update(overrides)
+ return mb.run_group_breakdowns(**kwargs)
+
+
+@pytest.fixture(scope="module")
+def world():
+ params, cola = invented.invented_parameters()
+ frames = _frames()
+ reexecution = _reexecute(frames, params, cola)
+ parent = common.CommittedArtifact(
+ document=common.json_normalized(dict(reexecution.result))
+ )
+ loader = mb.invented_group_attribute_loader(frames, seed=SIDE_FRAME_SEED)
+ calls: list[list[int]] = []
+
+ def spy(record):
+ def load(person_ids):
+ record.append(list(person_ids))
+ return loader(person_ids)
+
+ return load
+
+ state = SimpleNamespace(
+ params=params,
+ cola=cola,
+ frames=frames,
+ reexecution=reexecution,
+ parent=parent,
+ loader=loader,
+ calls=calls,
+ spy=spy,
+ )
+ state.document = _run(state, load_side_frame=spy(calls))
+ return state
+
+
+@pytest.fixture(scope="module")
+def pieces(world):
+ """The group work's objects, computed directly on the same
+ re-execution (the run returns only the document)."""
+
+ universe = [int(p) for p in world.reexecution.cohort.persons["person_id"]]
+ side_frame = world.loader(universe)
+ attributes = mb.person_attributes(
+ world.reexecution,
+ world.params,
+ side_frame,
+ tax_rates=g2.load_oasdi_tax_rates(),
+ interest_rates=g2.load_trust_fund_interest_rates(),
+ )
+ assignment = mb.assign(attributes)
+ breakdowns = mb.tabulate_breakdowns(
+ world.reexecution,
+ assignment,
+ data_provenance=INVENTED,
+ registration_pointer=None,
+ parent_sha256=None,
+ )
+ return SimpleNamespace(
+ side_frame=side_frame,
+ attributes=attributes,
+ assignment=assignment,
+ breakdowns=breakdowns,
+ )
+
+
+# =========================================================================
+# The document
+# =========================================================================
+def test_the_document_is_labelled_invented_post_hoc_and_report_only(world):
+ document = world.document
+ assert document["header"] == common.INVENTED_DATA_HEADER
+ assert document["data_provenance"] == INVENTED
+ assert document["blind"] is False
+ assert document["scored"] is False
+ assert document["publishes_regardless"] is True
+ labels = document["labels"]
+ assert labels[0] == g3.INVENTED_DATA_LABEL
+ assert labels[1 : 1 + len(OUTPUT_LABELS)] == list(OUTPUT_LABELS)
+ assert labels[-2:] == list(common.POST_HOC_LABELS)
+ assert document["breakdown"]["post_hoc_labels"] == list(
+ common.POST_HOC_LABELS
+ )
+ json.dumps(document, allow_nan=False)
+
+
+def test_every_row_option_and_d430_is_broken_down(world):
+ breakdowns = world.document["breakdowns"]
+ assert list(breakdowns) == [*REGISTERED_ROWS, mb.SENSITIVITY_KEY]
+ for options in breakdowns.values():
+ assert list(options) == [str(n) for n in TABLE6_OPTIONS]
+ for number, cell in options.items():
+ assert cell["indicator_column"] == f"receives_{number}"
+ assert cell["scored"] is False
+ assert [d["key"] for d in cell["dimensions"]] == [
+ d.key for d in mb.SCHEME.dimensions
+ ]
+
+
+def test_the_reproduction_is_exact_and_complete(world):
+ reproduction = world.document["reproduction"]
+ assert reproduction["identical"] is True
+ assert reproduction["comparison"]["tolerance"] is None
+ names = [check["name"] for check in reproduction["checks"]]
+ assert all(check["identical"] for check in reproduction["checks"])
+ assert all(
+ check["n_leaves_compared"] > 0 for check in reproduction["checks"]
+ )
+ expected = {
+ "parameters",
+ "inputs.source.psid_files_sha256",
+ "pipeline.rows",
+ "pipeline.cohort_structure",
+ "pipeline.sensitivities",
+ f"group_rows.{mb.SENSITIVITY_KEY}.tabulation",
+ f"group_rows.{mb.SENSITIVITY_KEY}.diagnostics",
+ }
+ for row in REGISTERED_ROWS:
+ expected |= {
+ f"group_rows.{row}.tabulation",
+ f"group_rows.{row}.diagnostics",
+ }
+ assert expected <= set(names)
+ # Every block the pipeline returns is compared, but the two the
+ # registered runner replaced.
+ pipeline_keys = {
+ n.split(".", 1)[1] for n in names if n.startswith("pipeline.")
+ }
+ assert pipeline_keys == set(world.reexecution.result) - {
+ "header",
+ "specification",
+ }
+
+
+def test_the_loader_is_called_once_with_the_universe(world):
+ assert len(world.calls) == 1
+ assert world.calls[0] == [
+ int(p) for p in world.reexecution.cohort.persons["person_id"]
+ ]
+
+
+def test_the_registered_cells_are_reproduced_by_g3(world):
+ consistency = world.document["consistency_with_registered_cells"]
+ assert consistency["identical"] is True
+ assert consistency["n_checks"] == (
+ (len(REGISTERED_ROWS) + 1) * len(TABLE6_OPTIONS) * 3
+ )
+
+
+def test_what_is_not_computed_says_why(world):
+ not_computed = world.document["not_computed"]
+ assert set(not_computed["dimensions"]) == {
+ "poverty_status",
+ "household_income_quintile",
+ }
+ assert all(
+ reason.startswith("not computed:")
+ for reason in not_computed["dimensions"].values()
+ )
+ statistics = not_computed["statistics"]
+ assert statistics["mint8_benefit_statistics"]["statistics"] == list(
+ g3.MINT_BENEFIT_STATISTICS
+ )
+ assert statistics["mint8_poverty_statistics"]["statistics"] == list(
+ g3.POVERTY_STATISTICS
+ )
+ dimensions = {d.key for d in mb.SCHEME.dimensions}
+ assert not dimensions & set(not_computed["dimensions"])
+
+
+def test_the_scheme_is_mint8_with_lifetime_less_what_is_not_computed():
+ full = [d.key for d in g3.MINT8_ANNUAL_WITH_LIFETIME_SCHEME.dimensions]
+ kept = [d.key for d in mb.SCHEME.dimensions]
+ assert kept == [k for k in full if k not in mb.NOT_COMPUTED_DIMENSIONS]
+ by_key = {
+ d.key: d for d in g3.MINT8_ANNUAL_WITH_LIFETIME_SCHEME.dimensions
+ }
+ for dimension in mb.SCHEME.dimensions:
+ assert [c.label for c in dimension.categories] == [
+ c.label for c in by_key[dimension.key].categories
+ ]
+ assert set(mb.COLUMNS_MAP) == set(kept) - {"total"}
+
+
+# =========================================================================
+# Invariants of the cells
+# =========================================================================
+def _result_cells(result: g3.GroupBreakdownResult):
+ for dimension in result.dimensions:
+ for cell in dimension.cells:
+ yield dimension, cell
+
+
+def test_categories_and_unclassified_partition_total(world, pieces):
+ assignment = pieces.assignment
+ for (key, number), result in pieces.breakdowns.items():
+ rows = (
+ world.reexecution.sensitivity_evaluation.rows
+ if key == mb.SENSITIVITY_KEY
+ else world.reexecution.evaluations[key].rows
+ )
+ weights = rows["weight"].to_numpy(dtype=float)
+ receiving = rows[f"receives_{number}"].to_numpy(dtype=bool)
+ total = result.cell("total", "total")
+ for dimension in result.dimensions:
+ if dimension.kind == g3.TOTAL:
+ continue
+ masks = [
+ assignment.mask(dimension.key, cell.category)
+ for cell in dimension.cells
+ ]
+ unclassified = assignment.unclassified_mask(dimension.key)
+ stacked = np.vstack([*masks, unclassified])
+ # disjoint and covering: every row in exactly one group
+ assert (stacked.sum(axis=0) == 1).all()
+ n = sum(c.counts["unweighted_n"] for c in dimension.cells)
+ assert (
+ n + dimension.unclassified["unweighted_n"]
+ == total.counts["unweighted_n"]
+ )
+ w = math.fsum(
+ [c.counts["weighted_n"] for c in dimension.cells]
+ + [dimension.unclassified["weighted_n"]]
+ )
+ assert math.isclose(w, total.counts["weighted_n"], rel_tol=1e-12)
+ numerators = [
+ math.fsum(weights[mask & receiving]) for mask in masks
+ ] + [math.fsum(weights[unclassified & receiving])]
+ assert math.isclose(
+ math.fsum(numerators),
+ math.fsum(weights[receiving]),
+ rel_tol=1e-12,
+ abs_tol=1e-9,
+ )
+
+
+def test_every_share_is_the_direct_formula(world, pieces):
+ """Differential: G3's cell against ``100 * fsum(w A) / fsum(w)``."""
+
+ assignment = pieces.assignment
+ compared = 0
+ for (key, number), result in pieces.breakdowns.items():
+ rows = (
+ world.reexecution.sensitivity_evaluation.rows
+ if key == mb.SENSITIVITY_KEY
+ else world.reexecution.evaluations[key].rows
+ )
+ weights = rows["weight"].to_numpy(dtype=float)
+ receiving = rows[f"receives_{number}"].to_numpy(dtype=bool)
+ for dimension, cell in _result_cells(result):
+ mask = (
+ np.ones(len(rows), dtype=bool)
+ if dimension.kind == g3.TOTAL
+ else assignment.mask(dimension.key, cell.category)
+ )
+ statistic = cell.statistic(g3.SHARE)
+ assert statistic.unweighted_n == int(mask.sum())
+ if not mask.any():
+ assert not statistic.defined
+ continue
+ expected = (
+ 100.0
+ * math.fsum(weights[mask & receiving].tolist())
+ / math.fsum(weights[mask].tolist())
+ )
+ assert statistic.defined
+ assert statistic.value.hex() == expected.hex()
+ compared += 1
+ assert compared > 0
+
+
+def test_total_and_sex_cells_equal_track_ms_cells(world, pieces):
+ """Differential: G3 against ``tabulate_track_m``, recomputed here."""
+
+ checks = mb.consistency_checks(world.reexecution, pieces.breakdowns)
+ assert checks and all(check.identical for check in checks)
+
+
+def test_shares_errors_and_floors_are_bounded(pieces):
+ for result in pieces.breakdowns.values():
+ for _, cell in _result_cells(result):
+ statistic = cell.statistic(g3.SHARE)
+ if not statistic.defined:
+ continue
+ assert 0.0 <= statistic.value <= 100.0
+ se = statistic.uncertainty.get("design_se")
+ if isinstance(se, dict):
+ se = se.get("value")
+ if se is not None:
+ assert se >= 0.0
+ floor = statistic.uncertainty.get("floor") or {}
+ for name, value in floor.items():
+ if isinstance(value, float) and name != "seeds":
+ assert value >= 0.0 or math.isnan(value)
+
+
+def test_the_document_carries_the_same_cells(world, pieces):
+ """The run's document holds the cells computed here, for every row."""
+
+ for (key, number), result in pieces.breakdowns.items():
+ document = result.as_dict()
+ recorded = world.document["breakdowns"][key][str(number)]
+ assert recorded["dimensions"] == common.json_normalized(
+ document["dimensions"]
+ )
+
+
+# =========================================================================
+# Attributes
+# =========================================================================
+def test_attributes_follow_the_rows(world, pieces):
+ frame = pieces.attributes.frame
+ rows = world.reexecution.evaluations["MS0"].rows
+ assert frame["person_id"].tolist() == rows["person_id"].tolist()
+ assert frame["weight"].tolist() == rows["weight"].tolist()
+ assert frame["sex"].tolist() == rows["sex"].tolist()
+ sensitivity = world.reexecution.sensitivity_evaluation.rows
+ assert sensitivity["person_id"].tolist() == rows["person_id"].tolist()
+
+
+def test_every_age_is_62_or_more_and_banded(pieces):
+ frame = pieces.attributes.frame
+ assert frame["age"].min() >= mb.YOUNGEST_AGE == 62
+ long = pieces.assignment.long
+ ages = long[long["dimension"] == "age"].sort_values("row_id")
+ for age, label in zip(frame["age"], ages["label"], strict=True):
+ band = (
+ "90 or older"
+ if age >= 90
+ else f"{age // 10 * 10}–{age // 10 * 10 + 9}"
+ )
+ assert label == band
+
+
+def test_codes_are_mint8s_or_declared_unclassified(pieces):
+ frame = pieces.attributes.frame
+ declared = pieces.attributes.unclassified_codes
+ for code in frame["benefit_type"]:
+ assert code in _BENEFIT_CODES or code in declared["benefit_type"]
+ for code in frame["marital_status"]:
+ assert code in common.MARITAL_STATUS_CODES or (
+ code in declared["marital_status"]
+ )
+ assert set(declared["marital_status"]) <= {
+ "unknown",
+ "no_marriage_history",
+ }
+
+
+def test_the_aime_identity_is_exercised(pieces):
+ record = pieces.attributes.provenance["initial_aime_at_62"]
+ assert record["convention"]["equals_mint8_initial_aime_convention"]
+ identity = record["identity_with_ms0_aime"]
+ assert identity["all_equal"] is True
+ assert identity["n_compared"] > 0
+
+
+def test_persons_without_an_own_record_have_no_aime(world, pieces):
+ frame = pieces.attributes.frame.set_index("person_id")
+ persons = world.reexecution.cohort.persons
+ without = [
+ str(pid)
+ for pid, own in zip(
+ persons["person_id"], persons["own_record_id"], strict=True
+ )
+ if own is None or pd.isna(own)
+ ]
+ assert frame.loc[without, "initial_aime_at_62"].isna().all()
+
+
+def test_quintiles_are_cut_within_each_birth_cohort(pieces):
+ frame = pieces.attributes.frame
+ long = pieces.assignment.long
+ for key in (
+ "initial_aime_quintile",
+ "lifetime_payroll_tax_quintile",
+ "lifetime_payroll_tax_quintile_shared",
+ ):
+ rows = long[long["dimension"] == key].sort_values("row_id")
+ values = frame[mb.COLUMNS_MAP[key]].to_numpy(dtype=float)
+ labels = rows["label"].to_numpy()
+ assert (np.isnan(values) == (labels == g3.UNCLASSIFIED)).all()
+ order = [c.label for c in mb.SCHEME.dimension(key).categories]
+ rank = {label: 5 - i for i, label in enumerate(order)}
+ for cohort_key in set(frame["birth_cohort_10y"]):
+ inside = (
+ frame["birth_cohort_10y"] == cohort_key
+ ).to_numpy() & ~np.isnan(values)
+ # A higher value never has a lower quintile in its cohort.
+ pairs = sorted(
+ zip(
+ values[inside],
+ [rank[x] for x in labels[inside]],
+ strict=True,
+ )
+ )
+ ranks = [r for _, r in pairs]
+ assert ranks == sorted(ranks)
+
+
+def test_attributes_are_deterministic(world, pieces):
+ again = mb.person_attributes(
+ world.reexecution,
+ world.params,
+ pieces.side_frame,
+ tax_rates=g2.load_oasdi_tax_rates(),
+ interest_rates=g2.load_trust_fund_interest_rates(),
+ )
+ pd.testing.assert_frame_equal(again.frame, pieces.attributes.frame)
+ assert common.compare_exact(
+ "provenance", again.provenance, pieces.attributes.provenance
+ ).identical
+
+
+def test_a_side_frame_for_other_persons_is_refused(world, pieces):
+ universe = [int(p) for p in world.reexecution.cohort.persons["person_id"]]
+ short = world.loader(universe[1:])
+ with pytest.raises(ValueError, match="exactly the universe"):
+ mb.person_attributes(
+ world.reexecution,
+ world.params,
+ short,
+ tax_rates=g2.load_oasdi_tax_rates(),
+ interest_rates=g2.load_trust_fund_interest_rates(),
+ )
+
+
+def test_side_frame_files_must_be_the_bytes_track_m_read(world, pieces):
+ files = {"IND2023ER.txt": "1" * 64}
+ reexecution = dataclasses.replace(
+ world.reexecution,
+ cohort_inputs=dataclasses.replace(
+ world.frames,
+ provenance={**world.frames.provenance, "psid_files_sha256": files},
+ ),
+ )
+ side = dataclasses.replace(
+ pieces.side_frame,
+ provenance={
+ **pieces.side_frame.provenance,
+ "inputs": {"psid_files_sha256": {"IND2023ER.txt": "2" * 64}},
+ },
+ )
+ with pytest.raises(ValueError, match="different bytes"):
+ mb.person_attributes(
+ reexecution,
+ world.params,
+ side,
+ tax_rates=g2.load_oasdi_tax_rates(),
+ interest_rates=g2.load_trust_fund_interest_rates(),
+ )
+
+
+def test_a_cohort_whose_2022_receipt_differs_is_refused(world):
+ built = world.reexecution.cohort
+ persons = built.persons.copy()
+ persons.loc[0, "types_2022"] = "other"
+ changed = dataclasses.replace(built, persons=persons)
+ with pytest.raises(ValueError, match="rebuilt from the anchor"):
+ mb._benefit_types(changed, world.frames, world.params.params)
+
+
+# =========================================================================
+# The order: nothing about groups before an exact reproduction
+# =========================================================================
+def _nudge(value: Any) -> Any:
+ if isinstance(value, bool):
+ return not value
+ if isinstance(value, int):
+ return value + 1
+ if isinstance(value, float):
+ return math.nextafter(value, math.inf)
+ if isinstance(value, str):
+ return value + "x"
+ return 0
+
+
+def _first_leaf(node: Any, kinds: tuple[type, ...]) -> list[Any] | None:
+ """The key path to the first leaf of one of ``kinds`` under ``node``."""
+
+ if isinstance(node, dict):
+ for key in node:
+ found = _first_leaf(node[key], kinds)
+ if found is not None:
+ return [key, *found]
+ return None
+ if isinstance(node, list):
+ for index, item in enumerate(node):
+ found = _first_leaf(item, kinds)
+ if found is not None:
+ return [index, *found]
+ return None
+ if (
+ isinstance(node, kinds)
+ and not isinstance(node, bool)
+ or (bool in kinds and isinstance(node, bool))
+ ):
+ return []
+ return None
+
+
+def _tampered(parent, prefix, kinds=(float,)):
+ document = copy.deepcopy(dict(parent.document))
+ node = document
+ for key in prefix:
+ node = node[key]
+ rest = _first_leaf(node, kinds)
+ assert rest is not None, prefix
+ path = [*prefix, *rest]
+ holder = document
+ for key in path[:-1]:
+ holder = holder[key]
+ holder[path[-1]] = _nudge(holder[path[-1]])
+ return common.CommittedArtifact(document=document), path
+
+
+def _cell(parent, row, key):
+ cells = parent.document["rows"][row]["tabulation"]["cells"]
+ index = next(i for i, c in enumerate(cells) if c["defined"])
+ return ["rows", row, "tabulation", "cells", index, key]
+
+
+TAMPERS: dict[str, Callable[[Any], tuple[list[Any], tuple[type, ...]]]] = {
+ "ms0_share": lambda p: (_cell(p, "MS0", "share_percent"), (float,)),
+ "ms4_design_se": lambda p: (_cell(p, "MS4", "design_se"), (float,)),
+ "ms6_weighted_n": lambda p: (_cell(p, "MS6", "weighted_n"), (float,)),
+ "ms2_unweighted_n": lambda p: (_cell(p, "MS2", "unweighted_n"), (int,)),
+ "ms1_floor": lambda p: (_cell(p, "MS1", "floor"), (float,)),
+ "ms3_diagnostics": lambda p: (["rows", "MS3", "diagnostics"], (int,)),
+ "cohort_structure": lambda p: (["cohort_structure"], (int,)),
+ "d430": lambda p: (["sensitivities", mb.SENSITIVITY_KEY], (float,)),
+ "d430_cohort_structure": lambda p: (
+ ["sensitivities", mb.SENSITIVITY_KEY, "cohort_structure"],
+ (int,),
+ ),
+ "labels": lambda p: (["labels"], (str,)),
+}
+
+
+@pytest.mark.parametrize("name", list(TAMPERS))
+def test_any_moved_committed_cell_refuses_before_any_group_work(
+ world, monkeypatch, name
+):
+ """The re-execution itself is the fixture's (deterministic, and
+ checked identical above); what is tested is that a parent differing
+ in one leaf stops the run before the side frame is requested."""
+
+ prefix, kinds = TAMPERS[name](world.parent)
+ bad, path = _tampered(world.parent, prefix, kinds)
+ monkeypatch.setattr(
+ mb, "reexecute_track_m", lambda *a, **k: world.reexecution
+ )
+ calls: list[Any] = []
+ with pytest.raises(common.ReproductionMismatchError) as error:
+ _run(world, bad, load_side_frame=world.spy(calls))
+ assert calls == []
+ message = str(error.value)
+ # The refusal names the differing block and path, never a value.
+ leaf = "$" + "".join(
+ f"[{key}]" if isinstance(key, int) else f".{key}" for key in path[1:]
+ )
+ assert f"pipeline.{path[0]}: 1 differing paths" in message
+ assert repr(leaf) in message
+ assert "refused before any group attribute or group cell" in message
+
+
+def test_one_ulp_refuses_through_a_real_reexecution(world):
+ """No stand-in: the run re-executes Track M on the INVENTED frames and
+ refuses a parent whose MS0 share moved by one unit in the last place,
+ before the side frame is requested."""
+
+ bad, path = _tampered(
+ world.parent, _cell(world.parent, "MS0", "share_percent")
+ )
+ calls: list[Any] = []
+ with pytest.raises(
+ common.ReproductionMismatchError, match="share_percent"
+ ):
+ _run(world, bad, load_side_frame=world.spy(calls))
+ assert calls == []
+
+
+def test_other_parameters_refuse_before_the_cohort_is_read(world):
+ document = copy.deepcopy(dict(world.parent.document))
+ document["parameters"]["oracle_pe_us_revision"] = "another"
+ reads: list[int] = []
+
+ def read():
+ reads.append(1)
+ return world.frames
+
+ with pytest.raises(common.ReproductionMismatchError, match="parameters"):
+ _run(
+ world,
+ common.CommittedArtifact(document=document),
+ load_cohort_inputs=read,
+ )
+ assert reads == []
+
+
+def test_other_psid_files_refuse_before_the_reexecution(world, monkeypatch):
+ document = copy.deepcopy(dict(world.parent.document))
+ document["inputs"]["source"]["psid_files_sha256"] = {"X.txt": "0" * 64}
+
+ def never(*args, **kwargs):
+ raise AssertionError("re-executed")
+
+ monkeypatch.setattr(mb, "reexecute_track_m", never)
+ with pytest.raises(common.ReproductionMismatchError, match="PSID files"):
+ _run(world, common.CommittedArtifact(document=document))
+
+
+@pytest.mark.parametrize(
+ ("overrides", "match"),
+ [
+ (
+ {"data_provenance": tabulation.REGISTERED_REAL},
+ "own issue #42 registration pointer",
+ ),
+ (
+ {
+ "data_provenance": tabulation.REGISTERED_REAL,
+ "registration_pointer": "https://example.org/42",
+ },
+ "own issue #42 registration pointer",
+ ),
+ (
+ {
+ "data_provenance": tabulation.REGISTERED_REAL,
+ "registration_pointer": POINTER,
+ },
+ "reads psid_files records",
+ ),
+ ({"provenance_kind": PSID_FILES}, "invented records"),
+ ({"data_provenance": "real"}, "unknown data_provenance"),
+ ],
+)
+def test_mismatched_provenance_refuses_before_anything_is_read(
+ world, overrides, match
+):
+ reads: list[int] = []
+ with pytest.raises(ValueError, match=match):
+ _run(
+ world,
+ load_cohort_inputs=lambda: reads.append(1),
+ **overrides,
+ )
+ assert reads == []
+
+
+def test_a_real_breakdown_needs_a_new_registration(world):
+ document = dict(world.parent.document)
+ parent = common.CommittedArtifact(
+ document={**document, "registration_pointer": POINTER}
+ )
+ with pytest.raises(ValueError, match="must not be the parent run's"):
+ _run(
+ world,
+ parent,
+ data_provenance=tabulation.REGISTERED_REAL,
+ registration_pointer=POINTER,
+ provenance_kind=PSID_FILES,
+ )
+
+
+# =========================================================================
+# The benefit type
+# =========================================================================
+_PARAMS = invented.invented_parameters()[0].params
+_ITEM = st.sampled_from([True, False, None])
+
+
+@settings(max_examples=400, deadline=None)
+@given(
+ st.fixed_dictionaries({name: _ITEM for name in ssr.SS_TYPES}),
+ st.booleans(),
+ st.sampled_from([None, cohort.BASIS_OLD_AGE, cohort.BASIS_DISABILITY]),
+ st.integers(min_value=1913, max_value=1960),
+ _ITEM,
+)
+def test_property_benefit_type(types, paid, basis, birth, other):
+ if paid and basis is None:
+ with pytest.raises(ValueError, match="own record"):
+ mb.benefit_type_2022(
+ types,
+ birth_year=birth,
+ paid_own_worker_benefit=paid,
+ own_record_basis=basis,
+ params=_PARAMS,
+ )
+ return
+
+ def code(items):
+ return mb.benefit_type_2022(
+ items,
+ birth_year=birth,
+ paid_own_worker_benefit=paid,
+ own_record_basis=basis,
+ params=_PARAMS,
+ )
+
+ result = code(types)
+ assert result in _BENEFIT_CODES or result.startswith("unclassified:")
+ survivor = types["survivor"] is True
+ dependent = (
+ types["dependent_of_disabled"] is True
+ or types["dependent_of_retired"] is True
+ )
+ if survivor and dependent:
+ assert result == "unclassified:survivor_and_dependent_mentioned"
+ elif survivor:
+ assert result == "widower"
+ elif dependent:
+ assert result == "spousal"
+ if result in ("retired_worker_only", "disabled_worker_only"):
+ # "only": no auxiliary type mentioned or unknown, own benefit paid
+ assert paid
+ assert all(types[name] is False for name in ssr.AUXILIARY_TYPES)
+ if result == "disabled_worker_only":
+ assert 12 * (2022 - birth) < _PARAMS.fra_months(birth)
+ # The 'other' item never decides the type.
+ assert code({**types, "other": other}) == result
+
+
+@pytest.mark.parametrize(
+ ("mentioned", "basis", "birth", "expected"),
+ [
+ ({"retirement"}, cohort.BASIS_OLD_AGE, 1950, "retired_worker_only"),
+ ({"retirement"}, cohort.BASIS_DISABILITY, 1958, "retired_worker_only"),
+ (
+ {"disability"},
+ cohort.BASIS_DISABILITY,
+ 1958,
+ "disabled_worker_only",
+ ),
+ # 2022 - 1950 = 72 is past the invented FRA: converted (MINT8)
+ ({"disability"}, cohort.BASIS_DISABILITY, 1950, "retired_worker_only"),
+ ({"disability"}, cohort.BASIS_OLD_AGE, 1958, "disabled_worker_only"),
+ # neither or both worker types: the record's basis decides
+ (set(), cohort.BASIS_DISABILITY, 1958, "disabled_worker_only"),
+ ({"other"}, cohort.BASIS_OLD_AGE, 1958, "retired_worker_only"),
+ (
+ {"retirement", "disability"},
+ cohort.BASIS_DISABILITY,
+ 1958,
+ "disabled_worker_only",
+ ),
+ ({"retirement", "survivor"}, cohort.BASIS_OLD_AGE, 1950, "widower"),
+ ({"dependent_of_retired"}, None, 1950, "spousal"),
+ ],
+)
+def test_benefit_type_examples(mentioned, basis, birth, expected):
+ types = {name: name in mentioned for name in ssr.SS_TYPES}
+ paid = basis is not None
+ assert (
+ mb.benefit_type_2022(
+ types,
+ birth_year=birth,
+ paid_own_worker_benefit=paid,
+ own_record_basis=basis,
+ params=_PARAMS,
+ )
+ == expected
+ )
+
+
+def test_an_unknown_auxiliary_item_leaves_only_unestablished():
+ types = {name: False for name in ssr.SS_TYPES}
+ types.update(retirement=True, survivor=None)
+ assert (
+ mb.benefit_type_2022(
+ types,
+ birth_year=1950,
+ paid_own_worker_benefit=True,
+ own_record_basis=cohort.BASIS_OLD_AGE,
+ params=_PARAMS,
+ )
+ == "unclassified:auxiliary_item_unknown"
+ )
+
+
+def test_every_combination_of_known_items_is_placed():
+ """Over all 64 known-item combinations, receipt and basis: the result
+ is a code, and unclassified only for the named reasons."""
+
+ reasons = set()
+ for values in itertools.product([True, False], repeat=len(ssr.SS_TYPES)):
+ types = dict(zip(ssr.SS_TYPES, values, strict=True))
+ for paid, basis in (
+ (True, cohort.BASIS_OLD_AGE),
+ (True, cohort.BASIS_DISABILITY),
+ (False, None),
+ ):
+ result = mb.benefit_type_2022(
+ types,
+ birth_year=1958,
+ paid_own_worker_benefit=paid,
+ own_record_basis=basis,
+ params=_PARAMS,
+ )
+ if result.startswith("unclassified:"):
+ reasons.add(result)
+ assert reasons == {
+ "unclassified:survivor_and_dependent_mentioned",
+ "unclassified:no_own_worker_benefit",
+ }
+
+
+# =========================================================================
+# Location fields: where a file was read is not what was read
+# =========================================================================
+_COLA_ELSEWHERE = "/elsewhere/checkout/" + mb.COLA_HISTORY_RELATIVE_PATH
+_COLA_HERE = "/here/worktree/" + mb.COLA_HISTORY_RELATIVE_PATH
+
+
+@pytest.mark.parametrize(
+ ("path", "expected"),
+ [
+ (_COLA_ELSEWHERE, mb.COLA_HISTORY_RELATIVE_PATH),
+ (mb.COLA_HISTORY_RELATIVE_PATH, mb.COLA_HISTORY_RELATIVE_PATH),
+ ("/elsewhere/data/external/other.json", None),
+ ("/elsewhere/xdata/external/ssa_cola_history.jsonx", None),
+ (None, None),
+ (3, None),
+ ],
+)
+def test_relative_location(path, expected):
+ assert mb.relative_location(path) == (
+ path if expected is None else expected
+ )
+
+
+def _with_cola(document: dict, record: dict) -> dict:
+ out = copy.deepcopy(document)
+ out.setdefault("inputs", {}).setdefault("source", {})[
+ "cola_history"
+ ] = record
+ return out
+
+
+def test_located_changes_only_the_location_field():
+ document = _with_cola({"x": 1.5}, {"path": _COLA_ELSEWHERE, "sha256": "a"})
+ before = copy.deepcopy(document)
+ out = mb.located(document)
+ assert document == before
+ assert out["inputs"]["source"]["cola_history"] == {
+ "path": mb.COLA_HISTORY_RELATIVE_PATH,
+ "sha256": "a",
+ }
+ assert out["x"] == 1.5
+ assert mb.located({"x": 1}) == {"x": 1}
+
+
+@pytest.mark.parametrize(
+ ("recomputed", "identical"),
+ [
+ ({"path": _COLA_HERE, "sha256": "a"}, True),
+ ({"path": _COLA_HERE, "sha256": "b"}, False),
+ ({"path": "/here/data/external/another.json", "sha256": "a"}, False),
+ ],
+)
+def test_the_cola_history_is_compared_by_bytes_not_location(
+ world, recomputed, identical
+):
+ committed = _with_cola(
+ dict(world.parent.document), {"path": _COLA_ELSEWHERE, "sha256": "a"}
+ )
+ reexecution = dataclasses.replace(
+ world.reexecution,
+ result=_with_cola(dict(world.reexecution.result), recomputed),
+ )
+ checks = {
+ c.name: c for c in mb.reproduction_checks(reexecution, committed)
+ }
+ assert checks["pipeline.inputs"].identical is identical
+ others = [c for name, c in checks.items() if name != "pipeline.inputs"]
+ assert all(c.identical for c in others)
+
+
+def test_the_document_records_the_location_rule(world):
+ record = world.document["reproduction"]["location_fields"]
+ assert record["fields"] == ["$.inputs.source.cola_history.path"]
+ assert record["rule"] == mb.LOCATION_RULE
diff --git a/tests/group_breakdowns/test_min_benefit_parent_artifact.py b/tests/group_breakdowns/test_min_benefit_parent_artifact.py
new file mode 100644
index 00000000..56c8b6a2
--- /dev/null
+++ b/tests/group_breakdowns/test_min_benefit_parent_artifact.py
@@ -0,0 +1,181 @@
+"""The committed exercise-4 run is the parent the breakdown can reproduce.
+
+Artifact tier: reads the committed ``runs/replication_urban2006_minimum_
+benefit_v1.json`` and its sidecar for their bindings, keys and
+environment only. No outcome value is read into an assertion or printed,
+no PSID file is opened and nothing is computed on real data.
+
+Checked here, before any host run: the parent's bytes are the pinned
+ones and its sidecar binds them; it is Registration 17's run on the M1
+specification at this commit; every block the adapter compares exists;
+the parameter pins the entry script checks are the specification's; the
+COLA history on disk is the file the parent recorded (bytes, not
+location); the environment fields the entry script compares are recorded;
+and ``scripts/make_nasi_repro_venv.sh`` pins exactly those versions.
+"""
+
+from __future__ import annotations
+
+import importlib.util
+import json
+import re
+from pathlib import Path
+
+import pytest
+
+from populace_dynamics.estimates.parameters import load_cola_history
+from populace_dynamics.group_breakdowns import common
+from populace_dynamics.group_breakdowns import min_benefit as mb
+from populace_dynamics.min_benefit_track_m.policy import REGISTERED_ROWS
+from populace_dynamics.min_benefit_track_m.specification import (
+ M1_SPECIFICATION_PATH,
+ m1_parameter_block,
+)
+
+ROOT = Path(__file__).resolve().parents[2]
+PARENT = ROOT / "runs" / "replication_urban2006_minimum_benefit_v1.json"
+VENV_SCRIPT = ROOT / "scripts" / "make_nasi_repro_venv.sh"
+
+
+def _script():
+ path = ROOT / "scripts" / "run_min_benefit_groups_registered.py"
+ spec = importlib.util.spec_from_file_location("_g4c_run_artifact", path)
+ module = importlib.util.module_from_spec(spec)
+ spec.loader.exec_module(module)
+ return module
+
+
+@pytest.fixture(scope="module")
+def parent() -> common.CommittedArtifact:
+ assert mb.PARENT_ARTIFACT_PATH == PARENT
+ return mb.load_parent_artifact()
+
+
+def test_the_parent_bytes_are_pinned_and_bound(parent):
+ assert parent.sha256 == mb.PARENT_ARTIFACT_SHA256
+ assert parent.sidecar["artifact"] == PARENT.name
+ assert parent.sidecar["artifact_sha256"] == mb.PARENT_ARTIFACT_SHA256
+
+
+def test_the_parent_is_registration_17s_run_on_this_specification(parent):
+ binding = mb.check_parent_binding(
+ parent,
+ specification_sha256=common.file_sha256(M1_SPECIFICATION_PATH),
+ )
+ assert binding["registration_pointer"] == mb.PARENT_REGISTRATION_POINTER
+ assert binding["specification_version"] == "m1-ratified-1"
+
+
+def test_every_block_the_adapter_compares_exists(parent):
+ document = parent.document
+ for key in (
+ "parameters",
+ "inputs",
+ "rows",
+ "sensitivities",
+ "cohort_structure",
+ "labels",
+ "data_provenance",
+ "registration_pointer",
+ ):
+ assert key in document, key
+ assert list(document["rows"]) == list(REGISTERED_ROWS)
+ for row in REGISTERED_ROWS:
+ assert {"tabulation", "diagnostics"} <= set(document["rows"][row])
+ sensitivity = document["sensitivities"][mb.SENSITIVITY_KEY]
+ assert {"tabulation", "diagnostics", "cohort_structure"} <= set(
+ sensitivity
+ )
+ files = document["inputs"]["source"]["psid_files_sha256"]
+ assert files and all(
+ re.fullmatch(r"[0-9a-f]{64}", digest) for digest in files.values()
+ )
+
+
+def test_the_parameter_pins_are_the_specifications(parent):
+ sources = m1_parameter_block()["sources"]
+ assert parent.document["checks"]["parameter_pins"] == {
+ "quarter_of_coverage": sources["quarter_of_coverage_amounts"][
+ "sha256"
+ ],
+ "census_thresholds": sources["census_thresholds"]["sha256"],
+ }
+
+
+def test_the_only_location_field_is_the_cola_history_path(parent):
+ """No other string in the parent names an absolute path."""
+
+ def walk(node, path="$"):
+ if isinstance(node, dict):
+ for key, value in node.items():
+ yield from walk(value, f"{path}.{key}")
+ elif isinstance(node, list):
+ for index, value in enumerate(node):
+ yield from walk(value, f"{path}[{index}]")
+ elif isinstance(node, str) and node.startswith("/"):
+ yield path
+
+ assert list(walk(parent.document)) == [
+ "$." + ".".join(keys) for keys in mb.LOCATION_FIELDS
+ ]
+
+
+def test_the_cola_history_on_disk_is_the_parents(parent):
+ """The re-execution will record this checkout's COLA history; by the
+ location rule it must equal the parent's record exactly."""
+
+ recorded = parent.document["inputs"]["source"]["cola_history"]
+ here = dict(load_cola_history().provenance)
+ check = common.compare_exact(
+ "cola_history",
+ mb.located({"inputs": {"source": {"cola_history": recorded}}}),
+ mb.located({"inputs": {"source": {"cola_history": here}}}),
+ )
+ assert check.identical, check.differences
+
+
+def test_the_environment_fields_are_recorded(parent):
+ script = _script()
+ environment = parent.sidecar["environment"]
+ for field in script.ENVIRONMENT_FIELDS:
+ value = environment
+ for key in field.split("."):
+ value = value[key]
+ assert isinstance(value, str) and value, field
+
+
+def _pins() -> dict[str, str]:
+ text = VENV_SCRIPT.read_text(encoding="utf-8")
+ found = dict(
+ re.findall(r'^([A-Z_]+_(?:VERSION|REVISION))="([^"]+)"', text, re.M)
+ )
+ return found
+
+
+def test_the_repro_venv_script_pins_the_parents_environment(parent):
+ environment = parent.sidecar["environment"]
+ assert _pins() == {
+ "PYTHON_VERSION": environment["python"],
+ "NUMPY_VERSION": environment["packages"]["numpy"],
+ "PANDAS_VERSION": environment["packages"]["pandas"],
+ "SCIPY_VERSION": environment["packages"]["scipy"],
+ "PE_US_REVISION": environment["policyengine_us_parameters"][
+ "revision"
+ ],
+ }
+
+
+def test_the_repro_venv_script_asks_for_the_gil_build():
+ text = VENV_SCRIPT.read_text(encoding="utf-8")
+ assert '--python "${PYTHON_VERSION}+gil"' in text
+ assert "Py_GIL_DISABLED" in text
+ assert "runs/replication_urban2006_minimum_benefit_v1.env.json" in text
+
+
+def test_the_parent_names_the_files_both_readers_record(parent):
+ """The side frame's reader and Track M's must agree on shared bytes;
+ the parent records Track M's by repository-relative path."""
+
+ files = parent.document["inputs"]["source"]["psid_files_sha256"]
+ assert all(not name.startswith("/") for name in files)
+ json.dumps(files)
diff --git a/tests/group_breakdowns/test_min_benefit_registered_script.py b/tests/group_breakdowns/test_min_benefit_registered_script.py
new file mode 100644
index 00000000..4ed5d088
--- /dev/null
+++ b/tests/group_breakdowns/test_min_benefit_registered_script.py
@@ -0,0 +1,567 @@
+"""The exercise-4 group-breakdown entry point refuses every other state.
+
+INVENTED DATA - NOT A COMPARISON. Unit tier: no PSID file, committed
+artifact or policyengine-us checkout is read. The "committed parent" is
+written to a temporary directory: the registered computation run once on
+``min_benefit_track_m.invented_psid``'s frames, marked as read from files
+(an INVENTED file name and hash, as ``tests/min_benefit_track_m/
+test_registered_run_script.py`` marks them), under Registration 17's
+pointer, with the bindings the preflight checks. ``git``, the Track M
+parameter loader, the environment resolver and the COLA history are
+stand-ins; every PSID loader is INVENTED.
+
+Stated and tested:
+
+* the preflight refuses a pointer that is not an issue #42 comment,
+ Registration 17's own pointer, a short commit, another ``HEAD``, a
+ dirty tree, an existing output or sidecar, other parent bytes, and a
+ parent that is not Registration 17's run on the M1 specification of
+ this commit;
+* before any PSID read: other parameter pins and another environment
+ (Python, numpy, pandas, scipy, the policyengine-us revision) refuse;
+* a parent with one committed cell moved by one unit in the last place
+ refuses after the re-execution, before the group-attribute loader is
+ called, and nothing is written;
+* the registered state writes the artifact and its sidecar exclusively,
+ with Track M's labels, the post hoc labels and the reproduction record.
+"""
+
+from __future__ import annotations
+
+import copy
+import dataclasses
+import importlib.util
+import json
+import math
+import re
+from pathlib import Path
+from types import SimpleNamespace
+
+import pytest
+from hypothesis import given, settings
+from hypothesis import strategies as st
+
+from populace_dynamics.group_breakdowns import common
+from populace_dynamics.group_breakdowns import min_benefit as mb
+from populace_dynamics.min_benefit_track_m import (
+ COVERED_EARNINGS_DISCLOSURE,
+ OUTPUT_LABELS,
+ invented,
+ invented_psid,
+)
+from populace_dynamics.min_benefit_track_m.evaluation import PSID_FILES
+from populace_dynamics.min_benefit_track_m.tabulation import REGISTERED_REAL
+
+ROOT = Path(__file__).resolve().parents[2]
+POINTER = (
+ "https://github.com/PolicyEngine/microcosm-dynamics/issues/42"
+ "#issuecomment-9"
+)
+COMMIT = "c" * 40
+PINS = {"pinned": "INVENTED"}
+#: An INVENTED environment, as Track A's resolver records one.
+ENVIRONMENT = {
+ "python": "3.0.0-INVENTED",
+ "platform": "INVENTED",
+ "packages": {"numpy": "0-INVENTED", "pandas": "0", "scipy": "0"},
+ "policyengine_us_parameters": {"revision": "INVENTED"},
+}
+
+
+def _script():
+ path = ROOT / "scripts" / "run_min_benefit_groups_registered.py"
+ spec = importlib.util.spec_from_file_location("_g4c_run", path)
+ module = importlib.util.module_from_spec(spec)
+ spec.loader.exec_module(module)
+ return module
+
+
+SCRIPT = _script()
+
+
+def _git(head: str = COMMIT, porcelain: str = ""):
+ def fake(*args: str) -> str:
+ if args == ("rev-parse", "HEAD"):
+ return head
+ if args == ("status", "--porcelain"):
+ return porcelain
+ raise AssertionError(args)
+
+ return fake
+
+
+class _Cola(dict):
+ provenance = {"kind": "INVENTED"}
+
+
+@pytest.fixture(scope="module")
+def stand_ins():
+ params, cola = invented.invented_parameters()
+ frames = invented_psid.invented_cohort_inputs(seed=5, n_family_units=30)
+ marked = dataclasses.replace(
+ frames,
+ provenance={
+ **frames.provenance,
+ "psid_files_sha256": {"INVENTED.txt": "0" * 64},
+ },
+ )
+ cola = _Cola(cola)
+ run = mb.reexecute_track_m(
+ marked,
+ params,
+ cola_rates=cola,
+ data_provenance=REGISTERED_REAL,
+ registration_pointer=mb.PARENT_REGISTRATION_POINTER,
+ provenance_kind=PSID_FILES,
+ source={"cola_history": dict(cola.provenance)},
+ )
+ document = common.json_normalized(dict(run.result))
+ document.update(
+ specification={
+ "sha256": mb.PARENT_SPECIFICATION_SHA256,
+ "version": "m1-ratified-1",
+ },
+ checks={"parameter_pins": dict(PINS)},
+ run={"registered_commit": mb.PARENT_REGISTERED_COMMIT},
+ )
+ return SimpleNamespace(
+ params=params, cola=cola, frames=marked, document=document
+ )
+
+
+def _write_parent(directory: Path, document: dict, environment=None):
+ path = directory / "parent_v1.json"
+ path.write_text(json.dumps(document, indent=2) + "\n")
+ digest = common.file_sha256(path)
+ path.with_suffix(".env.json").write_text(
+ json.dumps(
+ {
+ "artifact": path.name,
+ "artifact_sha256": digest,
+ "environment": environment or ENVIRONMENT,
+ }
+ )
+ )
+ return path, digest
+
+
+def _minimal_parent(**changes) -> dict:
+ document = {
+ "registration_pointer": mb.PARENT_REGISTRATION_POINTER,
+ "specification": {"sha256": mb.PARENT_SPECIFICATION_SHA256},
+ "run": {"registered_commit": mb.PARENT_REGISTERED_COMMIT},
+ }
+ document.update(changes)
+ return document
+
+
+def _preflight(tmp_path, document=None, **kwargs):
+ path, digest = _write_parent(tmp_path, document or _minimal_parent())
+ arguments = dict(
+ registration_pointer=POINTER,
+ registered_commit=COMMIT,
+ output=tmp_path / "out" / "groups_posthoc_v1.json",
+ git=_git(),
+ parent_path=path,
+ parent_sha256=digest,
+ )
+ arguments.update(kwargs)
+ return SCRIPT.preflight(**arguments)
+
+
+# =========================================================================
+# Preflight
+# =========================================================================
+def test_the_registered_state_passes_preflight(tmp_path):
+ state = _preflight(tmp_path)
+ assert state["head"] == COMMIT
+ assert state["binding"]["registered_commit"] == mb.PARENT_REGISTERED_COMMIT
+ assert not (tmp_path / "out").exists()
+
+
+@pytest.mark.parametrize(
+ ("kwargs", "match"),
+ [
+ ({"registration_pointer": "https://example.org/x"}, "issue #42"),
+ (
+ {
+ "registration_pointer": (
+ "https://github.com/PolicyEngine/microcosm-dynamics/"
+ "issues/420#issuecomment-1"
+ )
+ },
+ "issue #42",
+ ),
+ (
+ {
+ "registration_pointer": (
+ "https://github.com/PolicyEngine/microcosm-dynamics/"
+ "issues/42"
+ )
+ },
+ "issue #42",
+ ),
+ ({"registration_pointer": None}, "issue #42"),
+ (
+ {"registration_pointer": mb.PARENT_REGISTRATION_POINTER},
+ "not Registration 17's",
+ ),
+ ({"registered_commit": "abc123"}, "full 40-hex"),
+ ({"git": _git(head="d" * 40)}, "is not the registered commit"),
+ ({"git": _git(porcelain=" M src/x.py")}, "clean"),
+ ({"parent_sha256": "0" * 64}, "not the committed"),
+ ],
+)
+def test_non_registered_states_are_refused(tmp_path, kwargs, match):
+ with pytest.raises(ValueError, match=match):
+ _preflight(tmp_path, **kwargs)
+ assert not (tmp_path / "out").exists()
+
+
+@pytest.mark.parametrize("suffix", [".json", ".env.json"])
+def test_an_existing_output_or_sidecar_is_refused(tmp_path, suffix):
+ output = tmp_path / "out" / "groups_posthoc_v1.json"
+ output.parent.mkdir()
+ existing = output.with_suffix(suffix)
+ existing.write_text("{}")
+ with pytest.raises(ValueError, match="one-shot"):
+ _preflight(tmp_path, output=output)
+ assert existing.read_text() == "{}"
+
+
+@pytest.mark.parametrize(
+ ("changes", "match"),
+ [
+ (
+ {"registration_pointer": POINTER},
+ "not Registration 17's artifact",
+ ),
+ ({"run": {"registered_commit": "e" * 40}}, "not 2e4e08be"),
+ (
+ {"specification": {"sha256": "f" * 64}},
+ "another M1 specification",
+ ),
+ ],
+)
+def test_a_parent_other_than_registration_17s_is_refused(
+ tmp_path, changes, match
+):
+ with pytest.raises(ValueError, match=match):
+ _preflight(tmp_path, _minimal_parent(**changes))
+
+
+def test_a_changed_m1_specification_is_refused(tmp_path):
+ other = tmp_path / "other_specification.md"
+ other.write_text("not the M1 specification\n")
+ with pytest.raises(ValueError, match="changed since Registration 17"):
+ _preflight(tmp_path, specification_path=other)
+
+
+@settings(max_examples=60, deadline=None)
+@given(st.text(max_size=120))
+def test_property_any_other_pointer_is_refused(pointer):
+ if re.fullmatch(
+ r"https://github\.com/PolicyEngine/microcosm-dynamics/issues/42"
+ r"#issuecomment-\d+",
+ pointer,
+ ):
+ return
+ with pytest.raises(ValueError, match="issue #42"):
+ SCRIPT.preflight(
+ registration_pointer=pointer,
+ registered_commit=COMMIT,
+ output=Path("never-written-by-the-preflight.json"),
+ git=_git(),
+ )
+
+
+@settings(max_examples=60, deadline=None)
+@given(st.text(alphabet="0123456789abcdef", min_size=40, max_size=40))
+def test_property_any_other_head_is_refused(head):
+ if head == COMMIT:
+ return
+ with pytest.raises(ValueError, match="not the registered commit"):
+ SCRIPT.preflight(
+ registration_pointer=POINTER,
+ registered_commit=COMMIT,
+ output=Path("never-written-by-the-preflight.json"),
+ git=_git(head=head),
+ )
+
+
+@settings(max_examples=40, deadline=None)
+@given(st.text(min_size=1, max_size=40).filter(lambda t: t != ""))
+def test_property_any_dirty_tree_is_refused(porcelain):
+ with pytest.raises(ValueError, match="clean"):
+ SCRIPT.preflight(
+ registration_pointer=POINTER,
+ registered_commit=COMMIT,
+ output=Path("never-written-by-the-preflight.json"),
+ git=_git(porcelain=porcelain),
+ )
+
+
+# =========================================================================
+# Pre-PSID checks
+# =========================================================================
+def test_the_environment_fields_are_the_reproductions():
+ assert SCRIPT.ENVIRONMENT_FIELDS == (
+ "python",
+ "packages.numpy",
+ "packages.pandas",
+ "packages.scipy",
+ "policyengine_us_parameters.revision",
+ )
+
+
+def test_the_parent_environment_passes():
+ record = SCRIPT.check_environment(
+ copy.deepcopy(ENVIRONMENT), {"environment": ENVIRONMENT}
+ )
+ assert record["matches_parent"] is True
+ assert set(record["fields"]) == set(SCRIPT.ENVIRONMENT_FIELDS)
+ assert isinstance(record["gil_disabled_this_run"], bool)
+
+
+def _with(environment: dict, dotted: str, value) -> dict:
+ out = copy.deepcopy(environment)
+ node = out
+ keys = dotted.split(".")
+ for key in keys[:-1]:
+ node = node[key]
+ node[keys[-1]] = value
+ return out
+
+
+@pytest.mark.parametrize("field", SCRIPT.ENVIRONMENT_FIELDS)
+def test_any_other_environment_field_is_refused(field):
+ with pytest.raises(ValueError, match=re.escape(field)):
+ SCRIPT.check_environment(
+ _with(ENVIRONMENT, field, "other"), {"environment": ENVIRONMENT}
+ )
+ with pytest.raises(ValueError, match=re.escape(field)):
+ SCRIPT.check_environment(
+ ENVIRONMENT, {"environment": _with(ENVIRONMENT, field, None)}
+ )
+
+
+def test_the_platform_is_recorded_not_compared():
+ record = SCRIPT.check_environment(
+ {**ENVIRONMENT, "platform": "elsewhere"}, {"environment": ENVIRONMENT}
+ )
+ assert record["platform_this_run"] == "elsewhere"
+
+
+def test_other_parameter_pins_are_refused():
+ parent = common.CommittedArtifact(
+ document={"checks": {"parameter_pins": dict(PINS)}}
+ )
+ assert SCRIPT.check_parameter_pins(dict(PINS), parent) == PINS
+ with pytest.raises(ValueError, match="parameter_pins"):
+ SCRIPT.check_parameter_pins({"pinned": "other"}, parent)
+
+
+def test_every_composed_file_exists_and_is_hashed():
+ paths = SCRIPT.composed_paths()
+ assert "scripts/run_min_benefit_groups_registered.py" in paths
+ assert "src/populace_dynamics/group_breakdowns/min_benefit.py" in paths
+ assert "src/populace_dynamics/group_breakdowns/common.py" in paths
+ assert "src/populace_dynamics/min_benefit_track_m/cohort.py" in paths
+ assert "src/populace_dynamics/ss/statutory_aime.py" in paths
+ digests = SCRIPT.source_sha256()
+ assert set(digests) == set(paths)
+ assert all(re.fullmatch(r"[0-9a-f]{64}", d) for d in digests.values())
+
+
+def test_the_header_carries_every_label_and_the_disclosure():
+ for label in (*OUTPUT_LABELS, *common.POST_HOC_LABELS):
+ assert label in SCRIPT.REGISTERED_HEADER
+ assert COVERED_EARNINGS_DISCLOSURE in SCRIPT.REGISTERED_HEADER
+ assert "Publishes regardless of outcome" in SCRIPT.REGISTERED_HEADER
+
+
+def test_the_default_output_is_a_new_artifact_beside_the_parent():
+ assert SCRIPT.DEFAULT_OUTPUT == mb.OUTPUT_ARTIFACT_PATH
+ assert mb.OUTPUT_ARTIFACT_PATH.parent == mb.PARENT_ARTIFACT_PATH.parent
+ assert mb.OUTPUT_ARTIFACT_PATH.name == (
+ mb.PARENT_ARTIFACT_PATH.stem.removesuffix("_v1")
+ + "_groups_posthoc_v1.json"
+ )
+
+
+def test_main_refuses_without_a_pointer(capsys):
+ with pytest.raises(SystemExit):
+ SCRIPT.main(["--registered-commit", COMMIT])
+ assert "--registration-pointer" in capsys.readouterr().err
+
+
+# =========================================================================
+# execute(): the order and the writes
+# =========================================================================
+class _Spy:
+ def __init__(self, inner=None):
+ self.calls: list = []
+ self.inner = inner
+
+ def __call__(self, *args):
+ self.calls.append(args)
+ if self.inner is None:
+ raise AssertionError("called")
+ return self.inner(*args)
+
+
+def _execute(tmp_path, stand_ins, document, **overrides):
+ path, digest = _write_parent(tmp_path, document)
+ output = tmp_path / "out" / "groups_posthoc_v1.json"
+ arguments = dict(
+ registration_pointer=POINTER,
+ registered_commit=COMMIT,
+ output=output,
+ argv=["--registration-pointer", POINTER],
+ git=_git(),
+ track_m_script=SimpleNamespace(
+ check_runnable=lambda: None,
+ committed_parameters=lambda block: (stand_ins.params, dict(PINS)),
+ ),
+ environment=lambda **_: copy.deepcopy(ENVIRONMENT),
+ load_cola=lambda: stand_ins.cola,
+ load_cohort_inputs=lambda: stand_ins.frames,
+ load_side_frame=mb.invented_group_attribute_loader(
+ stand_ins.frames, seed=1
+ ),
+ parent_path=path,
+ parent_sha256=digest,
+ )
+ arguments.update(overrides)
+ return output, SCRIPT.execute(**arguments)
+
+
+def test_a_moved_cell_refuses_before_any_group_attribute(tmp_path, stand_ins):
+ """The orchestrator's order test: a parent whose MS0 share moved by
+ one unit in the last place is refused after a real re-execution of
+ the registered computation, before the group-attribute loader is
+ called, and neither the artifact nor its sidecar is written."""
+
+ document = copy.deepcopy(stand_ins.document)
+ cells = document["rows"]["MS0"]["tabulation"]["cells"]
+ index = next(i for i, cell in enumerate(cells) if cell["defined"])
+ cells[index]["share_percent"] = math.nextafter(
+ cells[index]["share_percent"], math.inf
+ )
+ side_frames = _Spy()
+ reads = _Spy(lambda: stand_ins.frames)
+ with pytest.raises(common.ReproductionMismatchError) as error:
+ _execute(
+ tmp_path,
+ stand_ins,
+ document,
+ load_side_frame=side_frames,
+ load_cohort_inputs=reads,
+ )
+ assert len(reads.calls) == 1
+ assert side_frames.calls == []
+ assert f"tabulation.cells[{index}].share_percent" in str(error.value)
+ assert not (tmp_path / "out").exists()
+
+
+def test_another_environment_refuses_before_the_cohort_is_read(
+ tmp_path, stand_ins
+):
+ reads = _Spy()
+ with pytest.raises(ValueError, match="packages.numpy"):
+ _execute(
+ tmp_path,
+ stand_ins,
+ stand_ins.document,
+ environment=lambda **_: _with(
+ ENVIRONMENT, "packages.numpy", "9.9.9"
+ ),
+ load_cohort_inputs=reads,
+ )
+ assert reads.calls == []
+ assert not (tmp_path / "out").exists()
+
+
+def test_other_parameter_pins_refuse_before_the_cohort_is_read(
+ tmp_path, stand_ins
+):
+ reads = _Spy()
+ with pytest.raises(ValueError, match="parameter_pins"):
+ _execute(
+ tmp_path,
+ stand_ins,
+ stand_ins.document,
+ track_m_script=SimpleNamespace(
+ check_runnable=lambda: None,
+ committed_parameters=lambda block: (
+ stand_ins.params,
+ {"pinned": "other"},
+ ),
+ ),
+ load_cohort_inputs=reads,
+ )
+ assert reads.calls == []
+
+
+def test_a_missing_component_refuses_before_the_parameters(
+ tmp_path, stand_ins
+):
+ def missing():
+ raise RuntimeError("the Track M pipeline is not built")
+
+ def never(block):
+ raise AssertionError("parameters read")
+
+ with pytest.raises(RuntimeError, match="not built"):
+ _execute(
+ tmp_path,
+ stand_ins,
+ stand_ins.document,
+ track_m_script=SimpleNamespace(
+ check_runnable=missing, committed_parameters=never
+ ),
+ )
+
+
+def test_the_registered_state_writes_the_artifact_and_sidecar(
+ tmp_path, stand_ins
+):
+ """INVENTED stand-ins through the whole registered path: the artifact
+ and its sidecar are created exclusively and carry every label."""
+
+ output, artifact = _execute(tmp_path, stand_ins, stand_ins.document)
+ written = json.loads(output.read_text())
+ assert written == common.json_normalized(artifact)
+ sidecar = json.loads(output.with_suffix(".env.json").read_text())
+ assert sidecar["artifact"] == output.name
+ assert sidecar["artifact_sha256"] == common.file_sha256(output)
+ assert sidecar["environment"] == ENVIRONMENT
+ assert next(iter(written)) == "header"
+ assert written["header"] == SCRIPT.REGISTERED_HEADER
+ assert written["data_provenance"] == REGISTERED_REAL
+ assert written["registration_pointer"] == POINTER
+ assert written["registration"]["pointer"] == POINTER
+ assert written["labels"] == [*OUTPUT_LABELS, *common.POST_HOC_LABELS]
+ assert written["breakdown"]["registration_pointer"] == POINTER
+ assert written["reads_comparator_values"] is False
+ assert written["publishes_regardless"] is True
+ assert written["blind"] is False and written["scored"] is False
+ assert written["reproduction"]["identical"] is True
+ assert written["reproduction"][
+ "reexecuted_under_registration_pointer"
+ ] == (mb.PARENT_REGISTRATION_POINTER)
+ assert written["consistency_with_registered_cells"]["identical"] is True
+ assert written["checks"]["parameter_pins"] == PINS
+ assert written["checks"]["environment"]["matches_parent"] is True
+ assert written["run"]["registered_commit"] == COMMIT
+ assert written["run"]["git_clean"] is True
+ assert set(written["code_sha256"]) == set(SCRIPT.composed_paths())
+ assert written["parent"]["registered_commit"] == (
+ mb.PARENT_REGISTERED_COMMIT
+ )
+ # One shot: a second run on the same output is refused by the
+ # preflight, and the first artifact is unchanged.
+ before = output.read_bytes()
+ with pytest.raises(ValueError, match="one-shot"):
+ _execute(tmp_path, stand_ins, stand_ins.document)
+ assert output.read_bytes() == before
diff --git a/tests/group_breakdowns/test_min_benefit_reproduction_psid.py b/tests/group_breakdowns/test_min_benefit_reproduction_psid.py
new file mode 100644
index 00000000..b0838ff7
--- /dev/null
+++ b/tests/group_breakdowns/test_min_benefit_reproduction_psid.py
@@ -0,0 +1,113 @@
+"""Host-only: the committed exercise-4 run reproduces exactly (NASI G4c).
+
+Integration tier (staged PSID under ``POPULACE_DYNAMICS_PSID_DIR`` or
+``~/PolicyEngine/psid-data``). Re-executes Registration 17's registered
+computation on the staged PSID exactly as the registered runner did
+(:func:`min_benefit.reproduce_parent`) and asserts that every recomputed
+block equals the committed artifact (floats bit for bit; the COLA
+history's location compared by its repository-relative path). That is
+the precondition of any group breakdown, checked here on its own: no
+side frame is loaded, no group attribute or group cell is computed, and
+nothing is written or printed (a mismatch names paths only). It
+recomputes the committed, published cells; it produces no new outcome.
+
+It skips unless all three hold: the staged PSID files exist; the oracle's
+policyengine-us checkout is at the parent's revision (``a03e82e503``);
+and Python, numpy, pandas and scipy are the versions the parent's sidecar
+records (build them with ``scripts/make_nasi_repro_venv.sh``).
+"""
+
+from __future__ import annotations
+
+import importlib.metadata
+import importlib.util
+import os
+import platform
+from pathlib import Path
+
+import pytest
+
+from populace_dynamics.group_breakdowns import common
+from populace_dynamics.group_breakdowns import min_benefit as mb
+
+ROOT = Path(__file__).resolve().parents[2]
+PSID_DIR = Path(
+ os.environ.get("POPULACE_DYNAMICS_PSID_DIR", "~/PolicyEngine/psid-data")
+).expanduser()
+
+
+def _environment_matches(sidecar: dict) -> list[str]:
+ recorded = sidecar["environment"]
+ here = {
+ "python": platform.python_version(),
+ **{
+ name: importlib.metadata.version(name)
+ for name in ("numpy", "pandas", "scipy")
+ },
+ }
+ wanted = {
+ "python": recorded["python"],
+ **{
+ name: recorded["packages"][name]
+ for name in ("numpy", "pandas", "scipy")
+ },
+ }
+ return [name for name in wanted if wanted[name] != here[name]]
+
+
+def _track_m_script():
+ path = ROOT / "scripts" / "run_track_m_registered.py"
+ spec = importlib.util.spec_from_file_location("_g4c_track_m", path)
+ module = importlib.util.module_from_spec(spec)
+ spec.loader.exec_module(module)
+ return module
+
+
+def test_the_registered_computation_reproduces_exactly():
+ if not PSID_DIR.is_dir():
+ pytest.skip("staged PSID files absent")
+ parent = mb.load_parent_artifact()
+ differ = _environment_matches(dict(parent.sidecar))
+ if differ:
+ pytest.skip(
+ f"environment differs from the parent run's in {differ} "
+ "(scripts/make_nasi_repro_venv.sh builds it)"
+ )
+ from populace_dynamics.estimates.parameters import load_cola_history
+ from populace_dynamics.min_benefit_track_m import cohort
+ from populace_dynamics.min_benefit_track_m.evaluation import PSID_FILES
+ from populace_dynamics.min_benefit_track_m.specification import (
+ m1_parameter_block,
+ )
+ from populace_dynamics.min_benefit_track_m.tabulation import (
+ REGISTERED_REAL,
+ )
+
+ track_m = _track_m_script()
+ try:
+ parameters, pins = track_m.committed_parameters(m1_parameter_block())
+ except FileNotFoundError:
+ pytest.skip("the oracle's policyengine-us checkout is absent")
+ revision = parent.sidecar["environment"]["policyengine_us_parameters"][
+ "revision"
+ ]
+ if parameters.params.pe_us_revision != revision:
+ pytest.skip(
+ f"policyengine-us is at {parameters.params.pe_us_revision}, not "
+ f"the parent's {revision}"
+ )
+ assert pins == parent.document["checks"]["parameter_pins"]
+ cola = load_cola_history()
+ _, checks = mb.reproduce_parent(
+ parameters=parameters,
+ parent=parent,
+ data_provenance=REGISTERED_REAL,
+ load_cohort_inputs=cohort.load_cohort_inputs,
+ cola_rates=cola,
+ provenance_kind=PSID_FILES,
+ source={"cola_history": dict(cola.provenance)},
+ )
+ assert all(check.identical for check in checks)
+ names = {check.name for check in checks}
+ assert {"pipeline.rows", "pipeline.cohort_structure"} <= names
+ common.require_identical(checks, stage="host reproduction")
From 72c0d78b19deb87eafae311bc943881c03b5fe58 Mon Sep 17 00:00:00 2001
From: Max Ghenis
Date: Sat, 3 Oct 2026 14:48:04 -0400
Subject: [PATCH 08/10] Reconcile the group-breakdown package and wire in the
report scheme
group_breakdowns/common.py is now one shared module for all four
adapters (cola, fra68, uniform_cut, min_benefit): one zero-tolerance
comparator, one registration preflight (issue #42 pointer, HEAD equal
to the registered commit, clean tree including untracked files, a new
runs/ destination), one paired exclusive write with rollback, and one
environment builder. Where the three builders' helpers differed, the
stricter behavior is kept. Every adapter still refuses before any group
attribute loads unless its committed artifact reproduces.
The three captures of SSA's MINT8 Table User Guide collapse to one
canonical capture (certified 2026-04-01) and one label file; tests show
the snapshots' main content and labels are identical.
The Butrica-Uccello report scheme now follows the cleared exercise-2
definitions extract: race rows, own and shared lifetime earnings
(wage-indexed earnings at ages 22-62 over a 41-year divisor, shared in
married years). The extract leaves the education and labor-force-
experience definitions open, so those rows stay unavailable, and the
unrecorded quintile conventions are named choices for the registration.
Adds the ten new src modules to the birth-evidence reducer's post-
review exclusions and the new tests to the tier manifest.
Co-Authored-By: Claude Opus 5.5
---
data/external/group_category_schemes_v1.json | 98 +-
.../lifetime_measure_sources.provenance.json | 4 +-
.../mint8_lifetime_quintile_definitions.json | 2 +-
.../mint8_row_categories.provenance.json | 66 +-
.../mint8_row_categories_2026.source.json | 1851 -----------------
.../mint8_table_user_guide.source.html | 18 +-
...t8_table_user_guide.source.provenance.json | 95 +-
.../ssa_mint8_payroll_option_row_labels.json | 1851 -----------------
...l_option_row_labels.source.provenance.json | 11 -
.../ssa_mint8_table_user_guide.source.html | 633 ------
...t8_table_user_guide.source.provenance.json | 14 -
.../ssa_mint8_user_guide_2026.source.html | 635 ------
.../exercise_2_uniform_cut/RESULTS.md | 17 +
.../track_u_groups_invented.env.json | 9 -
docs/design/lifetime_measures.md | 62 +-
docs/design/nasi_g1_group_attributes.md | 8 +-
docs/design/nasi_projection_groups_posthoc.md | 13 +-
scripts/extract_lifetime_measure_sources.py | 8 +-
scripts/first_estimates_birth_evidence.py | 13 +
scripts/make_nasi_repro_venv.sh | 243 ++-
scripts/run_min_benefit_groups_registered.py | 46 +-
scripts/run_projection_groups_registered.py | 61 +-
.../cohorts/group_attributes.py | 22 +-
.../estimates/group_breakdown.py | 7 +-
.../estimates/lifetime_measures.py | 215 +-
.../group_breakdowns/__init__.py | 9 +-
.../group_breakdowns/common.py | 896 +++++++-
.../group_breakdowns/uniform_cut.py | 159 +-
tests/cohorts/test_group_attributes.py | 20 +
tests/cohorts/test_group_category_schemes.py | 51 +-
.../estimates/test_birth_evidence_artifact.py | 10 +
.../estimates/test_group_breakdown_sources.py | 73 +-
.../test_lifetime_measure_sources.py | 25 +
tests/estimates/test_lifetime_measures.py | 213 +-
.../test_lifetime_measures_properties.py | 44 +-
tests/group_breakdowns/test_common.py | 364 ++++
.../test_min_benefit_parent_artifact.py | 69 +-
.../test_min_benefit_registered_script.py | 18 +-
.../test_min_benefit_reproduction_psid.py | 10 +
tests/group_breakdowns/test_registration.py | 20 +-
tests/group_breakdowns/test_uniform_cut.py | 35 +
.../test_uniform_cut_properties.py | 22 +-
tests/tier_counts.json | 6 +-
43 files changed, 2577 insertions(+), 5469 deletions(-)
delete mode 100644 data/external/mint8_row_categories_2026.source.json
delete mode 100644 data/external/ssa_mint8_payroll_option_row_labels.json
delete mode 100644 data/external/ssa_mint8_payroll_option_row_labels.source.provenance.json
delete mode 100644 data/external/ssa_mint8_table_user_guide.source.html
delete mode 100644 data/external/ssa_mint8_table_user_guide.source.provenance.json
delete mode 100644 data/external/ssa_mint8_user_guide_2026.source.html
create mode 100644 docs/analysis/nasi_group_breakdowns_invented_20261001/exercise_2_uniform_cut/RESULTS.md
delete mode 100644 docs/analysis/nasi_group_breakdowns_invented_20261001/exercise_2_uniform_cut/track_u_groups_invented.env.json
diff --git a/data/external/group_category_schemes_v1.json b/data/external/group_category_schemes_v1.json
index ed725078..97b02bd9 100644
--- a/data/external/group_category_schemes_v1.json
+++ b/data/external/group_category_schemes_v1.json
@@ -5,24 +5,26 @@
"mint8_user_guide": {
"title": "Table User Guide—Modeling Income in the Near Term (MINT) 8",
"url": "https://www.ssa.gov/policy/docs/projections/user-guide.html",
- "date_certified": "2025-10-01",
- "capture": "https://web.archive.org/web/20260419231110id_/https://www.ssa.gov/policy/docs/projections/user-guide.html",
- "committed_file": "data/external/ssa_mint8_table_user_guide.source.html",
- "sha256": "278d5d19c1b50f1d354db1ada515288af563c035a67eb16fb971d25700fb94e9",
+ "date_certified": "2026-04-01",
+ "capture": "Direct HTTPS GET, 2026-10-01T22:31:33Z; exact live SSA response body",
+ "committed_file": "data/external/mint8_table_user_guide.source.html",
+ "sha256": "8d5bc3f0de17831c2ed07383003257a63bb657df179d0c54868fda052f23f100",
"locator": "Definitions—Table Rows and Columns > Characteristic Subgroups—Table Rows"
},
"mint8_row_labels": {
"title": "MINT 8 policy-option tables, 'Increase payroll tax rate': row labels only",
"url": "https://www.ssa.gov/policy/docs/projections/policy-options/increase-payroll-tax-rate.html",
"capture": "http://web.archive.org/web/20260519190449id_/https://www.ssa.gov/policy/docs/projections/policy-options/increase-payroll-tax-rate.html",
- "committed_file": "data/external/ssa_mint8_payroll_option_row_labels.json",
+ "committed_file": "data/external/mint8_row_categories.json",
"sha256": "23fbfbc8dbc14b06144a83bf0583ea0a6c89505638f3f7dc5fd6b34a3a4ac650",
"locator": "Tables 1–3 (benefits), 7–9 (household income), 10–12 (official poverty), 13–16 (benefit/tax ratio), and 17–20 (initial replacement rate): captions and row labels only; no data cells."
},
"boomers2004_report_rows": {
- "title": "Butrica and Uccello (2004), How Will Boomers Fare at Retirement?, Tables 19 and 21 row labels as recorded in this repository",
- "repository_locator": "src/populace_dynamics/estimates/uniform_cut_tabulation.py NOT_COMPUTED_REPORT_ROWS['race_ethnicity'] and ['education']; docs/design/boomers2004_uniform_cut_comparison.md, 'Report rows not computed in v1 (named omissions)'",
- "note": "The repository records the Report's row labels as printed (cleared extract) and, for education, that the Report does not say whether 'High school graduate' includes some college. It records no other definition of these rows. The Report itself was not opened for this file."
+ "title": "Butrica and Uccello (2004), How Will Boomers Fare at Retirement?, cleared exercise-2 definitions",
+ "note": "Definition-only cleared extract, hash verified before reading. No Report PDF, restricted page or comparator was opened. Lines 231 and 327 explicitly leave Education, Labor Force Experience, quintile population and marital-status timing undefined.",
+ "extract_file": "exercise2-definitions-cleared-20260924.md",
+ "sha256": "a3978b683b4275424b6d12e9fe45f021fae277ccf9731a7883952564b6ed0384",
+ "locator": "cleared extract lines 58–66, 86–94, 223, 227, 229, 231 and 327"
}
},
"schemes": {
@@ -626,7 +628,7 @@
}
},
"boomers2004": {
- "title": "Butrica and Uccello (2004) Tables 19 and 21 rows (Track U's named omissions)",
+ "title": "Butrica and Uccello (2004) Tables 19 and 21 rows (cleared definitions)",
"dimensions": {
"race_ethnicity": {
"column": "race_ethnicity_report4",
@@ -642,9 +644,12 @@
},
"multiple_races": "other_non_hispanic",
"assumptions": [
- "The repository records only the row labels. Because the White and Black rows are marked non-Hispanic, this scheme reads 'Hispanic' as Hispanic of any race and 'Other' as every other non-Hispanic person; both readings are provisional, not Report definitions.",
- "Persons reporting two or more races are placed in 'Other', as in scheme mint8; the Report's rule is not recorded."
- ]
+ "report_hispanic_any_race: Hispanic takes precedence over race because the White and Black rows are non-Hispanic (lines 58 and 86); this precedence is a builder convention pending registration.",
+ "report_multiple_races_other: non-Hispanic persons reporting multiple races are placed in Other; the extract gives no multiple-race rule. Pending registration.",
+ "report_other_partition: after Hispanic, White non-Hispanic and Black non-Hispanic assignments, other known non-Hispanic race reports enter Other; its cited gloss includes Asian and Native American people (line 229). Pending registration for any remaining PSID categories."
+ ],
+ "locator": "cleared extract lines 58 and 86 (row labels), 229 (Other gloss)",
+ "definition": "Other comprises other minority groups, including Asian and Native American people (cleared definition-only paraphrase, line 229). White and Black rows are explicitly non-Hispanic."
},
"education": {
"column": "education_report3",
@@ -658,7 +663,7 @@
"min": 0,
"max": 17,
"status": "definition_not_recorded",
- "reason": "The cleared repository records only these three row labels and the uncertainty about some college. It records no numeric years-of-schooling definitions for any row, so no cutoff is inferred."
+ "reason": "report_education_mapping: no education definition is stated, including whether High school graduate includes some college (cleared extract lines 231 and 327). Leave assignments unavailable until registration supplies the mapping; numeric schooling boundaries are not inferred."
}
],
"assumptions": [],
@@ -666,8 +671,73 @@
"High school dropout",
"High school graduate",
"College graduate"
- ]
+ ],
+ "locator": "cleared extract lines 59 and 87 (row labels), 231 and 327 (definitions not stated)",
+ "builder_defaults": {
+ "report_education_mapping": "unavailable pending registration; no years-of-schooling or degree mapping is defined in the cleared extract (lines 231 and 327)"
+ }
+ }
+ },
+ "row_groups": [
+ {
+ "group": "Race/Ethnicity",
+ "labels": [
+ "White, non-hispanic",
+ "Black, non-hispanic",
+ "Hispanic",
+ "Other"
+ ],
+ "locator": "cleared extract lines 58 and 86"
+ },
+ {
+ "group": "Education",
+ "labels": [
+ "High school dropout",
+ "High school graduate",
+ "College graduate"
+ ],
+ "locator": "cleared extract lines 59 and 87"
+ },
+ {
+ "group": "Labor Force Experience",
+ "labels": [
+ "Less than 20 years",
+ "20 to 29 years",
+ "30 to 34 years",
+ "35 or more years"
+ ],
+ "locator": "cleared extract lines 60 and 88"
+ },
+ {
+ "group": "Lifetime Earnings (Own)",
+ "labels": [
+ "1st Quintile",
+ "2nd Quintile",
+ "3rd Quintile",
+ "4th Quintile",
+ "5th Quintile"
+ ],
+ "locator": "cleared extract lines 61 and 89"
+ },
+ {
+ "group": "Lifetime Earnings (Shared)",
+ "labels": [
+ "1st Quintile",
+ "2nd Quintile",
+ "3rd Quintile",
+ "4th Quintile",
+ "5th Quintile"
+ ],
+ "locator": "cleared extract lines 62 and 90"
}
+ ],
+ "builder_defaults": {
+ "report_education_mapping": "unavailable pending registration (lines 231 and 327)",
+ "report_labor_force_experience": "unavailable pending registration: years in the labor force is a label, not a career-based formula (lines 231 and 327); positive-earnings years are not substituted",
+ "report_quintile_population": "unavailable pending registration (lines 231 and 327)",
+ "report_quintile_order": "unavailable pending registration: rows name 1st through 5th Quintile but do not specify high-to-low order (lines 61–62 and 89–90)",
+ "report_marital_status_timing": "unavailable pending registration (lines 231 and 327)",
+ "report_quintile_ties": "unavailable pending registration; weighted ties are not defined in the cleared extract"
}
}
}
diff --git a/data/external/lifetime_measure_sources.provenance.json b/data/external/lifetime_measure_sources.provenance.json
index f5e21cf9..1a8c6786 100644
--- a/data/external/lifetime_measure_sources.provenance.json
+++ b/data/external/lifetime_measure_sources.provenance.json
@@ -42,7 +42,7 @@
"acquisition": "Direct HTTPS GET of the live ssa.gov page with curl -A 'Wget/1.21.4' on 2026-10-01 (HTTP 200, content-type text/html; charset=UTF-8; response Date header Thu, 01 Oct 2026 22:31:33-34 GMT). The body is committed byte for byte. A per-request Akamai mPulse script in changes between fetches; an earlier fetch the same day parsed to the identical tables."
},
"mint8_user_guide": {
- "committed_source_file": "data/external/ssa_mint8_user_guide_2026.source.html",
+ "committed_source_file": "data/external/mint8_table_user_guide.source.html",
"source_url": "https://www.ssa.gov/policy/docs/projections/user-guide.html",
"document": "MINT8 Table User Guide",
"source_sha256": "8d5bc3f0de17831c2ed07383003257a63bb657df179d0c54868fda052f23f100",
@@ -51,7 +51,7 @@
"acquisition": "Direct HTTPS GET of the live ssa.gov page with curl -A 'Wget/1.21.4' on 2026-10-01 (HTTP 200, content-type text/html; charset=UTF-8; response Date header Thu, 01 Oct 2026 22:31:33-34 GMT). The body is committed byte for byte. A per-request Akamai mPulse script in changes between fetches; an earlier fetch the same day parsed to the identical tables."
},
"mint8_table_row_labels": {
- "committed_source_file": "data/external/mint8_row_categories_2026.source.json",
+ "committed_source_file": "data/external/mint8_row_categories.json",
"source_url": "https://www.ssa.gov/policy/docs/projections/policy-options/increase-payroll-tax-rate.html",
"document": "MINT8 payroll-tax option table labels (labels only)",
"source_sha256": "23fbfbc8dbc14b06144a83bf0583ea0a6c89505638f3f7dc5fd6b34a3a4ac650",
diff --git a/data/external/mint8_lifetime_quintile_definitions.json b/data/external/mint8_lifetime_quintile_definitions.json
index 2e2df976..063b9d38 100644
--- a/data/external/mint8_lifetime_quintile_definitions.json
+++ b/data/external/mint8_lifetime_quintile_definitions.json
@@ -55,7 +55,7 @@
"date_certified": "2026-04-01",
"label_file": "mint8_row_categories.json (MINT-categories lane)",
"label_file_sha256": "23fbfbc8dbc14b06144a83bf0583ea0a6c89505638f3f7dc5fd6b34a3a4ac650",
- "committed_label_file": "data/external/mint8_row_categories_2026.source.json",
+ "committed_label_file": "data/external/mint8_row_categories.json",
"tables": "13-20 (benefit/tax ratios and initial replacement rates)",
"note": "Labels only. The lane's parser emitted th, caption and heading text and no data cell; the option table itself is not committed and supplies no input to this package."
},
diff --git a/data/external/mint8_row_categories.provenance.json b/data/external/mint8_row_categories.provenance.json
index 52abd516..7d6f60ab 100644
--- a/data/external/mint8_row_categories.provenance.json
+++ b/data/external/mint8_row_categories.provenance.json
@@ -10,6 +10,68 @@
"source_sha256": "23fbfbc8dbc14b06144a83bf0583ea0a6c89505638f3f7dc5fd6b34a3a4ac650",
"source_length_bytes": 37524,
"recorded_in": "~/microcosm-launch-evidence/dynasim-parity-20260909/RESTRICTED-FILES.md changelog entry 2026-10-01 22:20 (the label file and its SHA-256 23fbfbc8...)",
- "consumed_by": "src/populace_dynamics/estimates/group_breakdown.py (MINT8 row-group headings and row labels, verbatim and in document order; verify_mint8_sources checks the schemes against this file)",
- "note": "LABELS ONLY: no table value was extracted, read or committed. The payroll-tax-rate option is not one of the reforms the blind tests or the NASI follow-ups compare."
+ "consumed_by": [
+ "data/external/group_category_schemes_v1.json (cohorts/group_attributes.py)",
+ "src/populace_dynamics/estimates/group_breakdown.py",
+ "scripts/extract_lifetime_measure_sources.py (estimates/lifetime_measures.py)"
+ ],
+ "note": "LABELS ONLY: no table value was extracted, read or committed. The payroll-tax-rate option is not one of the reforms the blind tests or the NASI follow-ups compare.",
+ "consolidation": {
+ "decision": "All three committed labels-only extractions were byte-identical, including every caption, column heading, row-group heading and ordered row label in all 20 tables. Keep the original MINT-categories lane filename and bytes as the canonical extraction.",
+ "snapshots": [
+ {
+ "original_file": "data/external/mint8_row_categories.json",
+ "source_sha256": "23fbfbc8dbc14b06144a83bf0583ea0a6c89505638f3f7dc5fd6b34a3a4ac650",
+ "source_length_bytes": 37524,
+ "date_certified": "2026-04-01",
+ "original_provenance": {
+ "schema_version": "external_source_provenance.v1",
+ "committed_source_file": "data/external/mint8_row_categories.json",
+ "source_url": "https://www.ssa.gov/policy/docs/projections/policy-options/increase-payroll-tax-rate.html",
+ "document": "Social Security Administration, MINT8 projected-effects tables for the policy-option page increase-payroll-tax-rate.html (labels only)",
+ "date_certified": "2026-04-01",
+ "retrieval_date": "2026-10-01",
+ "fetch_method": "The MINT-categories lane (workflow wf_35b496b0-085, orchestrating session 95606380, 2026-10-01) fetched the Internet Archive raw capture http://web.archive.org/web/20260519190449id_/https://www.ssa.gov/policy/docs/projections/policy-options/increase-payroll-tax-rate.html (raw SHA-256 3cfb6eb636cad6d062eb22734c533fb349a0bc4761014eee8f2ccb0eb6e28d8c) and extracted only captions, column headings and row labels with a parser that emits no ordinary data cell. The committed file is that extract, byte-identical; the raw page is not committed.",
+ "raw_page_sha256": "3cfb6eb636cad6d062eb22734c533fb349a0bc4761014eee8f2ccb0eb6e28d8c",
+ "source_sha256": "23fbfbc8dbc14b06144a83bf0583ea0a6c89505638f3f7dc5fd6b34a3a4ac650",
+ "source_length_bytes": 37524,
+ "recorded_in": "~/microcosm-launch-evidence/dynasim-parity-20260909/RESTRICTED-FILES.md changelog entry 2026-10-01 22:20 (the label file and its SHA-256 23fbfbc8...)",
+ "consumed_by": "src/populace_dynamics/estimates/group_breakdown.py (MINT8 row-group headings and row labels, verbatim and in document order; verify_mint8_sources checks the schemes against this file)",
+ "note": "LABELS ONLY: no table value was extracted, read or committed. The payroll-tax-rate option is not one of the reforms the blind tests or the NASI follow-ups compare."
+ }
+ },
+ {
+ "original_file": "data/external/ssa_mint8_payroll_option_row_labels.json",
+ "source_sha256": "23fbfbc8dbc14b06144a83bf0583ea0a6c89505638f3f7dc5fd6b34a3a4ac650",
+ "source_length_bytes": 37524,
+ "date_certified": "2026-04-01",
+ "original_provenance": {
+ "schema_version": "external_source_provenance.v1",
+ "committed_source_file": "data/external/ssa_mint8_payroll_option_row_labels.json",
+ "source_url": "https://www.ssa.gov/policy/docs/projections/policy-options/increase-payroll-tax-rate.html",
+ "document": "SSA MINT 8 policy-option tables, 'Increase payroll tax rate' — row and column labels only",
+ "fetch_method": "Built by the MINT-categories lane of orchestrating session 95606380 from the Internet Archive capture http://web.archive.org/web/20260519190449id_/https://www.ssa.gov/policy/docs/projections/policy-options/increase-payroll-tax-rate.html (raw page SHA-256 3cfb6eb636cad6d062eb22734c533fb349a0bc4761014eee8f2ccb0eb6e28d8c, recorded inside the file) with a parser that emitted table captions, header cells and row labels and never an ordinary data cell. Copied byte-identically from the lane's label file.",
+ "source_sha256": "23fbfbc8dbc14b06144a83bf0583ea0a6c89505638f3f7dc5fd6b34a3a4ac650",
+ "acquired_by": "MINT-categories lane (workflow wf_35b496b0-085), 2026-10-01; committed by the G1 group-attributes builder",
+ "consumed_by": "data/external/group_category_schemes_v1.json (scheme mint8 attribute and lifetime-quintile row labels; annual and cohort table row-group orders)",
+ "note": "LABELS ONLY. No MINT data cell is in this file. The restricted-files list allows one payroll-tax option table to be read for its row labels only; this is that table's label record."
+ }
+ },
+ {
+ "original_file": "data/external/mint8_row_categories_2026.source.json",
+ "source_sha256": "23fbfbc8dbc14b06144a83bf0583ea0a6c89505638f3f7dc5fd6b34a3a4ac650",
+ "source_length_bytes": 37524,
+ "date_certified": "2026-04-01",
+ "original_provenance": {
+ "committed_source_file": "data/external/mint8_row_categories_2026.source.json",
+ "source_url": "https://www.ssa.gov/policy/docs/projections/policy-options/increase-payroll-tax-rate.html",
+ "document": "MINT8 payroll-tax option table labels (labels only)",
+ "source_sha256": "23fbfbc8dbc14b06144a83bf0583ea0a6c89505638f3f7dc5fd6b34a3a4ac650",
+ "source_length_bytes": 37524,
+ "locator": "label-only JSON: dateCertified; tables.13 through tables.20; groups whose group is Current-law initial AIME quintile, Lifetime payroll tax quintile or Lifetime payroll tax quintile (shared); their labels arrays",
+ "acquisition": "Copied byte for byte from the cleared MINT-categories lane's followup/inputs/mint/mint8_row_categories.json; label-only extraction of the archived source named in MINT8_TABLE_LABEL_PROVENANCE, never a table body."
+ }
+ }
+ ]
+ }
}
diff --git a/data/external/mint8_row_categories_2026.source.json b/data/external/mint8_row_categories_2026.source.json
deleted file mode 100644
index 887adf55..00000000
--- a/data/external/mint8_row_categories_2026.source.json
+++ /dev/null
@@ -1,1851 +0,0 @@
-{
- "source_url": "https://www.ssa.gov/policy/docs/projections/policy-options/increase-payroll-tax-rate.html",
- "retrieved_via": "http://web.archive.org/web/20260519190449id_/https://www.ssa.gov/policy/docs/projections/policy-options/increase-payroll-tax-rate.html",
- "raw_sha256": "3cfb6eb636cad6d062eb22734c533fb349a0bc4761014eee8f2ccb0eb6e28d8c",
- "dateCertified": "2026-04-01",
- "note": "LABELS ONLY. No data cells were extracted. Table numbers are document order on the page.",
- "tables": {
- "1": {
- "caption": "Projected Effects of Proposal on Social Security Benefits in 2030 POPULATION: Current-law beneficiaries aged 60 or older (characteristics)",
- "columns": [
- "Characteristic",
- "Percent of population with a—",
- "Percent change in Social Security benefits at the—",
- "Benefit decrease",
- "Benefit increase",
- "10th %ile",
- "Median",
- "90th %ile"
- ],
- "groups": [
- {
- "group": "Total",
- "labels": []
- },
- {
- "group": "Sex",
- "labels": [
- "Female",
- "Male"
- ]
- },
- {
- "group": "Race and ethnicity",
- "labels": [
- "Hispanic or Latino, any race",
- "White, non-Hispanic",
- "Black or African American, non-Hispanic",
- "All other races, non-Hispanic"
- ]
- },
- {
- "group": "Country of birth",
- "labels": [
- "United States",
- "Other countries"
- ]
- },
- {
- "group": "Age",
- "labels": [
- "60–69",
- "70–79",
- "80–89",
- "90 or older"
- ]
- },
- {
- "group": "Marital status",
- "labels": [
- "Married",
- "Divorced",
- "Widowed",
- "Never married"
- ]
- },
- {
- "group": "Highest education level",
- "labels": [
- "Graduate",
- "Bachelor",
- "Associate",
- "High school",
- "Less than high school"
- ]
- },
- {
- "group": "Current-law poverty status",
- "labels": [
- "Above poverty",
- "In poverty"
- ]
- },
- {
- "group": "Current-law household income quintile",
- "labels": [
- "Highest",
- "Second highest",
- "Middle",
- "Second lowest",
- "Lowest"
- ]
- },
- {
- "group": "Current-law benefit type",
- "labels": [
- "Retired worker only",
- "Widow(er) (includes dually entitled)",
- "Spousal (includes dually entitled)",
- "Disabled worker only"
- ]
- }
- ]
- },
- "2": {
- "caption": "Projected Effects of Proposal on Social Security Benefits in 2050 POPULATION: Current-law beneficiaries aged 60 or older (characteristics)",
- "columns": [
- "Characteristic",
- "Percent of population with a—",
- "Percent change in Social Security benefits at the—",
- "Benefit decrease",
- "Benefit increase",
- "10th %ile",
- "Median",
- "90th %ile"
- ],
- "groups": [
- {
- "group": "Total",
- "labels": []
- },
- {
- "group": "Sex",
- "labels": [
- "Female",
- "Male"
- ]
- },
- {
- "group": "Race and ethnicity",
- "labels": [
- "Hispanic or Latino, any race",
- "White, non-Hispanic",
- "Black or African American, non-Hispanic",
- "All other races, non-Hispanic"
- ]
- },
- {
- "group": "Country of birth",
- "labels": [
- "United States",
- "Other countries"
- ]
- },
- {
- "group": "Age",
- "labels": [
- "60–69",
- "70–79",
- "80–89",
- "90 or older"
- ]
- },
- {
- "group": "Marital status",
- "labels": [
- "Married",
- "Divorced",
- "Widowed",
- "Never married"
- ]
- },
- {
- "group": "Highest education level",
- "labels": [
- "Graduate",
- "Bachelor",
- "Associate",
- "High school",
- "Less than high school"
- ]
- },
- {
- "group": "Current-law poverty status",
- "labels": [
- "Above poverty",
- "In poverty"
- ]
- },
- {
- "group": "Current-law household income quintile",
- "labels": [
- "Highest",
- "Second highest",
- "Middle",
- "Second lowest",
- "Lowest"
- ]
- },
- {
- "group": "Current-law benefit type",
- "labels": [
- "Retired worker only",
- "Widow(er) (includes dually entitled)",
- "Spousal (includes dually entitled)",
- "Disabled worker only"
- ]
- }
- ]
- },
- "3": {
- "caption": "Projected Effects of Proposal on Social Security Benefits in 2070 POPULATION: Current-law beneficiaries aged 60 or older (characteristics)",
- "columns": [
- "Characteristic",
- "Percent of population with a—",
- "Percent change in Social Security benefits at the—",
- "Benefit decrease",
- "Benefit increase",
- "10th %ile",
- "Median",
- "90th %ile"
- ],
- "groups": [
- {
- "group": "Total",
- "labels": []
- },
- {
- "group": "Sex",
- "labels": [
- "Female",
- "Male"
- ]
- },
- {
- "group": "Race and ethnicity",
- "labels": [
- "Hispanic or Latino, any race",
- "White, non-Hispanic",
- "Black or African American, non-Hispanic",
- "All other races, non-Hispanic"
- ]
- },
- {
- "group": "Country of birth",
- "labels": [
- "United States",
- "Other countries"
- ]
- },
- {
- "group": "Age",
- "labels": [
- "60–69",
- "70–79",
- "80–89",
- "90 or older"
- ]
- },
- {
- "group": "Marital status",
- "labels": [
- "Married",
- "Divorced",
- "Widowed",
- "Never married"
- ]
- },
- {
- "group": "Highest education level",
- "labels": [
- "Graduate",
- "Bachelor",
- "Associate",
- "High school",
- "Less than high school"
- ]
- },
- {
- "group": "Current-law poverty status",
- "labels": [
- "Above poverty",
- "In poverty"
- ]
- },
- {
- "group": "Current-law household income quintile",
- "labels": [
- "Highest",
- "Second highest",
- "Middle",
- "Second lowest",
- "Lowest"
- ]
- },
- {
- "group": "Current-law benefit type",
- "labels": [
- "Retired worker only",
- "Widow(er) (includes dually entitled)",
- "Spousal (includes dually entitled)",
- "Disabled worker only"
- ]
- }
- ]
- },
- "4": {
- "caption": "Projected Effects of Proposal on Social Security Taxes Paid in 2030 POPULATION: Current-law payroll taxpayers aged 31 or older (characteristics)",
- "columns": [
- "Characteristic",
- "Percent of population with a—",
- "Percent change in Social Security taxes paid at the—",
- "Change in taxes paid (in 2024$) at the—",
- "Tax decrease",
- "Tax increase",
- "10th %ile",
- "Median",
- "90th %ile",
- "10th %ile",
- "Median",
- "90th %ile"
- ],
- "groups": [
- {
- "group": "Total",
- "labels": []
- },
- {
- "group": "Sex",
- "labels": [
- "Female",
- "Male"
- ]
- },
- {
- "group": "Race and ethnicity",
- "labels": [
- "Hispanic or Latino, any race",
- "White, non-Hispanic",
- "Black or African American, non-Hispanic",
- "All other races, non-Hispanic"
- ]
- },
- {
- "group": "Country of birth",
- "labels": [
- "United States",
- "Other countries"
- ]
- },
- {
- "group": "Age",
- "labels": [
- "31–39",
- "40–49",
- "50–59",
- "60–69",
- "70 or older"
- ]
- },
- {
- "group": "Marital status",
- "labels": [
- "Married",
- "Divorced",
- "Widowed",
- "Never married"
- ]
- },
- {
- "group": "Highest education level",
- "labels": [
- "Graduate",
- "Bachelor",
- "Associate",
- "High school",
- "Less than high school"
- ]
- },
- {
- "group": "Current-law household income quintile",
- "labels": [
- "Highest",
- "Second highest",
- "Middle",
- "Second lowest",
- "Lowest"
- ]
- },
- {
- "group": "Current-law payroll taxes quintile",
- "labels": [
- "Highest",
- "Second highest",
- "Middle",
- "Second lowest",
- "Lowest"
- ]
- }
- ]
- },
- "5": {
- "caption": "Projected Effects of Proposal on Social Security Taxes Paid in 2050 POPULATION: Current-law payroll taxpayers aged 31 or older (characteristics)",
- "columns": [
- "Characteristic",
- "Percent of population with a—",
- "Percent change in Social Security taxes paid at the—",
- "Change in taxes paid (in 2024$) at the—",
- "Tax decrease",
- "Tax increase",
- "10th %ile",
- "Median",
- "90th %ile",
- "10th %ile",
- "Median",
- "90th %ile"
- ],
- "groups": [
- {
- "group": "Total",
- "labels": []
- },
- {
- "group": "Sex",
- "labels": [
- "Female",
- "Male"
- ]
- },
- {
- "group": "Race and ethnicity",
- "labels": [
- "Hispanic or Latino, any race",
- "White, non-Hispanic",
- "Black or African American, non-Hispanic",
- "All other races, non-Hispanic"
- ]
- },
- {
- "group": "Country of birth",
- "labels": [
- "United States",
- "Other countries"
- ]
- },
- {
- "group": "Age",
- "labels": [
- "31–39",
- "40–49",
- "50–59",
- "60–69",
- "70 or older"
- ]
- },
- {
- "group": "Marital status",
- "labels": [
- "Married",
- "Divorced",
- "Widowed",
- "Never married"
- ]
- },
- {
- "group": "Highest education level",
- "labels": [
- "Graduate",
- "Bachelor",
- "Associate",
- "High school",
- "Less than high school"
- ]
- },
- {
- "group": "Current-law household income quintile",
- "labels": [
- "Highest",
- "Second highest",
- "Middle",
- "Second lowest",
- "Lowest"
- ]
- },
- {
- "group": "Current-law payroll taxes quintile",
- "labels": [
- "Highest",
- "Second highest",
- "Middle",
- "Second lowest",
- "Lowest"
- ]
- }
- ]
- },
- "6": {
- "caption": "Projected Effects of Proposal on Social Security Taxes Paid in 2070 POPULATION: Current-law payroll taxpayers aged 31 or older (characteristics)",
- "columns": [
- "Characteristic",
- "Percent of population with a—",
- "Percent change in Social Security taxes paid at the—",
- "Change in taxes paid (in 2024$) at the—",
- "Tax decrease",
- "Tax increase",
- "10th %ile",
- "Median",
- "90th %ile",
- "10th %ile",
- "Median",
- "90th %ile"
- ],
- "groups": [
- {
- "group": "Total",
- "labels": []
- },
- {
- "group": "Sex",
- "labels": [
- "Female",
- "Male"
- ]
- },
- {
- "group": "Race and ethnicity",
- "labels": [
- "Hispanic or Latino, any race",
- "White, non-Hispanic",
- "Black or African American, non-Hispanic",
- "All other races, non-Hispanic"
- ]
- },
- {
- "group": "Country of birth",
- "labels": [
- "United States",
- "Other countries"
- ]
- },
- {
- "group": "Age",
- "labels": [
- "31–39",
- "40–49",
- "50–59",
- "60–69",
- "70 or older"
- ]
- },
- {
- "group": "Marital status",
- "labels": [
- "Married",
- "Divorced",
- "Widowed",
- "Never married"
- ]
- },
- {
- "group": "Highest education level",
- "labels": [
- "Graduate",
- "Bachelor",
- "Associate",
- "High school",
- "Less than high school"
- ]
- },
- {
- "group": "Current-law household income quintile",
- "labels": [
- "Highest",
- "Second highest",
- "Middle",
- "Second lowest",
- "Lowest"
- ]
- },
- {
- "group": "Current-law payroll taxes quintile",
- "labels": [
- "Highest",
- "Second highest",
- "Middle",
- "Second lowest",
- "Lowest"
- ]
- }
- ]
- },
- "7": {
- "caption": "Projected Effects of Proposal on Household Income in 2030 POPULATION: Current-law beneficiaries aged 60 or older (characteristics)",
- "columns": [
- "Characteristic",
- "Percent of population with an—",
- "Percent change in household income at the—",
- "Income decrease",
- "Income increase",
- "10th %ile",
- "Median",
- "90th %ile"
- ],
- "groups": [
- {
- "group": "Total",
- "labels": []
- },
- {
- "group": "Sex",
- "labels": [
- "Female",
- "Male"
- ]
- },
- {
- "group": "Race and ethnicity",
- "labels": [
- "Hispanic or Latino, any race",
- "White, non-Hispanic",
- "Black or African American, non-Hispanic",
- "All other races, non-Hispanic"
- ]
- },
- {
- "group": "Country of birth",
- "labels": [
- "United States",
- "Other countries"
- ]
- },
- {
- "group": "Age",
- "labels": [
- "60–69",
- "70–79",
- "80–89",
- "90 or older"
- ]
- },
- {
- "group": "Marital status",
- "labels": [
- "Married",
- "Divorced",
- "Widowed",
- "Never married"
- ]
- },
- {
- "group": "Highest education level",
- "labels": [
- "Graduate",
- "Bachelor",
- "Associate",
- "High school",
- "Less than high school"
- ]
- },
- {
- "group": "Current-law poverty status",
- "labels": [
- "Above poverty",
- "In poverty"
- ]
- },
- {
- "group": "Current-law household income quintile",
- "labels": [
- "Highest",
- "Second highest",
- "Middle",
- "Second lowest",
- "Lowest"
- ]
- },
- {
- "group": "Current-law benefit type",
- "labels": [
- "Retired worker only",
- "Widow(er) (includes dually entitled)",
- "Spousal (includes dually entitled)",
- "Disabled worker only"
- ]
- }
- ]
- },
- "8": {
- "caption": "Projected Effects of Proposal on Household Income in 2050 POPULATION: Current-law beneficiaries aged 60 or older (characteristics)",
- "columns": [
- "Characteristic",
- "Percent of population with an—",
- "Percent change in household income at the—",
- "Income decrease",
- "Income increase",
- "10th %ile",
- "Median",
- "90th %ile"
- ],
- "groups": [
- {
- "group": "Total",
- "labels": []
- },
- {
- "group": "Sex",
- "labels": [
- "Female",
- "Male"
- ]
- },
- {
- "group": "Race and ethnicity",
- "labels": [
- "Hispanic or Latino, any race",
- "White, non-Hispanic",
- "Black or African American, non-Hispanic",
- "All other races, non-Hispanic"
- ]
- },
- {
- "group": "Country of birth",
- "labels": [
- "United States",
- "Other countries"
- ]
- },
- {
- "group": "Age",
- "labels": [
- "60–69",
- "70–79",
- "80–89",
- "90 or older"
- ]
- },
- {
- "group": "Marital status",
- "labels": [
- "Married",
- "Divorced",
- "Widowed",
- "Never married"
- ]
- },
- {
- "group": "Highest education level",
- "labels": [
- "Graduate",
- "Bachelor",
- "Associate",
- "High school",
- "Less than high school"
- ]
- },
- {
- "group": "Current-law poverty status",
- "labels": [
- "Above poverty",
- "In poverty"
- ]
- },
- {
- "group": "Current-law household income quintile",
- "labels": [
- "Highest",
- "Second highest",
- "Middle",
- "Second lowest",
- "Lowest"
- ]
- },
- {
- "group": "Current-law benefit type",
- "labels": [
- "Retired worker only",
- "Widow(er) (includes dually entitled)",
- "Spousal (includes dually entitled)",
- "Disabled worker only"
- ]
- }
- ]
- },
- "9": {
- "caption": "Projected Effects of Proposal on Household Income in 2070 POPULATION: Current-law beneficiaries aged 60 or older (characteristics)",
- "columns": [
- "Characteristic",
- "Percent of population with an—",
- "Percent change in household income at the—",
- "Income decrease",
- "Income increase",
- "10th %ile",
- "Median",
- "90th %ile"
- ],
- "groups": [
- {
- "group": "Total",
- "labels": []
- },
- {
- "group": "Sex",
- "labels": [
- "Female",
- "Male"
- ]
- },
- {
- "group": "Race and ethnicity",
- "labels": [
- "Hispanic or Latino, any race",
- "White, non-Hispanic",
- "Black or African American, non-Hispanic",
- "All other races, non-Hispanic"
- ]
- },
- {
- "group": "Country of birth",
- "labels": [
- "United States",
- "Other countries"
- ]
- },
- {
- "group": "Age",
- "labels": [
- "60–69",
- "70–79",
- "80–89",
- "90 or older"
- ]
- },
- {
- "group": "Marital status",
- "labels": [
- "Married",
- "Divorced",
- "Widowed",
- "Never married"
- ]
- },
- {
- "group": "Highest education level",
- "labels": [
- "Graduate",
- "Bachelor",
- "Associate",
- "High school",
- "Less than high school"
- ]
- },
- {
- "group": "Current-law poverty status",
- "labels": [
- "Above poverty",
- "In poverty"
- ]
- },
- {
- "group": "Current-law household income quintile",
- "labels": [
- "Highest",
- "Second highest",
- "Middle",
- "Second lowest",
- "Lowest"
- ]
- },
- {
- "group": "Current-law benefit type",
- "labels": [
- "Retired worker only",
- "Widow(er) (includes dually entitled)",
- "Spousal (includes dually entitled)",
- "Disabled worker only"
- ]
- }
- ]
- },
- "10": {
- "caption": "Projected Effects of Proposal on Official Poverty Measure in 2030 POPULATION: Current-law beneficiaries aged 60 or older (characteristics)",
- "columns": [
- "Characteristic",
- "Official poverty rate",
- "Number of population in poverty (in thousands)",
- "Percent change in the number in poverty",
- "Under current law",
- "With proposal",
- "Under current law",
- "With proposal",
- "Change"
- ],
- "groups": [
- {
- "group": "Total",
- "labels": []
- },
- {
- "group": "Sex",
- "labels": [
- "Female",
- "Male"
- ]
- },
- {
- "group": "Race and ethnicity",
- "labels": [
- "Hispanic or Latino, any race",
- "White, non-Hispanic",
- "Black or African American, non-Hispanic",
- "All other races, non-Hispanic"
- ]
- },
- {
- "group": "Country of birth",
- "labels": [
- "United States",
- "Other countries"
- ]
- },
- {
- "group": "Age",
- "labels": [
- "60–69",
- "70–79",
- "80–89",
- "90 or older"
- ]
- },
- {
- "group": "Marital status",
- "labels": [
- "Married",
- "Divorced",
- "Widowed",
- "Never married"
- ]
- },
- {
- "group": "Highest education level",
- "labels": [
- "Graduate",
- "Bachelor",
- "Associate",
- "High school",
- "Less than high school"
- ]
- },
- {
- "group": "Current-law poverty status",
- "labels": [
- "Above poverty",
- "In poverty"
- ]
- },
- {
- "group": "Current-law benefit type",
- "labels": [
- "Retired worker only",
- "Widow(er) (includes dually entitled)",
- "Spousal (includes dually entitled)",
- "Disabled worker only"
- ]
- }
- ]
- },
- "11": {
- "caption": "Projected Effects of Proposal on Official Poverty Measure in 2050 POPULATION: Current-law beneficiaries aged 60 or older (characteristics)",
- "columns": [
- "Characteristic",
- "Official poverty rate",
- "Number of population in poverty (in thousands)",
- "Percent change in the number in poverty",
- "Under current law",
- "With proposal",
- "Under current law",
- "With proposal",
- "Change"
- ],
- "groups": [
- {
- "group": "Total",
- "labels": []
- },
- {
- "group": "Sex",
- "labels": [
- "Female",
- "Male"
- ]
- },
- {
- "group": "Race and ethnicity",
- "labels": [
- "Hispanic or Latino, any race",
- "White, non-Hispanic",
- "Black or African American, non-Hispanic",
- "All other races, non-Hispanic"
- ]
- },
- {
- "group": "Country of birth",
- "labels": [
- "United States",
- "Other countries"
- ]
- },
- {
- "group": "Age",
- "labels": [
- "60–69",
- "70–79",
- "80–89",
- "90 or older"
- ]
- },
- {
- "group": "Marital status",
- "labels": [
- "Married",
- "Divorced",
- "Widowed",
- "Never married"
- ]
- },
- {
- "group": "Highest education level",
- "labels": [
- "Graduate",
- "Bachelor",
- "Associate",
- "High school",
- "Less than high school"
- ]
- },
- {
- "group": "Current-law poverty status",
- "labels": [
- "Above poverty",
- "In poverty"
- ]
- },
- {
- "group": "Current-law benefit type",
- "labels": [
- "Retired worker only",
- "Widow(er) (includes dually entitled)",
- "Spousal (includes dually entitled)",
- "Disabled worker only"
- ]
- }
- ]
- },
- "12": {
- "caption": "Projected Effects of Proposal on Official Poverty Measure in 2070 POPULATION: Current-law beneficiaries aged 60 or older (characteristics)",
- "columns": [
- "Characteristic",
- "Official poverty rate",
- "Number of population in poverty (in thousands)",
- "Percent change in the number in poverty",
- "Under current law",
- "With proposal",
- "Under current law",
- "With proposal",
- "Change"
- ],
- "groups": [
- {
- "group": "Total",
- "labels": []
- },
- {
- "group": "Sex",
- "labels": [
- "Female",
- "Male"
- ]
- },
- {
- "group": "Race and ethnicity",
- "labels": [
- "Hispanic or Latino, any race",
- "White, non-Hispanic",
- "Black or African American, non-Hispanic",
- "All other races, non-Hispanic"
- ]
- },
- {
- "group": "Country of birth",
- "labels": [
- "United States",
- "Other countries"
- ]
- },
- {
- "group": "Age",
- "labels": [
- "60–69",
- "70–79",
- "80–89",
- "90 or older"
- ]
- },
- {
- "group": "Marital status",
- "labels": [
- "Married",
- "Divorced",
- "Widowed",
- "Never married"
- ]
- },
- {
- "group": "Highest education level",
- "labels": [
- "Graduate",
- "Bachelor",
- "Associate",
- "High school",
- "Less than high school"
- ]
- },
- {
- "group": "Current-law poverty status",
- "labels": [
- "Above poverty",
- "In poverty"
- ]
- },
- {
- "group": "Current-law benefit type",
- "labels": [
- "Retired worker only",
- "Widow(er) (includes dually entitled)",
- "Spousal (includes dually entitled)",
- "Disabled worker only"
- ]
- }
- ]
- },
- "13": {
- "caption": "Projected Effects of Proposal on Benefit/Tax Ratios POPULATION: Workers born 1960–1969 with a benefit/tax ratio (characteristics)",
- "columns": [
- "Characteristic",
- "Percent of population with a—",
- "Percent change in benefit/tax ratio at the—",
- "Benefit/tax ratio under current law at the—",
- "Benefit/tax ratio with proposal at the—",
- "Ratio decrease",
- "Ratio increase",
- "10th %ile",
- "Median",
- "90th %ile",
- "10th %ile",
- "Median",
- "90th %ile",
- "10th %ile",
- "Median",
- "90th %ile"
- ],
- "groups": [
- {
- "group": "Total",
- "labels": []
- },
- {
- "group": "Sex",
- "labels": [
- "Female",
- "Male"
- ]
- },
- {
- "group": "Race and ethnicity",
- "labels": [
- "Hispanic or Latino, any race",
- "White, non-Hispanic",
- "Black or African American, non-Hispanic",
- "All other races, non-Hispanic"
- ]
- },
- {
- "group": "Country of birth",
- "labels": [
- "United States",
- "Other countries"
- ]
- },
- {
- "group": "Highest education level",
- "labels": [
- "Graduate",
- "Bachelor",
- "Associate",
- "High school",
- "Less than high school"
- ]
- },
- {
- "group": "Current-law initial AIME quintile",
- "labels": [
- "Highest",
- "Second highest",
- "Middle",
- "Second lowest",
- "Lowest"
- ]
- },
- {
- "group": "Lifetime payroll tax quintile",
- "labels": [
- "Highest",
- "Second highest",
- "Middle",
- "Second lowest",
- "Lowest"
- ]
- },
- {
- "group": "Lifetime payroll tax quintile (shared)",
- "labels": [
- "Highest",
- "Second highest",
- "Middle",
- "Second lowest",
- "Lowest"
- ]
- }
- ]
- },
- "14": {
- "caption": "Projected Effects of Proposal on Benefit/Tax Ratios POPULATION: Workers born 1980–1989 with a benefit/tax ratio (characteristics)",
- "columns": [
- "Characteristic",
- "Percent of population with a—",
- "Percent change in benefit/tax ratio at the—",
- "Benefit/tax ratio under current law at the—",
- "Benefit/tax ratio with proposal at the—",
- "Ratio decrease",
- "Ratio increase",
- "10th %ile",
- "Median",
- "90th %ile",
- "10th %ile",
- "Median",
- "90th %ile",
- "10th %ile",
- "Median",
- "90th %ile"
- ],
- "groups": [
- {
- "group": "Total",
- "labels": []
- },
- {
- "group": "Sex",
- "labels": [
- "Female",
- "Male"
- ]
- },
- {
- "group": "Race and ethnicity",
- "labels": [
- "Hispanic or Latino, any race",
- "White, non-Hispanic",
- "Black or African American, non-Hispanic",
- "All other races, non-Hispanic"
- ]
- },
- {
- "group": "Country of birth",
- "labels": [
- "United States",
- "Other countries"
- ]
- },
- {
- "group": "Highest education level",
- "labels": [
- "Graduate",
- "Bachelor",
- "Associate",
- "High school",
- "Less than high school"
- ]
- },
- {
- "group": "Current-law initial AIME quintile",
- "labels": [
- "Highest",
- "Second highest",
- "Middle",
- "Second lowest",
- "Lowest"
- ]
- },
- {
- "group": "Lifetime payroll tax quintile",
- "labels": [
- "Highest",
- "Second highest",
- "Middle",
- "Second lowest",
- "Lowest"
- ]
- },
- {
- "group": "Lifetime payroll tax quintile (shared)",
- "labels": [
- "Highest",
- "Second highest",
- "Middle",
- "Second lowest",
- "Lowest"
- ]
- }
- ]
- },
- "15": {
- "caption": "Projected Effects of Proposal on Benefit/Tax Ratios POPULATION: Workers born 2000–2009 with a benefit/tax ratio (characteristics)",
- "columns": [
- "Characteristic",
- "Percent of population with a—",
- "Percent change in benefit/tax ratio at the—",
- "Benefit/tax ratio under current law at the—",
- "Benefit/tax ratio with proposal at the—",
- "Ratio decrease",
- "Ratio increase",
- "10th %ile",
- "Median",
- "90th %ile",
- "10th %ile",
- "Median",
- "90th %ile",
- "10th %ile",
- "Median",
- "90th %ile"
- ],
- "groups": [
- {
- "group": "Total",
- "labels": []
- },
- {
- "group": "Sex",
- "labels": [
- "Female",
- "Male"
- ]
- },
- {
- "group": "Race and ethnicity",
- "labels": [
- "Hispanic or Latino, any race",
- "White, non-Hispanic",
- "Black or African American, non-Hispanic",
- "All other races, non-Hispanic"
- ]
- },
- {
- "group": "Country of birth",
- "labels": [
- "United States",
- "Other countries"
- ]
- },
- {
- "group": "Highest education level",
- "labels": [
- "Graduate",
- "Bachelor",
- "Associate",
- "High school",
- "Less than high school"
- ]
- },
- {
- "group": "Current-law initial AIME quintile",
- "labels": [
- "Highest",
- "Second highest",
- "Middle",
- "Second lowest",
- "Lowest"
- ]
- },
- {
- "group": "Lifetime payroll tax quintile",
- "labels": [
- "Highest",
- "Second highest",
- "Middle",
- "Second lowest",
- "Lowest"
- ]
- },
- {
- "group": "Lifetime payroll tax quintile (shared)",
- "labels": [
- "Highest",
- "Second highest",
- "Middle",
- "Second lowest",
- "Lowest"
- ]
- }
- ]
- },
- "16": {
- "caption": "Projected Effects of Proposal on Benefit/Tax Ratios POPULATION: Workers born 2020–2029 with a benefit/tax ratio (characteristics)",
- "columns": [
- "Characteristic",
- "Percent of population with a—",
- "Percent change in benefit/tax ratio at the—",
- "Benefit/tax ratio under current law at the—",
- "Benefit/tax ratio with proposal at the—",
- "Ratio decrease",
- "Ratio increase",
- "10th %ile",
- "Median",
- "90th %ile",
- "10th %ile",
- "Median",
- "90th %ile",
- "10th %ile",
- "Median",
- "90th %ile"
- ],
- "groups": [
- {
- "group": "Total",
- "labels": []
- },
- {
- "group": "Sex",
- "labels": [
- "Female",
- "Male"
- ]
- },
- {
- "group": "Race and ethnicity",
- "labels": [
- "Hispanic or Latino, any race",
- "White, non-Hispanic",
- "Black or African American, non-Hispanic",
- "All other races, non-Hispanic"
- ]
- },
- {
- "group": "Country of birth",
- "labels": [
- "United States",
- "Other countries"
- ]
- },
- {
- "group": "Highest education level",
- "labels": [
- "Graduate",
- "Bachelor",
- "Associate",
- "High school",
- "Less than high school"
- ]
- },
- {
- "group": "Current-law initial AIME quintile",
- "labels": [
- "Highest",
- "Second highest",
- "Middle",
- "Second lowest",
- "Lowest"
- ]
- },
- {
- "group": "Lifetime payroll tax quintile",
- "labels": [
- "Highest",
- "Second highest",
- "Middle",
- "Second lowest",
- "Lowest"
- ]
- },
- {
- "group": "Lifetime payroll tax quintile (shared)",
- "labels": [
- "Highest",
- "Second highest",
- "Middle",
- "Second lowest",
- "Lowest"
- ]
- }
- ]
- },
- "17": {
- "caption": "Projected Effects of Proposal on Initial Replacement Rates POPULATION: Current-law beneficiaries born 1960–1969 with a replacement rate (characteristics)",
- "columns": [
- "Characteristic",
- "Percent of population with a—",
- "Percent change in initial replacement rate at the—",
- "Initial replacement rate under current law at the—",
- "Initial replacement rate with proposal at the—",
- "Rate decrease",
- "Rate increase",
- "10th %ile",
- "Median",
- "90th %ile",
- "10th %ile",
- "Median",
- "90th %ile",
- "10th %ile",
- "Median",
- "90th %ile"
- ],
- "groups": [
- {
- "group": "Total",
- "labels": []
- },
- {
- "group": "Sex",
- "labels": [
- "Female",
- "Male"
- ]
- },
- {
- "group": "Race and ethnicity",
- "labels": [
- "Hispanic or Latino, any race",
- "White, non-Hispanic",
- "Black or African American, non-Hispanic",
- "All other races, non-Hispanic"
- ]
- },
- {
- "group": "Country of birth",
- "labels": [
- "United States",
- "Other countries"
- ]
- },
- {
- "group": "Highest education level",
- "labels": [
- "Graduate",
- "Bachelor",
- "Associate",
- "High school",
- "Less than high school"
- ]
- },
- {
- "group": "Current-law initial AIME quintile",
- "labels": [
- "Highest",
- "Second highest",
- "Middle",
- "Second lowest",
- "Lowest"
- ]
- },
- {
- "group": "Lifetime payroll tax quintile",
- "labels": [
- "Highest",
- "Second highest",
- "Middle",
- "Second lowest",
- "Lowest"
- ]
- },
- {
- "group": "Lifetime payroll tax quintile (shared)",
- "labels": [
- "Highest",
- "Second highest",
- "Middle",
- "Second lowest",
- "Lowest"
- ]
- }
- ]
- },
- "18": {
- "caption": "Projected Effects of Proposal on Initial Replacement Rates POPULATION: Current-law beneficiaries born 1980–1989 with a replacement rate (characteristics)",
- "columns": [
- "Characteristic",
- "Percent of population with a—",
- "Percent change in initial replacement rate at the—",
- "Initial replacement rate under current law at the—",
- "Initial replacement rate with proposal at the—",
- "Rate decrease",
- "Rate increase",
- "10th %ile",
- "Median",
- "90th %ile",
- "10th %ile",
- "Median",
- "90th %ile",
- "10th %ile",
- "Median",
- "90th %ile"
- ],
- "groups": [
- {
- "group": "Total",
- "labels": []
- },
- {
- "group": "Sex",
- "labels": [
- "Female",
- "Male"
- ]
- },
- {
- "group": "Race and ethnicity",
- "labels": [
- "Hispanic or Latino, any race",
- "White, non-Hispanic",
- "Black or African American, non-Hispanic",
- "All other races, non-Hispanic"
- ]
- },
- {
- "group": "Country of birth",
- "labels": [
- "United States",
- "Other countries"
- ]
- },
- {
- "group": "Highest education level",
- "labels": [
- "Graduate",
- "Bachelor",
- "Associate",
- "High school",
- "Less than high school"
- ]
- },
- {
- "group": "Current-law initial AIME quintile",
- "labels": [
- "Highest",
- "Second highest",
- "Middle",
- "Second lowest",
- "Lowest"
- ]
- },
- {
- "group": "Lifetime payroll tax quintile",
- "labels": [
- "Highest",
- "Second highest",
- "Middle",
- "Second lowest",
- "Lowest"
- ]
- },
- {
- "group": "Lifetime payroll tax quintile (shared)",
- "labels": [
- "Highest",
- "Second highest",
- "Middle",
- "Second lowest",
- "Lowest"
- ]
- }
- ]
- },
- "19": {
- "caption": "Projected Effects of Proposal on Initial Replacement Rates POPULATION: Current-law beneficiaries born 2000–2009 with a replacement rate (characteristics)",
- "columns": [
- "Characteristic",
- "Percent of population with a—",
- "Percent change in initial replacement rate at the—",
- "Initial replacement rate under current law at the—",
- "Initial replacement rate with proposal at the—",
- "Rate decrease",
- "Rate increase",
- "10th %ile",
- "Median",
- "90th %ile",
- "10th %ile",
- "Median",
- "90th %ile",
- "10th %ile",
- "Median",
- "90th %ile"
- ],
- "groups": [
- {
- "group": "Total",
- "labels": []
- },
- {
- "group": "Sex",
- "labels": [
- "Female",
- "Male"
- ]
- },
- {
- "group": "Race and ethnicity",
- "labels": [
- "Hispanic or Latino, any race",
- "White, non-Hispanic",
- "Black or African American, non-Hispanic",
- "All other races, non-Hispanic"
- ]
- },
- {
- "group": "Country of birth",
- "labels": [
- "United States",
- "Other countries"
- ]
- },
- {
- "group": "Highest education level",
- "labels": [
- "Graduate",
- "Bachelor",
- "Associate",
- "High school",
- "Less than high school"
- ]
- },
- {
- "group": "Current-law initial AIME quintile",
- "labels": [
- "Highest",
- "Second highest",
- "Middle",
- "Second lowest",
- "Lowest"
- ]
- },
- {
- "group": "Lifetime payroll tax quintile",
- "labels": [
- "Highest",
- "Second highest",
- "Middle",
- "Second lowest",
- "Lowest"
- ]
- },
- {
- "group": "Lifetime payroll tax quintile (shared)",
- "labels": [
- "Highest",
- "Second highest",
- "Middle",
- "Second lowest",
- "Lowest"
- ]
- }
- ]
- },
- "20": {
- "caption": "Projected Effects of Proposal on Initial Replacement Rates POPULATION: Current-law beneficiaries born 2020–2029 with a replacement rate (characteristics)",
- "columns": [
- "Characteristic",
- "Percent of population with a—",
- "Percent change in initial replacement rate at the—",
- "Initial replacement rate under current law at the—",
- "Initial replacement rate with proposal at the—",
- "Rate decrease",
- "Rate increase",
- "10th %ile",
- "Median",
- "90th %ile",
- "10th %ile",
- "Median",
- "90th %ile",
- "10th %ile",
- "Median",
- "90th %ile"
- ],
- "groups": [
- {
- "group": "Total",
- "labels": []
- },
- {
- "group": "Sex",
- "labels": [
- "Female",
- "Male"
- ]
- },
- {
- "group": "Race and ethnicity",
- "labels": [
- "Hispanic or Latino, any race",
- "White, non-Hispanic",
- "Black or African American, non-Hispanic",
- "All other races, non-Hispanic"
- ]
- },
- {
- "group": "Country of birth",
- "labels": [
- "United States",
- "Other countries"
- ]
- },
- {
- "group": "Highest education level",
- "labels": [
- "Graduate",
- "Bachelor",
- "Associate",
- "High school",
- "Less than high school"
- ]
- },
- {
- "group": "Current-law initial AIME quintile",
- "labels": [
- "Highest",
- "Second highest",
- "Middle",
- "Second lowest",
- "Lowest"
- ]
- },
- {
- "group": "Lifetime payroll tax quintile",
- "labels": [
- "Highest",
- "Second highest",
- "Middle",
- "Second lowest",
- "Lowest"
- ]
- },
- {
- "group": "Lifetime payroll tax quintile (shared)",
- "labels": [
- "Highest",
- "Second highest",
- "Middle",
- "Second lowest",
- "Lowest"
- ]
- }
- ]
- }
- }
-}
\ No newline at end of file
diff --git a/data/external/mint8_table_user_guide.source.html b/data/external/mint8_table_user_guide.source.html
index 889f4d81..e4348ad1 100644
--- a/data/external/mint8_table_user_guide.source.html
+++ b/data/external/mint8_table_user_guide.source.html
@@ -5,11 +5,11 @@
Table User Guide - Modeling Income in the Near Term (MINT) 8
-
+
-
+
-
+
@@ -19,13 +19,13 @@
-
-
+
+