You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
US default has 3.55M more children than PEP V2024: the 15–19 band fits, but ages 15–16 are 40% high and 18–19 are 40% low #1070
The certified US default populace-us-2024-spm-20260915 has 76.68M children under 18 in 2024, against 73.13M in PEP V2024 (+3.55M, +4.9%). #880 first reported this level (76.7M against about 72–73M); this issue locates it. The release meets all 18 national PEP age-band targets to within 0.05%. Almost all of the excess sits at ages 15–17, offset by missing 18–19-year-olds inside the 15–19 target band.
No population target sees 18. Population targets are PEP 5-year bands, so nothing separates 15–17 from 18–19 inside the 15–19 band, and no row targets the under-18 total. The only rows bounded at 18 are two SSI recipient counts (under 18 and 18–64).
The support does not force the excess. On the same 57,240 households, a feasibility linear program found weights that keep every published PEP fit (national and state), the household total and the 5× cap, and give exactly 73,132,720 children (details below).
The split starts in the initial weights and widens in calibration. The initial weights carry 81.56M children, with ages 15–16 already +0.91M and 18–19 −2.02M against PEP. Calibration widens these to +3.59M and −3.54M, while adding a net 0.82M to meet the 15–19 band.
In sensitivity re-solves, the move from 18–19 to 15–16 appears only with SOI tax targets. PEP targets alone land at 74.75M children and raise 18–19 (6.92M → 7.34M). Adding SOI families together raises the landing point: CTC and ACTC to 75.84M, CTC/ACTC plus return counts to 76.11M, and taxable income, income tax before credits and income tax together to 76.60M. Any one family alone (CTC, ACTC, return counts, or any single SOI income or tax variable) lands at 74.7–75.1M.
The household total (122.54M) is fixed before calibration. One rescale to a rounded 334.2M person benchmark sets it, and mass="conserve" then preserves it.
Downstream, policyengine.py 6.1.1 applies policyengine-us 2.2.1's household_weight uprating, which multiplies every 2024 weight by 1.016492 with no aging (6.1.2 and 6.2.0 pin the same file). That gives 77.95M children in 2026, against 72.02M in Census V2025 (July 2025).
Artifact
Release populace-us-2024-spm-20260915: populace_us_2024.h5 (sha256 6496cc43…) and populace_us_2024_calibration.npz (sha256 b6e9c067…).
Its household weights (initial and final) are identical to parent Build P, populace-us-2024-buildp-sparse-rmloss100-cae8640-20260728T011454Z. The SPM-role enrichment only appends a role column. So the calibration behaviour below is cae8640's.
Code citations are at cae8640f9e65e274aea65c7916cb37b956978e32. bfr = tools/build_us_fiscal_refresh_release.py; us_runtime/ = packages/populace-build/src/populace/build/us_runtime/; solve.py = packages/populace-calibrate/src/populace/calibrate/solve.py. On main (3601f64) these modules moved to packages/microcosm-*. The cited logic is unchanged, except that since Exact-k ladder launcher + release plumbing: pool → k ∈ {N, 57,240, 20,000} datasets (#578 inc 3) #607 a --pool-manifest build skips the base-population rescale in item 4; this release went through the --base-h5 path, which applies it.
Reproduction
This takes a few seconds with pandas, tables, numpy and huggingface_hub; no engine run is needed. The default path downloads the 827 MB h5.
repro_counts.py
"""Reproduce the child, age and household counts of populace_us_2024 (populace-us-2024-spm-20260915).Needs only pandas, tables, numpy and huggingface_hub; no engine run. python repro_counts.py # downloads the pinned release files from Hugging Face python repro_counts.py --h5 path/to/populace_us_2024.h5 --npz path/to/populace_us_2024_calibration.npzThe 2026 figures apply policyengine-us 2.2.1's uprating of household_weight, whichmultiplies every 2024 weight by the same factor. Pass --factor to change it; thedefault is the measured ratio of the 2026 file written by policyengine.py 6.1.1to the 2024 base."""importargparseimportnumpyasnpimportpandasaspdREPO, REV="policyengine/populace-us", "populace-us-2024-spm-20260915"ap=argparse.ArgumentParser()
ap.add_argument("--h5")
ap.add_argument("--npz")
ap.add_argument("--factor", type=float, default=1.016492)
a=ap.parse_args()
ifnot (a.h5anda.npz):
fromhuggingface_hubimporthf_hub_downloadget=lambdaf: hf_hub_download(REPO, f, repo_type="dataset", revision=REV) # noqa: E731a.h5=a.h5orget("populace_us_2024.h5")
a.npz=a.npzorget("populace_us_2024_calibration.npz")
h=pd.read_hdf(a.h5, "household")
p=pd.read_hdf(a.h5, "person", columns=["person_household_id", "age", "SPM_POOR"])
z=np.load(a.npz, allow_pickle=True)
assertnp.allclose(z["household_weight"], h["household_weight"].values)
ids=h["household_id"].valuesage=p["age"].valuesfinal=p["person_household_id"].map(pd.Series(z["household_weight"], ids)).valuesinit=p["person_household_id"].map(pd.Series(z["initial_household_weight"], ids)).valuesdefm(x):
returnf"{x/1e6:8.3f}M"print(" initial final final x factor (2026)")
forlabel, maskin [
("persons", np.ones_like(age, bool)),
("children < 18", age<18),
("ages 0-14", age<15),
("ages 15-17", (age>=15) & (age<18)),
("ages 18-19", (age>=18) & (age<20)),
("ages 15-19 (a target band)", (age>=15) & (age<20)),
]:
print(f"{label:27s}{m(init[mask].sum())}{m(final[mask].sum())}{m(final[mask].sum() *a.factor)}")
print(f"{'households':27s}{m(z['initial_household_weight'].sum())}{m(z['household_weight'].sum())} "f"{m(z['household_weight'].sum() *a.factor)}")
print(f"persons per household {init.sum() /z['initial_household_weight'].sum():9.3f} "f"{final.sum() /z['household_weight'].sum():10.3f}")
print("\nsingle year of age, 2024 (initial -> final weights)")
foryrinrange(0, 25):
k=age==yrprint(f" age {yr:2d}: {m(init[k].sum())} -> {m(final[k].sum())}")
poor=p["SPM_POOR"].valuesforlabel, maskin [("children", age<18), ("all people", np.ones_like(age, bool))]:
rate= (final[mask] *poor[mask]).sum() /final[mask].sum()
print(f"Census SPM_POOR flag on the records, calibrated weights, {label}: {rate:.2%}")
initial final final x factor (2026)
persons 334.200M 340.077M 345.686M
children < 18 81.557M 76.683M 77.948M
ages 0-14 66.914M 59.707M 60.692M
ages 15-17 14.642M 16.976M 17.256M
ages 18-19 6.920M 5.404M 5.493M
ages 15-19 (a target band) 21.562M 22.380M 22.749M
households 122.537M 122.537M 124.558M
persons per household 2.727 2.775
single year of age, 2024 (initial -> final weights)
age 0: 3.394M -> 2.751M
age 1: 4.078M -> 3.859M
age 2: 4.304M -> 3.754M
age 3: 4.309M -> 3.880M
age 4: 4.605M -> 4.359M
age 5: 4.389M -> 4.095M
age 6: 4.546M -> 4.029M
age 7: 4.622M -> 3.769M
age 8: 4.786M -> 4.419M
age 9: 4.502M -> 3.888M
age 10: 4.702M -> 4.370M
age 11: 4.912M -> 4.689M
age 12: 4.494M -> 4.093M
age 13: 4.552M -> 3.916M
age 14: 4.719M -> 3.835M
age 15: 5.111M -> 6.241M
age 16: 4.691M -> 6.244M
age 17: 4.840M -> 4.491M
age 18: 4.151M -> 3.634M
age 19: 2.769M -> 1.770M
age 20: 3.886M -> 4.590M
age 21: 4.202M -> 4.587M
age 22: 4.066M -> 4.243M
age 23: 4.328M -> 4.663M
age 24: 4.431M -> 4.339M
Census SPM_POOR flag on the records, calibrated weights, children: 16.50%
Census SPM_POOR flag on the records, calibrated weights, all people: 14.03%
Model vs PEP V2024 (July 1, 2024), millions
Ages
Initial weights
Final weights
Benchmark
Final / benchmark
0
3.394
2.751
3.616
0.76
0–14
66.914
59.707
59.698
1.000
15
5.111
6.241
4.364
1.43
16
4.691
6.244
4.533
1.38
17
4.840
4.491
4.538
0.99
18
4.151
3.634
4.484
0.81
19
2.769
1.770
4.457
0.40
15–19 (a target band)
21.562
22.380
22.376
1.000
Under 18
81.557
76.683
73.133
1.049
Households
122.537
122.537
132.216 (CPS HH-1), 132.737 (ACS B11001)
0.927 / 0.923
The 18 national census_pep.cy2024.national_resident_population_age.* targets equal PEP V2024 (nc-est2024-agesex-res.csv) exactly, and their final errors are at most 0.049%.
The 918 state band targets fit less tightly: 379 miss by more than 0.05%, and the worst is DC ages 35–39 at −14.9%.
Mechanism
1. No population target sees 18.
Population rows are PEP 5-year bands, materialized with lower-inclusive, upper-exclusive ages on household state (bfr:3704-3735). No population row has a bound at 17 or 18.
The state rows come from sc-est2024-alldata6.csv, which has single years of age, so a state split at 18 needs no new source.
Feasibility. We solved a feasibility linear program on the published support, using the rebuilt matrix rows described in item 3. It holds the published fitted values of all 936 PEP rows (18 national, 918 state), the household mass (122,537,091.9) and the 5× initial-weight cap. It keeps every weight at or above 10⁻⁶ of its initial value and adds one constraint: children = PEP V2024. HiGHS finds it feasible. The result has 73,132,720 children, and the maximum change to any fitted PEP count is 1.6e-7 persons. We checked the witness separately against the h5. It is a vertex, not a usable weight set: 45,085 of the 57,240 households sit at the floor and 11,235 at the cap. It does not check the fiscal targets.
2. The starting weights are far off on children.
Initial household weights are ASEC HSUP_WGT × 0.01 (us_runtime/asec_pool.py:38) times one constant per source year: 0.9084 (2022), 0.9024 (2023) and 0.8884 (2024). The same constant applies to the asec records and to their puf_tax_detail clones (asec_pool.py:397-407).
Each constant combines three factors: the pool's per-year population scale (asec_pool.py:113-123), the ÷2 split across the two channels (us_runtime/puf_support.py:481-490) and the ×5.3304 repair in item 4. Before the repair, the selection carried 22.99M households.
The selected support is Build M's frozen selection, which is Build I with nine households substituted. It has 57,240 households, drawn from 47,277 distinct source households, out of 337,704 candidates.
At initial weights, those households average 2.727 persons, and 24.4% of those persons are children. The full ASEC pool averages 2.420 persons, with 21.4% children.
So the initial weights start at 81.56M children (24.4% of 334.2M). That is 8.42M over PEP, mostly at ages 0–14 (66.9M against 59.7M).
At the start, 15–17 is 1.21M high and 18–19 is 2.02M low.
Calibration pulls 0–14 back to its bands but raises 15–17 to 17.0M and cuts 18–19 to 5.4M. Only 15–19 as a whole is targeted, so no population row resists that shift.
The solver starts from these weights. Warm start is disabled in the release diagnostics, so the solver starts at log(w0) (solve.py:676-677, :740). These weights also set the 5× cap and the conserved total (solve.py:736, :849, :861). No re-solve below starts from other weights.
3. In re-solves, SOI tax-unit targets move the within-band split.
We rebuilt the 2024 target matrix with policyengine-us 2.2.1 (the build used 1.764.6). We did not rebuild the 11 JCT tax-expenditure rows, because each needs its own reform simulation. They are left empty, and the re-solves give them zero weight. The reconstruction is close but not exact:
loss 0.3489 → 0.0209 with the empty JCT rows counted at full error, against the recorded 0.3479 → 0.0177;
254 of the 5,648 non-JCT rows differ from the build's final estimates by more than 1%.
We then re-solved it with the build's loss, epochs, learning rate, cap and mass="conserve", using a NumPy reimplementation of the Adam loop. Each ablation zeroes the loss weights of the dropped rows and keeps the rest as built; it does not rebuild the target registry. Treat the results as sensitivities of the solver's landing point, not an additive decomposition (millions):
Retained targets (JCT excluded throughout)
15–16
18–19
Under 18
Default weights
12.48
5.40
76.68
Full re-solve
12.34
5.56
76.52
PEP only
10.10
7.34
74.75
PEP + CTC/ACTC
11.89
6.22
75.84
PEP + CTC/ACTC + return counts
12.08
6.01
76.11
PEP + taxable income, income tax before credits, income tax
13.01
5.50
76.60
PEP + all SOI
12.36
5.70
76.37
Full minus CTC/ACTC, taxable income, income taxes, return counts
10.00
7.46
74.61
PEP V2024
8.90
8.94
73.13
Added to PEP, taxable income, income tax before credits and income tax together move the landing point furthest (76.60M, with 15–16 at 13.01M). CTC/ACTC comes next (75.84M), and return counts add 0.27M on top of CTC/ACTC. We have not traced why those three rows move 15–16 most.
Household weight mass, relative to the initial weights, moves the same way:
households with a 15/16-year-old: ×1.04 with PEP only, ×1.22 with PEP + CTC/ACTC + return counts, ×1.27 at the published weights;
households with an 18/19-year-old who is not a tax-unit dependent (policyengine-us 2.2.1 flags): ×1.01, ×0.56 and ×0.35.
Two families have materializer details that bear on age:
CTC and ACTC (national plus 51 states, claims and amounts).
The CTC rows use max(min(ctc, ctc_limiting_tax_liability), 0) (bfr:4163-4175).
The ACTC rows use refundable_ctc (us_runtime/fiscal_targets.py:301) through the generic tax-unit path (bfr:4184-4189), floored at 0 (bfr:3776-3790).
Claims rows count tax units in the slice with a positive value (bfr:3797-3798).
The national claims row is an unaged TY2022 count; the amount row is a TY2022 amount aged ×1.120. In the build diagnostics, claims go from 35.33M to 37.97M against a 38.07M target, and the amount goes from $82.0B to $89.2B against $92.8B.
The qualifying-child test is under 17 (dependency and SSN tests also apply), and ages 0–14 are pinned by their bands. So within the 15–19 band, CTC per head is highest at 15–16.
Age 17 is not worth zero to these rows. ctc includes the $500 credit for other dependents (policyengine-us ctc_individual_maximum adds ctc_adult_individual_maximum). So a 17-year-old dependent adds $500 and can put its tax unit on the claims row.
Even so, 17 is not inflated: 4.84M initial and 4.49M final, against 4.54M in PEP. Adding CTC/ACTC to a PEP-only re-solve lowers 17 from 4.94M to 4.26M and raises 15–16 from 10.10M to 11.89M.
Return counts (source_variable == "count").
These are mask.astype(np.float64) over every tax unit in the AGI slice (bfr:4205-4254).
That constant sums rounded bands in us_runtime/demographics.py:60-77, which the code labels Vintage 2023 and describes as "approximate round figures for a directional benchmark".
The repair's reason string calls it "the Census 2024 national person-population benchmark" (bfr:394-396).
Calibration then runs with mass="conserve": bfr:9036 on the dense path Build P used, :9063 on the L0 path. The solver rescales the weights to the input total every step and once more at the end (solve.py:823-835, :851-861).
So households = 334.2M / 2.7273 = 122.54M, set by the support's household size before any target is seen.
Persons then rise to 340.08M under the PEP targets while households stay fixed, giving 2.775 residents per household. PEP's 340.11M residents over ACS's 132.74M households give 2.56. ACS B25010's 2.50 covers household population only (its universe is occupied housing units).
No household-count target exists; the SNAP and LIHEAP rows count benefit units. The model has 122.54M households, against 132.74M in ACS B11001 and 132.22M in CPS HH-1.
Effect on a CTC reform (policyengine.py 6.1.1, policyengine-us 2.2.1, tax year 2026)
The example raises the CTC base amount from $2,200 to $3,000 for 2026 (static, against a current-law baseline). On this release it costs $31.19B, 19.2% of households gain, and 189,306 children leave SPM poverty. We have not produced a corrected estimate; that needs the same run on a recalibrated release. These diagnostics show how much of those figures rests on the population issues above:
The excess teens are CTC-age. The 2026 file has 12.69M children aged 15–16, against 8.62M in Census V2025 (July 2025). If each tax unit's gain is split equally across its qualifying children, 15–16-year-olds carry 19.3% of the cost.
The share gaining rests on households with a 0–16-year-old. Every tax unit whose income tax changes has a CTC-qualifying child. 31.9% of model households contain a 0–16-year-old, against 25.9% in the CPS ASEC 2026 file (137.1M households).
The children-lifted count is concentrated. It comes from 40 SPM units, and two of them carry half of it, so it moves with a handful of household weights.
A corrected release can be scored by rerunning the reform. The per-record values behind these figures do not depend on the weights. Rerunning with random per-household weight factors (0.2–5) left every income-tax, CTC, SPM and net-income output column bitwise identical. Only the income-decile rank and the Medicaid cost columns moved. The decile rank enters none of these figures, and the reform leaves the Medicaid cost columns unchanged.
Adding the boundary costs little elsewhere in calibration. We re-solved the rebuilt matrix (all 5,648 non-JCT targets) with PEP V2024 national rows added. These are calibration diagnostics, not fiscal figures. The re-solve does not reproduce the published weights (76.52M children with no added rows, against 76.68M):
Added rows
Children
15–16 / 17
Loss on the original surface
Targets within 1%
SOI CTC claims
SOI CTC amount
Default weights (reference)
76.68M
12.48M / 4.49M
0.0182
4,738
−0.2%
−3.7%
None
76.52M
12.34M / 4.48M
0.0162
5,005
−0.3%
−3.2%
15–17 and 18–19
73.14M
10.18M / 3.25M
0.0163
4,975
−0.3%
−5.5%
Single years 0–19
73.14M
8.90M / 4.54M
0.0164
4,932
−0.3%
−7.2%
PEP V2024
73.13M
8.90M / 4.54M
The split alone leaves 15–16 14% high and 17 28% low, across the CTC age line; single years fix both. Among the tracked rows, the tension is the CTC amount: with single-year PEP ages, dollars per claim fall further short of SOI. The claims target is an unaged TY2022 count, while the amount target is aged ×1.120. It is worth checking whether that shortfall is a per-child microdata issue or a target-vintage one.
The same records carry Census's own SPM_POOR flag, which lets the child rate be traced step by step. This is descriptive accounting, not a causal split.
Official 2022–24 rates at the selected people's source-year mix (MARSUPWT)
13.17%
12.76%
Distinct selected source people (MARSUPWT)
14.22%
11.41%
All retained records (MARSUPWT; 57,240 households from 47,277 source households, so some appear in both the asec and puf_tax_detail channels)
15.16%
11.96%
Initial weights
15.06%
11.90%
Final weights
16.50%
14.03%
policyengine-us 2.2.1 SPM, 2024
15.71%
13.02%
policyengine-us 2.2.1 SPM, 2026
17.42%
13.63%
Most of the child-specific elevation is present before the engine computes poverty.
As Pooled prior-year ASEC wages are not aged to the target year #944 found, wages are not aged. On the ASEC channel, employment_income_before_lsr is filled from the source year's WSAL_VAL (us_runtime/cps_carried.py:107; at main 3601f64, packages/microcosm-build/src/microcosm/build/us_runtime/cps_carried.py:172). This covers all 66,001 ASEC-channel persons.
puf_tax_detail wages also stay at source-year scale: the weighted model/raw WSAL_VAL ratio is 1.000 in each source year.
Rebuilding with 2022/2023-source wages aged to 2024 would show how much of the child-rate elevation this explains.
The annual builder uses anchor="frame", which multiplies each 2024 age cell by its projected growth. So it carries the 2024 child excess forward and does not replace fix 1 below.
Either split PEP 15–19 into 15–17 and 18–19, nationally and by state, or use national single years of age as policyengine-us-data's build_loss_matrix does.
Single years are the safer option: in the re-solves above, the split alone left 15–16 14% high and 17 28% low.
The feasibility LP (national under-18) and the re-solves (national 15–17/18–19 and single years 0–19) show that the national population side is satisfiable on the current support. State splits were not tested; their source file already has single years.
Under mass="conserve", a household-count target cannot move the total, because the solver projects household weights back to the repaired 122.54M. So the rescale has to change along with adding the targets (ACS B11001 132.74M and B25009 size shares). The rescale's benchmark is itself a rounded V2023 constant.
The person margin is PEP residents, while ACS household sizes exclude residents outside households: roughly 8M, since 340.11M − 2.50 × 132.74M ≈ 8.3M. Those two margins need reconciling.
A first try on the rebuilt matrix (ACS size rows plus a conserved 132.74M total) had not converged at 6,000 epochs (loss about 0.15 against 0.016) and pushed children to 83.4–84.5M. So this needs more than new rows.
Census V2025 has under-18 falling 0.73% from 2024 to 2025. The uniform factor (policyengine-us's CBO-based population series) adds 1.65% over two years.
Annual cuts use anchor="frame", so they carry the 2024 child excess forward unless fix 1 lands first.
Analysis scripts and outputs are available on request. They cover the reconstructed matrix, the re-solves, the feasibility LP, the SPM accounting and the per-record 2026 rerun.
Summary
The certified US default
populace-us-2024-spm-20260915has 76.68M children under 18 in 2024, against 73.13M in PEP V2024 (+3.55M, +4.9%). #880 first reported this level (76.7M against about 72–73M); this issue locates it. The release meets all 18 national PEP age-band targets to within 0.05%. Almost all of the excess sits at ages 15–17, offset by missing 18–19-year-olds inside the 15–19 target band.mass="conserve"then preserves it.Downstream, policyengine.py 6.1.1 applies policyengine-us 2.2.1's
household_weightuprating, which multiplies every 2024 weight by 1.016492 with no aging (6.1.2 and 6.2.0 pin the same file). That gives 77.95M children in 2026, against 72.02M in Census V2025 (July 2025).Artifact
populace-us-2024-spm-20260915:populace_us_2024.h5(sha2566496cc43…) andpopulace_us_2024_calibration.npz(sha256b6e9c067…).populace-us-2024-buildp-sparse-rmloss100-cae8640-20260728T011454Z. The SPM-role enrichment only appends a role column. So the calibration behaviour below iscae8640's.cae8640f9e65e274aea65c7916cb37b956978e32.bfr=tools/build_us_fiscal_refresh_release.py;us_runtime/=packages/populace-build/src/populace/build/us_runtime/;solve.py=packages/populace-calibrate/src/populace/calibrate/solve.py. On main (3601f64) these modules moved topackages/microcosm-*. The cited logic is unchanged, except that since Exact-k ladder launcher + release plumbing: pool → k ∈ {N, 57,240, 20,000} datasets (#578 inc 3) #607 a--pool-manifestbuild skips the base-population rescale in item 4; this release went through the--base-h5path, which applies it.Reproduction
This takes a few seconds with pandas, tables, numpy and huggingface_hub; no engine run is needed. The default path downloads the 827 MB h5.
repro_counts.pyModel vs PEP V2024 (July 1, 2024), millions
census_pep.cy2024.national_resident_population_age.*targets equal PEP V2024 (nc-est2024-agesex-res.csv) exactly, and their final errors are at most 0.049%.Mechanism
1. No population target sees 18.
bfr:3704-3735). No population row has a bound at 17 or 18.sc-est2024-alldata6.csv, which has single years of age, so a state split at 18 needs no new source.Feasibility. We solved a feasibility linear program on the published support, using the rebuilt matrix rows described in item 3. It holds the published fitted values of all 936 PEP rows (18 national, 918 state), the household mass (122,537,091.9) and the 5× initial-weight cap. It keeps every weight at or above 10⁻⁶ of its initial value and adds one constraint: children = PEP V2024. HiGHS finds it feasible. The result has 73,132,720 children, and the maximum change to any fitted PEP count is 1.6e-7 persons. We checked the witness separately against the h5. It is a vertex, not a usable weight set: 45,085 of the 57,240 households sit at the floor and 11,235 at the cap. It does not check the fiscal targets.
2. The starting weights are far off on children.
HSUP_WGT× 0.01 (us_runtime/asec_pool.py:38) times one constant per source year: 0.9084 (2022), 0.9024 (2023) and 0.8884 (2024). The same constant applies to the asec records and to their puf_tax_detail clones (asec_pool.py:397-407).asec_pool.py:113-123), the ÷2 split across the two channels (us_runtime/puf_support.py:481-490) and the ×5.3304 repair in item 4. Before the repair, the selection carried 22.99M households.log(w0)(solve.py:676-677,:740). These weights also set the 5× cap and the conserved total (solve.py:736,:849,:861). No re-solve below starts from other weights.3. In re-solves, SOI tax-unit targets move the within-band split.
We rebuilt the 2024 target matrix with policyengine-us 2.2.1 (the build used 1.764.6). We did not rebuild the 11 JCT tax-expenditure rows, because each needs its own reform simulation. They are left empty, and the re-solves give them zero weight. The reconstruction is close but not exact:
We then re-solved it with the build's loss, epochs, learning rate, cap and
mass="conserve", using a NumPy reimplementation of the Adam loop. Each ablation zeroes the loss weights of the dropped rows and keeps the rest as built; it does not rebuild the target registry. Treat the results as sensitivities of the solver's landing point, not an additive decomposition (millions):Two families have materializer details that bear on age:
max(min(ctc, ctc_limiting_tax_liability), 0)(bfr:4163-4175).refundable_ctc(us_runtime/fiscal_targets.py:301) through the generic tax-unit path (bfr:4184-4189), floored at 0 (bfr:3776-3790).bfr:3797-3798).ctcincludes the $500 credit for other dependents (policyengine-usctc_individual_maximumaddsctc_adult_individual_maximum). So a 17-year-old dependent adds $500 and can put its tax unit on the claims row.source_variable == "count").mask.astype(np.float64)over every tax unit in the AGI slice (bfr:4205-4254).4. The household total is set before calibration.
_with_base_population_mass_repair(bfr:4952-4993) multiplies every base household weight by 5.3304, so persons equalUS_BASE_PERSON_POPULATION_BENCHMARK= 334.2M (bfr:392). This repair came from Decide how US fiscal refresh should repair underweighted base H5 mass #94 and Repair US fiscal refresh base mass #128.us_runtime/demographics.py:60-77, which the code labels Vintage 2023 and describes as "approximate round figures for a directional benchmark".bfr:394-396).mass="conserve":bfr:9036on the dense path Build P used,:9063on the L0 path. The solver rescales the weights to the input total every step and once more at the end (solve.py:823-835,:851-861).Effect on a CTC reform (policyengine.py 6.1.1, policyengine-us 2.2.1, tax year 2026)
The example raises the CTC base amount from $2,200 to $3,000 for 2026 (static, against a current-law baseline). On this release it costs $31.19B, 19.2% of households gain, and 189,306 children leave SPM poverty. We have not produced a corrected estimate; that needs the same run on a recalibrated release. These diagnostics show how much of those figures rests on the population issues above:
The excess teens are CTC-age. The 2026 file has 12.69M children aged 15–16, against 8.62M in Census V2025 (July 2025). If each tax unit's gain is split equally across its qualifying children, 15–16-year-olds carry 19.3% of the cost.
The share gaining rests on households with a 0–16-year-old. Every tax unit whose income tax changes has a CTC-qualifying child. 31.9% of model households contain a 0–16-year-old, against 25.9% in the CPS ASEC 2026 file (137.1M households).
The children-lifted count is concentrated. It comes from 40 SPM units, and two of them carry half of it, so it moves with a handful of household weights.
A corrected release can be scored by rerunning the reform. The per-record values behind these figures do not depend on the weights. Rerunning with random per-household weight factors (0.2–5) left every income-tax, CTC, SPM and net-income output column bitwise identical. Only the income-decile rank and the Medicaid cost columns moved. The decile rank enters none of these figures, and the reform leaves the Medicaid cost columns unchanged.
Adding the boundary costs little elsewhere in calibration. We re-solved the rebuilt matrix (all 5,648 non-JCT targets) with PEP V2024 national rows added. These are calibration diagnostics, not fiscal figures. The re-solve does not reproduce the published weights (76.52M children with no added rows, against 76.68M):
The split alone leaves 15–16 14% high and 17 28% low, across the CTC age line; single years fix both. Among the tracked rows, the tension is the CTC amount: with single-year PEP ages, dollars per claim fall further short of SOI. The claims target is an unaged TY2022 count, while the amount target is aged ×1.120. It is worth checking whether that shortfall is a per-child microdata issue or a target-vintage one.
Related issues and PRs
US certified default undercounts households (122.5M) and overcounts married couples (81.5M): no household-type or marital-status targets in the registry #880 (households 122.5M, too many married couples). US certified default undercounts households (122.5M) and overcounts married couples (81.5M): no household-type or marital-status targets in the registry #880 already lists 76.7M children against about 72–73M; this issue locates that excess inside the 15–19 band. It also supplies the household mechanism: a total fixed by the pre-calibration rescale (Decide how US fiscal refresh should repair underweighted base H5 mass #94, Repair US fiscal refresh base mass #128) plus
mass="conserve". The child excess is an age-split problem, not only a household-type one.Child weights inflated ~46% in June 15 us_2024 build (0cdbb27) — breaks EITC/CTC, concentrated in EITC plateau #64 (closed) suggested making the under-18 total an active calibration constraint. Its closing comment still showed under-18 at +3.5% (75.7M against about 73.1M), on an earlier build.
Diagnose the child-vs-total SPM composition anomaly (poverty stays a permanent holdout) #646 (child-vs-total SPM anomaly; poverty stays a permanent holdout), Pooled prior-year ASEC wages are not aged to the target year #944 (pooled prior-year wages not aged) and Age donor dollar values to the base year per record, not through weights (SCF, SIPP, PUF) #981 (per-record aging of donor values).
SPM_POORflag, which lets the child rate be traced step by step. This is descriptive accounting, not a causal split.would_claim_wic, Published US releases store would_claim_wic, which policyengine-us 2.x ignores, so WIC take-up is 100% under policyengine.py 6.1.1 #1026), so WIC take-up is 100% in those engine rows.employment_income_before_lsris filled from the source year'sWSAL_VAL(us_runtime/cps_carried.py:107; at main 3601f64,packages/microcosm-build/src/microcosm/build/us_runtime/cps_carried.py:172). This covers all 66,001 ASEC-channel persons.WSAL_VALratio is 1.000 in each source year.Long-term CPS projections (2026–2100) have no populace replacement — decide rebuild / inherit / retire before #204 closes #333 / Add demographic static aging with aggregate-consistent income factors #957 / Export annual US projections and validate immutable release cuts #963 (annual static aging; Add demographic static aging with aggregate-consistent income factors #957 and Export annual US projections and validate immutable release cuts #963 are merged PRs).
anchor="frame", which multiplies each 2024 age cell by its projected growth. So it carries the 2024 child excess forward and does not replace fix 1 below.populace_us_2024 under-1 geography: CA carries 21.7% of US infants (~1.9x its population share), MI 1.0% (a third of its share) #520 / Preserve the WIC holdout and diagnose the infant/child eligibility offset #648 (infants).
Epic: publish the US default from the graph-native build #956 (graph-native US default). The fixes below need to land on the build that will publish the next default.
Proposed fixes
build_loss_matrixdoes.mass="conserve", a household-count target cannot move the total, because the solver projects household weights back to the repaired 122.54M. So the rescale has to change along with adding the targets (ACS B11001 132.74M and B25009 size shares). The rescale's benchmark is itself a rounded V2023 constant.anchor="frame", so they carry the 2024 child excess forward unless fix 1 lands first.Analysis scripts and outputs are available on request. They cover the reconstructed matrix, the re-solves, the feasibility LP, the SPM accounting and the per-record 2026 rerun.
🤖 Generated with Claude Code