Skip to content

Rebase SPI income draws to the FRS survey year - #532

Draft
MaxGhenis wants to merge 7 commits into
mainfrom
spi-income-rebase
Draft

MaxGhenis wants to merge 7 commits into
mainfrom
spi-income-rebase

Conversation

@MaxGhenis

@MaxGhenis MaxGhenis commented Oct 3, 2026 •

Copy link
Copy Markdown
Contributor

What was wrong

datasets/imputations/income.py trains its income QRF on SPI 2022-23 (SPI_FISCAL_YEAR = 2022) and draws into the FRS 2024-25 dataset (CURRENT_FRS_RELEASE.survey_year = 2024). The draws are the SPI-synthetic copy's six incomes plus Gift Aid and qualifying-investment gifts, and the FRS rows' dividends. Nothing moved them from 2022-23 to 2024-25 pounds:

  • Second stage. The second-stage QRF (frs_only.py) trains on FRS respondents in 2024-25 pounds, but predicts each SPI row's benefits, pension contributions and tax-free savings from that row's 2022-23 incomes.
  • Stacking. stack_datasets then put 2022-23 and 2024-25 amounts side by side.
  • Calibration. create_datasets.py uprates the stacked dataset as a whole (uprate_dataset, 2024 → 2025 for calibration and back), so the gap survived into calibration, and the weights had to absorb it.

PolicyEngine/microcosm#879 fixed the same gap on the microcosm side.

The change

rebase_spi_draws(draws, year) multiplies each drawn column by its variable's index ratio from SPI_FISCAL_YEAR to year. impute_over_incomes calls it on the model's output, with year = dataset.time_period, before anything is written. That puts it before the second stage, before stacking, and on the FRS half's dividend draw.

Which index, and which way the residual errors point

The index is the one uprate_dataset applies: storage/uprating_factors.csv, read through uprate_values. The same table then carries the whole dataset 2024 → 2025, so each variable has one index source across the build. I compared it with the two other places an index could come from, at the locked policyengine-uk 2.93.0 (analysis/index_compare.out):

column 2022→2024, CSV (used) policyengine-uk at load (data/uprating_indices.yaml, yoy growth) variable uprating attribute
employment_income 1.1189 1.1164 1.1164 (employment_income_before_lsr)
self_employment_income 1.0734 1.0211 1.0211
savings_interest_income 1.0897 1.0924 2.2692 (ons.household_interest_income)
dividend_income 1.0897 1.0924 1.0924
private_pension_income 1.1026 1.1025 1.1025
property_income 1.0897 1.0924 1.0924
gift_aid, charitable_investment_gifts none none none
  • Five of six agree within 0.3% with what policyengine-uk applies to a dataset at load.
  • Self-employment. The CSV is a stale snapshot of OBR indices, so SPI self-employment income ends up about 5% higher relative to the engine's index (×1.073 vs ×1.021). The CSV's 2024 → 2025 column is off for most incomes too. That is a defect in uprate_dataset's table, not in this rebase, and it is a separate follow-up with its own build and impact. Once the table is regenerated, this rebase follows it automatically.
  • Savings interest. The variable attribute (household interest income, ×2.27, covering the period of the rate rise) is not what policyengine-uk applies to a dataset (per-capita GDP). microcosm#879 used the attribute. If household interest grew as the attribute says, the SPI rows' savings draws stay low relative to the 2024 FRS respondents. This PR follows the load path and the CSV, which agree.
  • Gift Aid and qualifying-investment gifts. There is no index anywhere in policyengine-uk (no attribute, nothing in the load-time YAML, no CSV row) and no calibration target, so they keep their SPI amounts, exactly as uprate_dataset and policyengine-uk keep them.
    • SPI_NOMINAL_IMPUTATIONS names them.
    • Any other drawn column without an index raises (KeyError) instead of silently staying nominal.

Draws, not training data

  • A control run (analysis/train_vs_draw_control.out, 10k SPI sample, 20k FRS people, microimpute 1.8.1) gives:
    • Refitting agrees with the original fit on every draw, to within 1e-9.
    • Fitting on training outputs ×2, which is exact in floating point, agrees with the original draws ×2 to the same tolerance.
    • At the real factors, which are inexact in floating point, rebasing the training outputs changes about 9% of employment draws beyond that tolerance.
  • What that shows. The difference is consistent with numerical sensitivity of the refit: rounding at the inexact factors breaking near-tied splits is the plausible explanation. I didn't inspect the splits. It is not a difference in method.
  • Rebasing the draws is exact. Each written value is exactly factor × the current model's draw, so build comparisons isolate the rebase.
  • The cached model stays valid. income_spi_2022_23.pkl is keyed on the SPI release and output list. Rebasing the training data would tie it to the FRS year as well.
  • Same choice as microcosm#879.

Invariants (property-tested, tests/test_spi_income_rebasing.py)

For every input:

  • equal years are the identity, bit for bit;
  • within a column the map is weakly increasing and keeps signs;
  • zero stays zero and only zero becomes zero;
  • penny amounts keep their ranks, ties included;
  • missing draws stay missing (Draw SPI incomes within earnings groups set by FRS employment status #529 returns NaN for rows outside every earnings group);
  • the factor is the table's index ratio, so rebasing composes (a→b→c = a→c) and inverts;
  • the nominal columns are untouched.

Hypothesis found one counterexample to composition: the subnormal float 2.2250738585e-313, rebased 2022→2020→2021, misses the 1e-12 relative bound by 2.4e-11. (Rounded to 2.2e-313 it passes.)

  • Subnormals carry too few significant bits to hold a 1e-12 relative bound through two multiplications.
  • Normal floats compose within 6e-16 relative over all 3,375 year triplets and all six indexed columns.
  • No money amount is subnormal, so the money strategy excludes subnormals; the reason is commented in the test.
  • Composition is a property checked within tolerance over the strategy's domain, not exact equality.

The differential test checks rebase_spi_draws against uprate_dataset on a dataset stamped with the SPI year. Both read the same table, so it proves they cover the same columns (indexed vs nominal) with the same arithmetic. It is not an independent check of the factor values.

The call-site tests check:

  • impute_over_incomes writes draws in the dataset's year, whether time_period is an int or a string;
  • the second stage receives rebased SPI draws;
  • the FRS half's dividends are rebased;
  • nothing else moves: not the FRS respondents the second stage trains on, not the FRS rows' undrawn incomes, and not an undrawn indexed money column (employee pension contributions).

The call-site fakes carry an employment status, so the same tests pass on main and on a merge with #529.

14 tests pass, and a mutation check kills all 15 mutants (analysis/mutation_check.out, analysis/mutate.py). Eleven change the function: no call, reversed direction, fixed year, year + 1, additive shift, an extra nominal column, skipping unknown columns, abs, rounding, NaN → 0, wrong SPI year. Four move the rebase: after the second stage, onto the whole copy, onto the FRS half's other incomes, onto the second stage's training set.

Evidence from real builds

All builds use production settings: 512 epochs, PE_UK_DATA_OA_CLONES=1, torch seed 0, the same cached target downloads, and the locked environment (policyengine-uk 2.93.0, microimpute 1.8.1).

The design is paired under two SPI draws:

draw main (b45c373) this PR
seed 0 base branch
prediction seed 43 main43 branch43

The seed-43 pair sets the income QRF's prediction seed to 43 through a build wrapper, with no code change. main43 reproduces the earlier placebo exactly (same weights and fit). The "redraw" column below is main43 − base, i.e. how much a different SPI draw alone moves things.

Rebuild noise is zero. A second base build of the same commit and seed gives identical aggregates.

The builds did what the code says (analysis/build_rebase_check*.out):

  • In both pairs, every one of the eight drawn columns on every SPI-synthetic person matches its partner times the CSV factor: 43,052 people under seed 0, 43,053 under seed 43.
    • The check tolerance is 1e-6 relative. Review r2 measured the largest discrepancy at 3.9e-16, i.e. floating-point rounding.
    • The two nominal columns match bit for bit.
  • On the 34,966 plain FRS people, every undrawn income is identical row by row. Only dividends move.
  • Person counts differ only inside the 270 CGT band-donor households, which are chosen by total income: 583 vs 613 people under seed 0, 632 vs 622 under seed 43. Everyone else is identical.

Impact

policyengine-uk main 00fb451d6 (2.112.1). The effect is branch − main under each draw. GBP bn; people in thousands (impact/paired_compare.out).

2025 base effect, seed 0 effect, seed 43 redraw
employee NI (class 1) 49.11 −0.82 −0.65 −0.10
employer NI (class 1) 149.70 −1.34 −0.89 +0.06
employment income 1,217.65 −4.31 −3.31 −1.96
of which on SPI-synthetic rows +41.65 +33.89
total income 1,800.75 −3.49 −2.42 −0.90
income tax 295.14 −1.25 +0.71 +1.84
capital gains 39.40 −0.88 −1.58 +5.26
Gift Aid 1.95 −0.23 −0.31 +0.28
universal credit 75.86 −0.31 +0.69 −1.08
household net income 1,709.24 −2.19 −4.93 +10.54
HBAI household net income 1,519.57 +1.58 −3.70 +7.55
people in absolute poverty, BHC 10,782 +89 +221 −584
people in absolute poverty, AHC 14,093 −15 +437 −157
people in relative poverty, AHC 16,360 +90 −86 +2
children in relative poverty, AHC 4,919 +57 +4 +75
Gini, equivalised HBAI income 0.3151 −0.0009 +0.0044 −0.0025

In 2026 the same pattern holds:

  • employee NI −0.80 and −0.65;
  • employer NI −1.39 and −0.92;
  • income tax −1.29 and +0.79;
  • relative AHC poverty +588k and −3k people (redraw +99k), and for children +294k and −1k (redraw +152k).

How to read it:

  • The rule. An effect counts as robust when it has the same sign under both draws and each effect is larger than the redraw's. I applied it mechanically to every metric in both years (impact/robust_classification.out).

  • Robust in both years:

    • employee NI falls £0.65–0.82bn;
    • employer NI falls £0.89–1.40bn;
    • employment income falls £3.3–4.5bn and total income £2.4–3.5bn.

    The weighted employment income on SPI-synthetic rows rises £34–43bn, and on FRS rows it falls £37–48bn. Each FRS row's amounts are unchanged, so that fall is reweighting.

  • Robust in one year only, small or borderline:

    • 2025: class 4 NI (−£42m / −£9m, redraw +£9m) and Pension Credit (+£15m / +£29m, redraw −£3m). Both change sign in 2026.
    • 2026: benefit units on UC (+60k / +80k, redraw +59k) and children in absolute AHC poverty (+23k / +131k, redraw −13k).
  • Not robust: these change sign between the draws, or move less than the redraw itself:

    • income tax and UC spending;
    • household and HBAI income;
    • capital gains and Gift Aid;
    • the Gini;
    • every other poverty measure. In particular, 2026 relative AHC poverty is +588k under one draw and −3k under the other, and that is not an effect of this PR I can state.
  • Two draws are still few. They show which effects survive a redraw; they don't bound the noise.

Calibration fit and weights

National fit (analysis/national_fit_paired.out): the 636 national targets, recomputed for each saved file (uprated to 2025 with uprate_dataset, the build's own target matrix and cached targets).

base branch main43 branch43
mean squared relative error, all targets 0.2697 0.2723 0.2789 0.2732
HMRC income-band targets (156), mean relative error 0.1335 0.1270 0.1129 0.1136
bands from £150k 0.2687 0.2537 0.2152 0.2160
HMRC bands within 10% 80.1% 83.3% 87.2% 85.9%

The changes go both ways between the draws and are smaller than the redraw's own. The rebase leaves the national fit essentially where it was.

Review r1 pointed to analysis/calibration_fit_vs_placebo.out, which compared the branch with the placebo. That comparison mixes the rebase with a different SPI draw, so the paired table replaces it.

Weights (analysis/weights_summary.out):

  • The SPI-synthetic share of household weight rises from 27.7% to 28.5% (seed 0) and from 28.1% to 29.5% (seed 43).
  • Effective sample size on SPI-synthetic households falls from 650 to 520 and from 727 to 599.
  • Overall effective sample size falls from 1,208 to 1,122 under seed 0 and is unchanged (1,184) under seed 43.

CI

CI's reduced build (TESTING=1, 32 epochs) of 3d82532 failed test_cgt_band_donors.py::test_built_total_gains: £236.1bn of above-AEA gains against HMRC's £65.9bn, just past the reduced-build bound.

That is consistent with reduced-build under-convergence:

It does not rule out a contribution from the rebase. The capital-gains imputation conditions donor selection and gain draws on total income bands, and donor composition does change between builds. The full-build bound guards large level errors; it does not show a zero CGT effect.

604cb02 takes #529's version of that test file byte for byte (from 88ef576 and d39371d), so whichever PR lands second merges cleanly. Under TESTING the total must lie within a factor of 6 of HMRC's; full builds keep the 50% bound.

Locally, the full suite on this branch's production build passes, except the subnormal property fixed in 080e05d.

With #529, #498 and the rank-preserving draw

The hypothesis dev dependency is a cherry-pick of #529's 313468b (the same lines as #514 and #524).

Follow-ups (not in this PR)

  • Regenerate storage/uprating_factors.csv from the locked policyengine-uk's load-time path, and add a test that the table matches the engine. Self-employment 2022→2024 is the stale entry this PR touches; 2024→2025 is off for most incomes.
  • policyengine-uk's savings index is inconsistent: savings_interest_income.uprating says household interest income, while uprating_indices.yaml uprates it by per-capita GDP.

axiom: n/a: uk-data input construction, no policy rule encoded.

Lands with the batched uk-data release (d833). Do not merge outside it.

🤖 Generated with Claude Code

MaxGhenis and others added 7 commits October 2, 2026 22:46
Same lines as #514 and #524, so whichever lands second merges cleanly.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The income QRF is trained on SPI 2022-23 amounts and draws into the FRS
2024-25 dataset, but nothing moved the draws to 2024-25 before the
second-stage FRS-only QRF conditioned on them and the halves were
stacked. `uprate_dataset` then uprates the whole stacked dataset by the
same factor, so the SPI rows stayed about two years of earnings and price
growth behind the FRS rows, and calibration had to make up the gap through
weights.

`rebase_spi_draws` multiplies each draw by its variable's index ratio in
storage/uprating_factors.csv (the index `uprate_dataset` applies), from
SPI_FISCAL_YEAR to the dataset's time_period, right after the model
predicts. That covers the SPI-synthetic copy's six incomes and the FRS
half's dividends. Gift Aid and qualifying-investment gifts have no index in
policyengine-uk and keep their SPI amounts, as `uprate_dataset` keeps them.

Property tests (hypothesis) cover identity at equal years, monotonicity,
zeros, penny ranks, missing draws, composition and inversion, a
differential test against `uprate_dataset`, and the call sites: the
second stage and the FRS dividends see rebased draws.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
With #529's earnings-group draw, impute_over_incomes reads each person's
employment_status. The fakes now carry one (working-age employees), so the
same tests pass on main and on a merge with #529.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The stage-2 capture now also checks that the FRS respondents the FRS-only
QRF trains on, the FRS rows' undrawn incomes and an undrawn indexed money
column (employee pension contributions) keep their survey-year values, so
rebasing the whole copy, the training set or the FRS half's other incomes
fails the test, as does rebasing after the second stage.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Hypothesis found that 2.2e-313 rebased 2022 -> 2020 -> 2021 differs from a
direct 2022 -> 2021 rebase by 2e-11 relative: subnormal floats carry too
few significant bits to hold a 1e-12 relative bound through two
multiplications. The smallest normal floats compose exactly. No amount of
money is subnormal, so the money strategy excludes them rather than
loosening the bound for real amounts.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
CI's reduced build (TESTING=1, 32 epochs) of 3d82532 carried £236.1bn of
above-AEA gains against HMRC's £65.9bn, just past the reduced-build bound
(relative error 2.5). That is reduced-build calibration noise, not this
change: #529's seeded reduced builds put main itself at about £243bn
(relative error 2.7), and this branch's production build passes the strict
50% bound. This takes #529's version of the test file unchanged (from
88ef576 and d39371d): under TESTING the total must lie within a factor of 6
of HMRC's; full builds keep the 50% bound. Identical bytes, so whichever PR
lands second merges cleanly.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Review r2: the comment rounded Hypothesis's counterexample to 2.2e-313,
which passes; the logged value 2.2250738585e-313 is the one that fails
(2.4e-11 relative). Normal floats compose within 6e-16 relative over every
pair of rebasings between 2020 and 2034. Comment only.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant