Skip to content

Draw SPI incomes within earnings groups set by FRS employment status - #529

Merged
MaxGhenis merged 11 commits into
mainfrom
spi-imputation-employment-status
Oct 7, 2026
Merged

MaxGhenis merged 11 commits into
mainfrom
spi-imputation-employment-status

Conversation

@MaxGhenis

@MaxGhenis MaxGhenis commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Draft. This changes the published dataset, so merging it is a data release. uk-data PRs land together as one batched release when Max says go (d833). No dataset was uploaded or released from this branch. Fixes #504.

Problem

impute_income stacks a zero-weight copy of 10,000 FRS households (household_is_spi_synthetic), and calibration later gives them weight. Their six incomes (employment, self-employment, savings interest, dividends, private pension, property) are redrawn by a QRF trained on the SPI, whose only predictors are age, gender and region. Every other column, including employment_status and hours_worked, stays the FRS donor's. So a row's incomes do not depend on whether it is a child, an employee, self-employed, or out of work.

On main's build (b45c373, seed 0), the SPI rows and their capital-gains clones hold 8.63m of 31.14m households. Share of people with earnings, by employment status, at 2024-25 calibrated weights:

Status FRS rows: pay > 0 SPI rows: pay > 0 FRS rows: profit > 0 SPI rows: profit > 0
Child (FRS child table) 0% 99.7% 0% 1.0%
Employee 98.4% 84.0% 2.8% 7.3%
Self-employed 6.4% 81.3% 89.5% 8.6%
Out of work (unemployed, retired, student, carer, sick or disabled) under 10 households 46.3% 0% 7.4%

In money, SPI-row children carry £73.3bn of employment income (#504 found the same, from the file as released). SPI rows out of work carry £54.3bn of pay and £6.8bn of profit. SPI-row self-employed people carry £41.3bn of pay but only £3.4bn of profit.

Everything that reads status and income together inherits this:

  • Universal Credit hours and conditionality;
  • the minimum income floor (policyengine-uk#2081 reads any profit as gainful self-employment when Set UC gainful self-employment from the FRS main-job status #525's input is absent);
  • income tax and NI on children's pay;
  • the second-stage QRF, which draws benefit reports from these incomes.

What the SPI can support

From the SPI 2022-23 Public Use Tape documentation (UKDS SN 9422, Annex A):

  • No ILO employment status, hours or full/part-time split. Status cannot be a QRF predictor, because the SPI has nothing to train it on.
  • Income sources, which it does have:
    • PAY ("Pay from employment net of benefits and foreign earnings"), with EPB and TAXTERM;
    • PROFITS ("Gross profits assessable for all sources of self-employment income"). Its range starts at 0, so losses are not recorded;
    • SEINC_NUM, the "Indicator for self-employed cases (those submitting pages in their tax return for income from trades or partnerships)". This finds traders whatever their profit: 17,298 records file self-employment pages with zero PROFITS;
    • MAINSRCE, the main source of income: pay, occupational pension, sole trader, partnership, other, or claims case. It is the largest income, not the main job, so this PR does not use it.

Change

1. Draw SPI incomes within earnings groups (imputations/income.py)

The FRS and the SPI share one thing that matters here: whether a person has pay and whether they have a trade. That gives four earnings groups.

Group SPI records (2022-23, weighted) FRS rows
EMPLOYEE pay > 0, no trade (33.0m) employee main job (FT/PT_EMPLOYED), or recorded pay, and no trade
SELF_EMPLOYED SEINC_NUM = 1 or PROFITS > 0, no pay (3.9m; 93.5% with a profit) self-employed main job, or recorded profit, and no pay
EMPLOYEE_AND_SELF_EMPLOYED both (1.5m; 81% with a profit) both, e.g. a self-employed main job plus an employee second job
NO_EARNINGS neither (12.3m) everyone else: retired, unemployed, student, carer, sick or disabled, other inactive
NOT_IMPUTED n/a the FRS child table, and anyone under 16. They keep their own incomes; this FRS build records no earnings for them
  • One QRF per group, fitted on that group's SPI records only, still on age, gender and region. So:
    • an employee always draws pay;
    • a self-employed person always draws a trade, whose profit is zero where the SPI trader's is;
    • someone out of work draws pension, property and investment income from SPI people with no earnings;
    • a child takes no draw.
  • The training sample is still a weighted resample of the SPI (100k; 10k in TESTING). It is now drawn within each group, and each group gets at least 10% of the nominal sample size: 65.1k employee, 24.2k no earnings, 10k self-employed and 10k both. Floors are added on top, so the sample is 109,312 records (10,931 in TESTING).
  • The cached model (income_spi_2022_23.pkl, 2.1 GB as before) holds one QRF per group. A cache in the old single-model format is retrained.
  • Main's FRS-half dividend draw goes through the same groups, so FRS children keep their own dividends. Keep FRS-reported dividends and key them on person_id #498 removes that draw; the two merge cleanly in either order.

Why not the other options

  • Status as a QRF predictor. The SPI has no status to learn from. A proxy built from income sources is these groups. As a soft predictor, the forest would usually split on it but not always. Separate models make the agreement hold by construction, and the tests check it.
  • Re-deriving status from the imputed incomes. That throws away what the FRS knows and the SPI does not: who is a child, retired, a student, a carer, sick, or unemployed. It would also need hours, which the SPI does not have.

2. Calibrate to LFS employees and self-employed (targets/sources/ons_labour_market.py)

Making rows coherent exposed a conflict in the calibration. The local targets include HMRC's count of income-tax payers with employment income in each constituency and local authority (SPI table 3.15, 30.0m in total). That count is annual. It includes people who had pay for part of the year but whose FRS status at interview is out of work, and in the FRS those people have no pay.

On main, SPI-row children and out-of-work rows with pay supplied part of that count. Once they draw no pay, calibration met the count by moving weight from people out of work to employees.

Nothing tied employment status to an official total, so this adds national targets for the ONS LFS levels of employees (MGRN) and self-employed (MGRQ), annual averages for 2022-2025, counted from the FRS ILO main-job status. The LFS also counts working dependants aged 16-19, whom the FRS child table records without a job. On FRS 2024-25 grossing weights the FRS has 28.1m employees, against the LFS's 29.1m for 2024.

3. Report, not train on, the local HMRC employment-income counts (create_datasets.py)

The two national LFS targets could not outweigh about 1,000 local count targets: with both trained, employees stayed at 32.1m. VALIDATION_ONLY_LOCAL_TARGETS moves hmrc/employment_income/count to the calibrator's existing validation mode, so it is still logged but no longer trained. The local amounts of employment income, the self-employment counts, and the national counts by income band still train. c3cb237 fixes the logged validation loss, which was NaN when every validation target is local (an empty mean); training and weights are unchanged. This is a calibration-methodology change; it is a separate commit (fc46a56) so it can be reverted on its own.

Labour-market composition across builds

All builds are seeded production builds (512 epochs, PE_UK_DATA_OA_CLONES=1, seed 0) with the same cached web targets. People by FRS employment status at 2024-25 calibrated weights, FRS and SPI rows together:

Build Employees Self-employed Out of work (16+) Children
main (b45c373) 28.2m 4.8m 21.2m 15.1m
placebo: main with only the SPI draw's seed changed 28.3m 4.6m 21.3m 15.1m
change 1 only (708ab3e) 32.2m 4.0m 18.1m 14.9m
changes 1 and 2 (40fbfea) 32.1m 4.4m 17.9m 14.9m
changes 1, 2 and 3 (this head) 30.3m 4.4m 19.6m 14.9m
ONS LFS, 2024 (2025) 29.1m (29.6m) 4.3m (4.4m)
FRS 2024-25 grossing weights 28.1m 4.2m 21.3m 14.7m

The placebo is main rebuilt with a different random draw for the SPI incomes and nothing else changed. It shows what one redraw does. It is a single draw, so read it as the scale of noise, not a bound on it.

Target fit at the final epoch. Rows marked LA are from the local-authority calibration; the local sums are from the constituency calibration.

Target Placebo Changes 1 and 2 This head
LFS employees, 29.6m (LA) not a target (28.3m by status, −4.3%) +7.9% +2.9%
LFS self-employed, 4.40m (LA) not a target +0.7% +1.6%
OBR employer NI (LA) −0.2% +2.8% +2.0%
OBR employee NI (LA) −3.4% −0.3% −1.2%
HMRC local employment-income count, summed (30.0m) +0.4% +0.0% −8.3% (not trained)
HMRC local employment-income amount, summed +4.3% +3.3% +1.5%
National targets within 10% 85.7% of 637 84.8% of 639 85.4% of 639

The untrained count ends 8.3% under HMRC's in total. That is a post-calibration residual, not a measurement of the part-year-earner gap. Not training on the counts also removes the constraint on each area's number of earners, and with it on mean pay per earner:

  • across 636 constituencies the count residual has a median of −11.0%, with 33 areas below −20% and 9 above +20%;
  • the area amounts still fit (median +1.6%, none beyond ±20%).

Impact

Real policyengine-uk microsimulations, 2026, on policyengine-uk main (3c48247eb, 2.107.0). Base is main's seed-0 build; "this head" is the changes 1-3 build. "SPI rows" is the part of the change on SPI-synthetic households, which includes their reweighting. The placebo column is one redraw, for scale.

Base This head Change of which SPI rows Placebo change
Income tax (£bn) 314.5 317.1 +2.6 -6.1 +1.9
Employee NI (class 1) (£bn) 51.3 52.8 +1.5 -1.4 -0.1
Employer NI (class 1) (£bn) 156.1 160.7 +4.5 -5.3 +0.1
Self-employed NI (class 4) (£bn) 2.7 2.7 +0.0 +0.4 +0.0
Employment income (£bn) 1,263.8 1,216.9 -46.9 -122.2 -2.0
Self-employment income (£bn) 116.5 119.7 +3.2 +16.6 +0.6
Universal Credit (£bn) 78.0 76.9 -1.2 +18.2 -1.0
Housing Benefit (£bn) 14.0 14.3 +0.3 +0.5 +0.1
Council tax reduction (£bn) 2.3 2.4 +0.2 +0.4 +0.2
Child Benefit (£bn) 17.7 17.5 -0.2 +1.1 +0.1
Household net income (£bn) 1,760.8 1,716.5 -44.3 -42.5 +10.7
People with employment income (m) 35.17 30.52 -4.64 -6.51 +0.09
People in absolute poverty, BHC (k) 10,345 10,711 +367 +2,181 -551
People in absolute poverty, AHC (k) 13,222 14,236 +1,013 +3,147 -55
Children in absolute poverty, BHC (k) 2,627 3,263 +635 +1,376 -78
Children in absolute poverty, AHC (k) 3,714 4,546 +832 +1,853 -15
People in relative poverty, BHC (k) 13,568 12,259 -1,309 +1,980 -469
People in relative poverty, AHC (k) 15,902 14,743 -1,159 +2,992 +203
Gini, equivalised HBAI net income (people) 0.3127 0.3090 -0.0037 -0.0023

On policyengine-uk#2081's head (94b7129b4), where the minimum income floor applies only to claimants with all work-related requirements, people in UC-receiving benefit units whose floor applies:

Base This head Change Placebo change
2025 307k 321k +15k +15k
2026 302k 315k +13k +14k

The change is similar in size to the placebo's.

Which part does what (2026, change from main):

Change 1 only Changes 1 and 2 Changes 1-3 (this head) Placebo
Income tax (£bn) +2.0 +2.0 +2.6 +1.9
Employer NI (£bn) +6.2 +6.1 +4.5 +0.1
Employee NI (£bn) +1.9 +1.9 +1.5 -0.1
Universal Credit (£bn) -1.7 -1.7 -1.2 -1.0
People in absolute poverty, AHC (k) +371 +370 +1,013 -55
Children in absolute poverty, AHC (k) +717 +728 +832 -15

Readings:

  • Income tax moves +£2.6bn, similar to the placebo's +£1.9bn, so much of it may be redraw noise. Income tax is calibrated.
  • NI rises for real. About £73bn of pay sat on SPI-row children, £66.7bn of it on under-16s. policyengine-uk charges neither employee nor employer Class 1 NI on under-16s (ni_liable and ni_class_1_secondary_liable both require over_16), so moving that pay to adults raises both. Change 3 takes about £1.6bn back off employer NI by putting the employee count near the LFS. Employer NI ends 2.0% over its OBR target (placebo −0.2%) and employee NI 1.2% under it (placebo −3.4%).
  • Absolute poverty rises, mostly among children: SPI-row children no longer carry earnings that kept their households above the line. SPI income imputation gives every child in a synthetic household employment income (£67bn on under-16s) #504 predicted this with a zero-the-children sensitivity: +5.3 points AHC child poverty.
  • Change 3 adds about 0.64m to AHC absolute poverty over changes 1 and 2. It returns about 1.8m people from employment to out of work (employees 32.1m → 30.3m), and out-of-work households are poorer. That moves the labour market back towards the LFS and towards the FRS's own grossed split (employees 28.1m, out of work 21.3m).
  • Relative poverty falls because median income falls with that pay.
  • Universal Credit moves −£1.2bn, similar to the placebo's −£1.0bn.

Interplay with open uk-data PRs

Invariants (Hypothesis property tests)

tests/test_spi_income_earnings_groups.py:

  1. FRS rows. Children (FRS child table, or under 16) are NOT_IMPUTED. Every other row is in one group. That group has pay if and only if the main job is as an employee or the FRS records pay, and has a trade if and only if the main job is self-employment or the FRS records a profit. Every policyengine-uk EmploymentStatus is covered.

  2. SPI records. Pay if and only if PAY + EPB + TAXTERM > 0. A trade if and only if SEINC_NUM = 1 or PROFITS > 0.

  3. Both mappings are monotone: more of one income adds that source and leaves the other alone. Each gives the same answer elementwise as row by row.

  4. Sample allocation. Every group with weight gets at least 10% of the nominal sample size. Groups without weight get none. Groups above the floor get their weighted share. The total lies between the nominal size, less rounding, and the nominal size plus one floor per group.

  5. generate_spi_table resamples each group only from its own records, in those counts.

  6. Draws from a real fitted per-group QRF. The trade checks cover both trade groups.

    • Pay is positive exactly in the groups with pay.
    • Profit is zero in the groups without a trade.
    • Every draw is non-negative.
    • NOT_IMPUTED rows get no draw, and the index is preserved.
    • The trade groups draw both zero and positive profits.
  7. apply_income_draws overwrites drawn rows and leaves NOT_IMPUTED rows, other columns and its input untouched.

  8. The cache. A cache in the old format, or missing a group, is retrained; a current one round-trips.

  9. On a built enhanced FRS. On SPI rows:

    • every employee has pay;
    • every child has no earnings;
    • over 60% of the self-employed have a profit;
    • under 2% of people out of work have earnings.

    This check fails on main's build, at its first assertion.

tests/test_lfs_employment_targets.py:

  • the ONS values, pinned independently;
  • year resolution (the 2025 values are held for up to three later years, then the targets drop; add each new annual average);
  • discovery of the source module;
  • a property for the household count column;
  • on a built enhanced FRS, employees and the self-employed are within 8% of the LFS. This also catches the local counts going back into training.

Mutation check. I made 10 deliberate defects in income.py and the tests caught each one: status ignored for pay or for a trade, children drawn, SEINC_NUM ignored, child rows overwritten, NaN draws kept, one model for every group, resample ignoring groups, no group floor, and a cache accepting a missing group.

Full suite. On the production build of changes 1 and 2, the full test suite passes except the LFS employee check that change 3 fixes.

CI's reduced build (TESTING=1, 32 epochs). Three existing checks failed on 54f0974 that pass on other PRs, so I rebuilt main and this branch at seed 0 with TESTING=1:

Check main, reduced this branch, reduced CI run, reduced this branch, full build old reduced bound new reduced bound
Salary-sacrifice users (7.7m) 6.38m 5.22m 5.4m 7.96m 20% 40%
Below-cap users (4.3m) 3.78m 2.64m 2.7m 4.36m 25% 45%
Above-AEA gains (£65.9bn) £242.6bn £267.4bn £249.3bn £53.2bn relative error 2.5 within 6× either way
  • Salary sacrifice. The full build calibrates to OBR's user counts and gets there: 4.5m users at epoch 0, 5.5m at epoch 30, 7.9m at epoch 510. Thirty-two epochs don't. Main's reduced build passed only because about 1.35m of its 6.38m users were SPI-row children with imputed salary sacrifice; without them it held 5.0m, below this branch's 5.22m.
  • Gains. Main's seed-0 reduced build already fails the old bound, so that check depends on the run in reduced builds.
  • What changed (88ef576, d39371d). The reduced bounds widen as above, following the precedents in Sort FRS household frame by ID to fix scrambled weights (2024-25 population undercount) #436 and Sort the benunit table by benunit_id in create_frs #462. The gains bound is now two-sided in reduced builds: the total must be within a factor of 6 of HMRC's, either way. Every full-build bound is unchanged, and the full build passes all of them. The new bounds pass on both seeded reduced builds.
  • Headroom is thin, and the looser bounds apply to every future PR's reduced build.
    • On this branch's seeded reduced build, the salary-sacrifice checks pass with 0.60m (total) and 0.27m (below cap) to spare.
    • An uncalibrated starting point (4.52m and 2.24m, epoch 0 of the full build) would fail both, but only by 1.3 and 2.9 points.
    • The unchanged above-cap check sits at −21.8% against its 25% reduced bound.
    • Some run-to-run flakiness in CI's unseeded reduced build remains possible.

The new SPI-row and LFS checks pass in CI's reduced build.

Not in this PR

Checklist

  • Data. Aggregates only; no record-level values. Cells under 10 survey households are suppressed. Nothing was uploaded.
  • axiom: n/a: survey imputation and calibration targets, no policy rule.

🤖 Generated with Claude Code

MaxGhenis and others added 9 commits October 2, 2026 12:37
The SPI income model drew every SPI-synthetic row's six incomes from age,
gender and region alone, while the row kept its FRS donor's employment
status. Children drew pay, the unemployed and retired drew earnings, and
self-employed rows rarely drew a profit.

Fit one QRF per earnings group (employee, self-employed, both, neither) on
the SPI records in that group, using PAY and the SPI self-employment
indicator SEINC_NUM, and draw each FRS row from the group its employment
status and recorded earnings put it in. Children keep their own incomes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Same lines as #514 and #524, so whichever lands second merges cleanly.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Hypothesis properties for the FRS and SPI group mappings, the training
sample allocation, the per-group model's draws and the donor values kept
for children, plus a check on a built enhanced FRS. The cache tests write
the per-group format.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
With SPI incomes drawn by earnings group, the SPI-synthetic rows no longer
give pay to children and people out of work. Calibration then met the HMRC
counts of income-tax payers with employment income, which are annual and
include part-year earners, by moving about 4m people's weight from out of
work to employees (32.2m against the LFS's 29.1m for 2024). Nothing tied
employment status to an official count.

Add national targets for the ONS LFS employee and self-employed levels
(MGRN, MGRQ annual averages, 2022-2025), counted from the FRS ILO main-job
status. Move the status groups to utils/employment_status.py so the income
imputation and the targets share them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A loaded runner tripped Hypothesis's too_slow health check while the
properties themselves held. Drop the deadline and that health check.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
HMRC's counts of income-tax payers with employment income by area are
annual, so they include people with pay for part of the year whose FRS
status at interview is out of work. Trained on, they held employees at
32m against the LFS target of 29.6m. The area amounts and the national
counts by income band still train; the counts stay in the calibration
logs as validation targets.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Values pinned against the ONS release, year resolution, source discovery,
a Hypothesis property for the household column, and a check that a built
enhanced FRS lands near both counts.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…get is local

With VALIDATION_ONLY_LOCAL_TARGETS no national target is held out, so the
national validation term was the mean of an empty set and the logged
validation loss was NaN. An empty mask now contributes zero, as the local
term already did. Training and the saved weights are unchanged.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…onicity

- Widen the reduced (TESTING) bounds on salary-sacrifice users (40% total,
  45% below cap) and above-AEA gains (relative error 5). The 32-epoch build
  stops short of the OBR salary-sacrifice targets (4.5m at epoch 0, 5.5m at
  epoch 30, 7.9m at epoch 510 in the full build), and main's reduced build
  met the old bounds only through about 1.35m SPI-synthetic children with
  salary sacrifice. Seed-0 reduced builds of main and of this branch both
  exceed the old gains bound. Full builds keep the strict bounds and pass.
- State the training-sample floor as a share of the nominal size, with the
  total's bounds tested.
- Add an SPI monotonicity property and check mixed profits in both trade
  groups.
- ONS release date is 15 September 2026; note how later years resolve.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A relative-error bound of 5 could never fail low: a total of zero has a
relative error of 1. Reduced builds now require the total to lie within a
factor of 6 of HMRC's in either direction; full builds keep the 50% bound.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@MaxGhenis
MaxGhenis marked this pull request as ready for review October 3, 2026 22:27
@MaxGhenis

Copy link
Copy Markdown
Contributor Author

Hand-off to the UK hub (session local_185ef39c, 2026-10-03).

MaxGhenis added a commit that referenced this pull request Oct 4, 2026
CI's reduced build (TESTING=1, 32 epochs) of 3d82532 carried £236.1bn of
above-AEA gains against HMRC's £65.9bn, just past the reduced-build bound
(relative error 2.5). That is reduced-build calibration noise, not this
change: #529's seeded reduced builds put main itself at about £243bn
(relative error 2.7), and this branch's production build passes the strict
50% bound. This takes #529's version of the test file unchanged (from
88ef576 and d39371d): under TESTING the total must lie within a factor of 6
of HMRC's; full builds keep the 50% bound. Identical bytes, so whichever PR
lands second merges cleanly.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@MaxGhenis

Copy link
Copy Markdown
Contributor Author

Queued in release PR #544 for the 10/8 uk-data batch. It lands only on Max's go (d833).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

SPI income imputation gives every child in a synthetic household employment income (£67bn on under-16s)

1 participant