Repository navigation
Draw SPI incomes within earnings groups set by FRS employment status - #529
Merged
Merged
Conversation
The SPI income model drew every SPI-synthetic row's six incomes from age, gender and region alone, while the row kept its FRS donor's employment status. Children drew pay, the unemployed and retired drew earnings, and self-employed rows rarely drew a profit. Fit one QRF per earnings group (employee, self-employed, both, neither) on the SPI records in that group, using PAY and the SPI self-employment indicator SEINC_NUM, and draw each FRS row from the group its employment status and recorded earnings put it in. Children keep their own incomes. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Hypothesis properties for the FRS and SPI group mappings, the training sample allocation, the per-group model's draws and the donor values kept for children, plus a check on a built enhanced FRS. The cache tests write the per-group format. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
With SPI incomes drawn by earnings group, the SPI-synthetic rows no longer give pay to children and people out of work. Calibration then met the HMRC counts of income-tax payers with employment income, which are annual and include part-year earners, by moving about 4m people's weight from out of work to employees (32.2m against the LFS's 29.1m for 2024). Nothing tied employment status to an official count. Add national targets for the ONS LFS employee and self-employed levels (MGRN, MGRQ annual averages, 2022-2025), counted from the FRS ILO main-job status. Move the status groups to utils/employment_status.py so the income imputation and the targets share them. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A loaded runner tripped Hypothesis's too_slow health check while the properties themselves held. Drop the deadline and that health check. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
HMRC's counts of income-tax payers with employment income by area are annual, so they include people with pay for part of the year whose FRS status at interview is out of work. Trained on, they held employees at 32m against the LFS target of 29.6m. The area amounts and the national counts by income band still train; the counts stay in the calibration logs as validation targets. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Values pinned against the ONS release, year resolution, source discovery, a Hypothesis property for the household column, and a check that a built enhanced FRS lands near both counts. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…get is local With VALIDATION_ONLY_LOCAL_TARGETS no national target is held out, so the national validation term was the mean of an empty set and the logged validation loss was NaN. An empty mask now contributes zero, as the local term already did. Training and the saved weights are unchanged. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Oct 3, 2026
…onicity - Widen the reduced (TESTING) bounds on salary-sacrifice users (40% total, 45% below cap) and above-AEA gains (relative error 5). The 32-epoch build stops short of the OBR salary-sacrifice targets (4.5m at epoch 0, 5.5m at epoch 30, 7.9m at epoch 510 in the full build), and main's reduced build met the old bounds only through about 1.35m SPI-synthetic children with salary sacrifice. Seed-0 reduced builds of main and of this branch both exceed the old gains bound. Full builds keep the strict bounds and pass. - State the training-sample floor as a share of the nominal size, with the total's bounds tested. - Add an SPI monotonicity property and check mixed profits in both trade groups. - ONS release date is 15 September 2026; note how later years resolve. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
A relative-error bound of 5 could never fail low: a total of zero has a relative error of 1. Reduced builds now require the total to lie within a factor of 6 of HMRC's in either direction; full builds keep the 50% bound. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
MaxGhenis
marked this pull request as ready for review
October 3, 2026 22:27
Contributor
Author
|
Hand-off to the UK hub (session
|
MaxGhenis
added a commit
that referenced
this pull request
Oct 4, 2026
CI's reduced build (TESTING=1, 32 epochs) of 3d82532 carried £236.1bn of above-AEA gains against HMRC's £65.9bn, just past the reduced-build bound (relative error 2.5). That is reduced-build calibration noise, not this change: #529's seeded reduced builds put main itself at about £243bn (relative error 2.7), and this branch's production build passes the strict 50% bound. This takes #529's version of the test file unchanged (from 88ef576 and d39371d): under TESTING the total must lie within a factor of 6 of HMRC's; full builds keep the 50% bound. Identical bytes, so whichever PR lands second merges cleanly. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Oct 5, 2026
23 of 57 tasks
Contributor
Author
|
Queued in release PR #544 for the 10/8 uk-data batch. It lands only on Max's go (d833). |
21 of 52 tasks
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Draft. This changes the published dataset, so merging it is a data release. uk-data PRs land together as one batched release when Max says go (d833). No dataset was uploaded or released from this branch. Fixes #504.
Problem
impute_incomestacks a zero-weight copy of 10,000 FRS households (household_is_spi_synthetic), and calibration later gives them weight. Their six incomes (employment, self-employment, savings interest, dividends, private pension, property) are redrawn by a QRF trained on the SPI, whose only predictors are age, gender and region. Every other column, includingemployment_statusandhours_worked, stays the FRS donor's. So a row's incomes do not depend on whether it is a child, an employee, self-employed, or out of work.On main's build (b45c373, seed 0), the SPI rows and their capital-gains clones hold 8.63m of 31.14m households. Share of people with earnings, by employment status, at 2024-25 calibrated weights:
In money, SPI-row children carry £73.3bn of employment income (#504 found the same, from the file as released). SPI rows out of work carry £54.3bn of pay and £6.8bn of profit. SPI-row self-employed people carry £41.3bn of pay but only £3.4bn of profit.
Everything that reads status and income together inherits this:
What the SPI can support
From the SPI 2022-23 Public Use Tape documentation (UKDS SN 9422, Annex A):
PAY("Pay from employment net of benefits and foreign earnings"), withEPBandTAXTERM;PROFITS("Gross profits assessable for all sources of self-employment income"). Its range starts at 0, so losses are not recorded;SEINC_NUM, the "Indicator for self-employed cases (those submitting pages in their tax return for income from trades or partnerships)". This finds traders whatever their profit: 17,298 records file self-employment pages with zeroPROFITS;MAINSRCE, the main source of income: pay, occupational pension, sole trader, partnership, other, or claims case. It is the largest income, not the main job, so this PR does not use it.Change
1. Draw SPI incomes within earnings groups (
imputations/income.py)The FRS and the SPI share one thing that matters here: whether a person has pay and whether they have a trade. That gives four earnings groups.
EMPLOYEEFT/PT_EMPLOYED), or recorded pay, and no tradeSELF_EMPLOYEDSEINC_NUM= 1 orPROFITS> 0, no pay (3.9m; 93.5% with a profit)EMPLOYEE_AND_SELF_EMPLOYEDNO_EARNINGSNOT_IMPUTEDTESTING). It is now drawn within each group, and each group gets at least 10% of the nominal sample size: 65.1k employee, 24.2k no earnings, 10k self-employed and 10k both. Floors are added on top, so the sample is 109,312 records (10,931 inTESTING).income_spi_2022_23.pkl, 2.1 GB as before) holds one QRF per group. A cache in the old single-model format is retrained.Why not the other options
2. Calibrate to LFS employees and self-employed (
targets/sources/ons_labour_market.py)Making rows coherent exposed a conflict in the calibration. The local targets include HMRC's count of income-tax payers with employment income in each constituency and local authority (SPI table 3.15, 30.0m in total). That count is annual. It includes people who had pay for part of the year but whose FRS status at interview is out of work, and in the FRS those people have no pay.
On main, SPI-row children and out-of-work rows with pay supplied part of that count. Once they draw no pay, calibration met the count by moving weight from people out of work to employees.
Nothing tied employment status to an official total, so this adds national targets for the ONS LFS levels of employees (MGRN) and self-employed (MGRQ), annual averages for 2022-2025, counted from the FRS ILO main-job status. The LFS also counts working dependants aged 16-19, whom the FRS child table records without a job. On FRS 2024-25 grossing weights the FRS has 28.1m employees, against the LFS's 29.1m for 2024.
3. Report, not train on, the local HMRC employment-income counts (
create_datasets.py)The two national LFS targets could not outweigh about 1,000 local count targets: with both trained, employees stayed at 32.1m.
VALIDATION_ONLY_LOCAL_TARGETSmoveshmrc/employment_income/countto the calibrator's existing validation mode, so it is still logged but no longer trained. The local amounts of employment income, the self-employment counts, and the national counts by income band still train. c3cb237 fixes the logged validation loss, which was NaN when every validation target is local (an empty mean); training and weights are unchanged. This is a calibration-methodology change; it is a separate commit (fc46a56) so it can be reverted on its own.Labour-market composition across builds
All builds are seeded production builds (512 epochs,
PE_UK_DATA_OA_CLONES=1, seed 0) with the same cached web targets. People by FRS employment status at 2024-25 calibrated weights, FRS and SPI rows together:The placebo is main rebuilt with a different random draw for the SPI incomes and nothing else changed. It shows what one redraw does. It is a single draw, so read it as the scale of noise, not a bound on it.
Target fit at the final epoch. Rows marked LA are from the local-authority calibration; the local sums are from the constituency calibration.
The untrained count ends 8.3% under HMRC's in total. That is a post-calibration residual, not a measurement of the part-year-earner gap. Not training on the counts also removes the constraint on each area's number of earners, and with it on mean pay per earner:
Impact
Real policyengine-uk microsimulations, 2026, on policyengine-uk main (3c48247eb, 2.107.0). Base is main's seed-0 build; "this head" is the changes 1-3 build. "SPI rows" is the part of the change on SPI-synthetic households, which includes their reweighting. The placebo column is one redraw, for scale.
On policyengine-uk#2081's head (94b7129b4), where the minimum income floor applies only to claimants with all work-related requirements, people in UC-receiving benefit units whose floor applies:
The change is similar in size to the placebo's.
Which part does what (2026, change from main):
Readings:
ni_liableandni_class_1_secondary_liableboth requireover_16), so moving that pay to adults raises both. Change 3 takes about £1.6bn back off employer NI by putting the employee count near the LFS. Employer NI ends 2.0% over its OBR target (placebo −0.2%) and employee NI 1.2% under it (placebo −3.4%).Interplay with open uk-data PRs
uc_is_in_gainful_self_employment).income.pyauto-merges with Set UC gainful self-employment from the FRS main-job status #525.uv.lockconflicts only because thehypothesislines differ; whichever lands second relocks.SELF_EMPLOYED_STATUSESinfrs.py; this PR addsutils/employment_status.py. Whichever lands second can point Set UC gainful self-employment from the FRS main-job status #525's constant at the shared module.git merge-tree). Thehypothesislines anduv.lockhunk are identical to Stop SPI-synthetic rows carrying benefit claims nobody observed #514's and Supply is_claimant_or_partner from the FRS adult table #524's.income.pyand auto-merge cleanly.uv.lock.Invariants (Hypothesis property tests)
tests/test_spi_income_earnings_groups.py:FRS rows. Children (FRS child table, or under 16) are
NOT_IMPUTED. Every other row is in one group. That group has pay if and only if the main job is as an employee or the FRS records pay, and has a trade if and only if the main job is self-employment or the FRS records a profit. Every policyengine-ukEmploymentStatusis covered.SPI records. Pay if and only if PAY + EPB + TAXTERM > 0. A trade if and only if
SEINC_NUM= 1 orPROFITS> 0.Both mappings are monotone: more of one income adds that source and leaves the other alone. Each gives the same answer elementwise as row by row.
Sample allocation. Every group with weight gets at least 10% of the nominal sample size. Groups without weight get none. Groups above the floor get their weighted share. The total lies between the nominal size, less rounding, and the nominal size plus one floor per group.
generate_spi_tableresamples each group only from its own records, in those counts.Draws from a real fitted per-group QRF. The trade checks cover both trade groups.
NOT_IMPUTEDrows get no draw, and the index is preserved.apply_income_drawsoverwrites drawn rows and leavesNOT_IMPUTEDrows, other columns and its input untouched.The cache. A cache in the old format, or missing a group, is retrained; a current one round-trips.
On a built enhanced FRS. On SPI rows:
This check fails on main's build, at its first assertion.
tests/test_lfs_employment_targets.py:Mutation check. I made 10 deliberate defects in
income.pyand the tests caught each one: status ignored for pay or for a trade, children drawn,SEINC_NUMignored, child rows overwritten, NaN draws kept, one model for every group, resample ignoring groups, no group floor, and a cache accepting a missing group.Full suite. On the production build of changes 1 and 2, the full test suite passes except the LFS employee check that change 3 fixes.
CI's reduced build (
TESTING=1, 32 epochs). Three existing checks failed on 54f0974 that pass on other PRs, so I rebuilt main and this branch at seed 0 withTESTING=1:The new SPI-row and LFS checks pass in CI's reduced build.
Not in this PR
hours_workedand the rest of the donor's columns are unchanged.Checklist
🤖 Generated with Claude Code