Skip to content

Draw SPI incomes at each donor's FRS rank within their cell - #551

Draft
MaxGhenis wants to merge 5 commits into
mainfrom
spi-rank-preserving-draw
Draft

MaxGhenis wants to merge 5 commits into
mainfrom
spi-rank-preserving-draw

Conversation

@MaxGhenis

@MaxGhenis MaxGhenis commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

What this changes

PR #529 draws each SPI-synthetic person's incomes from SPI records in their own earnings group, so employees draw pay and the out-of-work draw none. Within the group the draw was still random: microimpute picks a random quantile, conditioned only on age, gender and region. The donor keeps their hours and FT/PT status, so a part-timer drew a full-timer's pay as often as a part-timer's. The SPI has no hours.

This PR draws each income at the person's FRS rank of that income within their cell instead of at a random quantile. This is the uk-data analogue of PolicyEngine/microcosm#840, "rank-preserving replacement". A part-timer on low FRS pay draws low SPI pay. A donor at the top of their cell draws from the top of the forest's distribution for that cell.

  • Rank variable: per component. Pay sets the pay draw, profit the profit draw, a private pension the pension draw, and so on.
    • For employees and the self-employed, this is the same as ranking on total earnings.
    • People with both pay and a trade keep each source's own position.
    • People with no earnings (21m weighted) get ranks from their pension, savings and property income. A total-earnings rank would leave their draws random.
    • The FRS does not record Gift Aid or charitable investment gifts, so those are drawn at a uniform random quantile, independently of each other and of every other output.
  • Ties and zeros. Within a cell, people are ordered by the income, with equal values in random order. Equal values are common: zeros, and the FRS's heaped pay, where only 16% of positive values are distinct. Each person takes a uniform point in their slice of the cell's weight.
    • The slices tile [0, 1), so the quantiles are exactly uniform in every cell, whatever the ties.
    • Zero is just the lowest tie. An employee with no recorded pay draws from the bottom of their cell and never above one with pay.
    • Mid-ranks are rejected: they put a whole tie, such as everyone without a private pension, at one quantile.
  • What uniform quantiles preserve, and what they don't.
    • For each group's first output, the draw keeps the forest's distribution for the cell, in expectation. The forest is queried at the person's exact age, but its training age is uniform noise within the SPI age band. FRS pay rank tracks age within a band, so the rank draw also picks up the forest's chance age pattern. On the full FRS, pay over £100k, summed over the three pay groups, is 0.802m at FRS pay ranks, against 0.845m with random within-cell ranks (SD 0.033m) and a forest expectation of 0.848m. For the EMPLOYEE group alone it is 0.780m against the forest's 0.827m.
    • Later outputs are conditioned on those drawn before, at correlated FRS ranks, so they carry more dependence between incomes than the SPI does. It is still less than microimpute's single random quantile per person gave (table under Evidence).
  • Cells: (earnings group, SPI age band, gender, region). These are the finest cells the SPI can tell apart, since it gives age only as a band. The forests are queried at the person's exact age and, after each group's first output, at the outputs drawn before it. So the cells match the forests' conditioning only up to that age noise and the chain.
    • Small cells are not pooled. Ranking a London donor nationally would map their higher national rank onto London's already-higher SPI distribution.
    • A one-donor cell gets a uniform random quantile, as before.
  • People with both pay and a trade are split by main source.
    • SPI: MAINSRCE 1 means pay is the main source; 3 (sole trader) or 4 (partnership) means a trade. Any other code (2 occupational pension, 5 other, 6 claims case, −1 not classified) leaves it to the larger income.
    • FRS: the main job's status.
    • Main source of income is one of the variables HMRC stratifies the self-assessment sample by (HMRC, Survey of Personal Incomes 2022-23: public use tape documentation, p. 3; MAINSRCE codes on p. 15). It is not the larger income: 57% of trade-main taxpayers have more pay than profit.
    • The FRS has 294 such people, 1–2 per rank cell, so their main source is most of what links their draw to their own jobs.
  • The draw reaches the tails.
    • Before: microimpute 1.8.1, the locked version, picks from ten forest quantiles between 1/11 and 10/11, so no one drew from the top or bottom 9% of the forest's conditional distribution at their predictors.
    • Now: outputs are drawn at the nearest of 1,000 midpoints of that distribution, from 0.05% to 99.95%.
    • The draw reimplements microimpute's chain (each output conditioned on those drawn before) at given quantiles. At microimpute's own random draws and grid, it reproduces predict bit for bit, which is tested. A model type it does not know raises.
  • The FRS rows' dividend draw uses the same ranks. Ranks are computed once, on the full FRS at grossing weights, and both draws use them.
    • Before: every FRS adult's dividends were drawn conditioned on a random drawn pay.
    • Now: they are conditioned on pay and savings drawn at the person's own ranks.
    • The FRS dividend column itself is empty until Keep FRS-reported dividends and key them on person_id #498 lands, because frs.py keys it on person.index. Until then the dividend quantile ties and is random.

Base and overlaps. Branched from main 4cbedbe (uk-data 1.58.0), which includes #529. Not part of the released 10/8 batch; it waits for the next batched release (d833). Open PRs that also edit income.py:

Evidence

Real builds pending. Production builds (base, branch, and a placebo redraw of main) have not run yet. Post-release rebuilds run one at a time through the UK hub's rebuild queue, and each needs more than 100 GB free on the build host's data volume. This PR is fourth in that queue, and the volume had 43 GB free on 10/9. Real-build figures for hours and pay, calibration and the policyengine-uk microsimulation will be added here before the PR leaves draft.

Offline, measured by review r1. These use this PR's code and forests on the full FRS at FRS weights, with single draws in SPI 2022-23 money. They are not a build: no calibration and no policyengine-uk.

Hourly pay below £5/hour moves from full-timers to part-timers, but the overall share does not fall:

Share of employees below £5/hour Status quo This PR
All employees 17.0% 16.9%
Full-time 18.1% 15.1%
Part-time 12.7% 24.2%
  • The FRS's own figure is 1.8%.
  • The SPI employee group is 33.0m weighted, against 27.6m FRS employees. 10.8% of it has annual pay under £4,160, from part-year and marginal pay.
  • Rank draws hand that low-pay mass to the employees with the lowest FRS pay, who are mostly part-timers. The group-definition mismatch predates this PR and is listed below as a follow-up.

Dependence between incomes (weighted Spearman):

Pair SPI Forest, independent quantiles This PR Status quo
Employees: pay–savings 0.146 0.150 0.204 0.343
No earnings: pension–savings 0.397 0.486 0.551 0.658

Totals on the full FRS:

Output Status quo This PR SPI-cell expectation
Dividends (the FRS-half draw, which carries real weights) £41.3bn £60.7bn £59.9bn
Gift Aid £7.29bn £2.54bn £2.52bn
Charitable investment gifts £2.80bn £0.05bn £0.05bn
Savings interest £7.16bn £9.03bn £7.77bn
Self-employment profit £87.0bn £101.2bn £90.1bn
  • Most of this undoes microimpute 1.8's shared quantile and its ten-point grid.
  • The profit and savings overshoots come from the forests' tails, not from the ranks: independent quantiles give £103.6bn and £9.18bn.
  • Property income for people with no earnings is £23.7bn, against £20.5bn with independent quantiles.
  • These move the calibration inputs (dividends, Gift Aid relief); the real builds will show by how much after calibration.

Over £100k the rank draw matches the SPI cells on the build's SPI copy: 0.512m at FRS ranks, against 0.511m from the SPI cells and 0.524m from the forest, counting every adult in the 10,000-household copy at FRS weights. Over £150k it still tracks the forest, but the forest does not track the SPI cells:

Copy adults with pay over £150k People
SPI-cell expectation 0.205m
microimpute 1.8.1's own draw on this PR's forests (drawn on the full FRS, then subset to the copy) 0.274m
This PR, at FRS ranks 0.306m
Forest expectation 0.325m

Over 1,000 copy subsamples the rank draw's excess over the SPI cells is +45% (SD 9%), so it is not subsample noise. The 0.274m row holds the forests fixed, so it is a controlled comparison, not the base build's own result: base has one forest for people with both pay and a trade, and drawing on the copy alone restarts microimpute's random sequence. Holding the forests fixed, the rank draw adds 32k weighted people over £150k (+11.9%), and the excess over the SPI cells rises from 33% to 49%. The excess belongs to the forests, and the 1,000-point grid exposes more of it than the ten-point grid did. The EMPLOYEE, NO_EARNINGS and SELF_EMPLOYED forests are the same in base and branch. Why the forest's tail exceeds the SPI cells' for these donors is not established. The real builds will report SPI-row pay over £100k and £150k after calibration to HMRC's income bands.

The forest's top tail also depends on its training resample: EMPLOYEE pay over £150k on the full FRS ranges from 0.26m to 0.50m across four resamples, and this PR's resample gives the 0.50m. The EMPLOYEE, NO_EARNINGS and SELF_EMPLOYED forests are identical between base and branch (same group index, so the same resample seed and size), so the base–branch delta isolates the draw rule for those groups.

Invariants (tested)

policyengine_uk_data/tests/test_spi_income_rank_draw.py (Hypothesis unless noted):

  1. Quantiles lie in [0, 1). Within a cell, a lower value never gets a higher quantile, and a weighted row always gets a strictly lower one. The rows' slices tile [0, 1) and each quantile lies in its row's slice, so the quantiles are uniform in every cell.
  2. Across seeds, a row's quantile is uniform on its slice, not a fixed point in it (KS test). A lone donor does not always draw its cell's median.
  3. With equal weights and m rows per grid point, every grid point is drawn exactly m times.
  4. Quantiles do not depend on row order, on strictly increasing transforms of the values, on other cells' values, or on the scale of a cell's weights.
    • Equal seeds give equal results.
    • A row whose value is unique in its cell stays in its slice whatever the seed.
  5. Missing values or weights, negative weights and duplicate ids are rejected.
  6. Differential test: draw_at_quantiles at microimpute 1.8's random draws and grid equals microimpute's predict exactly.
  7. The first output drawn never falls as the quantile rises, and every draw lies within the training range.
    • Each output is drawn at its own quantile: raising one output's quantile never lowers it and never changes an output drawn before it.
    • An unknown model type raises.
  8. The group model draws each group at the given quantiles, gives NOT_IMPUTED rows no draw, and without quantiles is deterministic.
  9. impute_over_incomes passes quantiles by person ID: a subsample drawn with the full data's quantiles gets exactly those, and more FRS income never draws lower in a cell.
  10. draw_quantiles ranks within rank_cells at household weights: in every cell, each person's quantile lies in their slice of the cell's household weight, for every output.
  11. Outputs everyone ties on (Gift Aid) get independent quantiles: no two quantile columns share an order (|Spearman| < 0.1 over 2,000 people in one cell).
  12. On a built eFRS, SPI-synthetic employees' pay rises with hours at least half as strongly as FRS employees' does. It is skipped without a build; it passed on CI's TESTING build.

test_imputation_source_flags.py: impute_income ranks once, on the full weighted FRS, and passes those same quantiles to both draws.

test_spi_income_earnings_groups.py gains the main-source split: SPI and FRS rows with pay and a trade fall in the employee-main group exactly by the rule above.

Mutation check. These are the eight mutants from review r1, run against the targeted tests:

Mutant (review r1's numbering) At df3a633 At c134a6b
M1: midpoint of the slice instead of a uniform point survived killed
M2: ties ordered by id instead of at random survived killed
M3: one seed for every output survived killed
M4: equal weights instead of household weights survived killed
M6: national ranks instead of cell ranks survived killed
M7: the copy's draw without the full-FRS quantiles survived killed
M9: every output at the first output's quantile survived killed
M10: value order reversed killed killed

Script and log: mutate_r1_fixes.py / .log (local evidence folder).

Tests. Six targeted files on the locked microimpute 1.8.1 (test_spi_income_rank_draw, test_spi_income_earnings_groups, test_imputation_source_flags, test_income_imputation_preserves_housing_costs, test_spi_build, test_uc_gainful_self_employment): 168 passed and 8 skipped. Every skip needs a built dataset, and the differential test runs. ruff format --check . reports 243 files formatted. CI at 6c26644: 2,217 passed, 3 skipped and 1 xfailed, with test_built_enhanced_frs_spi_pay_rises_with_hours passing on the TESTING build. CI at c134a6b: green (Test passed in 47m16s). At 5079ea6, the comment-only change below: the six targeted files give 168 passed and 8 skipped, and the differential test runs. CI at 5079ea6 is green: 2,226 passed, 3 skipped and 1 xfailed, and test_built_enhanced_frs_spi_pay_rises_with_hours passes on the TESTING build.

Not in this PR

  • SPI draws are in 2022-23 money and are not uprated to the 2024-25 FRS year before stacking. This is pre-existing; Rebase SPI income draws to the FRS survey year #532 addresses it.
  • microcosm#840's other bridge predictors (industry, State Pension age) are not included.
  • Per-component ranks plus the chain's conditioning overstate the dependence between incomes (table above). This PR measures it but does not correct it.
  • Follow-up: train the forests on the SPI age band rather than a noise age within it. Everyone in a rank cell would then face the same conditional distribution, and the rank draw would reproduce it exactly.
  • Follow-up: the SPI employee group carries part-year and marginal pay that FRS employees do not. That puts very low hourly pay on SPI-synthetic part-timers.
  • Not changed: draw_quantiles(dataset) and the FRS-half impute_over_incomes(dataset, …) each build the model inputs, so a Microsimulation is built twice. That costs seconds in a build that takes tens of minutes. The FRS-half call is also due to go with Keep FRS-reported dividends and key them on person_id #498.

Review

  • r1 (Opus, independent; subfleet job 20261008-171857-ukdata-rank-draw-review): REQUEST CHANGES at df3a633. It confirmed the core: rank_quantiles is unbiased, draw_at_quantiles matches microimpute 1.8.1 exactly, and the person-ID wiring is right. Its findings and the answers:
    1. CI failed at df3a633 (Set UC gainful self-employment from the FRS main-job status #525's test). Fixed in 6c26644; CI green there.
    2. 7 of 8 mutants survived the tests. New tests (invariants 2, 5, 7, 10, 11 and the impute_income wiring test) kill all eight; see the mutation check above.
    3. "SPI distribution preserved" went too far. The claim is now qualified in code, DESIGN and this body, and the dependence is measured above. The band-trained forest is listed as a follow-up.
    4. The offline tail result (b) is the forest and its resample, not a rank bias. Explained in DESIGN.
    5. The hourly-pay shift is disclosed above, with numbers. The group mismatch is a follow-up.
    6. Keep FRS-reported dividends and key them on person_id #498 and Rebase SPI income draws to the FRS survey year #532 were misdescribed. Corrected under "Base and overlaps".
    7. The aggregate changes are disclosed above.
    8. "Percentiles of each cell" now reads "of the forest's conditional distribution".
    9. An unknown model type now raises.
    10. Duplicate ids now raise.
    11. The duplicate Microsimulation build is kept, with the reason under "Not in this PR".
    12. The MAINSRCE claim now cites the SPI documentation. The "impact table" claim now waits for the real builds.
  • r2 (delta review of c134a6b; subfleet job 20261008-215049-ukdata-rank-draw-review-r2): REQUEST CHANGES at c134a6b, for written evidence only. Code and tests pass: 168 passed and 8 skipped, the differential test runs, all eight of r1's mutants are killed, and a mixed-type prediction probe matches microimpute exactly. It marked r1's findings 1, 2, 5, 6, 7, 9, 10 and 12 addressed, 11 an acceptable deferral, and 3, 4 and 8 partly addressed. Its findings and the answers:
    1. DESIGN labelled £100k tail figures as £150k. Relabelled. The £150k figures are now shown separately, here and in DESIGN: 0.306m at FRS ranks against 0.205m from the SPI cells.
    2. The 0.802m/0.845m/0.848m figures sum the three pay groups, not EMPLOYEE alone. Relabelled; the EMPLOYEE figures (0.780m against 0.827m) are added.
    3. Nits: the cells are now described as the SPI's resolution of the forests' conditioning, not their exact conditioning. The forests also take exact age and the earlier outputs. This is fixed in the income.py comment (5079ea6, comment only), in DESIGN and in this body. DESIGN's cutoffs now read as percentiles of the forest's conditional distribution.
  • r3 (delta review of 5079ea6, GPT-6.1 Sol; subfleet job 20261009-103633-ukdata-rank-draw-review-r3): APPROVE WITH NITS at 5079ea6. All three r2 findings are addressed. It confirmed every new figure against r1's JSON, confirmed the 14 changed code lines are comments (identical Python AST), and checked the new mechanism statements in the code. Its two nits, both on written evidence, are fixed in this body and in DESIGN:
    1. The £150k table's 0.274m is microimpute's own draw on this PR's forests, drawn on the full FRS and then subset to the copy. It is a controlled comparison, not base's own result, and only the EMPLOYEE, NO_EARNINGS and SELF_EMPLOYED forests are shared with base. Relabelled.
    2. DESIGN's 0.521m forest expectation uses the eval's resamples with split forests for people with both pay and a trade, so it compares with the eval's split variant (0.5728m), not the unsplit 0.573m. Relabelled, with the three forests that sum to it.

axiom: n/a: data-construction change (SPI income imputation), no policy rule

🤖 Generated with Claude Code

MaxGhenis and others added 5 commits October 8, 2026 17:05
The SPI-synthetic copy drew each income at a random quantile of its earnings
group's forest, conditioned on age, gender and region only, so a part-timer
drew a full-timer's pay as often as a part-timer's. Each output is now drawn at
the donor's weighted rank of the same income among FRS people in the same cell
(earnings group, SPI age band, gender, region): the cells the forests condition
on. Ties, mostly zeros, are put in random order and each person takes a uniform
point in their slice of the cell's weight, so the quantiles are exactly uniform
in every cell and the SPI's distribution within each cell is kept.

The draw reimplements microimpute's chained predict at given quantiles, on 1,000
midpoints instead of microimpute 1.8's ten points between 1/11 and 10/11, which
never drew from the top or bottom 9% of a cell. At microimpute's own random
draws and grid it reproduces predict exactly (tested). Ranks come from the full
FRS at grossing weights and also drive the FRS rows' dividend draw.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
People with both pay and a trade drew from one forest, so an FRS
self-employed person with a small PAYE job drew like a salaried person with a
side trade. They are now split into employee-main and self-employed-main
groups. In the SPI the split follows MAINSRCE: 1 means pay is the main
source, 3 (sole trader) or 4 (partnership) a trade, otherwise the larger
income. MAINSRCE is the self-assessment main source HMRC stratifies the SPI
by, not the larger income. In the FRS the split follows the main job's
status, otherwise the larger income. The FRS has only about 300 such people,
one or two per rank cell, so the main source is most of what links their
draw to their own jobs.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
#525's test runs impute_income on a fake dataset with the imputation
stubbed. impute_income now also computes the draw quantiles from the FRS,
which builds a Microsimulation the fake dataset cannot support, so the test
stubs draw_quantiles too and its impute_over_incomes stub takes the
quantiles argument, as test_imputation_source_flags does.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Review r1 found 7 of 8 design-breaking mutants survived the tests. New tests:
- a row's quantile is uniform on its slice across seeds (KS), so a lone
  donor does not always draw the median;
- draw_quantiles tiles each rank cell at household weights;
- outputs everyone ties on get independent quantiles;
- each output is drawn at its own quantile;
- impute_income passes one set of full-FRS quantiles to both draws.
All eight mutants now fail.

draw_at_quantiles raises on a model type it does not know, and
rank_quantiles raises on duplicate ids. The module docstring no longer
says the SPI distribution is kept: each group's first output keeps the
forest's distribution for the cell in expectation, and later outputs carry
extra dependence. The MAINSRCE comment cites the SPI 2022-23 public use
tape documentation and lists every code.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The forests are queried at exact age (trained on uniform noise within the
SPI age band) and, after each group's first output, at the outputs drawn
before it, so the rank cells match their conditioning only up to that age
noise and the chain. Comment only (review r2 nit).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant