Repository navigation
Conversation
The SPI-synthetic copy drew each income at a random quantile of its earnings group's forest, conditioned on age, gender and region only, so a part-timer drew a full-timer's pay as often as a part-timer's. Each output is now drawn at the donor's weighted rank of the same income among FRS people in the same cell (earnings group, SPI age band, gender, region): the cells the forests condition on. Ties, mostly zeros, are put in random order and each person takes a uniform point in their slice of the cell's weight, so the quantiles are exactly uniform in every cell and the SPI's distribution within each cell is kept. The draw reimplements microimpute's chained predict at given quantiles, on 1,000 midpoints instead of microimpute 1.8's ten points between 1/11 and 10/11, which never drew from the top or bottom 9% of a cell. At microimpute's own random draws and grid it reproduces predict exactly (tested). Ranks come from the full FRS at grossing weights and also drive the FRS rows' dividend draw. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
People with both pay and a trade drew from one forest, so an FRS self-employed person with a small PAYE job drew like a salaried person with a side trade. They are now split into employee-main and self-employed-main groups. In the SPI the split follows MAINSRCE: 1 means pay is the main source, 3 (sole trader) or 4 (partnership) a trade, otherwise the larger income. MAINSRCE is the self-assessment main source HMRC stratifies the SPI by, not the larger income. In the FRS the split follows the main job's status, otherwise the larger income. The FRS has only about 300 such people, one or two per rank cell, so the main source is most of what links their draw to their own jobs. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
#525's test runs impute_income on a fake dataset with the imputation stubbed. impute_income now also computes the draw quantiles from the FRS, which builds a Microsimulation the fake dataset cannot support, so the test stubs draw_quantiles too and its impute_over_incomes stub takes the quantiles argument, as test_imputation_source_flags does. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Review r1 found 7 of 8 design-breaking mutants survived the tests. New tests: - a row's quantile is uniform on its slice across seeds (KS), so a lone donor does not always draw the median; - draw_quantiles tiles each rank cell at household weights; - outputs everyone ties on get independent quantiles; - each output is drawn at its own quantile; - impute_income passes one set of full-FRS quantiles to both draws. All eight mutants now fail. draw_at_quantiles raises on a model type it does not know, and rank_quantiles raises on duplicate ids. The module docstring no longer says the SPI distribution is kept: each group's first output keeps the forest's distribution for the cell in expectation, and later outputs carry extra dependence. The MAINSRCE comment cites the SPI 2022-23 public use tape documentation and lists every code. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The forests are queried at exact age (trained on uniform noise within the SPI age band) and, after each group's first output, at the outputs drawn before it, so the rank cells match their conditioning only up to that age noise and the chain. Comment only (review r2 nit). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this changes
PR #529 draws each SPI-synthetic person's incomes from SPI records in their own earnings group, so employees draw pay and the out-of-work draw none. Within the group the draw was still random: microimpute picks a random quantile, conditioned only on age, gender and region. The donor keeps their hours and FT/PT status, so a part-timer drew a full-timer's pay as often as a part-timer's. The SPI has no hours.
This PR draws each income at the person's FRS rank of that income within their cell instead of at a random quantile. This is the uk-data analogue of PolicyEngine/microcosm#840, "rank-preserving replacement". A part-timer on low FRS pay draws low SPI pay. A donor at the top of their cell draws from the top of the forest's distribution for that cell.
predictbit for bit, which is tested. A model type it does not know raises.person.index. Until then the dividend quantile ties and is random.Base and overlaps. Branched from main 4cbedbe (uk-data 1.58.0), which includes #529. Not part of the released 10/8 batch; it waits for the next batched release (d833). Open PRs that also edit
income.py:person_idand deletes the FRS-half dividend draw, so FRS rows keep their reported dividends. After it, dividend ranks matter only for the SPI copy, where this code ranks the recorded dividends with no change. This PR's FRS-half rewiring is therefore transitional. Whichever of the two merges second drops the FRS-halfimpute_over_incomes(dataset, model, ["dividend_income"], quantiles)call. Keep FRS-reported dividends and key them on person_id #498's test stub ofimpute_over_incomesthen needs thequantilesargument, as Set UC gainful self-employment from the FRS main-job status #525's did here.output_df = model.predict(input_df)inimpute_over_incomes. This PR leaves that line as it is: the rank draw happens inside the model'spredict. The two compose. The rebase is a positive scaling per variable, so it keeps ranks.Evidence
Real builds pending. Production builds (base, branch, and a placebo redraw of main) have not run yet. Post-release rebuilds run one at a time through the UK hub's rebuild queue, and each needs more than 100 GB free on the build host's data volume. This PR is fourth in that queue, and the volume had 43 GB free on 10/9. Real-build figures for hours and pay, calibration and the policyengine-uk microsimulation will be added here before the PR leaves draft.
Offline, measured by review r1. These use this PR's code and forests on the full FRS at FRS weights, with single draws in SPI 2022-23 money. They are not a build: no calibration and no policyengine-uk.
Hourly pay below £5/hour moves from full-timers to part-timers, but the overall share does not fall:
Dependence between incomes (weighted Spearman):
Totals on the full FRS:
Over £100k the rank draw matches the SPI cells on the build's SPI copy: 0.512m at FRS ranks, against 0.511m from the SPI cells and 0.524m from the forest, counting every adult in the 10,000-household copy at FRS weights. Over £150k it still tracks the forest, but the forest does not track the SPI cells:
Over 1,000 copy subsamples the rank draw's excess over the SPI cells is +45% (SD 9%), so it is not subsample noise. The 0.274m row holds the forests fixed, so it is a controlled comparison, not the base build's own result: base has one forest for people with both pay and a trade, and drawing on the copy alone restarts microimpute's random sequence. Holding the forests fixed, the rank draw adds 32k weighted people over £150k (+11.9%), and the excess over the SPI cells rises from 33% to 49%. The excess belongs to the forests, and the 1,000-point grid exposes more of it than the ten-point grid did. The EMPLOYEE, NO_EARNINGS and SELF_EMPLOYED forests are the same in base and branch. Why the forest's tail exceeds the SPI cells' for these donors is not established. The real builds will report SPI-row pay over £100k and £150k after calibration to HMRC's income bands.
The forest's top tail also depends on its training resample: EMPLOYEE pay over £150k on the full FRS ranges from 0.26m to 0.50m across four resamples, and this PR's resample gives the 0.50m. The EMPLOYEE, NO_EARNINGS and SELF_EMPLOYED forests are identical between base and branch (same group index, so the same resample seed and size), so the base–branch delta isolates the draw rule for those groups.
Invariants (tested)
policyengine_uk_data/tests/test_spi_income_rank_draw.py(Hypothesis unless noted):draw_at_quantilesat microimpute 1.8's random draws and grid equals microimpute'spredictexactly.NOT_IMPUTEDrows no draw, and without quantiles is deterministic.impute_over_incomespasses quantiles by person ID: a subsample drawn with the full data's quantiles gets exactly those, and more FRS income never draws lower in a cell.draw_quantilesranks withinrank_cellsat household weights: in every cell, each person's quantile lies in their slice of the cell's household weight, for every output.test_imputation_source_flags.py:impute_incomeranks once, on the full weighted FRS, and passes those same quantiles to both draws.test_spi_income_earnings_groups.pygains the main-source split: SPI and FRS rows with pay and a trade fall in the employee-main group exactly by the rule above.Mutation check. These are the eight mutants from review r1, run against the targeted tests:
Script and log:
mutate_r1_fixes.py/.log(local evidence folder).Tests. Six targeted files on the locked microimpute 1.8.1 (
test_spi_income_rank_draw,test_spi_income_earnings_groups,test_imputation_source_flags,test_income_imputation_preserves_housing_costs,test_spi_build,test_uc_gainful_self_employment): 168 passed and 8 skipped. Every skip needs a built dataset, and the differential test runs.ruff format --check .reports 243 files formatted. CI at 6c26644: 2,217 passed, 3 skipped and 1 xfailed, withtest_built_enhanced_frs_spi_pay_rises_with_hourspassing on the TESTING build. CI at c134a6b: green (Test passed in 47m16s). At 5079ea6, the comment-only change below: the six targeted files give 168 passed and 8 skipped, and the differential test runs. CI at 5079ea6 is green: 2,226 passed, 3 skipped and 1 xfailed, andtest_built_enhanced_frs_spi_pay_rises_with_hourspasses on the TESTING build.Not in this PR
draw_quantiles(dataset)and the FRS-halfimpute_over_incomes(dataset, …)each build the model inputs, so a Microsimulation is built twice. That costs seconds in a build that takes tens of minutes. The FRS-half call is also due to go with Keep FRS-reported dividends and key them on person_id #498.Review
subfleetjob 20261008-171857-ukdata-rank-draw-review): REQUEST CHANGES at df3a633. It confirmed the core:rank_quantilesis unbiased,draw_at_quantilesmatches microimpute 1.8.1 exactly, and the person-ID wiring is right. Its findings and the answers:impute_incomewiring test) kill all eight; see the mutation check above.subfleetjob 20261008-215049-ukdata-rank-draw-review-r2): REQUEST CHANGES at c134a6b, for written evidence only. Code and tests pass: 168 passed and 8 skipped, the differential test runs, all eight of r1's mutants are killed, and a mixed-type prediction probe matches microimpute exactly. It marked r1's findings 1, 2, 5, 6, 7, 9, 10 and 12 addressed, 11 an acceptable deferral, and 3, 4 and 8 partly addressed. Its findings and the answers:income.pycomment (5079ea6, comment only), in DESIGN and in this body. DESIGN's cutoffs now read as percentiles of the forest's conditional distribution.subfleetjob 20261009-103633-ukdata-rank-draw-review-r3): APPROVE WITH NITS at 5079ea6. All three r2 findings are addressed. It confirmed every new figure against r1's JSON, confirmed the 14 changed code lines are comments (identical Python AST), and checked the new mechanism statements in the code. Its two nits, both on written evidence, are fixed in this body and in DESIGN:axiom: n/a: data-construction change (SPI income imputation), no policy rule
🤖 Generated with Claude Code