Repository navigation
Conversation
The income QRF is trained on SPI 2022-23 amounts and draws into the FRS 2024-25 dataset, but nothing moved the draws to 2024-25 before the second-stage FRS-only QRF conditioned on them and the halves were stacked. `uprate_dataset` then uprates the whole stacked dataset by the same factor, so the SPI rows stayed about two years of earnings and price growth behind the FRS rows, and calibration had to make up the gap through weights. `rebase_spi_draws` multiplies each draw by its variable's index ratio in storage/uprating_factors.csv (the index `uprate_dataset` applies), from SPI_FISCAL_YEAR to the dataset's time_period, right after the model predicts. That covers the SPI-synthetic copy's six incomes and the FRS half's dividends. Gift Aid and qualifying-investment gifts have no index in policyengine-uk and keep their SPI amounts, as `uprate_dataset` keeps them. Property tests (hypothesis) cover identity at equal years, monotonicity, zeros, penny ranks, missing draws, composition and inversion, a differential test against `uprate_dataset`, and the call sites: the second stage and the FRS dividends see rebased draws. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The stage-2 capture now also checks that the FRS respondents the FRS-only QRF trains on, the FRS rows' undrawn incomes and an undrawn indexed money column (employee pension contributions) keep their survey-year values, so rebasing the whole copy, the training set or the FRS half's other incomes fails the test, as does rebasing after the second stage. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Hypothesis found that 2.2e-313 rebased 2022 -> 2020 -> 2021 differs from a direct 2022 -> 2021 rebase by 2e-11 relative: subnormal floats carry too few significant bits to hold a 1e-12 relative bound through two multiplications. The smallest normal floats compose exactly. No amount of money is subnormal, so the money strategy excludes them rather than loosening the bound for real amounts. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
CI's reduced build (TESTING=1, 32 epochs) of 3d82532 carried £236.1bn of above-AEA gains against HMRC's £65.9bn, just past the reduced-build bound (relative error 2.5). That is reduced-build calibration noise, not this change: #529's seeded reduced builds put main itself at about £243bn (relative error 2.7), and this branch's production build passes the strict 50% bound. This takes #529's version of the test file unchanged (from 88ef576 and d39371d): under TESTING the total must lie within a factor of 6 of HMRC's; full builds keep the 50% bound. Identical bytes, so whichever PR lands second merges cleanly. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Review r2: the comment rounded Hypothesis's counterexample to 2.2e-313, which passes; the logged value 2.2250738585e-313 is the one that fails (2.4e-11 relative). Normal floats compose within 6e-16 relative over every pair of rebasings between 2020 and 2034. Comment only. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
21 of 52 tasks
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What was wrong
datasets/imputations/income.pytrains its income QRF on SPI 2022-23 (SPI_FISCAL_YEAR = 2022) and draws into the FRS 2024-25 dataset (CURRENT_FRS_RELEASE.survey_year = 2024). The draws are the SPI-synthetic copy's six incomes plus Gift Aid and qualifying-investment gifts, and the FRS rows' dividends. Nothing moved them from 2022-23 to 2024-25 pounds:frs_only.py) trains on FRS respondents in 2024-25 pounds, but predicts each SPI row's benefits, pension contributions and tax-free savings from that row's 2022-23 incomes.stack_datasetsthen put 2022-23 and 2024-25 amounts side by side.create_datasets.pyuprates the stacked dataset as a whole (uprate_dataset, 2024 → 2025 for calibration and back), so the gap survived into calibration, and the weights had to absorb it.PolicyEngine/microcosm#879 fixed the same gap on the microcosm side.
The change
rebase_spi_draws(draws, year)multiplies each drawn column by its variable's index ratio fromSPI_FISCAL_YEARtoyear.impute_over_incomescalls it on the model's output, withyear = dataset.time_period, before anything is written. That puts it before the second stage, before stacking, and on the FRS half's dividend draw.Which index, and which way the residual errors point
The index is the one
uprate_datasetapplies:storage/uprating_factors.csv, read throughuprate_values. The same table then carries the whole dataset 2024 → 2025, so each variable has one index source across the build. I compared it with the two other places an index could come from, at the locked policyengine-uk 2.93.0 (analysis/index_compare.out):data/uprating_indices.yaml, yoy growth)upratingattributeemployment_income_before_lsr)ons.household_interest_income)uprate_dataset's table, not in this rebase, and it is a separate follow-up with its own build and impact. Once the table is regenerated, this rebase follows it automatically.uprate_datasetand policyengine-uk keep them.SPI_NOMINAL_IMPUTATIONSnames them.KeyError) instead of silently staying nominal.Draws, not training data
analysis/train_vs_draw_control.out, 10k SPI sample, 20k FRS people, microimpute 1.8.1) gives:income_spi_2022_23.pklis keyed on the SPI release and output list. Rebasing the training data would tie it to the FRS year as well.Invariants (property-tested,
tests/test_spi_income_rebasing.py)For every input:
Hypothesis found one counterexample to composition: the subnormal float 2.2250738585e-313, rebased 2022→2020→2021, misses the 1e-12 relative bound by 2.4e-11. (Rounded to 2.2e-313 it passes.)
The differential test checks
rebase_spi_drawsagainstuprate_dataseton a dataset stamped with the SPI year. Both read the same table, so it proves they cover the same columns (indexed vs nominal) with the same arithmetic. It is not an independent check of the factor values.The call-site tests check:
impute_over_incomeswrites draws in the dataset's year, whethertime_periodis an int or a string;The call-site fakes carry an employment status, so the same tests pass on main and on a merge with #529.
14 tests pass, and a mutation check kills all 15 mutants (
analysis/mutation_check.out,analysis/mutate.py). Eleven change the function: no call, reversed direction, fixed year, year + 1, additive shift, an extra nominal column, skipping unknown columns,abs, rounding, NaN → 0, wrong SPI year. Four move the rebase: after the second stage, onto the whole copy, onto the FRS half's other incomes, onto the second stage's training set.Evidence from real builds
All builds use production settings: 512 epochs,
PE_UK_DATA_OA_CLONES=1, torch seed 0, the same cached target downloads, and the locked environment (policyengine-uk 2.93.0, microimpute 1.8.1).The design is paired under two SPI draws:
The seed-43 pair sets the income QRF's prediction seed to 43 through a build wrapper, with no code change. main43 reproduces the earlier placebo exactly (same weights and fit). The "redraw" column below is main43 − base, i.e. how much a different SPI draw alone moves things.
Rebuild noise is zero. A second base build of the same commit and seed gives identical aggregates.
The builds did what the code says (
analysis/build_rebase_check*.out):Impact
policyengine-uk main 00fb451d6 (2.112.1). The effect is branch − main under each draw. GBP bn; people in thousands (
impact/paired_compare.out).In 2026 the same pattern holds:
How to read it:
The rule. An effect counts as robust when it has the same sign under both draws and each effect is larger than the redraw's. I applied it mechanically to every metric in both years (
impact/robust_classification.out).Robust in both years:
The weighted employment income on SPI-synthetic rows rises £34–43bn, and on FRS rows it falls £37–48bn. Each FRS row's amounts are unchanged, so that fall is reweighting.
Robust in one year only, small or borderline:
Not robust: these change sign between the draws, or move less than the redraw itself:
Two draws are still few. They show which effects survive a redraw; they don't bound the noise.
Calibration fit and weights
National fit (
analysis/national_fit_paired.out): the 636 national targets, recomputed for each saved file (uprated to 2025 withuprate_dataset, the build's own target matrix and cached targets).The changes go both ways between the draws and are smaller than the redraw's own. The rebase leaves the national fit essentially where it was.
Review r1 pointed to
analysis/calibration_fit_vs_placebo.out, which compared the branch with the placebo. That comparison mixes the rebase with a different SPI draw, so the paired table replaces it.Weights (
analysis/weights_summary.out):CI
CI's reduced build (
TESTING=1, 32 epochs) of 3d82532 failedtest_cgt_band_donors.py::test_built_total_gains: £236.1bn of above-AEA gains against HMRC's £65.9bn, just past the reduced-build bound.That is consistent with reduced-build under-convergence:
It does not rule out a contribution from the rebase. The capital-gains imputation conditions donor selection and gain draws on total income bands, and donor composition does change between builds. The full-build bound guards large level errors; it does not show a zero CGT effect.
604cb02 takes #529's version of that test file byte for byte (from 88ef576 and d39371d), so whichever PR lands second merges cleanly. Under
TESTINGthe total must lie within a factor of 6 of HMRC's; full builds keep the 50% bound.Locally, the full suite on this branch's production build passes, except the subnormal property fixed in 080e05d.
With #529, #498 and the rank-preserving draw
impute_over_incomesbut keeps a singleoutput_df = model.predict(input_df)line. Whichever PR lands second wraps that line inrebase_spi_draws(..., int(dataset.time_period)).out/income_py_on_529.patch; it has more hunks because it also carries this PR's own non-conflicting additions.test_spi_income_earnings_groups.py,test_imputation_source_flags.py,test_income_imputation_preserves_housing_costs.py,test_cgt_band_donors.py) pass: 34 passed, 10 skipped for built data.apply_income_drawswrites zero for drawn rows and keeps undrawn rows' own 2024 values.test_second_stage_and_frs_dividends_see_rebased_draws(its docstring says so), and the FRS-dividend part of this PR no longer applies.The hypothesis dev dependency is a cherry-pick of #529's 313468b (the same lines as #514 and #524).
Follow-ups (not in this PR)
storage/uprating_factors.csvfrom the locked policyengine-uk's load-time path, and add a test that the table matches the engine. Self-employment 2022→2024 is the stale entry this PR touches; 2024→2025 is off for most incomes.savings_interest_income.upratingsays household interest income, whileuprating_indices.yamluprates it by per-capita GDP.axiom: n/a: uk-data input construction, no policy rule encoded.
Lands with the batched uk-data release (d833). Do not merge outside it.
🤖 Generated with Claude Code