Skip to content

Regenerate the uprating factor table from policyengine-uk's load-time uprating - #541

Open
MaxGhenis wants to merge 5 commits into
mainfrom
uprating-factors-v2
Open

MaxGhenis wants to merge 5 commits into
mainfrom
uprating-factors-v2

Conversation

@MaxGhenis

@MaxGhenis MaxGhenis commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Post-batch refresh (10/7)

New head: 31888a2d52aaf49bd860941a9d5e1ad111275d1d, a fast-forward descendant of c4c24a261e123d39f953201718551fcb0d17fa18. The merge commit has post-batch main 4cbedbecca352f52dffe47e849553c2759a00711 as a parent. The planned 10/8 batch landed on 10/7 as 1.58.0; this PR remains for the next batch, on Max's go (d833).

Each requested test file ran sequentially in its own foreground pytest process, with -p no:cacheprovider and no xdist:

File Passed Skipped Failed Errors
test_uprating_factors_table.py 30 0 0 0
test_income_projection.py 10 0 0 0
test_road_fuel_volume_uprating.py 15 0 0 0
test_salary_sacrifice_headcount.py 2 3 0 0
test_uprating_range.py 3 0 0 0
test_pension_credit_reported_capital.py 6 0 0 0
test_hmrc_salary_sacrifice_targets.py 22 1 0 0

Total: 88 passed, 4 skipped, 0 failed, 0 errors. The four skips require built enhanced FRS. The table/headcount pair remains 32 passed, 3 skipped, matching the recorded r3 run at the old head. The runbook's #532 SPI-rebasing and #506 boarder/lodger test files are absent because those drafts are not on main.

Impact remains pending a real rebuild after this refresh. No dataset build was run. The hub schedules that rebuild and the delta review. Evidence is untracked under _build/hub-evidence/ in the assigned workspace.

What was wrong

The diagnosis and numerical comparisons below record the original pre-batch evidence at policyengine-uk 2.93.0. The refreshed branch uses the post-batch lock (policyengine-uk 2.122.2 / core 3.32.13); current results are recorded in the "Post-batch refresh (10/7)" section.

storage/uprating_factors.csv was last built on 2026-05-20, from each variable's uprating attribute. uprate_dataset uses it to move the FRS 2024 build to the 2025 calibration year and back. uprate_values uses it to project the HMRC SPI income targets (incomes_projection.csv) and the CGT targets.

The original locked policyengine-uk (2.93.0) moves a saved dataset differently. Simulation calls extend_single_year_dataset, which projects the file to 2030. Each year, apply_single_year_uprating multiplies every variable listed in data/uprating_indices.yaml by one plus the year-on-year growth parameter it is listed under. Variables the file doesn't list are carried forward unchanged. Three load-time changes are not single indices:

  • council tax is uprated by country;
  • rent is uprated by region and tenure;
  • student loan plans are reassigned, zeroing the repayments of loans the engine writes off.

policyengine-core's attribute-based uprating (simulation.py, the variable.uprating branch) runs only for periods with no stored value. For a dataset column, that means after 2030.

So calibration fitted 2025 values the model does not run on. Growth from 2024 to 2025 (check/table_evidence.out):

variable main's table main's generator, rerun on 2.93.0 this PR policyengine-uk at load
employment income 3.73% no row 5.30% 5.30% (obr.average_earnings)
self-employment income 4.65% 0.76% 0.69% 0.69% (obr.per_capita.mixed_income)
savings interest 2.77% 3.50% 4.38% 4.38% (obr.per_capita.gdp)
dividends, property income, capital gains 2.77% 4.32% / no CGT row 4.38% 4.38% (obr.per_capita.gdp)
private pension income 4.74% 3.62% 3.59% 3.59% (obr.private_pension_index)
state pension 3.23% no row 3.40% 3.40% (obr.consumer_price_index)
salary-sacrificed pension contributions 3.73% 5.26% no row 0 (not listed)

Rerunning the old generator wouldn't fix it. On 2.93.0 it emits no row for employment_income (an adds formula with no attribute), capital_gains, state_pension, state_pension_reported, employee_pension_contributions or student_loan_repayments. A missing capital_gains row silently truncates the CGT target projection. It also indexes savings interest by household interest income (×2.27 for 2022→2024), where the engine uses GDP per head (×1.092). policyengine-uk#1862 tracks that attribute mismatch and the four declared-but-unlisted variables.

Over the 75 rows main and this PR share, the 2024→2025 factor moves by between −3.8% and +1.6%. Only 3 rows are unchanged to 0.1%.

The change

  • Generator (utils/uprating.py). policyengine_uk_load_time_index() compounds the growth parameter each variable is listed under in the locked policyengine-uk's uprating_indices.yaml. It raises if a variable is listed twice. build_uprating_factors_table() then applies the existing household-weight and road-fuel litre-proxy overrides. Levels are stored to six decimal places: at three, rounding alone moved a year-on-year factor by up to 0.07%.

  • Table. On the refreshed 2.122.2 lock, it drops five rows from post-batch main that the engine does not list for load-time uprating:

    • benunit_rent: an adds aggregate, so no stored column changes;
    • domestic_energy_consumption: absent from the locked engine's load-time index list;
    • housing_service_charges, pension_contributions_via_salary_sacrifice and water_and_sewerage_charges: policyengine-uk carries these unchanged, so calibration now does too.

    Relative to post-batch main it gains bus_fare_spending and bus_fare_spending_reported (CPI), private_pension_wealth, capital_gains_badr, capital_gains_carried_interest, capital_gains_residential_property, household_lifetime_isa_balance and lifetime_isa_balance (GDP per head), and domestic_rates (Northern Ireland domestic rates). Use the FRS benefit-unit capital for Pension Credit's capital test #513's pension_credit_reported_capital row is retained (GDP per head). The table has 84 rows.

  • incomes_projection.csv is regenerated (see the check below).

  • Salary sacrifice headcount targets (targets/compute/income.py). compute_ss_headcount used to deflate contributions to 2023 prices through the salary sacrifice row before applying the £2,000 cap. That row no longer exists. The loss matrix catches the resulting KeyError and skips the target (build_loss_matrix.py, "Skipping target"), so all three headcount targets would have dropped out.

    It now classifies the contributions the calibration-year simulation holds, which on the refreshed 2.122.2 lock are still the 2024-25 survey amounts. These are the values test_salary_sacrifice_headcount's built-data tests already classify. The cutoff moves. Main's uprating and deflation together compared survey amounts with an effective £2,092.95 (2,000 × 1.261 / 1.205, main's 2024 and 2023 rows), a leftover of the 2023-24 base year. This PR compares them with £2,000.

    There are two readings: survey-year contribution groups, or nominal calibration-year contributions. Today they coincide, because policyengine-uk carries this variable unchanged at load. They separate by a year of earnings growth once policyengine-uk#1863 makes the engine uprate it. A new test (test_calibration_year_contributions_are_the_survey_amounts) fails when the table gains that row, which forces the choice at that point. The choice is queued for Max as d979, recommending survey-year groups because the targets grow both groups at the same rate.

  • create_datasets.py: a comment only. rail_usage is not re-uprated at load, and bus_fare_spending now moves by the same CPI factor both ways.

incomes_projection.csv: the table is the only cause of the change

check/incomes_projection_check.py and its .out file record this:

  1. Main's table (b45c373), run through this repo's unchanged project_income_table on the committed incomes.csv, rebuilds main's incomes_projection.csv byte for byte.
  2. This PR's table, run through the same code on the same input, rebuilds this PR's incomes_projection.csv byte for byte. As a control, main's table does not reproduce this PR's file.
  3. The committed incomes.csv equals the live HMRC SPI table that create_income_projections() downloads.
  4. This PR changes neither incomes_projection.py, hmrc_spi.py, spi.py nor incomes.csv.
  5. No count cell and no 2024 cell moves. Every amount cell moves by exactly its variable's table ratio, to within £1 rounding.

All-band totals, this PR / main, 2025 and 2029:

variable 2025 2029
employment income 1.0152 1.0252
self-employment income 0.9622 0.9653
state pension 1.0016 1.0062
private pension income 0.9890 1.0154
property, savings and dividend income 1.0156 1.0122

Invariants, tested

tests/test_uprating_factors_table.py checks these for every row, and for every year or pair of years:

  • The committed table is exactly what the generator builds from the locked policyengine-uk. Its rows are exactly the variables policyengine-uk uprates at load. A policyengine-uk bump that changes a growth parameter or the YAML fails this test until the table is regenerated.
  • Outside the overridden rows, each row grows each year by one plus its engine growth parameter, within the six-decimal rounding bound.
  • Differential against the engine itself. extend_single_year_dataset projects a dataset of ones from 2020 to 2034 and agrees with the table for every row and year. Fuel spending × household weight agrees with the engine's.
  • Differential through the build path. uprate_dataset from the FRS base year gives what policyengine-uk's scalar YAML growth gives in each year from 2024 to 2030. That includes the four columns both paths carry unchanged.
  • The one load-time change outside the table that touches a table row. Repayments on a kept Plan 2 loan agree on both paths. Only the engine zeroes a written-off Plan 1 loan (test_student_loan_plan_reassignment_is_outside_the_table).
  • Every row is a level index: 1 in 2020, finite and above 0.5. So uprate_dataset there and back between any two of the 15 years is the identity (225 pairs).
  • Every variable a target projection uprates (the SPI income variables, household_weight, capital_gains) has a row.
  • A variable listed under two indices raises.

test_income_projection.py adds a check that incomes_projection.csv is the committed SPI table projected with the committed table. test_salary_sacrifice_headcount.py adds two tests:

  • a property test of the cap split on edge cases and 1,000 random amounts: below-cap and above-cap partition the users, and no amount is adjusted;
  • the guard described above, which fails if the table gains a salary sacrifice row.

Tests before the refresh

The pre-refresh author and reviewers recorded the following results. On Python 3.13 with the then-locked environment (uv sync --frozen --extra dev; no lock change):

  • This PR's head. test_uprating_factors_table, test_income_projection, test_road_fuel_volume_uprating, test_salary_sacrifice_headcount, test_uprating_range, test_spi_build, test_property_income_targets, test_frs_prerequisites, test_local_la_extras, test_la_loss_missing_sources and test_hmrc_cgt_targets give 116 passed, 4 skipped at cc5b4c1. The skips are the four tests that need the built enhanced FRS. At 8d8b608, before the two tests review r1 prompted, the same set gave 114 passed, 4 skipped. c4c24a2 changes only that fixture's ages, and its two test files give 32 passed, 3 skipped.
  • A trial merge with Rebase SPI income draws to the FRS survey year #532's head c02b070. Rebase SPI income draws to the FRS survey year #532's test_spi_income_rebasing (Hypothesis) and test_cgt_band_donors, plus this PR's test_uprating_factors_table and test_income_projection, give 56 passed, 5 skipped (built-data gated).
  • Lint. ruff format --check and ruff check on the changed files are clean.

How this interacts with #532 (SPI income rebase)

The pre-batch trial merge with #532 was clean. #532 did not enter the batch: it remains an open, conflicting draft. The intended interaction is by design. #532's rebase_spi_draws reads this same table through uprate_values, so once both are on main the SPI draws follow the regenerated table. That is the follow-up #532 lists. The pre-batch 2022→2024 factor comparison was:

variable main's table this PR
employment income 1.1189 1.1164
self-employment income 1.0734 1.0211
savings, dividend and property income 1.0897 1.0924
private pension income 1.1026 1.1025
Gift Aid, charitable investment gifts no row no row

These are the "policyengine-uk at load" column of #532's own table. Gift Aid and qualifying-investment gifts stay nominal, as #532's SPI_NOMINAL_IMPUTATIONS expects.

Order: whichever of #532 and this PR lands second must regenerate both uprating tables and incomes_projection.csv with the locked engine. #532 is still an open draft and was excluded from the batch. Its earlier measured impact used the old table; this PR's pending rebuild will measure the refreshed branch against post-batch main, which does not include #532.

Batch and later interactions

The planned 10/8 batch landed on 10/7 as uk-data 1.58.0 (4cbedbecca). #513 landed and introduced the pension-capital row conflict in both uprating CSVs. #501, #506 and #528 remain open drafts; their table-row interactions below are future dependencies. #533 and #529 landed, and their source/test changes are preserved in the refresh.

PR rows in policyengine-uk's YAML?
#513 pension_credit_reported_capital yes, under obr.per_capita.gdp (2.107.0 and main)
#506 rent_paid_as_boarder, rent_paid_as_lodger after PE-UK #2002
#501 private_pension_wealth yes, under obr.per_capita.gdp (already locked 2.93.0; this PR adds the row)
#501 cash_isa, directly_held_shares, stocks_and_shares_isa, unit_and_investment_trusts after PE-UK #2131 or #1974
#501 the three secured-debt columns after PE-UK #1969
#528 (next release) trading_loss after PE-UK #2085

The refresh resolves the two table conflicts by regenerating both files, never by hand-merging them. Each row then follows the engine the lock points to: rows whose variables the locked engine lists are kept, and the rest are carried unchanged, as the engine carries them.

The refresh retires #513's test_uprating_rows_match_the_model, which expected the attribute-based 3-decimal row, in favour of this PR's tests of every row against the engine. All other pension-capital tests are retained. #506's future test_rent_paid_is_uprated_* tests require the PE-UK #2002 rows in the lock; #506 is not on main.

Impact

Impact: pending a real rebuild after this refresh on post-batch main (disk rule). No dataset was built for this PR or during the refresh. Runbook: ~/reviews/uk-hub/jobs/rescue-uprating-factors/rebuild/README.md. With the refreshed branch and regenerated tables:

K=~/reviews/uk-hub/jobs/rescue-uprating-factors/rebuild
$K/run_build.sh main <post-batch main worktree> 0      # base: production settings, OA clones 1, torch seed 0
$K/run_build.sh placebo <post-batch main worktree> 1   # noise floor: same commit, torch seed 1
$K/run_build.sh branch <this branch's worktree> 0
for k in main placebo branch; do <pe-uk venv>/bin/python $K/dataset_impact.py $K/efrs_$k.h5 $K/impact_$k 2025,2026; done
<pe-uk venv>/bin/python $K/compare.py $K/impact_main.json $K/impact_branch.json $K/impact_placebo.json

run_build.sh holds the host-wide build lock and runs one build at a time. It needs more than 60 GB free.

Known gaps, unchanged here

These three load-time changes are not single indices, and neither this table nor main's covers them:

  • Council tax and rent. policyengine-uk uprates council_tax by country and rent by region and tenure at load. Calibration therefore sees their 2024 amounts in 2025 (NOT_SINGLE_INDICES in the test module).
  • Student loan plans. policyengine-uk reassigns plans each year at load and zeroes the repayments of written-off loans. The table grows every borrower's repayments by average earnings. test_student_loan_plan_reassignment_is_outside_the_table pins exactly where the two paths diverge.

Review

This PR was excluded from the batch that landed on 10/7 as 1.58.0. It remains for the next uk-data batch, on Max's go (d833).

axiom: n/a: uk-data input construction, no policy rule encoded.

🤖 Generated with Claude Code

MaxGhenis and others added 5 commits October 5, 2026 17:16
… uprating

storage/uprating_factors.csv was last built on 2026-05-20 from each
variable's `uprating` attribute. policyengine-uk 2.93.0 projects a saved
dataset with the year-on-year growth parameters in
data/uprating_indices.yaml instead, so the build calibrated 2025 values the
model does not run on: earnings were uprated 3.7% for 2024-25 against the
engine's 5.3%, self-employment income 4.7% against 0.7%, and GDP-indexed
incomes and wealth 2.8% against 4.4%. Regenerating with the old generator
would not fix it: on 2.93.0 it emits no employment_income row (an `adds`
formula with no attribute) and indexes savings interest by household
interest income where the engine uses GDP per head.

The generator now compounds the growth parameter each variable is listed
under in uprating_indices.yaml, keeps the household-weight and road-fuel
overrides, and stores six decimal places. The table loses four rows the
engine does not uprate at load (benunit_rent, housing_service_charges,
pension_contributions_via_salary_sacrifice, water_and_sewerage_charges)
and gains bus_fare_spending and private_pension_wealth.

incomes_projection.csv is regenerated: the table is its only input that
changed (rebuilding it from main's table reproduced main's file exactly).
compute_ss_headcount deflated contributions through the salary sacrifice
row, which no longer exists; the loss matrix would have skipped all three
headcount targets. It now classifies contributions as simulated in the
calibration year, the values test_salary_sacrifice_headcount checks.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Review r1 of #541: policyengine-uk also reassigns student loan plans at
load and zeroes the repayments of loans it writes off, which the table's
single average-earnings index cannot represent. The generator docstring
and the test module now name it beside council tax and rent, and a test
projects a kept Plan 2 loan and a written-off Plan 1 loan through both
paths: repayments agree for the first, and only the engine zeroes the
second. The docstring also states when policyengine-core applies a
variable's `uprating` attribute: only to periods with no stored value.

A second test fails if the table gains a salary sacrifice row, as
policyengine-uk#1863 would add. The headcount targets then move from
survey-year to nominal calibration-year contributions, a methodology
choice to make at that point.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Review r2 of #541: policyengine-uk carries ages unchanged at load, so ages
taken from the base year put the engine's university starts at 2013 and
1983 rather than the 2012 and 1982 the comment names. Deriving them from
the projected year makes the cohorts exact; 2012 is the first Plan 2 year.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Merge uk-data 1.58.0 (4cbedbe) into uprating-factors-v2. Resolve both
uprating CSV conflicts with the PR generator under policyengine-uk 2.122.2
and core 3.32.13, retaining pension_credit_reported_capital. Regenerate
the HMRC income projection and retire the obsolete attribute-based
three-decimal pension-capital assertion in favour of all-row engine checks.

Preserve restored salary-sacrifice relief, the direct 2,000 pound headcount
cutoff and d979 guard, and the batch pension-age mask pipeline wiring.
The release updated pyproject.toml but left the editable package version
at 1.57.4 in uv.lock, making uv lock --check fail. Update only that metadata
entry; dependency versions and hashes, including UK 2.122.2 and core
3.32.13, remain unchanged.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant