Skip to content

Seed the calibration weight dropout so builds are reproducible - #507

Draft
MaxGhenis wants to merge 1 commit into
mainfrom
deterministic-calibration
Draft

MaxGhenis wants to merge 1 commit into
mainfrom
deterministic-calibration

Conversation

@MaxGhenis

Copy link
Copy Markdown
Contributor

Problem

Two production builds of the same commit, with the same raw inputs, the same trained imputation models and the same locked environment, give different household weights. Four builds of main b45c373 on 30 Sep and 1 Oct had identical person and benunit tables, but nearly every household_weight differed (52,842 to 52,846 of 52,846 rows per pair), and so did the six columns create_datasets.py rescales after calibration. Headline 2026 outputs on policyengine-uk differed between them by up to about £0.8bn of income tax and £3bn of household net income.

So a rebuild cannot isolate the effect of a small data fix: the fix's effect is mixed with run-to-run noise of the same size or larger. That affects every open uk-data PR that changes inputs.

Cause

calibrate_local_areas replaces 5% of the log-weights with their mean each epoch, choosing them with torch.rand_like. That draws from torch's global generator. Nothing in the repo seeds it, and torch seeds it differently in every process: torch.initial_seed() differs between fresh interpreters, with or without importing policyengine-uk or this package.

Calibrating one saved pre-calibration dataset 24 epochs at a time, in fresh processes:

runs targets dropout weights
same seed, PYTHONHASHSEED 1 vs unset same same identical
one unseeded, one seeded (different masks) same different 52,844 of 52,846 differ
same seed, 1 vs 6 vs 12 torch threads same same 3,312 to 6,883 rows differ; total absolute difference about 1e-8 of total weight (summation order)

Change

The dropout masks come from their own torch.Generator, seeded from a new seed argument (default 0). Calibration no longer touches torch's global generator.

Invariants (tested)

  1. The saved weights depend only on the inputs and seed: not on the process, and not on how much of torch's global generator was consumed before calibration.
  2. Calibration does not consume torch's global generator.
  3. seed is used: different seeds give different weights.

test_calibrate_determinism.py checks 1 over a grid of seeds and global-generator states (and two fresh processes against the parent), 2 directly and 3 on two seeds. These are parametrized rather than Hypothesis-generated, because I didn't want to add a dev dependency while the enhanced FRS is in maintenance mode. On the old code, the two tests that don't pass seed fail: fresh processes give different weights, and calibration changes the global generator's state.

Measured with real builds

(Two builds of this commit and two of main, all sharing one frozen set of target downloads, are running; results to follow before this leaves draft.)

Effect on the published dataset

The next release after this merges uses seeded dropout masks. Its weights are one draw from the same dropout process as before, not a systematic change. Releases from then on are reproducible from their inputs.

Not in this PR (enhanced FRS is in maintenance mode ahead of the Microcosm migration)

These were measured on 1 Oct and are recorded here, not fixed:

  • Target downloads change the target set from build to build. Target sources are downloaded every time targets are collected, and failures fall back or drop targets with at most a log line. In 8 calibration runs on one dataset, the national target set came out three different ways:

    • in 2 runs, the gov.scot chargeable-dwellings workbook failed and the hard-coded Scottish band shares were used (Band H 15,739 instead of 14,481). A build logs this as one warning line;
    • in 2 runs, ONS returned 429 and the 10 ONS household-type targets dropped out. The tenure workbook also got a 429, but it succeeded on a later call in the same run, because failures aren't cached.

    With seeded dropout, the Scottish fallback alone moved weights by 5.6e-4 of total weight over 24 epochs. Two builds of the same commit are bit-identical only when their downloads agree.

  • The HMRC salary sacrifice CSV in sources.yaml returns 410 Gone in every build, including the release builds of 4 Sep and 25 Sep 2026. Its income tax and NICs relief targets have therefore been absent from published datasets; only the static contributions target is emitted.

  • Income-dependent draws reshuffle when one record changes. stack_cgt_band_donors uses Generator.choice(p=propensity), and impute_cg_to_doubled_dataset draws gain quantiles in row order. In Stop counting rent paid by boarders and lodgers as their property income #503, a change to 43 records' property income left 41 of 270 donor households in common and re-drew 23,669 people's capital gains. Per-record keyed draws (Efraimidis–Spirakis keys from a hash of the household or person id) fix both; they're worth carrying into Microcosm. The impute_income 10,000-household subsample does not depend on incomes: it zeroes weights first, so the draw is uniform over row positions.

  • Results also depend on the torch thread count (about 1e-8 relative) and on CPU architecture, so bit-identity holds on one machine, not between a Mac and a CI runner.

axiom: n/a: data build infrastructure (calibration RNG)

🤖 Generated with Claude Code

calibrate_local_areas replaced 5% of the log-weights with their mean each
epoch, choosing them with torch.rand_like. That draws from torch's global
generator, which torch seeds differently in every process and nothing in
the build seeded, so two builds of the same commit with the same inputs gave
different household weights. The dropout now has its own torch.Generator,
seeded from a new seed argument (default 0), and leaves the global
generator alone.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant