Repository navigation
Conversation
calibrate_local_areas replaced 5% of the log-weights with their mean each epoch, choosing them with torch.rand_like. That draws from torch's global generator, which torch seeds differently in every process and nothing in the build seeded, so two builds of the same commit with the same inputs gave different household weights. The dropout now has its own torch.Generator, seeded from a new seed argument (default 0), and leaves the global generator alone. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This was referenced Oct 2, 2026
21 of 52 tasks
This was referenced Oct 2, 2026
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
Two production builds of the same commit, with the same raw inputs, the same trained imputation models and the same locked environment, give different household weights. Four builds of main
b45c373on 30 Sep and 1 Oct had identical person and benunit tables, but nearly everyhousehold_weightdiffered (52,842 to 52,846 of 52,846 rows per pair), and so did the six columnscreate_datasets.pyrescales after calibration. Headline 2026 outputs on policyengine-uk differed between them by up to about £0.8bn of income tax and £3bn of household net income.So a rebuild cannot isolate the effect of a small data fix: the fix's effect is mixed with run-to-run noise of the same size or larger. That affects every open uk-data PR that changes inputs.
Cause
calibrate_local_areasreplaces 5% of the log-weights with their mean each epoch, choosing them withtorch.rand_like. That draws from torch's global generator. Nothing in the repo seeds it, and torch seeds it differently in every process:torch.initial_seed()differs between fresh interpreters, with or without importing policyengine-uk or this package.Calibrating one saved pre-calibration dataset 24 epochs at a time, in fresh processes:
PYTHONHASHSEED1 vs unsetChange
The dropout masks come from their own
torch.Generator, seeded from a newseedargument (default 0). Calibration no longer touches torch's global generator.Invariants (tested)
seed: not on the process, and not on how much of torch's global generator was consumed before calibration.seedis used: different seeds give different weights.test_calibrate_determinism.pychecks 1 over a grid of seeds and global-generator states (and two fresh processes against the parent), 2 directly and 3 on two seeds. These are parametrized rather than Hypothesis-generated, because I didn't want to add a dev dependency while the enhanced FRS is in maintenance mode. On the old code, the two tests that don't passseedfail: fresh processes give different weights, and calibration changes the global generator's state.Measured with real builds
(Two builds of this commit and two of main, all sharing one frozen set of target downloads, are running; results to follow before this leaves draft.)
Effect on the published dataset
The next release after this merges uses seeded dropout masks. Its weights are one draw from the same dropout process as before, not a systematic change. Releases from then on are reproducible from their inputs.
Not in this PR (enhanced FRS is in maintenance mode ahead of the Microcosm migration)
These were measured on 1 Oct and are recorded here, not fixed:
Target downloads change the target set from build to build. Target sources are downloaded every time targets are collected, and failures fall back or drop targets with at most a log line. In 8 calibration runs on one dataset, the national target set came out three different ways:
With seeded dropout, the Scottish fallback alone moved weights by 5.6e-4 of total weight over 24 epochs. Two builds of the same commit are bit-identical only when their downloads agree.
The HMRC salary sacrifice CSV in
sources.yamlreturns 410 Gone in every build, including the release builds of 4 Sep and 25 Sep 2026. Its income tax and NICs relief targets have therefore been absent from published datasets; only the static contributions target is emitted.Income-dependent draws reshuffle when one record changes.
stack_cgt_band_donorsusesGenerator.choice(p=propensity), andimpute_cg_to_doubled_datasetdraws gain quantiles in row order. In Stop counting rent paid by boarders and lodgers as their property income #503, a change to 43 records' property income left 41 of 270 donor households in common and re-drew 23,669 people's capital gains. Per-record keyed draws (Efraimidis–Spirakis keys from a hash of the household or person id) fix both; they're worth carrying into Microcosm. Theimpute_income10,000-household subsample does not depend on incomes: it zeroes weights first, so the draw is uniform over row positions.Results also depend on the torch thread count (about 1e-8 relative) and on CPU architecture, so bit-identity holds on one machine, not between a Mac and a CI runner.
axiom: n/a: data build infrastructure (calibration RNG)
🤖 Generated with Claude Code