Skip to content

Rebuild a subsampled simulation from its inputs, not its calculated values - #573

Merged
MaxGhenis merged 16 commits into
masterfrom
fix-subsample-inputs-only
Oct 10, 2026
Merged

MaxGhenis merged 16 commits into
masterfrom
fix-subsample-inputs-only

Conversation

@MaxGhenis

@MaxGhenis MaxGhenis commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Subsampling rebuilds from surviving source inputs recorded by the holder, including explicit overrides of formula-backed variables. Calculated values stay calculated; deleted or cache-replaced inputs cannot become source inputs again. Loaded IDs, memberships and roles provide any missing structural leaves, and pre-sample calculation branches are recreated against the retained populations.

Merged master at a934dc3465cd3284b6d21fb6fd7a0506b3e6581b with merge commit a19b8eaffab026ce18cfea81b00e60bb572d12e0, incorporating #558, #561, #571 and #581. The subsequent normalization fix is d5e775675dd9fe5c3cc1fd268bfae07d4b5d7605. The dump/restore conflict uses master's _restore_input implementation, so restored values enter #561's per-tier input record through the holder lifecycle. Dataframe export, dictionary export and subsample rebuilds share that record and read surviving inputs from their recorded tiers. An unrecorded memory cache cannot replace a recorded disk source input during export. The existing structural-only default public export and include_computed_variables=True behavior are retained.

d1026 Proposal A: each recorded weight column preserves its own entity's source total. Normalization groups by the exported membership column that the rebuild reads, addressing review r2 P3-1 when membership is overridden after dataset load. Source totals also come from recorded input tiers. Zero-total inputs stay zero; an unquantized sample retaining only zero weights for a positive source total raises before replacing the dataset.

Requested-period sampling weights still use available inputs, formulas, uprating or carry-over, falling back to recorded dataset-period weights when no requested-period source exists. Missing both sources raises a clear error. Legacy dumps without provenance may still restore calculated values as inputs.

Authorization and hub gates

Max gave go under d1118 on 2026-10-09, subject to the hub gates. The hub's current-head US and UK microsimulations, marginal-tax-rate check and partner-output check are in "Hub validation" below. Existing #570/#572 holds and d1024's fresh-parameter-lookup contract remain in force.

Master contains the earlier d899 sequence, including #561, #571 and #581, with #558 alongside. #601 still overlaps structural rebuilding; #560 composition must retain master's restored-input record and provenance lifecycle. Historical composition results do not validate the updated head.

Master merge (2026-10-10): Merge commit 372b328638fcc35f6f65ca607e44717534ce02fd includes current master (78a6c503ca6623f2735410412b8ca9e93816ec0d) and #560 (b5beeda37c9e79b50b0310ec0944f00a831cd71b), with no conflicts or additional code changes. Combination validation at this merge: 430 passed across 24 files — 151 in all 12 subsample files, test_dump_restore.py and the selected subsample case of test_clone_reform_baseline.py, plus 279 in all 10 #560 test files; its supporting fixture module collected no tests. Each file ran separately in the foreground with uv run pytest -q under Python 3.13.9, with Hypothesis integration enabled. Source review found no composition defect: export and invalidation share surviving per-tier input provenance, deleted overrides stay excluded, subsampling resets store history and recreates branches on the sample, and restore loads inputs before calculated arrays. Formatting, lint and git diff --check pass. The hub's microsimulations and downstream checks are below.

Validation

Focused validation: 278 passed across 19 files — 172 subsample, dump/restore and clone/baseline checks, plus 106 checks across all five requested #561 input-record files. These include all ten added regression cases and property tests configured for up to 150 input-rebuild examples, 35 sampling differential examples and 200 memory/disk provenance sequences.

The required single tests/core run completed with 2,298 passed, 1 skipped, 3 xfailed and 457 warnings, exit code 0. Ruff formatting and lint checks and git diff --check pass for the final PR diff. All thirteen changed Python files parse successfully.

Focused tests ran one file at a time in the foreground with uv run pytest -q -p no:cacheprovider, using Python 3.13 development dependencies, checkout-pinned core imports and explicit Hypothesis pytest integration. The broad command was uv run pytest -q -p no:cacheprovider tests/core, run once after the focused files. Unrelated pytest and Hypothesis plugin discovery was disabled; property tests remained enabled. An initial deleted-input invocation with startup diagnostics exited 139 before reporting any cases. The rerun without diagnostic dumps passed all 21 cases, and the subsequent broad suite passed.

Documentation review

The public export and subsample docstrings describe recorded input tiers, sampling fallback, structural leaves, native entity totals and branch recreation; the existing API autodoc exposes them. The existing single towncrier fragment covers the behavior change. No generated documentation was edited. A documentation build is outside this validation run because the generator and templates are unchanged. Impact: medium. Confidence rests on the hub validation below.

Remaining review notes

Review r2 P3-2 (datasets without time_period) and P3-3 (parameter-based abolition in sampling-weight source selection) are unchanged and outside this merge reconciliation. The hub's validation of that change's downstream effects is below.

Hub validation (2026-10-10, at this head 372b328638)

All runs use the default PolicyEngine-US dataset on PE-US main 75cdd8019e, one at a time under the machine's heavy-job lock.

Full-sample control (no subsample), 2026. Core master 78a6c503ca against this head: household_net_income, income_tax, state_income_tax and household_weight are identical in every record.

Marginal tax rates, 2026. Core a934dc3465 (master before #560) against this head, which contains #560 and this PR: marginal_tax_rate and household_net_income are identical in every record. Wall time was 543 s against 566 s.

Subsample A/B, 2026, n = 10,000 households, three seeds. Core master 78a6c503ca against this head. Ratios are the subsample total over the full-sample total.

Seed Weights Side household_weight person_weight tax_unit_weight household_net_income income_tax zero/NaN weights*
0 quantized base 1.0012 1.0000 1.0133 0.9934 0.9732 0
0 unquantized base 0.9702 1.0000 0.9809 1.0150 1.0828 0
0 quantized PR 1.0000 0.9988 1.0122 0.9923 0.9721 0
0 unquantized PR 1.0000 1.0307 1.0110 1.0462 1.1161 0
1 quantized base 0.9993 1.0000 1.0024 1.0016 0.9803 0
1 unquantized base 0.9909 1.0000 1.0014 0.9962 0.9664 0
1 quantized PR 1.0000 1.0007 1.0031 1.0024 0.9810 0
1 unquantized PR 1.0000 1.0092 1.0106 1.0053 0.9753 0
2 quantized base 1.0044 1.0000 1.0012 1.0130 1.0089 0
2 unquantized base 1.0090 1.0000 1.0062 1.0435 1.0870 0
2 quantized PR 1.0000 0.9956 0.9968 1.0085 1.0045 0
2 unquantized PR 1.0000 0.9911 0.9972 1.0343 1.0774 0

*family_weight is zero on both sides (no data and no formula) and is excluded.

  • The dataset records only household_weight. In PolicyEngine-US, person_weight, tax_unit_weight, spm_unit_weight and marital_unit_weight are formulas that copy the household's weight.
  • This head preserves the recorded column's own household total exactly in all six runs, which is d1026 Proposal A. Master preserved the repeated person-row sum, so its person total is exact instead.
  • The total that is not preserved varies with the sample, and its sign changes across seeds. Neither side is systematically closer to the full sample on dollar aggregates.
  • No run produced a zero or NaN weight.

Partner contract tests. The PolicyEngine-US partner files (policyengine_us/tests/policy/baseline/partners, 145 files) pass with this head: 623 passed, 0 failed.

PolicyEngine-UK. PE-UK 2.122.2 with a cached populace_uk_2023.h5 (a July 2026 build, not the current production file), 2026:

subsample exported every stored value, calculated ones included, and
loaded them all back as inputs. A formula result calculated before
subsampling then replaced its formula for good: it was carried over
past the formula's end and survived apply_reform and later set_input
calls, so results after subsampling depended on what had been
calculated before.

It now exports the values the simulation was given (loaded from the
dataset or passed to set_input), for variables with a formula too, so
formula-backed structural IDs from the dataset are still kept.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@MaxGhenis

Copy link
Copy Markdown
Contributor Author

#601 supersedes this PR. It keeps the same rule (keep inputs, including those of formula-backed variables, and drop calculated values), with three differences:

  • It reads the storage's own input mark (Holder.get_input_periods) rather than _user_input_keys, so it doesn't depend on Keep the record of set_input values in step with storage #561.
  • It calculates a formula-backed id the dataset lacks, since the rebuild needs it.
  • It tests the reform simulation's baseline arm, where the leak showed: the baseline read the reform's values after subsampling.

#601 carries this PR's tests unchanged, and they pass there. Run against this branch merged locally onto current master, #601's tests fail in three cases: a value calculated where an input was deleted, and a formula-backed household_id the dataset lacks, with and without calculating it first. I suggest closing this one when #601 merges.

MaxGhenis and others added 11 commits October 7, 2026 16:53
Preserve current storage provenance and baseline binding while retaining the original inputs-only export.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Preserve the baseline policy while recreating calculation branches on the sampled population. Extend the existing differential property to formulas that create branches.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Keep restored input provenance in the input export registry while excluding values listed in derived_periods.txt. Add a differential restore and subsample regression without relying on PR #560.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Calculate sampling probabilities from the requested period independently of the input export. Keep existing input-column normalization and recompute derived weights on the sample. Add focused differential and property coverage.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Use the flat-file loader\x27s recorded membership labels when an input-only household ID is absent. Keep calculated default IDs out of the rebuilt dataset and verify complete household partitions with shuffled, nonconsecutive labels.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Extend differential sampling examples and properties to monthly requests. Verify annual FLOW weights divide by twelve while normalized probabilities, source inputs and the default calculation period stay consistent.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@MaxGhenis
MaxGhenis marked this pull request as ready for review October 9, 2026 15:56
@MaxGhenis
MaxGhenis merged commit 4ba5bed into master Oct 10, 2026
19 checks passed
@MaxGhenis

Copy link
Copy Markdown
Contributor Author

US + core hub merge audit, core#573 at 372b328638fcc35f6f65ca607e44717534ce02fd (squash): a subsampled simulation is rebuilt from its inputs, and each recorded weight column keeps its own entity's total.

  • Max's rulings: d1026 Proposal A (each weight column preserves its native entity total) and d1118 (merge go, 2026-10-09).
  • Review: five independent rounds. GPT-6.1 Sol r5 approves this exact head. It reasoned through seven scenarios for the interaction with Make set_input on a branch drop values calculated from the input it replaces #560 (merged as b5beeda37c) and found no defect. Its sandbox could not run tests, so execution evidence is CI at this head and the build's 430 passes across both PRs' test files.
  • Evidence at this head (in the body's "Hub validation" section):
    • Full-sample control: identical in every record.
    • Marginal tax rates: identical in every record against core before Make set_input on a branch drop values calculated from the input it replaces #560.
    • Subsample A/B over three seeds: the recorded household weight total is exact in all six runs, with no zero or NaN weights.
    • PolicyEngine-US partner contract tests: 623 of 623 pass.
    • PolicyEngine-UK: full-sample totals are identical. subsample() fails identically on master and here (policyengine-uk#2254), so nothing changes for PE-UK.
  • Gates checked live at merge:
    • gh pr checks exits 0 (19/19);
    • the latest Pull request run at the head concluded success;
    • MERGEABLE, not a draft, no CHANGES_REQUESTED;
    • master is still 78a6c503ca, the base of the A/B;
    • head pinned.
  • Follow-up: core#601 has the same title and addresses the same defect; the hub checks whether it holds anything this PR lacks, then closes or salvages it.
  • --admin: used only because a required approving review is the sole block.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant