Skip to content

Rebuild a subsampled simulation from its inputs, not its calculated values - #601

Closed
MaxGhenis wants to merge 1 commit into
masterfrom
fix-subsample-computed-export
Closed

MaxGhenis wants to merge 1 commit into
masterfrom
fix-subsample-computed-export

Conversation

@MaxGhenis

@MaxGhenis MaxGhenis commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Simulation.subsample rebuilt the population from to_input_dataframe(include_computed_variables=True): every value the simulation had stored, calculated ones included, reloaded as inputs. It then rebuilds a reform simulation's baseline arm with self.get_branch("baseline") from that population. Anything calculated under the reform before subsampling therefore became an input of the baseline arm too, and the baseline read the reform's values.

Observed on policyengine-us main (two earners at $50,000 and $80,000, spm={"geography_kind": "national"}, reform: gov.irs.deductions.standard.amount.SINGLE = 0 for 2024, taxable_income calculated for 2024, then subsample(2, seed=0)):

taxable_income 2024 reform arm baseline arm
before subsample 48,754.84 / 76,595.86 35,400 / 65,400
after subsample, master 5d68130 48,754.84 / 76,595.86 48,754.84 / 76,595.86
after subsample, this PR 48,754.84 / 76,595.86 35,400 / 65,400

The rule

subsample now rebuilds from _to_subsample_dataframe():

  • Inputs, from the storage's own mark. For each variable, it exports the periods Holder.get_input_periods(self.branch_name) reports. That is the storage's input/derived mark, which put_in_cache(..., derived=True) sets for every calculated, carried, uprated or default value.
    • It includes inputs of formula-backed variables. A dataset can supply them, and they override the formula in the original simulation. In policyengine-us these are person_id (formula np.arange(len(age))) and is_spm_independent_minor_role (DATASET_SOURCE_INPUTS in spm.py). Both are the reason 56c8b82 switched to the full export.
    • It does not read _user_input_keys. That record can outlive the value it names. policyengine-us system.py moves employment_income into employment_income_before_lsr with Holder.delete_arrays and leaves the employment_income keys behind. A calculated employment_income (which includes the labour-supply response) at such a key would otherwise pass as an input.
  • Structural fallback. A structural variable is one build_from_dataset reads to place people in entities: the person id, each group's id, and each person's membership id and role in it. If one has a formula but no input, it is calculated at the dataset period and kept. Before this PR it was exported only if something had calculated it earlier.
  • Nothing else that was calculated is kept.

to_input_dataframe and its include_computed_variables flag are unchanged.

Invariants

Stated for every population, reform (none, parametric, structural formula, changed default, added variable), set of calculations on either arm, sample size and seed. Hypothesis tests in tests/core/test_subsample_computed_values_property.py, 200 examples plus 5 explicit ones:

  1. The export depends only on inputs. Calculating anything leaves _to_subsample_dataframe() unchanged.
  2. History independence (differential). Subsampling after the calculations rebuilds the same inputs, and both arms return the same values, as subsampling without them.
  3. Baseline fidelity (differential). The baseline arm returns what a simulation without the reform, subsampled the same way, returns.
  4. Restriction. Every sampled person reads, in each arm, what they read in the same simulation not subsampled.

From #573 (tests/core/test_subsample_inputs_only_property.py, carried here unchanged): whatever was calculated before subsampling, with auto-carry-over on or off, the subsampled simulation stores exactly what one subsampled straight after loading stores, and every later request agrees.

Tests

File What On master 5d68130
test_subsample_computed_values.py 9 examples: parametric and structural reform values (income_tax, disposable_income) calculated before subsample; a default the reform changes; a value calculated where an input was deleted (stale record); a dataset value for a formula-backed variable is kept; a formula-backed household_id the dataset lacks, calculated first or not 7 fail (each at its intended assertion), 2 pass (they pin kept behaviour)
test_subsample_computed_values_property.py the four invariants above its explicit examples fail (run with master's export as the hook)
test_subsample_inputs_only.py, ..._property.py, fixtures/subsample_inputs.py #573's tests, unchanged: carry-over past a formula's end, stored values, following a new input, following a reform, dataset values for a formula variable, formulas recomputed on the sample 4 of 6 fail (carry-over past a formula's end, stored values, following a new input, following a reform)

Existing subsample tests (formula-backed IDs, fast cache, baseline branch) pass.

Relation to other PRs

Impact

  • Unchanged when nothing was calculated first. On fresh simulations, master's export and this PR's are identical column for column and value for value: 7 columns on the country template with and without a reform, and 181 on a policyengine-us reform simulation of synthetic data. So subsample called straight after construction samples the same households with the same inputs and weights as before.
  • Callers. I searched code in PolicyEngine and TheAxiomFoundation (gh search code) and read each call site. No caller of Simulation.subsample on a default branch calculates anything before subsampling:
    • policyengine-app-v2 build-parameter-dependency-map.py, axiom-oracles populace_us.py, policyengine-us docs and microsimulation tests, and core's docs and tests all subsample straight after construction.
    • policyengine.py, policyengine-api and microcosm don't call it; microcosm and policyengine-uk-data use their own samplers.
    • So no published baseline-versus-reform comparison was affected. Interactive use that calculated first was: notebooks, analyses, and tests such as #9994's property test.
  • policyengine-uk can't use core's subsample. On policyengine-uk main 9dd771a05, Microsimulation(...).subsample(2) raises TypeError: Simulation.build_from_dataset() missing 1 required positional argument: 'dataset'. Core's subsample calls self.build_from_dataset() with no argument (on master and here), UK overrides it as build_from_dataset(self, dataset), and the UK baseline is a separate Simulation, not a branch. This PR changes nothing there.
  • Sampling. The weights used to sample, and how they are rescaled, are unchanged. Only weight columns that are inputs are exported and rescaled now. A formula-backed or uprated weight (person_weight, later years' household_weight) is recalculated from the rescaled inputs, as it would be after subsampling a fresh simulation.

Real data. I used a 500-household, 1,421-person slice of the default policyengine-us dataset (default_population_sample(500) from policyengine_us/tests/microsimulation/populace_fixture.py), on policyengine-us main 19c240d0ff.

  • The reform sets the standard deduction to 0 for every filing status. income_tax and household_net_income were calculated for 2026, then subsample(200, seed=0) ran. "Master" means this branch with subsample's export restored to master's, the only behavioural change. The script is below.
  • A fresh reform simulation stores 0 values that are not inputs, so subsampling straight after construction is unchanged.
After subsample(200), 2026 Impact read: master Impact read: this PR True impact
income_tax, weighted sum $0.000m $1,822.419m $1,822.419m
household_net_income, weighted sum $0.000m −$1,772.836m −$1,772.836m

The true impact is the reform subsample minus an unreformed subsample with the same seed.

  • On master, the baseline arm is the reform for both variables: the largest household difference from the unreformed subsample is $8,694 (income_tax) and $9,765.94 (household_net_income).
  • With this PR, the baseline arm equals the unreformed subsample exactly, and the reform arm equals a reform subsample that calculated nothing first.
Script (run from a policyengine-us checkout with this core; argument fix or master)
"""Real-data check of subsample on a policyengine-us reform simulation.

Mode "fix": the branch's subsample. Mode "master": the branch with
subsample's export restored to master's full export (the only line the fix
changes in subsample's behaviour).
"""
import sys, time, traceback
import numpy as np
from policyengine_core.reforms import Reform
from policyengine_core.simulations.simulation import Simulation as CoreSimulation
from policyengine_us import Microsimulation
from policyengine_us.tests.microsimulation.populace_fixture import default_population_sample

mode = sys.argv[1]
if mode == "master":
    CoreSimulation._to_subsample_dataframe = lambda self: self.to_input_dataframe(
        include_computed_variables=True
    )

YEAR = 2026
reform = Reform.from_dict(
    {
        f"gov.irs.deductions.standard.amount.{status}": {"2024-01-01.2035-12-31": 0}
        for status in ["SINGLE", "JOINT", "HEAD_OF_HOUSEHOLD", "SEPARATE", "SURVIVING_SPOUSE"]
    },
    country_id="us",
)
t0 = time.time()
dataset = default_population_sample(500)
print(f"[{mode}] slice: {len(dataset.household)} households, {len(dataset.person)} people ({time.time()-t0:.0f}s)")

def build(with_reform):
    return Microsimulation(dataset=dataset, reform=reform if with_reform else None)

fresh = build(True)
not_inputs = sorted(
    (name, str(period))
    for population in fresh.populations.values()
    for name, holder in population._holders.items()
    for period in holder.get_known_periods()
    if period not in set(holder.get_input_periods())
)
print(f"[{mode}] fresh reform sim: {len(not_inputs)} stored values that are not inputs: {not_inputs[:20]}")

VARS = ["income_tax", "household_net_income"]
sim = build(True)
full = {v: (float(sim.calc(v, YEAR).sum()), float(sim.baseline.calc(v, YEAR).sum())) for v in VARS}
for v, (r, b) in full.items():
    print(f"[{mode}] before subsample {v} {YEAR}: reform {r/1e6:,.3f}m baseline {b/1e6:,.3f}m impact {(r-b)/1e6:,.3f}m (weighted, this slice)")
try:
    sim.subsample(200, seed=0)
    unreformed = build(False).subsample(200, seed=0)
    uncalculated = build(True).subsample(200, seed=0)
    for v in VARS:
        base_after = np.asarray(sim.baseline.calculate(v, YEAR))
        base_true = np.asarray(unreformed.calculate(v, YEAR))
        ref_after = np.asarray(sim.calculate(v, YEAR))
        ref_true = np.asarray(uncalculated.calculate(v, YEAR))
        w = np.asarray(unreformed.calculate("household_weight", YEAR))
        print(
            f"[{mode}] after subsample {v}: baseline arm == unreformed subsample: {np.allclose(base_after, base_true)} "
            f"(max abs diff {np.abs(base_after-base_true).max():,.2f}); reform arm == uncalculated reform subsample: {np.allclose(ref_after, ref_true)}; "
            f"impact read after subsample {float(sim.calc(v, YEAR).sum() - sim.baseline.calc(v, YEAR).sum())/1e6:,.3f}m vs true {float(uncalculated.calc(v, YEAR).sum() - unreformed.calc(v, YEAR).sum())/1e6:,.3f}m"
        )
except Exception:
    print(f"[{mode}] subsample after calculating raised:")
    traceback.print_exc()
print(f"[{mode}] done in {time.time()-t0:.0f}s")

Checks run

Documentation review

subsample's docstring now states what the rebuild keeps; docs/python_api is generated from docstrings by autodoc. No other docs describe the export. Impact: medium (public method's behaviour after calculations). Confidence: high.

axiom: n/a: core engine, no policy encoding

🤖 Generated with Claude Code

…alues

subsample exported to_input_dataframe(include_computed_variables=True),
every stored value, and loaded it back as inputs of the rebuilt population.
The baseline arm of a reform simulation is rebuilt from that population, so
anything calculated under the reform before subsampling became an input of
the baseline arm too: after calculating taxable_income on a policyengine-us
reform simulation, the baseline arm read the reform's taxable income.

subsample now rebuilds from _to_subsample_dataframe: for each variable, the
periods Holder.get_input_periods reports as inputs (the storage's own
input/derived mark, so a stale _user_input_keys entry cannot pass a
calculated value off as an input), including inputs of formula-backed
variables such as policyengine-us's dataset-supplied person_id and
is_spm_independent_minor_role. A formula-backed id, membership or role
variable that has no input is calculated at the dataset period, as the
rebuild needs it.

Tests: examples and a Hypothesis property (history independence, baseline
arm equals an unreformed subsample, restriction, export depends only on
inputs). Also carries the tests from #573, which this supersedes, and which
pass here.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@MaxGhenis

Copy link
Copy Markdown
Contributor Author

US + core hub: closing as superseded.

  • Production changes: Rebuild a subsampled simulation from its inputs, not its calculated values #573 merged as 4ba5bed9a5. It fixes the same defect: a subsample rebuilt from calculated values. Master now exports only recorded inputs, restores baseline policy on the rebuilt baseline arm, and removes deleted input provenance. The hub's supersession check went through each behaviour this PR fixes and found master covers all of them.
  • Tests: the coverage only this PR held was the reform-baseline regressions and the calculated-before-subsample property. It was ported onto master in Preserve reform-baseline subsampling regression coverage #606, merged as ee841192d5, and every ported test passes on unchanged master.
  • Not carried over, on purpose:

Nothing further is needed from this branch.

@MaxGhenis MaxGhenis closed this Oct 10, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant