diff --git a/changelog.d/1063-identity-kernel-residential-split-branch.fixed.md b/changelog.d/1063-identity-kernel-residential-split-branch.fixed.md new file mode 100644 index 000000000..fe014f7d1 --- /dev/null +++ b/changelog.d/1063-identity-kernel-residential-split-branch.fixed.md @@ -0,0 +1 @@ +The UK geography identity kernel now lists `cgt_residential_split` among its structural branches, between the CGT incidence clone and the geographic pool expansion. The population node had keyed residential arms (microcosm#1063 item 7) on that branch since the stage landed, but the kernel's roster never learned it, so `--release-role dense` refused every spine carrying a residential arm with "unknown structural branch" (the main-head spine of 2026-10-04 carried 2,735). Arms now key their declared path; every other household's key is unchanged. diff --git a/changelog.d/1115-dense-battery-register-gates.fixed.md b/changelog.d/1115-dense-battery-register-gates.fixed.md new file mode 100644 index 000000000..9304a9528 --- /dev/null +++ b/changelog.d/1115-dense-battery-register-gates.fixed.md @@ -0,0 +1 @@ +The dense build's terminal gate battery now evaluates the register gates the way the release certifier does. The spine checkpoint publishes its build state (household weight kind, period, mass log) as a `spine_build_state` artifact and the terminal gate hands it to `uk_release_input_coverage` as the build-state frame, so the stage families are checked against the spine's importance weights rather than the calibrated release frame (every required family had failed on weight kind alone). `uk_degenerate_release_surface` skips the columns the release boundary drops before writing (`UK_RELEASE_EXPORT_DROPPED_COLUMNS`) and records them, instead of reporting a column the exported H5 never carries. The dense gate context computes the CGT projection artifact the `uk_cgt_projection_entrants` fence needs, as the national calibration kernel does. diff --git a/changelog.d/1115-engine-blocks-represent-the-pool.changed.md b/changelog.d/1115-engine-blocks-represent-the-pool.changed.md new file mode 100644 index 000000000..873db785d --- /dev/null +++ b/changelog.d/1115-engine-blocks-represent-the-pool.changed.md @@ -0,0 +1 @@ +Per-clone engine resolution (`--engine-blocks K`) now scales every block's engine household weights by pool mass / block mass before the policyengine-uk engine loads it, so formulas that allocate a national total by weighted share (corporate land value, shareholding, consumption share) see the pool's denominator instead of handing each block the whole aggregate (the K-times land-value artefact behind the #736 erratum). The measures receipt records the representation (`engine_population_representation`: the factor per block and whether the blocks are identical copies of one another, which makes the scaling exact), and the block-sensitivity caveat is dropped when it is. `release_verdict` gains `engine_population_exact`; a multi-block run with an exact representation is releasable, and the dense release-candidate posture no longer refuses `--engine-blocks` greater than one. The resolver's own frame keeps the block's true mass; the scaling lives on the engine scratch frame's mass log. diff --git a/changelog.d/1115-export-allow-list-graph-columns.changed.md b/changelog.d/1115-export-allow-list-graph-columns.changed.md new file mode 100644 index 000000000..422b5f4bc --- /dev/null +++ b/changelog.d/1115-export-allow-list-graph-columns.changed.md @@ -0,0 +1 @@ +The UK export-surface allow-list (`uk/gates.json`, `uk_export_surface`) names the graph export's columns: `household.household_clone_index`, `benunit.benunit_clone_index` and `person.person_clone_index` replace the rowwise tool's `household.clone_index`; `household.constituency_code`, `household.local_authority_code` and `household.region_code` replace the `_oa` aliases; the ITL1–3 and ward codes the atomic geography derives are allowed. The gate-battery vintage pins move with the spec. diff --git a/changelog.d/1115-export-area-code-names.changed.md b/changelog.d/1115-export-area-code-names.changed.md new file mode 100644 index 000000000..14e284023 --- /dev/null +++ b/changelog.d/1115-export-area-code-names.changed.md @@ -0,0 +1 @@ +The UK single-year export writes the three area codes under the names every consumer reads (microcosm#1114): `constituency_code_oa`, `la_code_oa` and `region_code_oa` replace `constituency_code`, `local_authority_code` and `region_code` at the export boundary (`graph_terminal._tables`, `uk_release_export_frame`), with no alias kept. The graph keeps the ladder names in memory; `uk_export_candidate_columns` translates them so the export-surface gate compares the artifact's names, the allow-list names the `_oa` columns again, the export read-back refuses a ladder name or an empty consumer column, and the export descriptor and the rowwise candidate manifest record `area_codes` (the columns and the code frames behind them, read from the committed support provenance). The gate-battery vintage pins move with the spec. diff --git a/changelog.d/1115-review-round-1.fixed.md b/changelog.d/1115-review-round-1.fixed.md new file mode 100644 index 000000000..4367bd04a --- /dev/null +++ b/changelog.d/1115-review-round-1.fixed.md @@ -0,0 +1 @@ +Review round 1 of microcosm#1115: a per-block engine resolution is `exact` only when the blocks are identical copies with a source identity and the allocation keys of policyengine-uk's weight-share formulas (`UK_WEIGHT_SHARE_FORMULA_INPUTS`: corporate land value, shareholding, consumption share) carry the same `sum(x * w)` in every block, recorded per block in the measures receipt; a weights-only comparison is never exact. The `release_input_coverage` binding requires the spine build state (evidence absent without it) and refuses to judge the calibrated frame. The diagnostics schema stays strict: the terminal node's converter hands the holdout in the schema's shape and the loosening that admitted the graph artifact's fields is withdrawn. diff --git a/changelog.d/1115-warm-start-size-search.added.md b/changelog.d/1115-warm-start-size-search.added.md new file mode 100644 index 000000000..f3e5679ed --- /dev/null +++ b/changelog.d/1115-warm-start-size-search.added.md @@ -0,0 +1 @@ +`--selection-initial-lambda` warm-starts the UK size selection's L0 budget search at a known penalty (`select_uk_dataset_size(initial_lambda=...)`, the size-search node's `initial_lambda` parameter, recorded in the run parameters and the search receipt). The shared search probes the given penalty first and stops there when the draw is feasible and within tolerance; a miss is followed by one probe half a decade away on the side it steered to, then interpolation between the measured ends (regula falsi with a progress safeguard) instead of re-bisecting the whole bracket. A stale hint costs probes, never feasibility; a warm and a cold search may settle on different penalties inside the budget window. The cold path (no warm start) is unchanged. diff --git a/changelog.d/932-collect-each-engine-block.fixed.md b/changelog.d/932-collect-each-engine-block.fixed.md new file mode 100644 index 000000000..19fc89d59 --- /dev/null +++ b/changelog.d/932-collect-each-engine-block.fixed.md @@ -0,0 +1 @@ +The per-clone measure resolution (`--engine-blocks` greater than one) now collects garbage after every engine block. A policyengine simulation is a large cyclic object graph that `del` alone does not reclaim, so the first K=25 dense build (2026-10-05) retained all 25 per-clone engines, reached a 140 GB memory footprint on a 24 GiB machine and was killed before the solve. diff --git a/changelog.d/932-engine-scratch-drops-native-alias-codes.fixed.md b/changelog.d/932-engine-scratch-drops-native-alias-codes.fixed.md new file mode 100644 index 000000000..6980389ed --- /dev/null +++ b/changelog.d/932-engine-scratch-drops-native-alias-codes.fixed.md @@ -0,0 +1 @@ +The scratch dataset a UK measure engine loads now leaves the nation-native alias codes of the derived geography layers (`output_area_code`, `data_zone_code`, `intermediate_zone_code`, `super_data_zone_code`, `district_electoral_area_code`) behind, exactly as the release export does. They are NA outside their own nation by design, and policyengine-uk refuses a single-year dataset with any NaN column, so every `--release-role dense` build under the default atomic assignment failed at `uk.full.measures` with "Column 'data_zone_code' contains NaN values" (the 28 Sept dense rungs ran the legacy assignment and never met it). One helper, `without_uk_native_alias_columns`, now serves both boundaries; the resolver receipt lists the columns it dropped. diff --git a/changelog.d/932-national-problem-per-engine-block.changed.md b/changelog.d/932-national-problem-per-engine-block.changed.md new file mode 100644 index 000000000..cbcf76091 --- /dev/null +++ b/changelog.d/932-national-problem-per-engine-block.changed.md @@ -0,0 +1 @@ +The dense role's measures node now compiles the national constraint problem per engine block and stitches the per-block matrices into the pool's household order (`resolve_uk_full_national_problem`), instead of materializing every national target as a column on the whole K-clone pool and copying the prepared pool twice more to compile and check it. On a 1.6 million-household pool that step needed tens of gigabytes: the first K=25 dense build (2026-10-05) reached a 140 GB footprint on a 24 GiB machine and was killed before the solve. Materialization is row-wise (band edges come from the compiled register), so the stitched problem is the pool's problem; the measures receipt records `national_materialization: per_engine_block`. `resolve_uk_full_measures` keeps its contract for the national seam and the evaluation tools. diff --git a/changelog.d/932-terminal-gates-skipped-holdout.fixed.md b/changelog.d/932-terminal-gates-skipped-holdout.fixed.md new file mode 100644 index 000000000..2d87a86b8 --- /dev/null +++ b/changelog.d/932-terminal-gates-skipped-holdout.fixed.md @@ -0,0 +1 @@ +The dense role's terminal gate kernel now hands the calibration diagnostics the rotated holdout in the schema's shape: a skipped holdout as the bare `{"skipped": true}` marker and a measured one without the graph's own `graph_binding` provenance, which the diagnostics records (no undeclared fields) refuse. The `uk.full.holdout` artifact keeps the full report. The first K=25 dense build (2026-10-06) reached its terminal gates after 28 hours with `--skip-holdout` and the battery refused the graph's marker with nineteen validation errors; the solve stayed in the content store and resumed. diff --git a/changelog.d/calibrate-compiled-row-axis-guard.fixed.md b/changelog.d/calibrate-compiled-row-axis-guard.fixed.md new file mode 100644 index 000000000..4505b4299 --- /dev/null +++ b/changelog.d/calibrate-compiled-row-axis-guard.fixed.md @@ -0,0 +1 @@ +A decoded problem's compiled rows (`OrderedProblem.to_target_set`) now guard the entity axis with one shared array and a vectorised comparison instead of rebuilding the axis as a Python tuple from the frame on every row. On the first K=25 UK dense build (2026-10-05) that rebuild made compiling 22,053 local rows over a 1.6 million-household pool cost ~3.5e10 scalar conversions, roughly ten hours per compile; the guard's refusal of a different or reordered axis is unchanged. diff --git a/docs/evidence/spec-engine/us-f0-coverage.json b/docs/evidence/spec-engine/us-f0-coverage.json index f99b29cdf..0fabd95a8 100644 --- a/docs/evidence/spec-engine/us-f0-coverage.json +++ b/docs/evidence/spec-engine/us-f0-coverage.json @@ -1656,13 +1656,13 @@ "compiler_ir.node_slices" ], "expected": { - "map_sha256": "0fa116d9f5dae4ba0378f190d6c822b92f3dcf1e3fe80c9880bcaf87431c283c", - "protocol_sha256": "b4afa376b62fbc073fd10d74e40921f5ab60c1efe366e2c8504872091ea69d08" + "map_sha256": "b4d6ec9c11ce323d83319a3d0289e0cf7216a7936d6c6f08cef0a463ad369641", + "protocol_sha256": "114217a2918329b0601d94deb8ed83f73919584b85ca8aabe45a40665afb6560" }, "failures": [], "observed": { - "map_sha256": "0fa116d9f5dae4ba0378f190d6c822b92f3dcf1e3fe80c9880bcaf87431c283c", - "protocol_sha256": "b4afa376b62fbc073fd10d74e40921f5ab60c1efe366e2c8504872091ea69d08" + "map_sha256": "b4d6ec9c11ce323d83319a3d0289e0cf7216a7936d6c6f08cef0a463ad369641", + "protocol_sha256": "114217a2918329b0601d94deb8ed83f73919584b85ca8aabe45a40665afb6560" }, "status": "covered" }, @@ -1677,7 +1677,7 @@ "compiler_ir.seed_stream_map" ], "expected": { - "implementation_sha256": "b4afa376b62fbc073fd10d74e40921f5ab60c1efe366e2c8504872091ea69d08", + "implementation_sha256": "114217a2918329b0601d94deb8ed83f73919584b85ca8aabe45a40665afb6560", "protocol": "legacy-v1", "streams": [ "build_model", @@ -1698,7 +1698,7 @@ }, "failures": [], "observed": { - "implementation_sha256": "b4afa376b62fbc073fd10d74e40921f5ab60c1efe366e2c8504872091ea69d08", + "implementation_sha256": "114217a2918329b0601d94deb8ed83f73919584b85ca8aabe45a40665afb6560", "protocol": "legacy-v1", "streams": [ "build_model", @@ -2599,7 +2599,7 @@ "country": "us", "schema_id": "country_spec", "schema_version": 1, - "spec_sha256": "d1df6b316e4937e809f5f020ea2d3a7bd38bccc1845f3ae6e104f7eed6a73323" + "spec_sha256": "f76f921d4f74392b106b8c0729ce250494e0a55455eaba7b910b19cdf7825178" } }, "report_schema_version": 3, @@ -2609,7 +2609,7 @@ "country": "us", "schema_id": "country_spec", "schema_version": 1, - "spec_sha256": "d1df6b316e4937e809f5f020ea2d3a7bd38bccc1845f3ae6e104f7eed6a73323" + "spec_sha256": "f76f921d4f74392b106b8c0729ce250494e0a55455eaba7b910b19cdf7825178" }, "status": "pass" } diff --git a/docs/uk-dense-release-assembly-runbook-762.md b/docs/uk-dense-release-assembly-runbook-762.md index 59c399a06..4e6f378ff 100644 --- a/docs/uk-dense-release-assembly-runbook-762.md +++ b/docs/uk-dense-release-assembly-runbook-762.md @@ -51,8 +51,12 @@ solve defaults (K=15, seed 42, 1,500 epochs, learning rate 0.15, `grain_equal`), names the outputs `microcosm_uk_2024_25_local.h5` and `microcosm_uk_2024_25_local.local_gates.json`, records itself in the manifest, and refuses the national role's flags. `--release-candidate` pins -the doctrine (bound 10, `grain_equal`, K=15, 1500 epochs), resolves the -engine in a single block, and runs the rotated holdout. Best-effort staging +the doctrine (bound 10, `grain_equal`, K=15, 1500 epochs) and runs the +rotated holdout. The engine may resolve in a single block or per clone +(`--engine-blocks K`): each block's engine weights are scaled to the pool so +weight-share formulas see the pool's denominator, and the measures receipt +records whether the blocks were identical copies (an exact representation); +a per-block run is releasable only when it was. Best-effort staging telemetry uploads every 300 s by default on this driver (the Hub allows about 128 commits per hour per repository). Expect about 3.5 hours and 10 GB at K=15 (measured on the rowwise tool; the @@ -73,7 +77,8 @@ uv run --no-sync python tools/score_uk_local_candidate.py ... --output-json [calibrate]`): the spine -with the build_930 recipe (the five NTS tabs, `--was-person-tab`, `--no-staging`), then the national -calibration (local staging only) against the Chronicle artifact 505e0e7 and the 2026-09-30 build as the -incumbent. Nothing is released from an arm. The engine is policyengine-uk at the repo's pin. - -The arms, in order: - -- C (e0b41a59b, the walk): all 27 spine gates pass. Residential gains gap £38.8m, 0.3% of the £12.2bn - target (main: £469m, 3.8%); count gap 1,428 against a bound of 10,331; the ten largest flagged stakes - carry 18.5% of residential gains (22.7% on the 2026-09-30 build); largest flagged stake £557m; the - largest stake in the pool £10.9bn, 89% of the target. The open band (£5m and over) holds 84 rows with - £47.7bn of gains, 5 flagged, achieved £1.27bn against £1.44bn expected: the deterministic bound of a - band whose heaviest row weighs 254 households and whose largest gain is £522m is vacuous (£132.6bn), - which is why the walk was replaced (R3). Table 7 by claimant status: claimants realise 87.2 / 5.3 / - 7.5 against the same targets to five decimals; non-claimants within 0.4 points of their restored - shares. Calibration: loss 0.00853 (main 0.00836), ESS 5,217 (5,198), max-to-median weight 1,115 - (1,079; bound 1,151), projection entrants 6,715 (6,711; bound 73,000), 1,090 targets. -- B (b561270df, WAS and identities): all 28 spine gates pass including the new - `uk_stage_was_wealth_coherence`: zero mortgage debt off mortgaged tenure, zero main-residence value - off owner tenures, zero owners without one, zero identity violations. (Superseded after review - round 1: Arm B's rule dropped the donor's off-tenure mortgage mass, £127bn of £1,266bn, all - other-property mortgages. Since 558325016 only the main-residence mortgage is tenure-stratified and - the other-property mortgage, `HMortGR8` less `TotMortR8`, is drawn on every tenure; on the final - build's spine the donor holds £127.3bn of mortgage debt off the mortgaged tenure, £111.6bn of it - other-property mortgages, and the recipient frame holds £273.1bn on 1,386 households, with zero - main-residence mortgage off the mortgaged tenure and zero `mortgage_debt` identity violations.) - Mortgaged - households owing more than their home: 5.4% against the donor's 2.7% (26% on the 2026-09-30 build). - Renters with property wealth: private 16.5% against the donor's 10.3%, social 4.6% against 1.1% - (10.7% before). The calibration stopped before an H5 on - `hmrc/state_pension_income_band_50_000_to_70_000@2025`, a target-fit exclusion inside its bound on - this spine, which the pensions branch retired at b19f14d98; dense loss at epoch 1,500: 0.0087. -- C2 (0ce1a5235, the residential split, on the arm C spine): all 28 spine gates pass. 2,276 - households carry a liable gainer, none more than one, so 2,276 residential arms; the residential - count and gains are identities on Table 8a in every band (solve error 1e-16 on the count, 7e-10 on - the gains, identity error 0); 106 arms weigh less than one household (smallest 0.14, largest - 1,263); the ten largest residential arms carry 9.4% of residential gains (18.5% on the walk, 22.7% - on the 2026-09-30 build); the largest residential arm stake is £176m (£557m flagged on the walk); - the largest liable stake in the pool stays £10.9bn, 89% of the target. Calibration: dense loss - 0.00909 (C 0.00853, main 0.00836); on the family fold the max-to-median weight is 610 (C 618, main - 609) and the ESS 5,207 (C 5,209, main 5,191), with 7,028 rows folded (4,752 support copies and the - 2,276 arms); the unfolded ESS fraction falls from 0.0895 to 0.0862 because the arms are extra rows. - The seam battery blocked the H5 on `uk_target_fit`: the North West £12,570–15,000 income-tax cell - (`hmrc.spi_region.income_tax_by_region_12570_15000@E12000002@2025`) at +25.3% against the 25% bound - (23.2% on main, 22.6% on C; its South East twin 22.6% and 23.3%: the regional band cells hover just - under the bound on every arm and the solver moves a point among them), and the stale - `hmrc/state_pension_income_band_50_000_to_70_000@2025` exclusion back inside the bound (24.9%; the - pensions branch retired it at b19f14d98, so the stack does not carry it). - On the calibrated weights (review item 4, `tools/uk_residential_readout.py` on the C2 solution): the - two bound national totals hold (count 202,628 against 202,630, gains £12.246bn against £12.242bn), - and within them the solver moves mass between bands: £250–500k 3,680 → 4,256 taxpayers, £500k–1m - 2,004 → 2,173 (gains £1.25bn → £1.67bn), £1–2m 730 → 552 (£0.94bn → £0.77bn), £2–5m 427 → 274 - (£1.24bn → £0.87bn), the open band 134 → 139; the bands below £250k move by under 3%. So the - by-band identities are a design-weight property; the release carries the national totals. Row - level with the arms in place: the smallest positive calibrated weight is 0.012 (an arm, 0.050; the - smallest design arm 0.144), the positive-weight median 26.6 (arms 11.8), no row at zero. -- Stack (ebff84125, the final-build candidate): the spine passed its gates in six minutes; the first - calibration refused the Chronicle artifact 505e0e7 because the stack pins the pensions feed - (facts 28b7105…, artifact 825406f). Against that feed the dense loss is 0.00818 (main 0.00836), but - the seam battery blocked the H5 on one cell with no support at all: - `dwp/uc_payment_dist/SINGLE_annual_payment_28_800_to_30_000@2025` (single, no children, monthly - award £2,400–2,500; target 1,654), initial estimate 0 against 763 at design weights on every - main-based arm and on every pensions arm. Traced: the cell is supported by one FRS household (a - disabled single council tenant in Yorkshire, design weight 760, with its incidence clone at 3.3); on - the stack its UC award is £678 a year lower, which moves it into the 27,600–28,800 band that the - uk-data#452 class already excludes. Nothing on the household's own inputs changed; its consumption, - energy, land and savings draws all moved at once, because the SPI income band donors now insert - 1,920 rows where 480 stood and every stage drawn after them that is not identity-keyed reads its - rows in a new order. The cell joins the measure exclusions beside its sibling (approved 2026-10-02, - expiring 2026-11-26 with the class, for her signature). -- Stack 2 (558325016, after review round 1): the spine passes; dense loss 0.00769; the battery blocked - on two further cells. `dwp/uc_payment_dist/COUPLE_NO_CHILDREN_annual_payment_22_800_to_24_000` - has two supporting households on every main-based spine (design weights 1,025 and 1,999 against - 2,761) and both awards moved out of the band when the stages drawn after the reworked WAS chain read - their rows in a new order; it joins the uk-data#452 measure exclusions. `hmrc.spi_region. - income_tax_by_region_12570_15000@E12000008@2025` (South East) sits on its target at design weights - is 11% short at design weights (69.6m against 78.3m; 62.4m on main, 58.3m on stack 1) and the - solver pulls it to +32.0%. The pull comes from the national employment-income bands: at design - weights the £12,570–15,000 and £15,000–20,000 employment-income bands are 20% and 31% short (the - earners gap the 2026-09-30 build found), the solver fills them by raising the weights of low-paid - employees everywhere (three-fold on the 200 households that carry most of this cell's gain), and - the South East income-tax cell, carried by the same households, overshoots while its total-income - sibling lands at +11.2% (main pulls the same cell from −20% to +22.6%). The cell is deferred four - weeks in the target-fit register (it was deferred on 2026-09-18 and retired on microcosm#1012 at - +24.5%); the fix is on the employment-income side, not on the draws. Separately, every regional - £12,570–15,000 income-tax cell moves 10–30% at design weights between spines that share the same - incomes (North West 63.0m, 44.5m, 52.0m; East 41.3m, 53.5m, 53.2m) because the salary-sacrifice and - pension draws after the SPI chain are positional; identity-keying those draws would make the - design values stable between spines but would not move where the solver lands. -- Stack 3 (2fe7d8bbb, the Table 3.6 amounts projected by HMRC's band growth): all 7 seam gates pass - and the H5 is written; dense loss 0.00713 (main 0.00836). The three lowest employment rows fit at - −3.9%, −11.9% and −10.0% (stack 2: −8.9%, −23.1%, −18.9%; main −9.9%, −23.2%, −19.6%). The South - East £12,570–15,000 income-tax cell lands at +30.3% under its deferral; the North West twin at - +14.4%. Folded weights: max-to-median 1,028 (main 609, stack 1 710, stack 2 1,116; bound 1,151), - ESS 4,953 (main 5,191, stack 1 5,222), median positive weight 40.0 (main 55.8), heaviest household - 41,085 (an East Midlands household in its thirties carrying 10–12% of that region's 30–39 - population, £20–30k taxpayers and band-B council tax rows; on stack 1 the heaviest was an 85–89 - Pension Credit unit at 29,690). The rise in the fold between stack 1 and stack 2 came with the WAS - mortgage rework and the first thin-cell exclusion; the margin to the bound is now thin and belongs - in the review. -- Final build (2fe7d8bbb): running at the time of writing; the readings follow. - -The South East cell, with the count side read: the frame has 0.278m South East taxpayers with total -income in £12,570–15,000 at design weights against HMRC's projected 0.364m (nationally 2.31m against -2.97m; North West 0.240m against 0.333m, London 0.235m against 0.340m, East 0.256m against 0.273m), -and the ones it has pay £250 of income tax each against HMRC's £215. The bound count rows make the -solver scale these households up by a third to hit the count (final error 0.000), the total-income -sibling overshoots to +11% and the tax cell to +30%. So root 2 is a deficit of taxpayers just above -the frozen allowance, concentrated in the South East, London and the North West, not an employment -problem: the fix is the lowest band's support in the frame (the SPI support channel's coverage of -that band, which the pension-age prior share of 0.2 reduced, and the FRS's small incomes around the -allowance), a spine change for the follow-up. - -Child Benefit children in payment (review item 5): the trial's 11.1m is at design weights on the -1 October spine, where the eligible-child base is 14.04m; the claim rate (86.7%) and the opted-out -family share (9.0%) match HMRC, and the opted-out families carry 1.82 children each, so claimed -children are 12.17m and 1.09m of them are in opted-out families. On the 2026-09-30 build's calibrated -weights the eligible base is 14.79m, which at the same rates gives 11.7m in payment against HMRC's -11.73m: the gap is the design-weight child base, not the rates or the opt-out ages. - -Child Benefit (c8, trial on the 1 October spine at design weights): 86.7% of eligible children claimed -for (HMRC 86.6%; the frame's age mix implies 87.1%), 9.0% of claiming families opted out, all from -families charged the whole benefit, 11.1m children in payment against 13.5m before. - -## The certifier rehearsal (R5, 2026-09-30 build, 15 passed / 5 failed of 20) - -- `uk_export_surface`: label `microcosm_uk_2024`; `household.household_weight` reported missing and - every id column reported as an extra (the certifier read every frame column); 49 columns outside the - enhanced-FRS surface without an allow-list entry; `incapacity_benefit_reported` to be dropped. -- `uk_degenerate_release_surface`: `incapacity_benefit_reported` and `is_in_approved_training` - all-zero. FRS 2024-25 (SN 9563) carries `TRAIN2` (codes 1–8 and 10 are schemes, 9 none of these) and - `TRAINEE` in place of `TRAIN`; 28 adults report a scheme, so the column carries signal once read. -- `uk_nonnegative_columns`: `housing_water_and_electricity_consumption` 250 rows below zero (minimum - −14,165), inherited from LCFS donor rows whose diary nets refunds below zero; the support clip took - its floor from the donor's realised range. -- `uk_input_mass_parity` against the enhanced FRS: `adult_ema` +786% (38.9m vs 4.4m), - `dfe_education_spending` +187,061% (£98.8bn vs £52.8m). -- `uk_qrf_tail_concentration`: `charitable_investment_gifts` top-100 rows carry 100% of the mass over - 168 carriers. - -c9 addresses each at its source (see the commit); the certifier is to be re-run on the final build. - -## The employment-income band shortfall (ruling 2026-10-02: fix, do not defer) - -The three lowest `hmrc/employment_income_income_band_*` rows are SPI 2023-24 Table 3.6 amounts: the -employment income received by taxpayers whose total income falls in the band, projected to 2025. The -frame's taxpayer counts and total income in those bands sit within a few percent of HMRC; what was -short was the share of the band's income that is pay, 20–31% at design weights. Two roots: - -- The registry projected the Table 3.6 amount rows with flat national indices (average earnings for - employment, mixed income for self-employment, the pension index for private pensions) while the - count, total-income and tax rows by band use HMRC's own band-specific growth (Income Tax - Liabilities Table 2.5). An amount held in a fixed nominal band moves with the band's membership: - HMRC projects the £12,570–15,000 band's total income +5.0% from 2023-24 to calendar 2025, the - £15,000–20,000 band's −2.3% and the £20,000–30,000 band's +0.3%, against the flat +10.7%. The - 39 amount references now use the band growth; the three lowest employment targets move from - £18.7bn, £53.8bn and £190.4bn to £17.7bn, £47.5bn and £172.5bn, and the frame's design-weight gaps - from −19.8%, −31.2% and −25.1% to −15.5%, −22.1% and −17.4%. Dividend and savings-interest amounts - keep their own indices (they move with rates, not membership). The SPI tape re-banded under the - engine's own indices agrees with HMRC's direction in every band. -- The remaining gap is the frame's composition of the low bands. Banding the frame on core income - only (pay, profits, private pensions, State Pension), its earners in the £12,570–15,000 band are - 1.36m, HMRC's 1.36m exactly, with £17.8bn of pay against the band-aware £17.7bn; adding the - non-employment leaves (savings interest, dividends, property, miscellaneous) lifts 0.14m of them and - £2.2bn of pay out of the band, and the engine's broader banding concept a further £0.6bn. In the - £15,000–20,000 band the core-only comparison still leaves pay 8.6% and earners 14% short, with State - Pensioners 7% above HMRC's count. The places to look, in order: the SPI stage-2 rewrite of dividend, - property and savings income onto FRS rows (`hmrc_spi_income_spine`; on the tape the low earners who - sit in these bands carry about £350 of such income, the frame's FRS-channel low earners show 7% with - dividends averaging £17,500), the total-income concept the band measure uses against SPI's, and the - SPI channel's age mix in the low bands (58% of the band on main, 43% on the stack). Spine work, a - follow-up. - -The South East £12,570–15,000 income-tax cell's +32% is this pull: the solver raised the weights of -the low-paid households that carry it (three-fold on the 200 that carry most of its gain) to fill the -under-projected employment rows; its own target is already band-aware. - -## Review round 1 (Vahid, 2026-10-02) - -- The five internal disability carriers leave the export allow-list and are dropped at the release - boundary with `incapacity_benefit_reported`; the input-mass evidence pin and the gate digests follow - the c9 register entries. -- WAS mortgages: only the main-residence mortgage (`TotMortR8`) is tenure-stratified; the mortgages - on other property (`HMortGR8` less `TotMortR8`) are drawn on every tenure without a stratum, and - `mortgage_debt` is derived as their sum, so a renter's or an outright owner's buy-to-let keeps its - mortgage beside the property. The chain runs eight segments; the coherence gate holds the - main-residence mortgage off the mortgaged tenure at zero and the mortgage identity on every row. -- The split's docstring and notes say what calibration binds: the two national residential rows. -- The plan-2 renewal stays on this PR (her ruling). - -## Residential split (c5b) - -Each liable gainer's residential probability `p` (the logistic solved to Table 8a's individuals-basis -count and gains, with the stock shift) is carried as weight: a household with one liable gainer becomes -a residential arm at `p·w` beside the incumbent at `(1 − p)·w`; with `k` gainers, `2^k` arms at -product weights (`k > 3` refused). At design weights the residential count and gains are identities on -Table 8a in every gain band (no draw, no seed, no error bound); the asset-type stage types the arms and -keeps the BADR claims and the Table 7 fit. The £522m copy (weight 20.8, `p` 0.2%) becomes a residential -arm at 0.04 households beside a non-residential arm at 20.76. Family folds, the geography identity -kernel, the export surface, national sampling, the student-loans receipt and the E8 identity receipt -carry the new layer. - -## Final build (2fe7d8bbb on main fe4e38dfa, 2026-10-02) - -Spine `spine-1063-2fe7d8bbb.h5` (sha256 ccda40bd…), national candidate `runs/uk-national-1063-2fe7d8bbb/ -microcosm_uk_2024_25.h5` (sha256 df3ca7cc…, staged), 7/7 seam gates, loss 0.00713 (1,090 targets; the -South East £12,570–15,000 income-tax cell +30.3% under its reviewed deferral). Evaluation against the -2026-09-30 incumbent PASSED: loss 0.00817 vs 0.327, 1,020 wins vs 30, 118 pruned. The release-cut -certifier stopped at its first part, `uk_ledger_compile_parity_incumbent_2025`: the signed receipt still -pinned the 39 Table 3.6 amount references at their flat-index values; re-signed against the pinned -ledger (2af5ad3e5, only those 39 rows move). E7 passed. E8 failed with three symptoms, all diagnosed on -the built spine: - -- `anchor_recompute` 55,622 households (max 286) and `support_split_selection_stored` extra 1 / - missing 2: the recompute banded the carriers on post-sacrifice pay. Since microcosm#1069 c9 (on main - as #1084) `salary_sacrifice` lowers a converted record's `employment_income` and zeroes its employee - contribution in place, after the support split (stage 30) and the anchor (33) banded on them; - 2,797 clone carriers carry salary sacrifice, 82 of them moved down a Table 3 band, so the anchor - derived loss 153,816 / sub-exempt 40,723 against the stage's 153,919 / 40,709 at the same liable mass, - and band 100,000 lost households 13421 and 2000006588 to band 0 while band 125,140 gained 6477. - The 30 September main spine predates the pay drop, which is why E8 was exact there. Fix: the stage - keeps a converted record's pre-conversion pay on `salary_sacrifice_pre_conversion_pay` (zero - elsewhere; dropped at the release boundary), and every receipt of a stage before `salary_sacrifice` - reverses the conversion from it (`_pre_salary_sacrifice_frame`). -- `residential_split_permutation`: the stored arms were reproduced to 3.4e-13, but two person orders - of the same frame differed by 1e-13 on 2,999 arm weights — the logistic solve's bisection sums are - pairwise float sums. Fix: the split walks the liable gainers in `person_id` order (the identity on - the spine's sorted tables) and breaks a tie for the exact-total correction row by household id; the - order test is now bitwise under a reversed and a random person order. - -The weight-shape warning stands (folded max-to-median 1,028 vs main's 609; bound 1,151). Both fixes -change the spine (a new carrier column; the arms bitwise unchanged), so the final build and the -certifier run again from the fixed head. - -## Final build, second run (64c460b66, 2026-10-03) - -Spine `spine-1063-64c460b66.h5` (sha256 b7202ef6…): bitwise identical to the first run's in every shared -column and every household weight; the only difference is the new carrier, set on exactly the 7,843 -converted records the stage reports. Candidate `runs/uk-national-1063-64c460b66/microcosm_uk_2024_25.h5` -(sha256 aa31bdf6…): the same calibration (loss 0.00713, 7/7 seam gates) and the same evaluation -(0.00817 vs 0.327; 1,020 wins vs 30). E7 and E8 pass; E8 reports the restored pre-sacrifice pay, anchor -targets equal to the stage's and the residential arms identical under permutation. The certifier -reaches all 20 gates: 17 pass, three fail. - -- `uk_nonnegative_columns` and `uk_qrf_tail_concentration` looked for declared stage outputs on the - release candidate, where the export now drops them (the five disability carriers and the - pre-conversion pay carrier). Both bindings now read an export-dropped declared column from the - spine frame the certifier supplies, at the spine's weights; without that frame the absence still - fails. -- `uk_input_mass_parity` against the uk-data 1.56.16 enhanced FRS (±452%): `access_fund` +535%, - `jsa_income_reported` +1,058%, `working_tax_credit_reported` +747%, and a stale `adult_ema` - exclusion (+74% now; entry retired). Measured on the candidate against the spine's design weights: - 98% of the access-fund mass is one FRS household (sernum 14443, South East) and its two CGT copies, - whose 19-year-old reports an access fund that the FRS weeklyises to £1,035.62 (`ACCSSAMT`, period - code 5) and the spine annualises to £54,037, eleven times the next largest value; the calibration - takes the household to its 10x bound (design 896, calibrated 8,927). The JSA and tax-credit masses sit - in the SPI synthetic channel (76 and 274 of the nonzero rows), where the stage-2 QRF imputes these - legacy-benefit leaves onto taxpayer rows at FRS rates, and in a handful of rows at or near the - bound (ratios 6–10). None of this is new: at design weights the stack spine carries less JSA mass - than the 30 September main spine (3.1e8 vs 4.7e8) and similar tax-credit and access-fund masses, - and the 30 September main candidate sat at +421% and +428% on the same two columns, just under the - tolerance, with the same 0.5% of households at the 10x bound. The three findings are the - calibration's weight concentration and the SPI channel's legacy-benefit leaves, not the stack's - stages; the rulings on them (an FRS outlier fence on `access_fund`, the SPI channel's treatment of - `jsa_income_reported` and `working_tax_credit_reported`, or reviewed exclusions) are hers. - -Certifier rerun from the gate-fix head (1e58ecb3a) on the same artifacts, 2026-10-03 00:23Z: 19 of 20 pass; -`uk_input_mass_parity` alone fails, on the three columns above (`final/certify2.err`, -`microcosm_uk_2024_25.release_cut_gates.rerun.json`). - -The access-fund record, checked against the FRS 2024-25 variable listing (review round 2, item 7): -`ACCSSPD` 5 is "Calendar month", and £1,035.62 a week is exactly £4,500 a calendar month -(1,035.616438 × 365/7 ÷ 12). A £4,500 access fund a month is not a plausible award; £4,500 a year sits -beside the next FRS values (£4,803, £4,654, £3,602). So the record is an annual or termly award under -the wrong period code, a data fix for the spine's period handling rather than an outlier fence, and -the household's weight is the calibration's: on the 30 September main build the same record sat at -design weight (ratio 1.00) because its incidence clone drew £2,355 of gains, under the exempt amount, -and the anchor returned the clone's mass; on this build the Table 3 redraw's positional stream, shifted -by the 1,920 inserted donor households, gives the clone £3,776, a liable South East gainer with a -residential arm, and the calibration takes the household and both copies to the 10x bound. - -Ruling 2026-10-04 (in session): the deferrals needed to build and publish the national dataset are -signed. `input_mass_reviewed_exclusions.json` gains `access_fund` (expiring 2027-01-04, with the -period rule owed), `jsa_income_reported` and `working_tax_credit_reported` (2027-04-04), each entry -carrying the diagnosis above; the certifier reruns on the same artifacts from the signed head. - -Certifier rerun from the signed head (427bec73f), 2026-10-04 16:23Z: all 20 gates pass. The run then -stopped composing the certificate: "spine: signature does not authenticate under the release signing -key". Not the key (every report, this run's own included, carries the same key fingerprint) but the -canonical bytes: the gate battery signs its reports over `json.dumps(sort_keys, compact)` while the -certifier verified the parts, and signed the certificate, over the logbook's #628 form, which renders -`1.0` as `1` and `1e-05` as `0.00001`. The two agree on the float-free synthetic fixtures and on no real -report, so the first certification to reach composition was the first to see it. The certifier now -verifies and signs over the battery's canonical bytes (`canonical_report_bytes`, the form the data -contract's verifier already uses), with a regression test that puts such floats in a part. A latent -defect on main, reached here for the first time. - -The engine-free lane's remaining failure, `test_uk_national_graph.py`'s seam-scope test, is the -committed target-fit register against the test's fixed review date: the South East deferral is -approved 2026-10-02 and the graph was reviewed at 2026-09-29, so the gate refused it as not yet in -force. Path-tiered CI had not selected the file before round 2 widened the diff. The test now reviews -at the latest approval across the committed registers, inside every live window. - -Certifier rerun from 5720b80dd on the same artifacts, 2026-10-04 16:55Z: all 20 gates pass and the signed -certification composes (`shippable: true`; spine, calibration-seam and release-cut parts, 31 entries; -candidate aa31bdf6…): `final/microcosm_uk_2024_25.release_certification.rerun.json`. - -## Still owed - -- The access-fund period rule in the spine (`ACCSSPD` 5 on sernum 14443), owed before the - `access_fund` exclusion expires on 2027-01-04. -- Follow-ups outside this PR: key the Table 3 redraw's within-band draws and the other positional - draws by identity, so an upstream row insertion no longer moves a thin UC cell or an incidence - clone across the exempt amount (review round 2, item 8); the SPI channel's treatment of the - legacy-benefit leaves (`jsa_income_reported`, `working_tax_credit_reported`); running - `salary_sacrifice` before the capital-gains stages, which would retire the pre-conversion pay - carrier. -- Her calls: Child Benefit opt-out pool order (fully charged first, then the taper, as built; or one - pool over every charged family); `other_residential_property_value` (WAS `DVHseVal`) excludes - buy-to-let (`DVBltVal`), which sits in the drawn remainder; the Table 8a rows stay fenced - (promotion to bound targets is the #970 leverage question, after the final build). -- Findings outside this PR: SPI-channel rows receive `child_benefit_reported` from the stage-2 QRF - without regard to children (1.13m weighted reporter units with no eligible child); the frame has - 0.67m reporter families with an adult above the £60,000 charge start against HMRC's 0.31m–0.44m - liable individuals; the E8 donor recompute takes the licensed tape (`--spi-tab`) rather than storing - the propensity table in the receipt (small cells). -- policyengine-uk: #1996 (VAT) filed; `is_renting` omitting `RENT_FROM_HA` and the engine ignoring - `child_benefit_opts_out` to file on her go. diff --git a/packages/microcosm-build/src/microcosm/build/spec_engine/inventory_coverage.py b/packages/microcosm-build/src/microcosm/build/spec_engine/inventory_coverage.py index 3fddff951..673ab08d3 100644 --- a/packages/microcosm-build/src/microcosm/build/spec_engine/inventory_coverage.py +++ b/packages/microcosm-build/src/microcosm/build/spec_engine/inventory_coverage.py @@ -359,8 +359,8 @@ "late_schedule": "e59c019d3d454eac99ac0ac209b6c5b6faaf9bdfcaeee18c36a25be19bf7da2f", "ownership": "5f64f0aac49e2313177564f71876bffc8c81b3ded4df701e70930e60e9c98356", "primary_tuples": "fdf23da429f8c501e198a2f5b719763d523a41565c2153fdf17b0344f21bcd7c", - "seed_map": "0fa116d9f5dae4ba0378f190d6c822b92f3dcf1e3fe80c9880bcaf87431c283c", - "seed_protocol": "b4afa376b62fbc073fd10d74e40921f5ab60c1efe366e2c8504872091ea69d08", + "seed_map": "b4d6ec9c11ce323d83319a3d0289e0cf7216a7936d6c6f08cef0a463ad369641", + "seed_protocol": "114217a2918329b0601d94deb8ed83f73919584b85ca8aabe45a40665afb6560", "source_manifest": "f4f8bfeb79268a4043a843b47cec1d1c0e58e9f5dd6bc23417ada95446dd7b8e", "take_up": "9522ce40f7dea569312dd7a7beb474e5a5afadd2f1e10cb5876534c3ef623d35", "tail": "ac92829c88a1a4fb6460d61190918d5d99c6c377fc8dd8f62f02b332d09bf59c", diff --git a/packages/microcosm-build/src/microcosm/build/uk/gates.json b/packages/microcosm-build/src/microcosm/build/uk/gates.json index e2b38403b..d87d0f2b2 100644 --- a/packages/microcosm-build/src/microcosm/build/uk/gates.json +++ b/packages/microcosm-build/src/microcosm/build/uk/gates.json @@ -854,7 +854,7 @@ "household.cgt_support_copy_index", "household.household_is_cgt_residential_clone", "household.cgt_residential_clone_index", - "household.clone_index", + "household.household_clone_index", "household.constituency_code_oa", "household.consumer_debt", "household.data_zone_code", @@ -972,7 +972,13 @@ "person.person_source_id", "person.person_support_channel", "person.person_support_clone_index", - "person.srp_regular_code5" + "person.srp_regular_code5", + "benunit.benunit_clone_index", + "person.person_clone_index", + "household.itl1_code", + "household.itl2_code", + "household.itl3_code", + "household.ward_code" ], "reviewed_exclusions": { "person.incapacity_benefit_reported": "The enhanced FRS stores this legacy reported-benefit input as an all-zero layer; the candidate must drop dead zero layers." diff --git a/packages/microcosm-build/src/microcosm/build/uk_runtime/atomic_area_support.py b/packages/microcosm-build/src/microcosm/build/uk_runtime/atomic_area_support.py index 32872b430..8efeb97f2 100644 --- a/packages/microcosm-build/src/microcosm/build/uk_runtime/atomic_area_support.py +++ b/packages/microcosm-build/src/microcosm/build/uk_runtime/atomic_area_support.py @@ -42,6 +42,59 @@ "super_data_zone_code", "district_electoral_area_code", ) +#: The committed provenance of the three support artifacts: per-column source +#: and vintage of every derived layer, with the artifacts' digests. +PROVENANCE_RESOURCE = "uk_atomic_area_supports.provenance.json" + + +def uk_area_code_frames() -> dict[str, dict[str, str]]: + """The code frames behind the exported area codes, per support system. + + microcosm#1114: consumers check their own area lists against the frames a + release carries. The frames are the per-column vintages the three support + artifacts declare (``2024_pcon`` for Westminster constituencies of July + 2024; ``2023_april_lad``, ``2019_council_area`` and ``2014_lgd`` for the + April 2023 local-authority code set; ``2024_rgn`` for the English regions + with the FRS sentinel codes for the other nations), read from the + committed provenance rather than restated here. Keyed by the exported + column name, then by support system. + """ + + import json + from importlib.resources import files + + from .geography_ladder import UK_EXPORT_AREA_CODE_COLUMNS + + payload = json.loads( + files("microcosm.build.uk").joinpath(PROVENANCE_RESOURCE).read_text() + ) + supports = payload["supports"] + _require(set(supports) == set(SYSTEMS), "provenance covers the three systems") + frames: dict[str, dict[str, str]] = {} + for ladder_column, export_column in UK_EXPORT_AREA_CODE_COLUMNS.items(): + frames[export_column] = {} + for system in SYSTEMS: + vintage = supports[system]["column_metadata"][ladder_column]["vintage"] + _require(_text(vintage), f"{system} {ladder_column} vintage is text") + frames[export_column][system] = str(vintage) + return frames + + +def without_uk_native_alias_columns(household): + """The household table without the nation-native alias columns. + + The aliases are NA outside their own nation by design, so no single-year + dataset may carry them: not the release export and not the scratch + dataset a measure engine loads (policyengine-uk refuses any NaN column). + Returns the same object when the table carries none of them. + """ + + aliases = [c for c in UK_NATIVE_ALIAS_COLUMNS if c in household.columns] + if not aliases: + return household + return household.drop(columns=aliases) + + _INPUT_COLUMNS = ( "oa_code", "population", diff --git a/packages/microcosm-build/src/microcosm/build/uk_runtime/atomic_household_identity.py b/packages/microcosm-build/src/microcosm/build/uk_runtime/atomic_household_identity.py index 7f8ad95aa..a36af7985 100644 --- a/packages/microcosm-build/src/microcosm/build/uk_runtime/atomic_household_identity.py +++ b/packages/microcosm-build/src/microcosm/build/uk_runtime/atomic_household_identity.py @@ -15,11 +15,13 @@ #: Structural branches in spine order: the SPI support channel, the CGT #: support split (copy ``k`` of a household, microcosm#1045), the incidence -#: clone and the geographic pool expansion. +#: clone, the residential split of a gaining clone (arm ``k`` of the clone, +#: microcosm#1063 item 7) and the geographic pool expansion. _BRANCHES = ( "spi_support_channel", "cgt_support_split", "cgt_incidence_clone", + "cgt_residential_split", "geographic_support", ) diff --git a/packages/microcosm-build/src/microcosm/build/uk_runtime/battery_bindings.py b/packages/microcosm-build/src/microcosm/build/uk_runtime/battery_bindings.py index af15e80d2..ca573aef0 100644 --- a/packages/microcosm-build/src/microcosm/build/uk_runtime/battery_bindings.py +++ b/packages/microcosm-build/src/microcosm/build/uk_runtime/battery_bindings.py @@ -78,7 +78,10 @@ load_uk_local_geography_contract, metric_names, ) -from microcosm.build.uk_runtime.national_frame import _uk_gate_surface +from microcosm.build.uk_runtime.national_frame import ( + UK_RELEASE_EXPORT_DROPPED_COLUMNS, + _uk_gate_surface, +) from microcosm.build.uk_runtime.release_input_coverage import ( assert_uk_release_input_coverage_build_stages, assert_uk_release_input_coverage_manifest_current, @@ -122,6 +125,7 @@ uk_qrf_tail_concentration_gate, ) from microcosm.calibrate.registry import TargetSpec +from microcosm.frame import Frame __all__ = [ "UK_GATE_REGISTRY", @@ -226,14 +230,31 @@ def _evaluate_release_input_coverage( # The release cut supplies the spine frame the stages produced, so the # family build-state half reads importance weights and stage receipts # where they live; the coverage halves read the release frame. + # The spine frame itself (the certifier's ``--spine-h5``), or the spine + # checkpoint's published build state (weight kind, period, mass log: all + # the build-state half reads) when a graph build hands it over. spine_frame = context.artifacts.get("spine_frame") + if spine_frame is None: + # Never the calibrated frame: its weights are the solve's, so the + # build-state half would judge the wrong frame (the first K=25 build + # failed all 17 families on weight kind that way). The binding's + # artifact selector marks the gate evidence_absent before this point; + # this refusal is the backstop. + raise ValueError( + "release_input_coverage needs the spine build state: the certifier " + "supplies --spine-h5 and a graph build the checkpoint's " + "spine_build_state artifact; none arrived, so the family " + "build-state half cannot be evaluated." + ) + if isinstance(spine_frame, Frame): + build_state = _uk_gate_surface(spine_frame) + else: + build_state = spine_frame return uk_release_input_coverage_gate( _uk_gate_surface(context.frame), engine, manifest=manifest, - build_state_frame=None - if spine_frame is None - else _uk_gate_surface(spine_frame), + build_state_frame=build_state, ) @@ -241,6 +262,16 @@ def _coverage_requires_frame(parameters: Mapping[str, Any]) -> bool: return parameters.get("check") != "manifest_current" +def _coverage_required_artifacts(parameters: Mapping[str, Any]) -> frozenset[str]: + # The preflight check reads the engine and the manifest only; the + # evaluation needs the spine build state beside them (microcosm#1115 + # review): without it the gate is evidence_absent, never a verdict on + # the calibrated frame. + if parameters.get("check") == "manifest_current": + return frozenset({"coverage_engine"}) + return frozenset({"coverage_engine", "spine_frame"}) + + def _evaluate_source_coverage( context: EvidenceContext, parameters: Mapping[str, Any] ) -> GateResult: @@ -934,6 +965,10 @@ def _evaluate_degenerate_release_surface( _uk_gate_surface(context.frame), reviewed_exclusions=resolved, now=_exclusion_clock(context), + # The release boundary drops these before writing; the certifier never + # sees them on the exported H5, so a gate fed the pre-export frame + # must not report them (found by the first graph dense build). + dropped_at_export=UK_RELEASE_EXPORT_DROPPED_COLUMNS, **kwargs, ) @@ -1647,6 +1682,7 @@ def _ledger_compile_parity_required_artifacts( evaluator=_evaluate_release_input_coverage, parameter_keys=frozenset({"check"}), artifact_keys=frozenset({"coverage_engine"}), + artifact_selector=_coverage_required_artifacts, frame_predicate=_coverage_requires_frame, legacy_name="uk_release_input_coverage", ), diff --git a/packages/microcosm-build/src/microcosm/build/uk_runtime/dataset_size.py b/packages/microcosm-build/src/microcosm/build/uk_runtime/dataset_size.py index d0e72b863..71a264b8d 100644 --- a/packages/microcosm-build/src/microcosm/build/uk_runtime/dataset_size.py +++ b/packages/microcosm-build/src/microcosm/build/uk_runtime/dataset_size.py @@ -52,6 +52,7 @@ class UKSizeSelection: learning_rate: float seed: int search_pi_hi: float + initial_lambda: float | None = None @dataclass(frozen=True) @@ -180,6 +181,20 @@ def _check_baseline_pi_floor(value: object) -> float: return floor +def _check_initial_lambda(value: object) -> float | None: + """``None`` (cold search) or a positive finite warm-start penalty.""" + if value is None: + return None + if ( + isinstance(value, bool) + or not isinstance(value, int | float) + or not np.isfinite(value) + or float(value) <= 0.0 + ): + raise ValueError("initial_lambda must be None or a positive finite number.") + return float(value) + + def _check_pi_hi(pi_hi: object) -> float: if not isinstance(pi_hi, float | int) or isinstance(pi_hi, bool): raise ValueError("pi_hi must be a number in (0, 1].") @@ -214,6 +229,7 @@ def select_uk_dataset_size( learning_rate: float, seed: int, pi_hi: float = 1.0, + initial_lambda: float | None = None, progress_callback: ProgressCallback | None = None, ) -> UKSizeSelection: """Run the informed L0 budget search for an exact-count draw of ``households``. @@ -223,9 +239,18 @@ def select_uk_dataset_size( the gates' open-probability mass until a draw of ``households`` at ``pi_hi`` is feasible on the learned probabilities (or the probe budget is spent; then the draw refuses with the measurement). + + ``initial_lambda`` warm-starts the search at a known penalty (the one a + previous search on the same pool and targets selected, microcosm#1115): + the search probes it first and stops there when the draw is feasible and + within tolerance, saving the full bisection. The search still verifies + every probe, so a stale hint costs probes, never feasibility; any penalty + whose draw lands inside the budget window is a valid stop, so a warm and a + cold search can settle on different penalties and select different rows. """ n = _check_size_inputs(frame, dense, households) pi_hi = _check_pi_hi(pi_hi) + initial_lambda = _check_initial_lambda(initial_lambda) if households == n: raise ValueError("a full-pool size needs no selection.") problem = dense.problem @@ -262,6 +287,7 @@ def select_uk_dataset_size( mass_reason=dense.options["mass_reason"], budget_basis=BUDGET_BASIS_OPEN_PROBABILITY_MASS, feasible_draw_pi_hi=pi_hi, + l0_lambda=0.0 if initial_lambda is None else initial_lambda, progress_callback=_phased(progress_callback, "size_search"), **_solver_common(dense, epochs=epochs, learning_rate=learning_rate, seed=seed), ) @@ -275,6 +301,7 @@ def select_uk_dataset_size( learning_rate=learning_rate, seed=seed, search_pi_hi=pi_hi, + initial_lambda=initial_lambda, ) @@ -287,6 +314,7 @@ def refit_uk_dataset_size( learning_rate: float, seed: int, pi_hi: float = 1.0, + initial_lambda: float | None = None, baseline_pi_floor: float = 0.0, selection: UKSizeSelection | None = None, draw: UKSizeDraw | None = None, @@ -350,6 +378,7 @@ def refit_uk_dataset_size( learning_rate=learning_rate, seed=seed, pi_hi=pi_hi, + initial_lambda=initial_lambda, progress_callback=progress_callback, ) reused = False diff --git a/packages/microcosm-build/src/microcosm/build/uk_runtime/full_build_cli.py b/packages/microcosm-build/src/microcosm/build/uk_runtime/full_build_cli.py index ffbd03c4b..11d1d413d 100644 --- a/packages/microcosm-build/src/microcosm/build/uk_runtime/full_build_cli.py +++ b/packages/microcosm-build/src/microcosm/build/uk_runtime/full_build_cli.py @@ -358,6 +358,20 @@ def parse_args(argv: list[str] | None = None) -> argparse.Namespace: parser.add_argument("--seed", type=int, help="Defaults to the role's seed.") parser.add_argument("--selection-seed", type=int) parser.add_argument("--selection-pi-hi", type=float, default=1.0) + parser.add_argument( + "--selection-initial-lambda", + type=float, + default=None, + help=( + "Warm-start the informed L0 budget search at this penalty (the " + "selected_l0_lambda of a previous search on the same pool and " + "targets, microcosm#1115). The search probes it first and stops " + "there when the draw is feasible and within tolerance; a stale " + "value costs probes, not feasibility (a warm and a cold search may " + "settle on different penalties inside the budget window). Requires " + "--dataset-households." + ), + ) parser.add_argument( "--baseline-pi-floor", type=float, @@ -733,6 +747,7 @@ def prepare_full_build( dataset_households=args.dataset_households, selection_seed=args.selection_seed, selection_pi_hi=args.selection_pi_hi, + selection_initial_lambda=args.selection_initial_lambda, baseline_pi_floor=args.baseline_pi_floor, target_weight_rule=args.target_weight_rule, ), diff --git a/packages/microcosm-build/src/microcosm/build/uk_runtime/full_gates.py b/packages/microcosm-build/src/microcosm/build/uk_runtime/full_gates.py index 692f43089..1bc886834 100644 --- a/packages/microcosm-build/src/microcosm/build/uk_runtime/full_gates.py +++ b/packages/microcosm-build/src/microcosm/build/uk_runtime/full_gates.py @@ -18,7 +18,11 @@ from microcosm.build.country_spec import GatesManifest, load_country_spec from microcosm.build.gate_battery import EvidenceContext -from microcosm.build.uk_runtime.calibration_run import uk_aggregate_admin_totals +from microcosm.build.uk_runtime.calibration_run import ( + uk_aggregate_admin_totals, + uk_cgt_projection_artifact, +) +from microcosm.build.uk_runtime.cgt_projection import UK_CGT_PROJECTION_ARTIFACT_KEY from microcosm.build.uk_runtime.graph_evidence import uk_spine_gate_artifacts from microcosm.build.uk_runtime.national_frame import uk_release_export_frame from microcosm.build.uk_runtime.parity_reference import load_efrs_parity_reference @@ -247,6 +251,20 @@ def build_full_gate_context( manifest = uk_full_gate_manifest(selection_receipt) admin_totals, admin_receipt = uk_aggregate_admin_totals(frame, manifest) artifacts = dict(supporting_evidence) + # The spine checkpoint's build state stands in for the certifier's spine + # frame: the input-coverage gate's build-state half reads the stages' + # importance weights and mass receipts from it, never from the calibrated + # release frame (found by the first graph dense build: every required + # family failed on weight kind alone). + spine_build_state = artifacts.pop("spine_build_state", None) + if spine_build_state is not None: + artifacts.setdefault("spine_frame", spine_build_state) + # The entrants fence needs the projection; the national seam computes it + # in its calibration kernel, the dense battery computes it here. + if UK_CGT_PROJECTION_ARTIFACT_KEY not in artifacts: + artifacts[UK_CGT_PROJECTION_ARTIFACT_KEY] = uk_cgt_projection_artifact( + frame, manifest + ) artifacts.update( { "stage_evidence": dict(stage_evidence), diff --git a/packages/microcosm-build/src/microcosm/build/uk_runtime/full_measure.py b/packages/microcosm-build/src/microcosm/build/uk_runtime/full_measure.py index 0dac211c3..fb82aa4a2 100644 --- a/packages/microcosm-build/src/microcosm/build/uk_runtime/full_measure.py +++ b/packages/microcosm-build/src/microcosm/build/uk_runtime/full_measure.py @@ -6,12 +6,14 @@ from __future__ import annotations +import gc from collections.abc import Mapping from pathlib import Path from typing import Any import numpy as np import pandas as pd +from scipy import sparse from microcosm.build.target_materialization import resolve_target_measures from microcosm.build.uk_runtime import ( @@ -24,6 +26,7 @@ materialize_uk_ledger_targets, ) from microcosm.build.uk_runtime.measure_simulation import UKMeasureResolver +from microcosm.calibrate.matrix import CalibrationProblem, build_constraint_matrix from microcosm.frame import Frame, MassChangeRecord UK_BLOCK_SENSITIVE_MEASURE_COLUMNS = ( @@ -32,6 +35,38 @@ "slc/student_loan_repayment/england", ) +#: policyengine-uk's weight-share formulas (a national total or a share +#: allocated by ``x * w / sum(x * w)``) and the frame columns their allocation +#: key ``x`` reads (microcosm#1115 review). Scaling a block's engine weights by +#: pool mass / block mass makes such a formula exact only when every block +#: carries the same ``sum(x * w)``; identical (source household, weight) +#: multisets guarantee that only while ``x`` is copied into every clone, which +#: holds for these keys (spine wealth and consumption inputs) and would not for +#: anything drawn per clone. The representation check therefore records each +#: block's ``sum(x * w)`` for every key here and requires them to agree; a +#: formula of this shape that reads a per-clone input must be added here so +#: the check sees it, never assumed exact. +UK_WEIGHT_SHARE_FORMULA_INPUTS: Mapping[str, tuple[str, ...]] = { + # corporate_land_value and shareholding allocate by corporate_sector_wealth. + "corporate_land_value": ("corporate_wealth", "private_pension_wealth"), + "shareholding": ("corporate_wealth", "private_pension_wealth"), + # consumption_shareholding allocates by the engine's consumption total. + "consumption_shareholding": ( + "food_and_non_alcoholic_beverages_consumption", + "alcohol_and_tobacco_consumption", + "clothing_and_footwear_consumption", + "housing_water_and_electricity_consumption", + "household_furnishings_consumption", + "health_consumption", + "transport_consumption", + "communication_consumption", + "recreation_consumption", + "education_consumption", + "restaurants_and_hotels_consumption", + "miscellaneous_consumption", + ), +} + def _without_scratch_paths(value, scratch_dir: Path): """Record scratch-relative paths in a receipt, never the scratch root. @@ -62,87 +97,226 @@ def _without_scratch_paths(value, scratch_dir: Path): return value -def resolve_uk_full_measures( - frame, - national_registry, - *, - period: int, - scratch_dir: Path, - band_edge_registry=None, - resolver_factory=UKMeasureResolver, - blocks: int = 1, - local_grains: tuple[str, ...] = ("constituency", "la"), -) -> tuple[Any, Any, UKRowwiseNationalRows, dict[str, pd.DataFrame], dict[str, Any]]: - """Resolve national inputs and local metrics on the cloned frame. - - ``blocks=1`` uses one scratch-mode engine for the whole clone. The - reviewed escape hatch ``blocks=K`` resolves each clone index separately, - then rejoins every entity-level prepared column by its stable entity id so - the full-frame target materialization and single solve retain frame order. - """ +def _engine_block_frames(frame, blocks: int) -> list[tuple[int | None, Frame]]: + """The engine blocks: the whole clone, or one scratch frame per clone index.""" household = frame.table("household") if blocks < 1: raise ValueError("engine resolution blocks must be positive.") if blocks == 1: - block_frames = [(None, frame)] - else: - clone_column = ladder_clone_index_column("household") - if clone_column not in household.columns: - raise ValueError(f"per-clone engine resolution requires {clone_column}.") - clone_indices = tuple(sorted(household[clone_column].unique().tolist())) - if len(clone_indices) != blocks: + return [(None, frame)] + clone_column = ladder_clone_index_column("household") + if clone_column not in household.columns: + raise ValueError(f"per-clone engine resolution requires {clone_column}.") + clone_indices = tuple(sorted(household[clone_column].unique().tolist())) + if len(clone_indices) != blocks: + raise ValueError( + "engine resolution blocks must match the realized clone indices: " + f"requested {blocks}, found {clone_indices}." + ) + person = frame.table("person") + block_frames = [] + for clone_index in clone_indices: + household_ids = set( + household.loc[ + household[clone_column] == clone_index, + "household_id", + ].tolist() + ) + person_mask = person["person_household_id"].isin(household_ids) + block = frame.select(person_mask) + # The block carries a K-th of the cloned mass while its log still + # ends on the full-clone record, and the scratch export validates + # the chain. Declare the subset explicitly: old = the cloned + # total, new = the block total, reason naming the block. The block + # frame is engine scratch and is discarded after resolution. + block_weights = block.weights_for("household") + full_total = float(frame.weights_for("household").total) + block_total = float(block_weights.total) + subset_record = MassChangeRecord( + entity="household", + old_total=full_total, + new_total=block_total, + declared_factor=block_total / full_total, + reason=( + f"engine resolution block {clone_index} of {blocks}: " + "scratch subset of the cloned frame for measure " + "resolution only, discarded after resolution" + ), + ) + block = Frame( + { + **{name: block.table(name) for name in block.entities}, + **{name: block.link(name) for name in block.links}, + }, + block.schema, + {entity: block.weights_for(entity) for entity in block.weighted_entities}, + block.strata, + mass_log=(*block.mass_log, subset_record), + metadata=block.metadata, + ) + block_frames.append((clone_index, block)) + + return block_frames + + +#: Columns that identify the spine household a cloned row was copied from, in +#: order of preference; every clone of a spine household shares the value. +_SOURCE_IDENTITY_COLUMNS = ( + "household_source_id", + "source_household_id", + "source_household_key", +) + + +def _engine_population_representation(frame, block_frames) -> dict[str, Any]: + """How each engine block stands for the pool, and whether exactly. + + A weight-share formula (a national total allocated by ``x*w / sum(x*w)``, + policyengine-uk's corporate land value among them) sees only its block's + denominator, so each block would reproduce the whole national total: the + ``K`` times artefact behind the #736 erratum. Every block's engine weights + are therefore scaled by ``pool mass / block mass``. For blocks that are + identical copies of one another (the clone expansion's contract) every + weighted sum then equals the pool's and the formula is exact; the checks + record whether the blocks are such copies. A pool that is not leaves the + scaling approximate, and the measures receipt keeps its caveat. + """ + + if len(block_frames) == 1 and block_frames[0][0] is None: + return {"mode": "single_block", "exact": True, "blocks": 1} + pool_mass = float(frame.weights_for("household").total) + signatures: list[tuple[np.ndarray, np.ndarray]] = [] + identity_key: str | None = None + identity_present = True + household_counts: list[int] = [] + person_counts: list[int] = [] + factors: dict[str, float] = {} + # sum(x * w) per block for every weight-share allocation key, and the key + # columns a block lacks (an absent column leaves the formula unverified). + input_sums: dict[str, dict[str, float]] = { + formula: {} for formula in UK_WEIGHT_SHARE_FORMULA_INPUTS + } + missing_inputs: dict[str, list[str]] = {} + for clone_index, block in block_frames: + household = block.table("household") + weights = np.asarray(block.weights_for("household").values, dtype=np.float64) + block_mass = float(weights.sum()) + if not block_mass > 0.0: raise ValueError( - "engine resolution blocks must match the realized clone indices: " - f"requested {blocks}, found {clone_indices}." - ) - person = frame.table("person") - block_frames = [] - for clone_index in clone_indices: - household_ids = set( - household.loc[ - household[clone_column] == clone_index, - "household_id", - ].tolist() - ) - person_mask = person["person_household_id"].isin(household_ids) - block = frame.select(person_mask) - # The block carries a K-th of the cloned mass while its log still - # ends on the full-clone record, and the scratch export validates - # the chain. Declare the subset explicitly: old = the cloned - # total, new = the block total, reason naming the block. The block - # frame is engine scratch and is discarded after resolution. - block_weights = block.weights_for("household") - full_total = float(frame.weights_for("household").total) - block_total = float(block_weights.total) - subset_record = MassChangeRecord( - entity="household", - old_total=full_total, - new_total=block_total, - declared_factor=block_total / full_total, - reason=( - f"engine resolution block {clone_index} of {blocks}: " - "scratch subset of the cloned frame for measure " - "resolution only, discarded after resolution" - ), + f"engine block {clone_index} carries no household mass; it cannot " + "represent the pool." ) - block = Frame( - { - **{name: block.table(name) for name in block.entities}, - **{name: block.link(name) for name in block.links}, - }, - block.schema, - { - entity: block.weights_for(entity) - for entity in block.weighted_entities - }, - block.strata, - mass_log=(*block.mass_log, subset_record), - metadata=block.metadata, - ) - block_frames.append((clone_index, block)) + factors[str(clone_index)] = pool_mass / block_mass + # Blocks are compared as multisets of (source household, weight): the + # clone expansion copies every spine household into every clone, so + # identical copies have identical multisets whatever geography each + # clone drew. Without a source identity the comparison is weights + # only, which cannot establish copies, so the result is never exact. + key = next((c for c in _SOURCE_IDENTITY_COLUMNS if c in household), None) + if key is None: + identity_present = False + source = np.zeros(len(weights), dtype=np.int64) + else: + identity_key = key + source = household[key].to_numpy() + order = np.lexsort((weights, source)) + signatures.append((source[order], weights[order])) + household_counts.append(int(len(household))) + person_counts.append(int(len(block.table("person")))) + for formula, columns in UK_WEIGHT_SHARE_FORMULA_INPUTS.items(): + absent = [c for c in columns if c not in household.columns] + if absent: + missing_inputs.setdefault(formula, []) + for column in absent: + if column not in missing_inputs[formula]: + missing_inputs[formula].append(column) + continue + key_values = np.zeros(len(weights), dtype=np.float64) + for column in columns: + key_values += np.asarray(household[column], dtype=np.float64) + input_sums[formula][str(clone_index)] = float(np.sum(key_values * weights)) + first_source, first_weights = signatures[0] + identical = all( + len(source) == len(first_source) + and np.array_equal(source, first_source) + and np.array_equal(weights, first_weights) + for source, weights in signatures + ) + equal_households = len(set(household_counts)) == 1 + equal_persons = len(set(person_counts)) == 1 + blocks = len(block_frames) + shares = np.array([1.0 / factor for factor in factors.values()], dtype=np.float64) + weight_share_inputs: dict[str, dict[str, object]] = {} + inputs_match = True + for formula, columns in UK_WEIGHT_SHARE_FORMULA_INPUTS.items(): + sums = input_sums[formula] + if formula in missing_inputs or len(sums) != blocks: + inputs_match = False + weight_share_inputs[formula] = { + "columns": list(columns), + "missing_columns": list(missing_inputs.get(formula, [])), + "sum_by_block": dict(sums), + "match": False, + } + continue + values = np.array(list(sums.values()), dtype=np.float64) + scale = max(float(np.max(np.abs(values))), 1.0) + max_rel_diff = float((np.max(values) - np.min(values)) / scale) + match = bool(max_rel_diff <= 1e-9) + inputs_match = inputs_match and match + weight_share_inputs[formula] = { + "columns": list(columns), + "missing_columns": [], + "sum_by_block": dict(sums), + "max_rel_diff": max_rel_diff, + "match": match, + } + return { + "mode": "block_weights_scaled_to_pool", + "exact": bool( + identity_present + and identical + and equal_households + and equal_persons + and inputs_match + ), + "blocks": blocks, + "factor_by_block": factors, + "checks": { + "identity_key": identity_key, + "source_identity_present": identity_present, + "equal_household_counts": equal_households, + "equal_person_counts": equal_persons, + "identical_source_weight_multisets": identical, + "weight_share_inputs_match": inputs_match, + "weight_share_inputs": weight_share_inputs, + "max_abs_mass_share_deviation": float( + np.max(np.abs(shares - 1.0 / blocks)) + ), + }, + } + + +def _run_engine_blocks( + block_frames, + national_registry, + *, + period: int, + scratch_dir: Path, + resolver_factory, + local_grains: tuple[str, ...], + on_block, + engine_weight_scales: Mapping[str, float] | None = None, +) -> tuple[ + dict[str, list[pd.DataFrame]], + list[Mapping[str, Any]], + list[dict[str, Any]], + set[tuple[str, str]], +]: + """Resolve every block's measures, hand each block's inputs to ``on_block`` + while its engine is alive, and release the engine before the next loads.""" - measure_parts: dict[tuple[str, str], list[pd.Series]] = {} metric_parts: dict[str, list[pd.DataFrame]] = {grain: [] for grain in local_grains} resolver_receipts: list[Mapping[str, Any]] = [] resolution_receipts: list[dict[str, Any]] = [] @@ -151,12 +325,17 @@ def resolve_uk_full_measures( block_scratch = ( scratch_dir if clone_index is None else scratch_dir / f"clone-{clone_index}" ) - resolver = resolver_factory( - simulation_source=None, - scratch_dir=block_scratch, - year=period, - frame=block_frame, - ) + resolver_kwargs: dict[str, Any] = { + "simulation_source": None, + "scratch_dir": block_scratch, + "year": period, + "frame": block_frame, + } + if engine_weight_scales is not None and clone_index is not None: + resolver_kwargs["engine_weight_scale"] = float( + engine_weight_scales[str(clone_index)] + ) + resolver = resolver_factory(**resolver_kwargs) resolution = resolve_target_measures( lambda block_frame=block_frame: CalibrationFrameAdapter(block_frame), national_registry, @@ -173,15 +352,7 @@ def resolve_uk_full_measures( raise RuntimeError( "per-clone engine resolution returned inconsistent national inputs." ) - for (entity, variable), values in resolution.measure_inputs.items(): - entity_table = block_frame.table(entity) - entity_id = f"{entity}_id" - measure_parts.setdefault((entity, variable), []).append( - pd.Series( - np.asarray(values), - index=entity_table[entity_id].tolist(), - ) - ) + on_block(block_frame, resolution.measure_inputs) block_household_ids = block_frame.table("household")["household_id"].tolist() for area_type in metric_parts: metric_parts[area_type].append( @@ -195,30 +366,27 @@ def resolve_uk_full_measures( resolver_receipts.append( _without_scratch_paths(dict(resolver.receipt()), scratch_dir) ) - del resolver + # A policyengine simulation is a large cyclic object graph; the + # collector does not reclaim it on `del`. Collect before the next + # block loads so one engine is alive at a time. + del resolution, resolver + gc.collect() simulation_input = block_scratch / "simulation-input.h5" simulation_input.unlink(missing_ok=True) try: block_scratch.rmdir() except OSError: pass + return ( + metric_parts, + resolver_receipts, + resolution_receipts, + set() if national_input_keys is None else national_input_keys, + ) - measure_inputs: dict[tuple[str, str], np.ndarray] = {} - for (entity, variable), parts in measure_parts.items(): - combined = pd.concat(parts) - if combined.index.has_duplicates: - raise RuntimeError( - f"per-clone engine resolution duplicated {entity} ids for {variable}." - ) - ordered_ids = frame.table(entity)[f"{entity}_id"] - ordered = combined.reindex(ordered_ids.tolist()) - if ordered.isna().any(): - raise RuntimeError( - f"per-clone engine resolution missed {entity} rows for {variable}." - ) - measure_inputs[(entity, variable)] = ordered.to_numpy() - full_household_ids = household["household_id"].tolist() +def _rejoin_local_metrics(frame, metric_parts) -> dict[str, pd.DataFrame]: + full_household_ids = frame.table("household")["household_id"].tolist() local_metrics = {} for area_type, parts in metric_parts.items(): combined = pd.concat(parts) @@ -232,29 +400,10 @@ def resolve_uk_full_measures( f"per-clone engine resolution missed {area_type} household rows." ) local_metrics[area_type] = ordered + return local_metrics - adapter = CalibrationFrameAdapter(frame) - # Injected engine inputs are scratch state for materialization only: - # they must be dropped before the prepared frame is assembled, or the - # flattening rule refuses columns that now exist on two entities - # (region, esa_* on the live spine). Same lifecycle as the national stage. - original_columns = { - entity: set(table.columns) for entity, table in adapter.tables.items() - } - inject_measure_inputs(adapter, measure_inputs) - materialized = materialize_uk_ledger_targets( - adapter, - national_registry, - period=period, - band_edge_registry=( - national_registry if band_edge_registry is None else band_edge_registry - ), - ) - if materialized.skipped: - raise RuntimeError( - "candidate national target materialization skipped row(s): " - f"{[skip.__dict__ for skip in materialized.skipped]}." - ) + +def _block_provenance(resolver_receipts) -> tuple[Any, Any, Any]: modes = {receipt.get("mode") for receipt in resolver_receipts} versions = {receipt.get("policyengine_uk_version") for receipt in resolver_receipts} if len(modes) != 1 or len(versions) != 1: @@ -265,19 +414,35 @@ def resolve_uk_full_measures( for block_receipt in resolver_receipts[1:] ): raise RuntimeError("per-clone CGT period contract is inconsistent.") + return next(iter(modes)), next(iter(versions)), cgt_period_contract + + +def _measures_receipt( + frame, + *, + mode, + engine_version, + cgt_period_contract, + national_input_keys: set[tuple[str, str]], + local_metrics: dict[str, pd.DataFrame], + blocks: int, + materialization_report, + resolution_receipts, + representation: Mapping[str, Any] | None = None, +) -> dict[str, Any]: receipt = { - "mode": next(iter(modes)), - "engine_version": next(iter(versions)), + "mode": mode, + "engine_version": engine_version, "households": len(frame.table("household")), "persons": len(frame.table("person")), "benunits": len(frame.table("benunit")), - "national_inputs": len(measure_inputs), + "national_inputs": len(national_input_keys), "local_metrics": { area_type: len(metrics.columns) for area_type, metrics in local_metrics.items() }, "blocks": blocks, - "target_materialization": materialized.report(), + "target_materialization": materialization_report, # The resolution loop's own receipt per block (which measure came from # which provider, the rounds, the provider's receipt): the seam # manifest's ``measure_resolution`` block, one per engine block. @@ -286,34 +451,162 @@ def resolve_uk_full_measures( if cgt_period_contract is not None: receipt["cgt_period_contract"] = cgt_period_contract if blocks > 1: + # The single-block receipt keeps its shape; a per-block resolution says + # how its blocks stood for the pool. + if representation is not None: + receipt["engine_population_representation"] = dict(representation) + exact = bool(representation and representation.get("exact")) receipt["deviation"] = "per_clone_block_engine_resolution" present = sorted( column for column in UK_BLOCK_SENSITIVE_MEASURE_COLUMNS - if column in {variable for _, variable in measure_inputs} + if column in {variable for _, variable in national_input_keys} ) receipt["block_sensitivity"] = { "known_population_normalised_measures": list( UK_BLOCK_SENSITIVE_MEASURE_COLUMNS ), "present_in_this_run": present, + "weight_share_formulas": { + formula: list(columns) + for formula, columns in UK_WEIGHT_SHARE_FORMULA_INPUTS.items() + }, + "mitigation": ( + "each block's engine household weights are scaled by pool mass / " + "block mass, so weight-share formulas see the pool's denominator; " + "exact only while the blocks are identical copies and every " + "allocation key named in weight_share_formulas carries the same " + "sum(x * w) in every block" + ), + "exact": exact, "caveat": ( - "per-block engine resolution mis-measures population-normalised " - "formulas (each block reproduces a national aggregate); rows " - "on these measures are not evidence for adjudication from this " - "run. Resolve in a single block before ruling on them." + None + if exact + else ( + "the engine blocks are not identical copies of one another, so " + "the scaling is approximate and per-block engine resolution " + "may mis-measure population-normalised formulas; rows on these " + "measures are not evidence for adjudication from this run. " + "Resolve in a single block before ruling on them." + ) ), } + return receipt + + +def _national_rows(national_registry, targets) -> UKRowwiseNationalRows: + return UKRowwiseNationalRows( + targets=targets, + registry=national_registry, + families=tuple(sorted({spec.family for spec in national_registry.specs})), + ) + + +def resolve_uk_full_measures( + frame, + national_registry, + *, + period: int, + scratch_dir: Path, + band_edge_registry=None, + resolver_factory=UKMeasureResolver, + blocks: int = 1, + local_grains: tuple[str, ...] = ("constituency", "la"), +) -> tuple[Any, Any, UKRowwiseNationalRows, dict[str, pd.DataFrame], dict[str, Any]]: + """Resolve national inputs and local metrics on the cloned frame. + + ``blocks=1`` uses one scratch-mode engine for the whole clone. The + reviewed escape hatch ``blocks=K`` resolves each clone index separately, + then rejoins every entity-level prepared column by its stable entity id so + the full-frame target materialization and single solve retain frame order. + The dense role's measures node uses :func:`resolve_uk_full_national_problem` + instead, which never materializes the whole pool. + """ + + block_frames = _engine_block_frames(frame, blocks) + representation = _engine_population_representation(frame, block_frames) + measure_parts: dict[tuple[str, str], list[pd.Series]] = {} + + def collect(block_frame, block_inputs): + for (entity, variable), values in block_inputs.items(): + entity_table = block_frame.table(entity) + entity_id = f"{entity}_id" + measure_parts.setdefault((entity, variable), []).append( + pd.Series( + np.asarray(values), + index=entity_table[entity_id].tolist(), + ) + ) + + metric_parts, resolver_receipts, resolution_receipts, _keys = _run_engine_blocks( + block_frames, + national_registry, + period=period, + scratch_dir=scratch_dir, + resolver_factory=resolver_factory, + engine_weight_scales=representation.get("factor_by_block"), + local_grains=local_grains, + on_block=collect, + ) + + measure_inputs: dict[tuple[str, str], np.ndarray] = {} + for (entity, variable), parts in measure_parts.items(): + combined = pd.concat(parts) + if combined.index.has_duplicates: + raise RuntimeError( + f"per-clone engine resolution duplicated {entity} ids for {variable}." + ) + ordered_ids = frame.table(entity)[f"{entity}_id"] + ordered = combined.reindex(ordered_ids.tolist()) + if ordered.isna().any(): + raise RuntimeError( + f"per-clone engine resolution missed {entity} rows for {variable}." + ) + measure_inputs[(entity, variable)] = ordered.to_numpy() + + local_metrics = _rejoin_local_metrics(frame, metric_parts) + + adapter = CalibrationFrameAdapter(frame) + # Injected engine inputs are scratch state for materialization only: + # they must be dropped before the prepared frame is assembled, or the + # flattening rule refuses columns that now exist on two entities + # (region, esa_* on the live spine). Same lifecycle as the national stage. + original_columns = { + entity: set(table.columns) for entity, table in adapter.tables.items() + } + inject_measure_inputs(adapter, measure_inputs) + materialized = materialize_uk_ledger_targets( + adapter, + national_registry, + period=period, + band_edge_registry=( + national_registry if band_edge_registry is None else band_edge_registry + ), + ) + if materialized.skipped: + raise RuntimeError( + "candidate national target materialization skipped row(s): " + f"{[skip.__dict__ for skip in materialized.skipped]}." + ) + mode, engine_version, cgt_period_contract = _block_provenance(resolver_receipts) + receipt = _measures_receipt( + frame, + mode=mode, + engine_version=engine_version, + cgt_period_contract=cgt_period_contract, + national_input_keys=set(measure_inputs), + local_metrics=local_metrics, + blocks=blocks, + representation=representation, + materialization_report=materialized.report(), + resolution_receipts=resolution_receipts, + ) try: scratch_dir.rmdir() except OSError: pass drop_injected_measure_inputs(adapter, measure_inputs, original_columns) - national_rows = UKRowwiseNationalRows( - targets=national_registry.to_target_set(), - registry=national_registry, - families=tuple(sorted({spec.family for spec in national_registry.specs})), - ) + national_rows = _national_rows(national_registry, national_registry.to_target_set()) return ( adapter.prepared_frame(), adapter.restore, @@ -321,3 +614,197 @@ def resolve_uk_full_measures( local_metrics, receipt, ) + + +def _materialize_national_problem( + block_frame, + block_inputs, + national_registry, + targets, + *, + period: int, + band_edge_registry, +) -> tuple[CalibrationProblem | None, dict[str, Any]]: + """Materialize the national targets on one engine block and compile its + columns of the national matrix, leaving the block untouched.""" + + adapter = CalibrationFrameAdapter(block_frame) + original_columns = { + entity: set(table.columns) for entity, table in adapter.tables.items() + } + inject_measure_inputs(adapter, block_inputs) + materialized = materialize_uk_ledger_targets( + adapter, + national_registry, + period=period, + band_edge_registry=band_edge_registry, + ) + if materialized.skipped: + raise RuntimeError( + "candidate national target materialization skipped row(s): " + f"{[skip.__dict__ for skip in materialized.skipped]}." + ) + drop_injected_measure_inputs(adapter, block_inputs, original_columns) + prepared = adapter.prepared_frame() + problem = None + if len(targets): + problem = build_constraint_matrix(prepared, targets, "household") + if problem.skipped: + failures = "; ".join( + f"{item.target.key}: {item.reason}" for item in problem.skipped + ) + raise ValueError( + "Selected national constraints failed to compile: " + failures + ) + # The materialization is scratch state: restoring the prepared frame must + # give the block's own tables back, column for column. + clean = adapter.restore(prepared) + for entity in block_frame.entities: + if not clean.table(entity).equals(block_frame.table(entity)): + raise RuntimeError( + f"national target materialization altered the {entity} table." + ) + return problem, materialized.report() + + +def _stitch_national_problems(frame, block_problems) -> CalibrationProblem: + """One national problem over the pool from the per-block problems, columns + in the pool's household order and the pool's weights as the start.""" + + first, _ = block_problems[0] + for problem, _ in block_problems[1:]: + if ( + problem.names != first.names + or problem.weight_entity != first.weight_entity + or not np.array_equal(problem.target_vector, first.target_vector) + ): + raise RuntimeError( + "per-clone national problems compiled different target rows." + ) + household_ids = frame.table("household")["household_id"].to_numpy() + position = pd.Index(household_ids) + if not position.is_unique: + raise RuntimeError("the pool's household ids are not unique.") + stacked_ids = np.concatenate([ids for _, ids in block_problems]) + columns = position.get_indexer(stacked_ids) + if ( + len(columns) != len(household_ids) + or (columns < 0).any() + or len(np.unique(columns)) != len(columns) + ): + raise RuntimeError( + "per-clone national problems do not cover the pool's households " + "exactly once." + ) + stacked = sparse.hstack( + [problem.matrix for problem, _ in block_problems], format="csc" + ) + ordered = sparse.csr_array(stacked[:, np.argsort(columns)]) + return CalibrationProblem( + matrix=ordered, + target_vector=first.target_vector, + names=first.names, + initial_weights=frame.resolve_weights(first.weight_entity), + weight_entity=first.weight_entity, + targets=first.targets, + skipped=first.skipped, + ) + + +def resolve_uk_full_national_problem( + frame, + national_registry, + *, + period: int, + scratch_dir: Path, + band_edge_registry=None, + resolver_factory=UKMeasureResolver, + blocks: int = 1, + local_grains: tuple[str, ...] = ("constituency", "la"), +) -> tuple[ + CalibrationProblem | None, + UKRowwiseNationalRows, + dict[str, pd.DataFrame], + dict[str, Any], +]: + """Resolve the national constraint problem block by block on the cloned frame. + + The dense role's measures node calls this instead of + :func:`resolve_uk_full_measures`: materializing every national target as a + column on a K-clone pool, then copying the prepared pool to compile it and + again to check it was left untouched, needs tens of gigabytes on a + 1.6 million-household pool (the first K=25 build, 2026-10-05, reached a + 140 GB footprint and was killed before the solve). Here each engine block + materializes its own targets while its engine is alive, compiles its own + columns of the national matrix and is released; the per-block matrices are + stitched into the pool's household order. Materialization is row-wise + (band edges come from the compiled register), so the stitched problem is + the pool's problem. ``None`` when the registry selects no national target. + """ + + block_frames = _engine_block_frames(frame, blocks) + representation = _engine_population_representation(frame, block_frames) + targets = national_registry.to_target_set() + band_edges = national_registry if band_edge_registry is None else band_edge_registry + block_problems: list[tuple[CalibrationProblem, np.ndarray]] = [] + reports: list[dict[str, Any]] = [] + + def materialize(block_frame, block_inputs): + problem, report = _materialize_national_problem( + block_frame, + block_inputs, + national_registry, + targets, + period=period, + band_edge_registry=band_edges, + ) + reports.append(report) + if problem is not None: + block_problems.append( + (problem, block_frame.table("household")["household_id"].to_numpy()) + ) + + metric_parts, resolver_receipts, resolution_receipts, national_input_keys = ( + _run_engine_blocks( + block_frames, + national_registry, + period=period, + scratch_dir=scratch_dir, + resolver_factory=resolver_factory, + engine_weight_scales=representation.get("factor_by_block"), + local_grains=local_grains, + on_block=materialize, + ) + ) + if any(report != reports[0] for report in reports[1:]): + raise RuntimeError( + "per-clone national target materialization differs across blocks." + ) + local_metrics = _rejoin_local_metrics(frame, metric_parts) + national_problem = ( + _stitch_national_problems(frame, block_problems) if block_problems else None + ) + mode, engine_version, cgt_period_contract = _block_provenance(resolver_receipts) + receipt = _measures_receipt( + frame, + mode=mode, + engine_version=engine_version, + cgt_period_contract=cgt_period_contract, + national_input_keys=national_input_keys, + local_metrics=local_metrics, + blocks=blocks, + representation=representation, + materialization_report=reports[0], + resolution_receipts=resolution_receipts, + ) + receipt["national_materialization"] = "per_engine_block" + try: + scratch_dir.rmdir() + except OSError: + pass + return ( + national_problem, + _national_rows(national_registry, targets), + local_metrics, + receipt, + ) diff --git a/packages/microcosm-build/src/microcosm/build/uk_runtime/geography_ladder.py b/packages/microcosm-build/src/microcosm/build/uk_runtime/geography_ladder.py index deba2a2fb..b8708598f 100644 --- a/packages/microcosm-build/src/microcosm/build/uk_runtime/geography_ladder.py +++ b/packages/microcosm-build/src/microcosm/build/uk_runtime/geography_ladder.py @@ -84,6 +84,7 @@ from collections.abc import Mapping from dataclasses import dataclass from pathlib import Path +from types import MappingProxyType from typing import Any import numpy as np @@ -121,7 +122,8 @@ #: name is either a policyengine-uk household input (``local_authority``, the #: enum member name resolved from ``local_authority_code``; ``region`` is #: pre-assigned, so it is not rewritten) or a plain ``*_code`` data column -#: carrying an ONS GSS code. +#: carrying an ONS GSS code. Three of them leave every single-year artifact +#: under the consumers' names (:data:`UK_EXPORT_AREA_CODE_COLUMNS`). UK_GEOGRAPHY_LADDER_COLUMNS = ( "oa_code", "lsoa_code", @@ -136,6 +138,54 @@ "itl1_code", ) +#: The three area codes consumers read under uk-data's names (microcosm#1114): +#: policyengine.py's constituency and local-authority filters and impact +#: outputs, the simulation API's geographic reports and the enhanced FRS all +#: carry ``constituency_code_oa``, ``la_code_oa`` and ``region_code_oa``. The +#: graph keeps the ladder names above everywhere in memory (gates, diagnostics, +#: target compilation, the engine's scratch dataset); the single-year export +#: boundary (``graph_terminal._tables`` and ``uk_release_export_frame``) +#: renames these three and only these three, so every artifact carries one +#: name per code and no alias. The suffix is accurate here: every code derives +#: from the assigned output area or data zone. +UK_EXPORT_AREA_CODE_COLUMNS: Mapping[str, str] = MappingProxyType( + { + "constituency_code": "constituency_code_oa", + "local_authority_code": "la_code_oa", + "region_code": "region_code_oa", + } +) + + +def export_area_code_columns(household: pd.DataFrame) -> pd.DataFrame: + """Rename the three area codes to their consumer names at the export boundary. + + Renames only the ladder columns present, so a table without geography (the + national line) passes unchanged, and refuses a table that already carries + a consumer name: a stale ``*_oa`` column riding through would shadow the + ladder's own code. + """ + + stale = [ + export + for export in UK_EXPORT_AREA_CODE_COLUMNS.values() + if export in household.columns + ] + if stale: + raise ValueError( + "household table already carries consumer area-code column(s) " + f"{stale!r}; the export boundary writes them from the ladder columns." + ) + present = { + ladder: export + for ladder, export in UK_EXPORT_AREA_CODE_COLUMNS.items() + if ladder in household.columns + } + if not present: + return household + return household.rename(columns=present) + + #: The nine English regions carry real ``E12`` GSS codes; Wales rides the #: ``W92000004`` country code in ONS lookups but the FRS-calibrated region #: dimension uses the ``W99999999`` pseudo-code, so the artifact builder @@ -1209,6 +1259,7 @@ def _region_code_for_value(value: Any) -> str | None: "GEOGRAPHY_LADDER_ARTIFACT_SHA256_ATTR", "GEOGRAPHY_LADDER_VINTAGES_ATTR", "UK_ENGLAND_WALES_REGION_CODES", + "UK_EXPORT_AREA_CODE_COLUMNS", "UK_GEOGRAPHY_LADDER_COLUMNS", "UK_LONDON_REGION_CODE", "UK_OA_LADDER_DERIVED_LAYERS", @@ -1221,4 +1272,5 @@ def _region_code_for_value(value: Any) -> str | None: "uk_region_mix", "uk_geography_ladder_assignment_summary", "uk_geography_ladder_gate", + "export_area_code_columns", ] diff --git a/packages/microcosm-build/src/microcosm/build/uk_runtime/graph_build.py b/packages/microcosm-build/src/microcosm/build/uk_runtime/graph_build.py index 55faff2ed..c8702af71 100644 --- a/packages/microcosm-build/src/microcosm/build/uk_runtime/graph_build.py +++ b/packages/microcosm-build/src/microcosm/build/uk_runtime/graph_build.py @@ -44,6 +44,7 @@ register_uk_calibration_kernels, uk_calibration_nodes, ) +from .graph_evidence import SPINE_BUILD_STATE_TYPE from .graph_kernels import UKClaimKernel, UKIdentityKernel, _normalize_create_frame from .graph_population import ( UKPoolCheckpointKernel, @@ -298,12 +299,34 @@ def run(self, context: KernelContext) -> KernelResult: provenance = calibration_run.strict_spine_provenance_from_sidecar( sidecar_path, sidecar, gate_report_path=context.sources["uk_spine_gates"] ) + created = _normalize_create_frame(frame, context) return KernelResult( - frame=_normalize_create_frame(frame, context), - artifacts={"spine_provenance": canonical_json(provenance)}, + frame=created, + artifacts={ + "spine_provenance": canonical_json(provenance), + "spine_build_state": canonical_json(spine_build_state(created)), + }, ) +def spine_build_state(frame: Frame) -> dict[str, object]: + """The checkpoint's build state the terminal gates read as the spine frame. + + The input-coverage gate's family build-state half checks each stage + family's typed importance weights, period and mass receipt on the spine + the stages produced (the certifier's ``--spine-h5``); a graph build has + no second population to hand a kernel, so the checkpoint publishes the + three things that half reads. + """ + + return { + "schema": "microcosm.uk.spine-build-state.v1", + "household_weight_kind": frame.weights_for("household").kind.value, + "time_period": str(frame.metadata["time_period"]), + "mass_log": [asdict(record) for record in frame.mass_log], + } + + def bound_spine_graph(frame: Frame) -> Graph: """Declare a checkpoint schema; the CREATE kernel verifies bound evidence.""" @@ -366,6 +389,7 @@ def token(dtype): outputs=outputs, artifact_outputs=( ArtifactOutput("spine_provenance", SPINE_PROVENANCE_TYPE), + ArtifactOutput("spine_build_state", SPINE_BUILD_STATE_TYPE), ), description="Resume a canonical spine with its bound lineage and source evidence.", ), diff --git a/packages/microcosm-build/src/microcosm/build/uk_runtime/graph_calibration.py b/packages/microcosm-build/src/microcosm/build/uk_runtime/graph_calibration.py index 6d2c3bc56..a6881d49d 100644 --- a/packages/microcosm-build/src/microcosm/build/uk_runtime/graph_calibration.py +++ b/packages/microcosm-build/src/microcosm/build/uk_runtime/graph_calibration.py @@ -79,6 +79,7 @@ class UKGraphCalibrationConfig: dataset_households: int | None = None selection_seed: int | None = None selection_pi_hi: float = 1.0 + selection_initial_lambda: float | None = None baseline_pi_floor: float = 0.0 target_weight_rule: str = "uniform" @@ -95,6 +96,7 @@ def __post_init__(self): ): raise ValueError("Dataset household count must be a positive integer.") dataset_size._check_pi_hi(self.selection_pi_hi) + dataset_size._check_initial_lambda(self.selection_initial_lambda) dataset_size._check_baseline_pi_floor(self.baseline_pi_floor) @@ -810,6 +812,10 @@ def uk_calibration_nodes( else config.selection_seed, "households": config.dataset_households, "pi_hi": config.selection_pi_hi, + # microcosm#1115: a warm-start penalty for the budget search; the + # search verifies it like any probe, so it is a parameter of the + # search node and of the nodes that reuse its selection. + "initial_lambda": config.selection_initial_lambda, } dense_input = ArtifactInput("dense", dense_id, "result", RESULT_TYPE) search_id, draw_id, refit_id = ( diff --git a/packages/microcosm-build/src/microcosm/build/uk_runtime/graph_evidence.py b/packages/microcosm-build/src/microcosm/build/uk_runtime/graph_evidence.py index 8eaf1b63c..772048c1a 100644 --- a/packages/microcosm-build/src/microcosm/build/uk_runtime/graph_evidence.py +++ b/packages/microcosm-build/src/microcosm/build/uk_runtime/graph_evidence.py @@ -46,6 +46,10 @@ from .cgt_asset_type import CGT_ASSET_TYPE_DOMAIN from .frs_relationships import CHRONICLE_ONS_HOUSEHOLD_TYPE_VALUE_IDS +#: The spine checkpoint's build state the terminal gates read as the spine +#: frame: household weight kind, period and mass log (microcosm#1115). +SPINE_BUILD_STATE_TYPE = ArtifactType("microcosm.uk.spine-build-state", 1) + SPINE_GATE_REPORT_TYPE = ArtifactType("microcosm.gate-phase-report", 1) diff --git a/packages/microcosm-build/src/microcosm/build/uk_runtime/graph_targets.py b/packages/microcosm-build/src/microcosm/build/uk_runtime/graph_targets.py index f52ae4104..6711e80d8 100644 --- a/packages/microcosm-build/src/microcosm/build/uk_runtime/graph_targets.py +++ b/packages/microcosm-build/src/microcosm/build/uk_runtime/graph_targets.py @@ -45,7 +45,7 @@ from microcosm.graph.codecs import SOURCE_CODECS from . import full_measure, full_problem, ladder_targets, ledger_targets, local_doctrine -from .full_measure import resolve_uk_full_measures +from .full_measure import resolve_uk_full_national_problem from .full_problem import build_uk_full_local_problem from .full_targets import CHRONICLE_SOURCE_CODEC, load_chronicle_source_bytes from .geography_ladder import load_uk_oa_ladder, uk_area_region_codes @@ -381,33 +381,22 @@ def run(self, context: KernelContext) -> KernelResult: with tempfile.TemporaryDirectory( prefix="microcosm-uk-full-measures-" ) as scratch: - prepared, restore, rows, metrics, evidence = resolve_uk_full_measures( - frame, - national, - period=int(full["calibration_year"]), - scratch_dir=Path(scratch), - band_edge_registry=registry_from_payload(full["band_edge_registry"]), - blocks=int(context.params["engine_blocks"]), - local_grains=grains, - ) - # Compile national measures while temporary columns exist. The - # immutable pool, rather than that evaluation Frame, goes forward. - national_problem = ( - build_constraint_matrix(prepared, rows.targets, "household") - if len(rows.targets) - else None - ) - if national_problem is not None and national_problem.skipped: - failures = "; ".join( - f"{item.target.key}: {item.reason}" - for item in national_problem.skipped + # The national problem is compiled per engine block and stitched + # in the pool's household order; the whole pool is never + # materialized (microcosm#932 follow-up, 2026-10-05). + national_problem, rows, metrics, evidence = ( + resolve_uk_full_national_problem( + frame, + national, + period=int(full["calibration_year"]), + scratch_dir=Path(scratch), + band_edge_registry=registry_from_payload( + full["band_edge_registry"] + ), + blocks=int(context.params["engine_blocks"]), + local_grains=grains, ) - raise ValueError( - "Selected national constraints failed to compile: " + failures - ) - clean = restore(prepared) - for entity in frame.entities: - pd.testing.assert_frame_equal(clean.table(entity), frame.table(entity)) + ) arrays = { f"metrics_{grain}": metrics[grain].to_numpy(dtype=np.float64) for grain in grains diff --git a/packages/microcosm-build/src/microcosm/build/uk_runtime/graph_terminal.py b/packages/microcosm-build/src/microcosm/build/uk_runtime/graph_terminal.py index 75a4bf45f..dcead4a3b 100644 --- a/packages/microcosm-build/src/microcosm/build/uk_runtime/graph_terminal.py +++ b/packages/microcosm-build/src/microcosm/build/uk_runtime/graph_terminal.py @@ -16,12 +16,14 @@ from dataclasses import asdict, replace from importlib import metadata from pathlib import Path +from types import SimpleNamespace +from typing import Any import numpy as np import pandas as pd from microcosm.calibrate import TargetRegistry -from microcosm.frame import Frame, engine_tables +from microcosm.frame import Frame, MassChangeRecord, WeightKind, engine_tables from microcosm.graph import ( ArtifactInput, ArtifactOutput, @@ -45,8 +47,12 @@ from ..artifact_files import file_artifact from . import geography_ladder, national_frame -from .atomic_area_support import UK_NATIVE_ALIAS_COLUMNS -from .geography_ladder import uk_geography_ladder_gate +from .atomic_area_support import uk_area_code_frames, without_uk_native_alias_columns +from .geography_ladder import ( + UK_EXPORT_AREA_CODE_COLUMNS, + export_area_code_columns, + uk_geography_ladder_gate, +) from .graph_population import context_frame, population_columns, population_slices from .national_frame import ( UK_RELEASE_EXPORT_DROPPED_COLUMNS, @@ -68,7 +74,12 @@ EXPORT_SOURCE_CODEC = "uk-single-year-h5@1" -def _tables(frame: Frame) -> dict[str, pd.DataFrame]: +def _ladder_tables(frame: Frame) -> dict[str, pd.DataFrame]: + """The export tables before the area codes take their consumer names. + + The geography ladder gate reads the ladder names, so it runs on these; + :func:`_tables` is what the artifact carries. + """ tables = engine_tables(frame, weighted_entities=("household",)) renamed = {} for entity in ("person", "benunit", "household"): @@ -81,11 +92,9 @@ def _tables(frame: Frame) -> dict[str, pd.DataFrame]: columns={column: ARTIFACT_CLONE_INDEX_COLUMN} ) # Nation-native aliases of derived layers are NA outside their own nation; - # the single-year artifact carries the ten ladder columns plus the + # the single-year artifact carries the ladder columns plus the # identity-keyed assignment columns, never the aliases. - aliases = [c for c in UK_NATIVE_ALIAS_COLUMNS if c in renamed["household"]] - if aliases: - renamed["household"] = renamed["household"].drop(columns=aliases) + renamed["household"] = without_uk_native_alias_columns(renamed["household"]) # The reviewed export exclusions leave at the same boundary (microcosm#1063 c9). for entity, columns in UK_RELEASE_EXPORT_DROPPED_COLUMNS.items(): dropped = [c for c in columns if c in renamed[entity]] @@ -94,6 +103,58 @@ def _tables(frame: Frame) -> dict[str, pd.DataFrame]: return renamed +def _tables(frame: Frame) -> dict[str, pd.DataFrame]: + tables = _ladder_tables(frame) + # The three area codes leave under the consumers' names (microcosm#1114). + tables["household"] = export_area_code_columns(tables["household"]) + return tables + + +def _area_codes_block(household: pd.DataFrame) -> dict[str, object]: + """The exported area-code columns and the code frames behind them.""" + columns = { + ladder: export + for ladder, export in UK_EXPORT_AREA_CODE_COLUMNS.items() + if export in household.columns + } + if not columns: + return {"columns": {}, "frames": {}} + frames = uk_area_code_frames() + return { + "columns": columns, + "frames": {export: dict(frames[export]) for export in columns.values()}, + } + + +def _area_code_failures( + household: pd.DataFrame, descriptor: Mapping[str, object] +) -> list[str]: + """microcosm#1114: the written table carries the consumer names, filled, and no ladder name.""" + failures = [] + for ladder, export in UK_EXPORT_AREA_CODE_COLUMNS.items(): + if ladder in household.columns: + failures.append( + f"Exported household table carries the ladder area code {ladder!r}; " + f"the artifact name is {export!r}." + ) + declared = dict(descriptor.get("area_codes", {}).get("columns", {})) + for export in declared.values(): + if export not in household.columns: + failures.append( + f"Exported household table lacks the declared area code {export!r}." + ) + continue + values = household[export] + empty = int(values.isna().sum()) + int( + (values.astype(str).str.len() == 0).sum() + ) + if empty: + failures.append( + f"Exported area code {export!r} is empty on {empty} household(s)." + ) + return failures + + def _table_description(table: pd.DataFrame) -> dict[str, object]: descriptor = { "rows": len(table), @@ -137,14 +198,15 @@ def describe_uk_export( ) -> dict[str, object]: """Validate and describe the maintained H5 layout without serializing it.""" validate_uk_national_frame(frame) - tables = _tables(frame) + ladder = _ladder_tables(frame) gate = uk_geography_ladder_gate( - tables["household"], frame.weights_for("household").values + ladder["household"], frame.weights_for("household").values ) if not gate.passed: raise ValueError( "UK export geography integrity failed: " + "; ".join(gate.failures) ) + tables = {**ladder, "household": export_area_code_columns(ladder["household"])} return { "schema_version": 1, "kind": "uk_full_build_export", @@ -156,6 +218,7 @@ def describe_uk_export( ), "bindings": dict(bindings), "geography_integrity": {"passed": gate.passed, "failures": list(gate.failures)}, + "area_codes": _area_codes_block(tables["household"]), "graph_only_metadata": [ "strata", "metadata_other_than_time_period", @@ -226,6 +289,7 @@ def validate_uk_export( for key in ("time_period", "weight_kind", "mass_log", "hash_environment"): if actual[key] != descriptor[key]: failures.append(f"Exported {key} differs from its graph descriptor.") + failures.extend(_area_code_failures(payload["household"], descriptor)) if file_artifact(path) != dataset: raise ValueError("UK exported file changed during graph readback validation.") return { @@ -659,6 +723,15 @@ def run(self, context: KernelContext) -> KernelResult: raise ValueError("UK full gate manifest differs from its declared binding.") stage_evidence, fit_weight_records = _spine_gate_evidence(context) supporting = _source_gate_evidence(context, self.engine) + if "spine_build_state" in context.artifacts: + state = json.loads(context.artifacts["spine_build_state"].payload) + supporting["spine_build_state"] = SimpleNamespace( + household_weight_kind=WeightKind(str(state["household_weight_kind"])), + time_period=str(state["time_period"]), + mass_log=tuple( + MassChangeRecord(**record) for record in state["mass_log"] + ), + ) phase = str(context.params["phase"]) diagnostics = [] support = [] @@ -776,7 +849,7 @@ def run(self, context: KernelContext) -> KernelResult: }, target_registry=target_registry, local_area_support=support_frame, - rotated_holdout=holdout, + rotated_holdout=diagnostics_rotated_holdout(holdout), build={ "build_kind": "uk_full_build", "target_scope": selection["selector"], @@ -843,7 +916,7 @@ def append_uk_full_gate_nodes( from ..gate_battery import _gates_manifest_payload from ..stage_evidence import STAGE_EVIDENCE_TYPE from .full_gates import uk_full_gate_manifest - from .graph_evidence import SPINE_GATE_REPORT_TYPE + from .graph_evidence import SPINE_BUILD_STATE_TYPE, SPINE_GATE_REPORT_TYPE from .graph_targets import TARGET_SELECTION_TYPE, TARGET_SURFACE_TYPE if not engine_identity: @@ -915,6 +988,26 @@ def append_uk_full_gate_nodes( if any(source.name == "uk_input_mass_reference" for source in graph.sources) else () ) + # The spine checkpoint's build state, when the bound checkpoint publishes + # it: the input-coverage gate reads the stages' importance weights and + # mass receipts from it rather than from the calibrated release frame. + build_state: tuple[ArtifactInput, ...] = () + if spine_provenance is not None: + producer = next( + (node for node in graph.nodes if node.id == spine_provenance.producer), + None, + ) + if producer is not None and any( + output.name == "spine_build_state" for output in producer.artifact_outputs + ): + build_state = ( + ArtifactInput( + "spine_build_state", + spine_provenance.producer, + "spine_build_state", + SPINE_BUILD_STATE_TYPE, + ), + ) final = Node( "uk.full.gates.calibrated", UKFullGateKernel.ref, @@ -924,6 +1017,7 @@ def append_uk_full_gate_nodes( params={**params, "phase": "terminal"}, artifact_inputs=( *common, + *build_state, prerequisite, ArtifactInput( "problem", calibration.problem_producer, "problem", PROBLEM_TYPE @@ -1109,6 +1203,23 @@ def uk_full_holdout_node( ) +def diagnostics_rotated_holdout(report: Mapping[str, Any]) -> dict[str, Any]: + """The graph's holdout report as the diagnostics schema declares it. + + The ``uk.full.holdout`` artifact carries the kernel's own binding + (``graph_binding``) and, when the holdout was skipped, the reason. The + diagnostics records forbid undeclared fields and state a skipped holdout + as the bare marker, so the first K=25 build (2026-10-06) reached its + terminal gates after 28 hours and the battery refused the marker with + nineteen validation errors. The holdout artifact keeps the full report; + the diagnostics receive the schema's shape. + """ + + if report.get("skipped") is True: + return {"skipped": True} + return {key: value for key, value in report.items() if key != "graph_binding"} + + def materialize_uk_terminal_artifacts( manifest, store, *, directory: str | Path, stem: str ) -> dict[str, dict[str, object]]: @@ -1326,12 +1437,18 @@ def rowwise_candidate_manifest_from_graph( unenforced_failures = [ line for line in blocking_lines if line[1:].split("]")[0] in unenforced ] + representation = dict(bindings.get("measure_resolution", {})).get( + "engine_population_representation" + ) releasable, release_posture = release_verdict( sample_fraction=args.sample_fraction, engine_blocks=args.engine_blocks, release_blocking_gates_passed=bool( enforcement["release_blocking_gates_passed"] ), + engine_population_exact=bool( + isinstance(representation, Mapping) and representation.get("exact") + ), ) materialization = { str(problem.problem.targets[index].row_name): str(row.get("materialization")) @@ -1411,6 +1528,16 @@ def rowwise_candidate_manifest_from_graph( "binding_adjudications": dict(bindings.get("binding_adjudications", {})), "cross_grain": dict(bindings.get("cross_geography", {})), "ladder_assignment_provenance": dict(ladder_provenance), + # microcosm#1114: the area-code columns the artifact carries and the + # code frames behind them, from the export descriptor. + "area_codes": dict( + ( + _optional_graph_json( + final_manifest, store, "uk.full.export.prepare", "export_descriptor" + ) + or {} + ).get("area_codes", {"columns": {}, "frames": {}}) + ), "household_dispersion": dict(surface.get("household_dispersion", {})), "parameters": rowwise_parameters(args, source_year=source_year), "inputs": { diff --git a/packages/microcosm-build/src/microcosm/build/uk_runtime/measure_simulation.py b/packages/microcosm-build/src/microcosm/build/uk_runtime/measure_simulation.py index 40e624b71..c4ccbfef6 100644 --- a/packages/microcosm-build/src/microcosm/build/uk_runtime/measure_simulation.py +++ b/packages/microcosm-build/src/microcosm/build/uk_runtime/measure_simulation.py @@ -14,6 +14,9 @@ import numpy as np import pandas as pd +from microcosm.build.uk_runtime.atomic_area_support import ( + without_uk_native_alias_columns, +) from microcosm.build.uk_runtime.national_frame import ( load_uk_national_frame, write_uk_national_frame, @@ -35,6 +38,7 @@ exclusion_evaluation_date, ) from microcosm.calibrate import TargetRegistry +from microcosm.frame import Frame, MassChangeRecord, Weights from microcosm.frame.adapters.policyengine_uk import validate_uc_claimant_input _ENTITY_LINK = {"benunit": "person_benunit_id", "household": "person_household_id"} @@ -342,6 +346,78 @@ def compute_uc_paid_diagnostic_masks( return masks +ENGINE_WEIGHT_SCALE_REASON = ( + "engine population representation: household weights scaled from the " + "engine block's mass to the pool's so weight-share formulas (a national " + "total allocated by x*w / sum(x*w)) see the pool's denominator; engine " + "scratch only, discarded after resolution" +) + + +def _engine_scratch_frame( + frame: Any, *, engine_weight_scale: float | None = None +) -> tuple[Any, tuple[str, ...]]: + """The frame a scratch-mode engine loads, and the columns it leaves behind. + + The nation-native alias codes of the derived geography layers + (microcosm#931) are NA outside their own nation by design and are never + engine inputs; the single-year dataset policyengine-uk loads refuses any + NaN column, so they leave here exactly as they leave at the export + boundary. + + ``engine_weight_scale`` multiplies the household weights the engine sees + (and only those: the resolver's own frame keeps the block's true mass). + A per-clone engine block carries a K-th of the pool, so a formula that + allocates a national total by weighted share would hand every block the + whole total; scaled to the pool's mass the block's weighted sums equal + the pool's and the formula is exact for identical clone copies. The + scaling is declared on the engine frame's mass log. + + A frame carrying no alias columns and no scale is returned as is. + """ + + household = frame.table("household") + kept = without_uk_native_alias_columns(household) + scale = None if engine_weight_scale is None else float(engine_weight_scale) + if scale is not None and not (np.isfinite(scale) and scale > 0.0): + raise ValueError("engine weight scale must be a positive finite factor.") + if kept is household and (scale is None or scale == 1.0): + return frame, () + dropped = tuple(c for c in household.columns if c not in kept.columns) + tables = {name: frame.table(name) for name in frame.entities} + tables["household"] = kept + weights = {entity: frame.weights_for(entity) for entity in frame.weighted_entities} + mass_log = frame.mass_log + if scale is not None and scale != 1.0: + if "household" not in weights: + raise ValueError("engine weight scaling needs explicit household weights.") + household_weights = weights["household"] + scaled = Weights( + np.asarray(household_weights.values, dtype=np.float64) * scale, + kind=household_weights.kind, + ) + mass_log = ( + *mass_log, + MassChangeRecord( + entity="household", + old_total=float(household_weights.total), + new_total=float(scaled.total), + declared_factor=scale, + reason=ENGINE_WEIGHT_SCALE_REASON, + ), + ) + weights["household"] = scaled + engine_frame = Frame( + {**tables, **{name: frame.link(name) for name in frame.links}}, + frame.schema, + weights, + frame.strata, + mass_log=mass_log, + metadata=frame.metadata, + ) + return engine_frame, dropped + + class UKMeasureResolver: """B2 measure provider backed by a policyengine-uk Microsimulation.""" @@ -353,8 +429,12 @@ def __init__( year: int, frame: Any, microsimulation_factory: Any | None = None, + engine_weight_scale: float | None = None, ): self.year = int(year) + self._engine_weight_scale = ( + None if engine_weight_scale is None else float(engine_weight_scale) + ) policyengine_uk = _policyengine_uk_module() factory = ( microsimulation_factory @@ -366,14 +446,24 @@ def __init__( raise ValueError("scratch-mode UKMeasureResolver requires a frame.") scratch_dir.mkdir(parents=True, exist_ok=True) source_path = scratch_dir / "simulation-input.h5" - write_uk_national_frame(frame, source_path) + engine_frame, dropped = _engine_scratch_frame( + frame, engine_weight_scale=self._engine_weight_scale + ) + write_uk_national_frame(engine_frame, source_path) mode = "scratch_frame_export" else: + if self._engine_weight_scale is not None: + raise ValueError( + "engine weight scaling needs scratch-mode resolution; a direct " + "H5 source carries its own weights." + ) + dropped = () source_path = Path(simulation_source) mode = "direct_h5" if frame is None: frame, _provenance = load_uk_national_frame(source_path) self.frame = frame + self._engine_scratch_dropped_columns = tuple(dropped) self._factory = factory self._source_path = source_path self._counterfactuals: dict[tuple[str, str | None], Any] = {} @@ -596,6 +686,12 @@ def _counterfactual_simulation(self, zeroed: str, folded: str | None) -> Any: def receipt(self) -> dict[str, Any]: receipt = dict(self._receipt) + if getattr(self, "_engine_weight_scale", None) is not None: + receipt["engine_weight_scale"] = float(self._engine_weight_scale) + if getattr(self, "_engine_scratch_dropped_columns", ()): + receipt["engine_scratch_dropped_columns"] = list( + self._engine_scratch_dropped_columns + ) counterfactual_measures = getattr(self, "_counterfactual_measures", None) if counterfactual_measures: receipt["counterfactual_measures"] = { diff --git a/packages/microcosm-build/src/microcosm/build/uk_runtime/national_frame.py b/packages/microcosm-build/src/microcosm/build/uk_runtime/national_frame.py index d068dc835..c5db4c442 100644 --- a/packages/microcosm-build/src/microcosm/build/uk_runtime/national_frame.py +++ b/packages/microcosm-build/src/microcosm/build/uk_runtime/national_frame.py @@ -481,7 +481,14 @@ def _write_uk_single_year_tables( def uk_release_export_frame(frame: Frame) -> Frame: - """The frame the release boundary writes: reviewed export exclusions dropped.""" + """The frame the release boundary writes: reviewed export exclusions dropped. + + The three area codes leave under the consumers' names (microcosm#1114, + ``geography_ladder.UK_EXPORT_AREA_CODE_COLUMNS``); a national frame + carries none and passes unchanged. + """ + + from .geography_ladder import export_area_code_columns tables: dict[str, pd.DataFrame] = {} changed = False @@ -495,6 +502,11 @@ def uk_release_export_frame(frame: Frame) -> Frame: if dropped: table = table.drop(columns=dropped) changed = True + if entity == "household": + renamed = export_area_code_columns(table) + if renamed is not table: + table = renamed + changed = True tables[entity] = table if not changed: return frame diff --git a/packages/microcosm-build/src/microcosm/build/uk_runtime/rowwise_cli.py b/packages/microcosm-build/src/microcosm/build/uk_runtime/rowwise_cli.py index 28d9f1236..b57cbd489 100644 --- a/packages/microcosm-build/src/microcosm/build/uk_runtime/rowwise_cli.py +++ b/packages/microcosm-build/src/microcosm/build/uk_runtime/rowwise_cli.py @@ -227,6 +227,10 @@ def rowwise_parameters(args: argparse.Namespace, *, source_year: int) -> dict[st "selection_pi_hi": None if args.dataset_households is None else float(args.selection_pi_hi), + "selection_initial_lambda": None + if args.dataset_households is None + or getattr(args, "selection_initial_lambda", None) is None + else float(args.selection_initial_lambda), "baseline_pi_floor": None if args.dataset_households is None else float(args.baseline_pi_floor), @@ -557,8 +561,6 @@ def validate_cli_args(args: argparse.Namespace) -> None: refused.append("--measure-exclusions") if args.skip_holdout: refused.append("--skip-holdout") - if args.engine_blocks > 1: - refused.append("--engine-blocks > 1") if args.sample_fraction != 1.0: refused.append("--sample-fraction != 1.0") if _geography_assignment(args) == "legacy": @@ -798,23 +800,37 @@ def release_verdict( sample_fraction: float, engine_blocks: int, release_blocking_gates_passed: bool, + engine_population_exact: bool | None = None, ) -> tuple[bool, dict[str, bool]]: - """``releasable`` needs the full rung, a single-block engine resolution and - every release-blocking gate passed. - - Per-block engine resolution mis-measures population-normalised formulas - (each block reproduces a national aggregate: the ×K land-value artefact - behind the #736 erratum), so a run resolved in more than one block is - diagnostic-only whatever its gates say. The posture is written beside the - verdict so a reader sees which leg failed. + """``releasable`` needs the full rung, an engine resolution that measures the + pool exactly and every release-blocking gate passed. + + A per-block engine resolution used to be diagnostic-only: population- + normalised formulas (a national total allocated by weighted share) saw one + block's denominator and each block reproduced the whole aggregate, the + ``K`` times land-value artefact behind the #736 erratum. The measures loop + now scales every block's engine weights to the pool and records whether the + blocks are identical copies with identical allocation keys for the named + weight-share formulas (``engine_population_representation.exact``; the + formulas and their inputs are ``UK_WEIGHT_SHARE_FORMULA_INPUTS``); a single + block is exact by construction. ``engine_population_exact`` is that + record; left unset, a multi-block run is treated as inexact. The posture is + written beside the verdict so a reader sees which leg failed. """ + single_block = int(engine_blocks) == 1 posture = { "full_rung": float(sample_fraction) == 1.0, - "single_block_engine": int(engine_blocks) == 1, + "single_block_engine": single_block, + "engine_population_exact": single_block or bool(engine_population_exact), "release_blocking_gates_passed": bool(release_blocking_gates_passed), } - return all(posture.values()), posture + releasable = ( + posture["full_rung"] + and posture["engine_population_exact"] + and posture["release_blocking_gates_passed"] + ) + return releasable, posture _release_verdict = release_verdict diff --git a/packages/microcosm-build/src/microcosm/build/uk_runtime/size_evaluation.py b/packages/microcosm-build/src/microcosm/build/uk_runtime/size_evaluation.py index 12a27f78e..546d44d3e 100644 --- a/packages/microcosm-build/src/microcosm/build/uk_runtime/size_evaluation.py +++ b/packages/microcosm-build/src/microcosm/build/uk_runtime/size_evaluation.py @@ -335,7 +335,13 @@ def add( ) add("epochs", epoch_pass, epoch_values if is_size else epochs, epoch_expected) measure = _mapping(solve.get("measure_resolution")) - add("engine_blocks", measure.get("blocks") == 1, measure.get("blocks"), 1) + representation = _mapping(measure.get("engine_population_representation")) + add( + "engine_blocks", + measure.get("blocks") == 1 or representation.get("exact") is True, + measure.get("blocks"), + "1, or per-block with an exact pool representation", + ) stretch_reference = weights_manifest.get( "stretch_reference", "pool_design" if not is_size else None ) diff --git a/packages/microcosm-build/src/microcosm/build/uk_runtime/terminal_gates.py b/packages/microcosm-build/src/microcosm/build/uk_runtime/terminal_gates.py index 46e0deb37..04b8cf1f7 100644 --- a/packages/microcosm-build/src/microcosm/build/uk_runtime/terminal_gates.py +++ b/packages/microcosm-build/src/microcosm/build/uk_runtime/terminal_gates.py @@ -184,7 +184,11 @@ def __post_init__(self) -> None: # #1063: the residential split's arm flag and index. "household.household_is_cgt_residential_clone", "household.cgt_residential_clone_index", - "household.clone_index", + "household.household_clone_index", + # microcosm#1114: the three area codes leave every artifact under the + # consumers' names (geography_ladder.UK_EXPORT_AREA_CODE_COLUMNS); the + # ladder names never reach an artifact and uk_export_candidate_columns + # translates them before this list is compared. "household.constituency_code_oa", "household.consumer_debt", # microcosm#932: the five nation-native aliases of the atomic assignment @@ -331,6 +335,12 @@ def __post_init__(self) -> None: "person.person_support_channel", "person.person_support_clone_index", "person.srp_regular_code5", + "benunit.benunit_clone_index", + "person.person_clone_index", + "household.itl1_code", + "household.itl2_code", + "household.itl3_code", + "household.ward_code", ) UK_KNOWN_MISSING_REFERENCE_EXPORT_COLUMNS: tuple[str, ...] = ( @@ -384,12 +394,20 @@ def uk_export_candidate_columns(frame: Any) -> set[str]: does. """ + from microcosm.build.uk_runtime.geography_ladder import ( + UK_EXPORT_AREA_CODE_COLUMNS, + ) + columns: set[str] = set() for entity in frame.entities: structural = _STRUCTURAL_COLUMNS.get(str(entity), frozenset()) for column in frame.table(entity).columns: if column in structural: continue + # microcosm#1114: the frame holds the ladder names; the artifact + # carries the consumers' names, which is what the surface compares. + if str(entity) == "household": + column = UK_EXPORT_AREA_CODE_COLUMNS.get(str(column), column) columns.add(f"{entity}.{column}") columns.add(f"household.{_WEIGHT_COLUMN}") return columns @@ -457,6 +475,7 @@ def uk_degenerate_release_surface_gate( *, reviewed_exclusions: Mapping[str, UKReviewedExclusion] | None = None, now: date | None = None, + dropped_at_export: Mapping[str, Iterable[str]] | None = None, ) -> GateResult: """Reject every all-null, all-zero, or constant nonstructural column. @@ -471,6 +490,14 @@ def uk_degenerate_release_surface_gate( exclusions = coerce_reviewed_exclusions( reviewed_exclusions, label="UK degenerate-surface" ) + # Columns the release boundary drops before writing: a gate fed the + # pre-export frame skips them, exactly as the certifier never sees them + # on the exported H5; the skipped names are recorded in the details. + dropped = { + entity: frozenset(str(column) for column in columns) + for entity, columns in (dropped_at_export or {}).items() + } + skipped_at_export: list[str] = [] present: set[str] = set() live: dict[str, dict[str, object]] = {} excluded: dict[str, dict[str, object]] = {} @@ -481,6 +508,9 @@ def uk_degenerate_release_surface_gate( for column in table.columns: if column in structural: continue + if column in dropped.get(entity, frozenset()): + skipped_at_export.append(f"{entity}.{column}") + continue checked += 1 name = f"{entity}.{column}" present.add(name) @@ -580,6 +610,7 @@ def uk_degenerate_release_surface_gate( failures=tuple(failures), details={ "columns_checked": checked, + "dropped_at_export": sorted(skipped_at_export), "findings": dict(sorted(live.items())), "all_null_columns": by_kind["all_null"], "all_zero_columns": by_kind["all_zero"], diff --git a/packages/microcosm-build/tests/engine/uk/test_uk_engine_population_representation.py b/packages/microcosm-build/tests/engine/uk/test_uk_engine_population_representation.py new file mode 100644 index 000000000..00bc6e228 --- /dev/null +++ b/packages/microcosm-build/tests/engine/uk/test_uk_engine_population_representation.py @@ -0,0 +1,110 @@ +"""A per-clone engine block scaled to the pool measures weight-share formulas exactly. + +policyengine-uk allocates the ONS aggregate corporate land value across households by +``corporate_sector_wealth * weight / sum(corporate_sector_wealth * weight)``. One engine over +a block that carries a K-th of the pool's mass hands that block the whole national total (the +K-times artefact behind the #736 erratum); the same block with its engine weights scaled by K +reproduces the pool's values. A formula without a weighted denominator is untouched. +""" + +from __future__ import annotations + +import numpy as np +import pandas as pd +import pytest + +from microcosm.build.uk_runtime.measure_simulation import UKMeasureResolver +from microcosm.build.uk_runtime.national_frame import uk_national_frame +from microcosm.frame import WeightKind + + +def _frame(weights): + person = pd.DataFrame( + { + "person_id": [1, 2, 3, 4], + "person_benunit_id": [10, 10, 20, 20], + "person_household_id": [100, 100, 200, 200], + "age": [35, 8, 40, 38], + "is_benunit_head": [True, False, True, False], + "is_parent": [True, False, False, False], + "is_uc_claimant": [False, False, False, False], + } + ) + benunit = pd.DataFrame( + { + "benunit_id": [10, 20], + "benunit_source_id": [10, 20], + "dependent_children": [1, 0], + "is_married": [False, False], + } + ) + household = pd.DataFrame( + { + "household_id": [100, 200], + "household_source_id": [100, 200], + "region": ["LONDON", "WALES"], + "council_tax": [1800.0, 900.0], + "rent": [0.0, 6000.0], + "tenure_type": ["OWNED_OUTRIGHT", "RENT_PRIVATELY"], + "corporate_wealth": [120_000.0, 15_000.0], + "private_pension_wealth": [80_000.0, 0.0], + "property_wealth": [650_000.0, 0.0], + } + ) + return uk_national_frame( + person=person, + benunit=benunit, + household=household, + time_period=2024, + household_weights=np.asarray(weights, dtype=float), + weight_kind=WeightKind.IMPORTANCE, + ) + + +def _values(resolver, variable): + return np.asarray(resolver.simulation.calculate(variable, 2025), dtype=np.float64) + + +def test_block_engine_scaled_to_the_pool_reproduces_the_pool_land_value(tmp_path): + pytest.importorskip("tables") + pool = _frame([1000.0, 500.0]) + # one of four identical clones: the same households at a quarter of the mass + block = _frame([250.0, 125.0]) + + pool_resolver = UKMeasureResolver( + simulation_source=None, frame=pool, year=2025, scratch_dir=tmp_path / "pool" + ) + naive_resolver = UKMeasureResolver( + simulation_source=None, frame=block, year=2025, scratch_dir=tmp_path / "naive" + ) + scaled_resolver = UKMeasureResolver( + simulation_source=None, + frame=block, + year=2025, + scratch_dir=tmp_path / "scaled", + engine_weight_scale=4.0, + ) + + pool_value = _values(pool_resolver, "corporate_land_value") + assert pool_value.sum() > 0.0 + # the artefact: the block alone reproduces the whole national aggregate + np.testing.assert_allclose( + _values(naive_resolver, "corporate_land_value"), 4.0 * pool_value, rtol=1e-6 + ) + # the representation: scaled to the pool, the block measures the pool's values + np.testing.assert_allclose( + _values(scaled_resolver, "corporate_land_value"), pool_value, rtol=1e-6 + ) + np.testing.assert_allclose( + _values(scaled_resolver, "land_value"), + _values(pool_resolver, "land_value"), + rtol=1e-6, + ) + # a formula without a weighted denominator never depended on the block's mass + np.testing.assert_allclose( + _values(naive_resolver, "household_land_value"), + _values(pool_resolver, "household_land_value"), + rtol=1e-9, + ) + assert scaled_resolver.receipt()["engine_weight_scale"] == 4.0 + assert "engine_weight_scale" not in naive_resolver.receipt() diff --git a/packages/microcosm-build/tests/engine/uk/test_uk_full_target_graph.py b/packages/microcosm-build/tests/engine/uk/test_uk_full_target_graph.py index a44c902ec..6d0c740f7 100644 --- a/packages/microcosm-build/tests/engine/uk/test_uk_full_target_graph.py +++ b/packages/microcosm-build/tests/engine/uk/test_uk_full_target_graph.py @@ -114,7 +114,8 @@ def test_default_all_has_direct_matrix_and_solver_parity_and_replays( source_year=2023, expected_constituency_vintage="2024_pcon", ) - prepared_frame, _, national_rows, metrics, _ = target_inputs["measures"]( + prepared_frame = target_inputs["prepared"](assignment.frame) + _, national_rows, metrics, _ = target_inputs["measures"]( assignment.frame, target_inputs["national"], local_grains=("constituency", "la") ) surface, cross = target_inputs["surface"]() diff --git a/packages/microcosm-build/tests/engine/uk/test_uk_weight_share_formula_inputs.py b/packages/microcosm-build/tests/engine/uk/test_uk_weight_share_formula_inputs.py new file mode 100644 index 000000000..fe9f62702 --- /dev/null +++ b/packages/microcosm-build/tests/engine/uk/test_uk_weight_share_formula_inputs.py @@ -0,0 +1,71 @@ +"""``UK_WEIGHT_SHARE_FORMULA_INPUTS`` mirrors policyengine-uk's allocation keys. + +microcosm#1115 review round 2: the constant names, by hand, the frame columns +each weight-share formula's allocation key reads. This test reads the same +dependencies from the installed engine, so an engine bump that changes a key's +``adds`` list, or drops a formula, fails here instead of letting the +per-block exactness check go stale. +""" + +from __future__ import annotations + +import pytest + +from microcosm.build.uk_runtime.full_measure import UK_WEIGHT_SHARE_FORMULA_INPUTS + +#: The allocation-key variable each weight-share formula divides by, as the +#: engine's formula reads it (``x * w / sum(x * w)``). +ALLOCATION_KEYS = { + "corporate_land_value": "corporate_sector_wealth", + "shareholding": "corporate_sector_wealth", + "consumption_shareholding": "consumption", +} + + +@pytest.fixture(scope="module") +def system(): + from policyengine_uk import CountryTaxBenefitSystem + + return CountryTaxBenefitSystem() + + +def test_constant_names_every_weight_share_formula_and_its_engine_inputs(system): + assert set(UK_WEIGHT_SHARE_FORMULA_INPUTS) == set(ALLOCATION_KEYS) + for formula, key in ALLOCATION_KEYS.items(): + variable = system.variables[formula] + assert variable.entity.key == "household", formula + assert not variable.is_input_variable(), f"{formula} is no longer a formula" + key_variable = system.variables[key] + adds = tuple(key_variable.adds) + assert adds, f"{key} no longer declares its inputs through `adds`" + assert adds == UK_WEIGHT_SHARE_FORMULA_INPUTS[formula], ( + f"{formula}: the engine's {key} adds {adds}; the constant names " + f"{UK_WEIGHT_SHARE_FORMULA_INPUTS[formula]}" + ) + for column in adds: + leaf = system.variables[column] + assert leaf.entity.key == "household", column + assert leaf.is_input_variable(), f"{column} is not an input column" + + +def test_named_columns_are_release_surface_household_columns(): + """The columns the check reads leave with the release: every one is an + enhanced-FRS household input (the parity reference) or a reviewed extra + household column of the export surface (the allow-list).""" + from microcosm.build.uk_runtime import load_efrs_parity_reference + from microcosm.build.uk_runtime.terminal_gates import ( + UK_ALLOWED_EXTRA_EXPORT_COLUMNS, + ) + + reference = load_efrs_parity_reference().input_entities + household = { + column for column, entity in reference.items() if entity == "household" + } + household |= { + name.split(".", 1)[1] + for name in UK_ALLOWED_EXTRA_EXPORT_COLUMNS + if name.startswith("household.") + } + for formula, columns in UK_WEIGHT_SHARE_FORMULA_INPUTS.items(): + missing = [c for c in columns if c not in household] + assert not missing, f"{formula}: {missing} are not release household columns" diff --git a/packages/microcosm-build/tests/engine_free/shared/test_spec_engine_loader.py b/packages/microcosm-build/tests/engine_free/shared/test_spec_engine_loader.py index 718f4ff56..17cd0b715 100644 --- a/packages/microcosm-build/tests/engine_free/shared/test_spec_engine_loader.py +++ b/packages/microcosm-build/tests/engine_free/shared/test_spec_engine_loader.py @@ -12,7 +12,7 @@ def test_semantic_hash_has_golden_vector_and_surface_separation(tmp_path) -> Non # Pin the domain separator, normalization rules, schema-set receipt, and # exact normative projection as one reviewable golden vector. assert first.spec_sha256 == ( - "6273f82f0e4899e97d1bede4db0662829b3df6f420dba32c0e67018d010c2c91" + "69de1ffc232c3a4da09925658c9d105af55bdaae90852824ac6283306932aab9" ) second_root = _rich_minimal(tmp_path / "xy", note="second", store="local:b") diff --git a/packages/microcosm-build/tests/engine_free/uk/test_uk_atomic_area_support.py b/packages/microcosm-build/tests/engine_free/uk/test_uk_atomic_area_support.py index 3e4f01f05..9029e31fe 100644 --- a/packages/microcosm-build/tests/engine_free/uk/test_uk_atomic_area_support.py +++ b/packages/microcosm-build/tests/engine_free/uk/test_uk_atomic_area_support.py @@ -403,3 +403,32 @@ def test_duplicate_or_unknown_household_identity_is_refused_by_shared_assignment households.loc[1, "region"] = "UNKNOWN" with pytest.raises(ValueError, match="exactly one"): geo.assign_atomic(households, spec, supports) + + +def test_area_code_frames_come_from_the_committed_support_provenance(): + """microcosm#1114: the frames a release records are the supports' declared vintages.""" + import json + from importlib.resources import files + + from microcosm.build.uk_runtime.atomic_area_support import ( + PROVENANCE_RESOURCE, + SYSTEMS, + uk_area_code_frames, + ) + from microcosm.build.uk_runtime.geography_ladder import ( + UK_EXPORT_AREA_CODE_COLUMNS, + ) + + frames = uk_area_code_frames() + assert set(frames) == set(UK_EXPORT_AREA_CODE_COLUMNS.values()) + provenance = json.loads( + files("microcosm.build.uk").joinpath(PROVENANCE_RESOURCE).read_text() + ) + for ladder, export in UK_EXPORT_AREA_CODE_COLUMNS.items(): + assert set(frames[export]) == set(SYSTEMS) + for system in SYSTEMS: + declared = provenance["supports"][system]["column_metadata"][ladder] + assert frames[export][system] == declared["vintage"] + assert frames["constituency_code_oa"] == dict.fromkeys(SYSTEMS, "2024_pcon") + assert frames["la_code_oa"]["uk_ew_output_area_2021"] == "2023_april_lad" + assert frames["region_code_oa"]["uk_ew_output_area_2021"] == "2024_rgn" diff --git a/packages/microcosm-build/tests/engine_free/uk/test_uk_atomic_household_lineage.py b/packages/microcosm-build/tests/engine_free/uk/test_uk_atomic_household_lineage.py index 5180b7cd3..8e07d41a4 100644 --- a/packages/microcosm-build/tests/engine_free/uk/test_uk_atomic_household_lineage.py +++ b/packages/microcosm-build/tests/engine_free/uk/test_uk_atomic_household_lineage.py @@ -581,3 +581,41 @@ def test_actual_expand_filter_receipts_feed_shared_atomic_graph_and_replay(tmp_p ) + "\n" ) + + +def _invented_key(path, source_household_id=1): + return household_draw_key( + source="invented-frs", + source_vintage="invented-2024-25", + source_household_id=source_household_id, + clone_path=path, + ) + + +def test_residential_split_branch_follows_the_incidence_clone(): + # microcosm#1063 item 7: a residential arm is an EXPAND of a gaining clone + # after the incidence clone and before the geographic pool. + a, b, c, s, _ = chain() + arm = expansion("cgt_residential_split", s.after_ids, ((51, 21),)) + pool = expansion("geographic_support", arm.after_ids, ((101, 1), (151, 51))) + actual = by_id(project(steps=(a, b, c, s, arm, pool), ids=pool.after_ids)) + assert actual[51] == _invented_key( + (("cgt_incidence_clone", 1), ("cgt_residential_split", 1)) + ) + assert actual[151] == _invented_key( + ( + ("cgt_incidence_clone", 1), + ("cgt_residential_split", 1), + ("geographic_support", 1), + ) + ) + assert actual[51] != actual[21] + assert actual.is_unique + + +def test_residential_split_before_the_incidence_clone_is_refused(): + a, b, c, _, _ = chain() + early = expansion("cgt_residential_split", b.after_ids, ((51, 11),)) + late_clone = expansion("cgt_incidence_clone", early.after_ids, ((21, 1), (31, 11))) + with pytest.raises(ValueError, match="repeated or reordered structural branch"): + project(steps=(a, b, early, late_clone), ids=late_clone.after_ids) diff --git a/packages/microcosm-build/tests/engine_free/uk/test_uk_atomic_identity_graph.py b/packages/microcosm-build/tests/engine_free/uk/test_uk_atomic_identity_graph.py index d9851d9ae..568bd9245 100644 --- a/packages/microcosm-build/tests/engine_free/uk/test_uk_atomic_identity_graph.py +++ b/packages/microcosm-build/tests/engine_free/uk/test_uk_atomic_identity_graph.py @@ -71,7 +71,11 @@ ] -def lineage_frame(rows=LINEAGE): +def _residential_index(row): + return int(row[7]) if len(row) > 7 else 0 + + +def lineage_frame(rows=LINEAGE, mutate=None): ids = np.asarray([r[0] for r in rows], dtype=np.int64) household = pd.DataFrame( { @@ -87,11 +91,18 @@ def lineage_frame(rows=LINEAGE): "household_is_capital_gains_clone": np.asarray([r[4] for r in rows]), "household_is_cgt_support_copy": np.asarray([r[5] != 0 for r in rows]), "cgt_support_copy_index": np.asarray([r[5] for r in rows], dtype=np.int64), - # microcosm#1063: a residential arm index (0 on every row here). - "household_is_cgt_residential_clone": np.zeros(len(rows), dtype=bool), - "cgt_residential_clone_index": np.zeros(len(rows), dtype=np.int64), + # microcosm#1063: the residential arm index of a gaining clone, an + # optional eighth field (0 on every seven-field row). + "household_is_cgt_residential_clone": np.asarray( + [_residential_index(r) != 0 for r in rows] + ), + "cgt_residential_clone_index": np.asarray( + [_residential_index(r) for r in rows], dtype=np.int64 + ), } ) + if mutate is not None: + mutate(household) person = pd.DataFrame( { "person_id": ids * 10, @@ -165,7 +176,15 @@ def run_identity(frame, tmp_path, *, k=2, endpoint="uk.full.identity"): return manifest, graph -def expected_key(source_id, channel, support_index, cgt, copy_index, clone_index): +def expected_key( + source_id, + channel, + support_index, + cgt, + copy_index, + clone_index, + residential_index=0, +): path = [] if channel == "spi": path.append(("spi_support_channel", int(support_index))) @@ -173,6 +192,8 @@ def expected_key(source_id, channel, support_index, cgt, copy_index, clone_index path.append(("cgt_support_split", int(copy_index))) if cgt: path.append(("cgt_incidence_clone", 1)) + if residential_index: + path.append(("cgt_residential_split", int(residential_index))) if clone_index: path.append(("geographic_support", int(clone_index))) return household_draw_key( @@ -331,3 +352,74 @@ def test_identity_inputs_are_the_declared_lineage_columns(): "cgt_residential_clone_index", CLONE, ) + + +# microcosm#1063 item 7: the residential split clones a gaining household into +# residential arms; the branch sits between the incidence clone and the +# geographic pool. The main-head spine of 2026-10-04 carried 2,735 such arms and +# the dense build refused every one of them as an unknown structural branch. +RESIDENTIAL_ROWS = sorted( + [ + *LINEAGE, + (1101, 2, "frs", 0, True, 1, "SCOTLAND"), # gaining clone of copy 1001 + (201, 1, "frs", 0, True, 0, "LONDON", 1), # arm 1 of gaining clone 101 + (211, 1, "spi", 1, True, 0, "LONDON", 1), # arm 1 of gaining clone 111 + (1201, 2, "frs", 0, True, 1, "SCOTLAND", 1), # arm 1 of gaining clone 1101 + ], + key=lambda row: row[0], # the frame requires ascending entity ids +) + + +@pytest.mark.parametrize("k", [1, 2]) +def test_residential_arm_households_key_their_declared_branch(tmp_path, k): + manifest, _ = run_identity(lineage_frame(RESIDENTIAL_ROWS), tmp_path, k=k) + household = manifest.population("uk.full.expand").table("household") + assert len(household) == k * len(RESIDENTIAL_ROWS) + assert household[IDENTITY_COLUMN].is_unique + # Every pool copy keeps its spine household's lineage columns, so the + # expected path is read from them, never from the remapped household id. + for ( + source_id, + channel, + support_index, + cgt, + copy_index, + arm, + clone_index, + key, + ) in zip( + household["source_household_id"].to_numpy(), + household["household_support_channel"].to_numpy(), + household["household_support_clone_index"].to_numpy(), + household["household_is_capital_gains_clone"].to_numpy(), + household["cgt_support_copy_index"].to_numpy(), + household["cgt_residential_clone_index"].to_numpy(), + household[CLONE].to_numpy(), + household[IDENTITY_COLUMN].to_numpy(), + strict=True, + ): + assert key == expected_key( + int(source_id), + str(channel), + int(support_index), + bool(cgt), + int(copy_index), + int(clone_index), + int(arm), + ) + arms = household[household["cgt_residential_clone_index"] != 0] + assert len(arms) == 3 * k + assert "cgt_residential_split" in arms[IDENTITY_COLUMN].iloc[0] + + +def test_residential_arm_flag_and_index_must_agree(tmp_path): + def clear_flag(household): + household.loc[ + household["household_id"] == 201, "household_is_cgt_residential_clone" + ] = False + + frame = lineage_frame(RESIDENTIAL_ROWS, mutate=clear_flag) + with pytest.raises( + NodeRejectedError, match="residential-arm flag and arm index disagree" + ): + run_identity(frame, tmp_path, k=1) diff --git a/packages/microcosm-build/tests/engine_free/uk/test_uk_battery_bindings.py b/packages/microcosm-build/tests/engine_free/uk/test_uk_battery_bindings.py index f4daf66f6..51de977b4 100644 --- a/packages/microcosm-build/tests/engine_free/uk/test_uk_battery_bindings.py +++ b/packages/microcosm-build/tests/engine_free/uk/test_uk_battery_bindings.py @@ -1057,18 +1057,24 @@ def test_terminal_mode_threads_frame_engine_and_manifest(self) -> None: family_coverage={}, ) binding = UK_GATE_REGISTRY["release_input_coverage"] + # The binding needs the spine build state (microcosm#1115 review); + # here the spine frame is the frame itself. result = binding.evaluate( EvidenceContext( frame=frame, artifacts={ "coverage_engine": engine, "coverage_manifest": manifest, + "spine_frame": frame, }, ), {}, ) direct = uk_release_input_coverage_gate( - _uk_gate_surface(frame), engine, manifest=manifest + _uk_gate_surface(frame), + engine, + manifest=manifest, + build_state_frame=_uk_gate_surface(frame), ) assert direct.name == "uk_release_input_coverage" diff --git a/packages/microcosm-build/tests/engine_free/uk/test_uk_full_build_cli.py b/packages/microcosm-build/tests/engine_free/uk/test_uk_full_build_cli.py index 96a19b1e0..a53eb8161 100644 --- a/packages/microcosm-build/tests/engine_free/uk/test_uk_full_build_cli.py +++ b/packages/microcosm-build/tests/engine_free/uk/test_uk_full_build_cli.py @@ -3,6 +3,7 @@ # ruff: noqa: F403, F405 from types import SimpleNamespace +import test_support.microcosm_build.uk_full_build_cli as support # noqa: E402 from test_support.microcosm_build.uk_full_build_cli import * @@ -1243,11 +1244,12 @@ def test_blocked_gates_partition_failures_by_criticality(tmp_path, monkeypatch, assert spool_rows(out)[0].disposition == "failed" -def test_multi_block_engine_run_is_never_releasable(tmp_path, monkeypatch): - """End to end: ``--engine-blocks K`` on f100 writes ``releasable: false``. - - Every release-blocking gate passes here; the posture alone withholds the - verdict, and the manifest names the leg (``single_block_engine``). +def test_multi_block_engine_run_is_releasable_when_its_blocks_represent_the_pool( + tmp_path, monkeypatch +): + """End to end: ``--engine-blocks K`` on f100 with every block an identical + copy scaled to the pool writes ``releasable: true``; the manifest names the + legs (``single_block_engine`` false, ``engine_population_exact`` true). """ pytest.importorskip("tables") status, out = run_dense_main( @@ -1257,10 +1259,33 @@ def test_multi_block_engine_run_is_never_releasable(tmp_path, monkeypatch): manifest = json.loads((out / cli.MANIFEST_FILENAME).read_text()) assert manifest["parameters"]["engine_blocks"] == 2 assert manifest["blocking_failures"] == [] - assert manifest["releasable"] is False + assert manifest["releasable"] is True posture = manifest["release_posture"] assert posture["full_rung"] is True assert posture["single_block_engine"] is False + assert posture["engine_population_exact"] is True + assert posture["release_blocking_gates_passed"] is True + + +def test_multi_block_engine_run_without_an_exact_representation_is_never_releasable( + tmp_path, monkeypatch +): + """The blocks were not identical copies: the scaling is approximate, the + measures receipt keeps its caveat, and the posture alone withholds the + verdict with every release-blocking gate passed (#736 erratum). + """ + pytest.importorskip("tables") + monkeypatch.setattr(support, "ENGINE_POPULATION_EXACT", False) + status, out = run_dense_main( + tmp_path, monkeypatch, "--n-clones", "2", "--engine-blocks", "2" + ) + assert status == 0 + manifest = json.loads((out / cli.MANIFEST_FILENAME).read_text()) + assert manifest["blocking_failures"] == [] + assert manifest["releasable"] is False + posture = manifest["release_posture"] + assert posture["single_block_engine"] is False + assert posture["engine_population_exact"] is False assert posture["release_blocking_gates_passed"] is True diff --git a/packages/microcosm-build/tests/engine_free/uk/test_uk_full_calibration_graph.py b/packages/microcosm-build/tests/engine_free/uk/test_uk_full_calibration_graph.py index 38b9e4f13..75ecec963 100644 --- a/packages/microcosm-build/tests/engine_free/uk/test_uk_full_calibration_graph.py +++ b/packages/microcosm-build/tests/engine_free/uk/test_uk_full_calibration_graph.py @@ -508,3 +508,51 @@ def test_size_kernels_forward_phased_epochs_to_the_registered_observer( replay.population(endpoints.population).weights_for("household").values, first.population(endpoints.population).weights_for("household").values, ) + + +def test_select_warm_starts_the_search_and_refuses_bad_hints(monkeypatch): + """microcosm#1115: ``initial_lambda`` reaches the shared search as its + warm-start penalty, is recorded on the selection and refused when invalid.""" + from microcosm.build.uk_runtime.graph_calibration import UKGraphCalibrationConfig + + frame = source_frame() + dense = calibrate(frame, targets(), epochs=2) + seen = {} + real = dataset_size.calibrate + + def spy(*args, **kwargs): + seen.update(kwargs) + return real(*args, **kwargs) + + monkeypatch.setattr(dataset_size, "calibrate", spy) + options = dict(households=2, epochs=2, learning_rate=0.02, seed=7) + cold = dataset_size.select_uk_dataset_size(frame, dense, **options) + assert seen["l0_lambda"] == 0.0 + assert cold.initial_lambda is None + assert cold.selection.options["budget_search"]["initial_lambda"] is None + warm = dataset_size.select_uk_dataset_size( + frame, dense, initial_lambda=1e-3, **options + ) + assert seen["l0_lambda"] == 1e-3 + assert warm.initial_lambda == 1e-3 + assert warm.selection.options["budget_search"]["initial_lambda"] == 1e-3 + assert warm.selection.options["budget_search"]["probes"][0]["l0_lambda"] == 1e-3 + for bad in (0.0, -1e-3, float("nan"), float("inf"), True): + with pytest.raises(ValueError, match="initial_lambda"): + dataset_size.select_uk_dataset_size( + frame, dense, initial_lambda=bad, **options + ) + with pytest.raises(ValueError, match="initial_lambda"): + UKGraphCalibrationConfig(selection_initial_lambda=bad) + assert ( + UKGraphCalibrationConfig(selection_initial_lambda=2e-6).selection_initial_lambda + == 2e-6 + ) + # The refit reuses a warm-searched selection like any other. + draw = dataset_size.draw_uk_dataset_size( + frame, dense, selection=warm, households=2, seed=7 + ) + refit = dataset_size.refit_uk_dataset_size( + frame, dense, selection=warm, draw=draw, initial_lambda=1e-3, **options + ) + assert refit.receipt["selection_l0_lambda"] == warm.selection.l0_lambda diff --git a/packages/microcosm-build/tests/engine_free/uk/test_uk_full_gates.py b/packages/microcosm-build/tests/engine_free/uk/test_uk_full_gates.py index c11c462f0..5a4ec666c 100644 --- a/packages/microcosm-build/tests/engine_free/uk/test_uk_full_gates.py +++ b/packages/microcosm-build/tests/engine_free/uk/test_uk_full_gates.py @@ -131,6 +131,11 @@ def evidence(monkeypatch): "uk_aggregate_admin_totals", lambda frame, manifest: ({"admin": 2.0}, [{"measured": 2.0}]), ) + monkeypatch.setattr( + runtime, + "uk_cgt_projection_artifact", + lambda frame, manifest: {"stub_projection": True}, + ) return ( frame, ordered, @@ -315,3 +320,46 @@ def test_missing_evidence_uses_declared_development_policy(): )["artifact_permitted"] is False ) + + +def test_gate_context_hands_the_spine_build_state_to_the_coverage_gate(evidence): + """The checkpoint's build state is the battery's ``spine_frame``: the + input-coverage gate reads the stages' importance weights from it, never + from the calibrated release frame; the raw key does not leak through.""" + frame, ordered, solution, supporting = evidence + state = SimpleNamespace( + household_weight_kind="importance", time_period="2024", mass_log=() + ) + context = runtime.build_full_gate_context( + frame, + ordered_problem=ordered, + solution=solution, + selection_receipt=_selection(), + stage_evidence={"frs_spine": {}}, + supporting_evidence={**supporting, "spine_build_state": state}, + ) + assert context.artifacts["spine_frame"] is state + assert "spine_build_state" not in context.artifacts + # without a published build state the certifier's own spine frame (if any) stands + plain = _context(evidence) + assert "spine_frame" not in plain.artifacts + + +def test_gate_context_computes_the_cgt_projection_for_the_entrants_fence(evidence): + from microcosm.build.uk_runtime.cgt_projection import UK_CGT_PROJECTION_ARTIFACT_KEY + + context = _context(evidence) + assert context.artifacts[UK_CGT_PROJECTION_ARTIFACT_KEY] == { + "stub_projection": True + } + # a projection the caller already supplies is kept as given + frame, ordered, solution, supporting = evidence + supplied = runtime.build_full_gate_context( + frame, + ordered_problem=ordered, + solution=solution, + selection_receipt=_selection(), + stage_evidence={"frs_spine": {}}, + supporting_evidence={**supporting, UK_CGT_PROJECTION_ARTIFACT_KEY: "given"}, + ) + assert supplied.artifacts[UK_CGT_PROJECTION_ARTIFACT_KEY] == "given" diff --git a/packages/microcosm-build/tests/engine_free/uk/test_uk_full_measure.py b/packages/microcosm-build/tests/engine_free/uk/test_uk_full_measure.py index d6abd74ff..9be6c4204 100644 --- a/packages/microcosm-build/tests/engine_free/uk/test_uk_full_measure.py +++ b/packages/microcosm-build/tests/engine_free/uk/test_uk_full_measure.py @@ -102,7 +102,28 @@ def test_full_measure_resolves_real_per_clone_blocks( second_cgt_period, toy_ladder, ) -> None: - frame = source_frame() + # The weight-share allocation keys ride into every clone, so the + # representation can verify them (microcosm#1115 review). + from microcosm.build.uk_runtime.full_measure import ( + UK_WEIGHT_SHARE_FORMULA_INPUTS, + ) + from microcosm.frame import Frame + + original = source_frame() + tables = {e: original.table(e).copy() for e in original.entities} + for inputs in UK_WEIGHT_SHARE_FORMULA_INPUTS.values(): + for column in inputs: + tables["household"][column] = ( + tables["household"]["household_id"].astype("float64") * 10.0 + ) + frame = Frame( + tables, + original.schema, + {"household": original.weights_for("household")}, + original.strata, + mass_log=original.mass_log, + metadata=original.metadata, + ) ladder, _ = toy_ladder clone = clone_uk_dataset_with_ladder_geography( frame, @@ -165,6 +186,17 @@ def resolve(): prepared, restore, _, metrics, receipt = resolve() assert len(constructions) == 2 + # Each block carries half the pool, so its engine weights are scaled by two + # and the receipt records an exact (identical-copy) representation. + assert [call["engine_weight_scale"] for call in constructions] == pytest.approx( + [2.0, 2.0] + ) + representation = receipt["engine_population_representation"] + assert representation["mode"] == "block_weights_scaled_to_pool" + assert representation["exact"] is True + assert representation["checks"]["identical_source_weight_multisets"] is True + assert receipt["block_sensitivity"]["exact"] is True + assert receipt["block_sensitivity"]["caveat"] is None assert [len(call["frame"].table("household")) for call in constructions] == [ frame.n("household"), frame.n("household"), @@ -185,7 +217,11 @@ def resolve(): assert set(sensitivity["present_in_this_run"]) <= set( sensitivity["known_population_normalised_measures"] ) - assert "not evidence for adjudication" in sensitivity["caveat"] + # Identical clone copies scaled to the pool: the known formulas see the + # pool's denominator, so the receipt names the mitigation and no caveat. + assert "scaled by pool mass / block mass" in sensitivity["mitigation"] + assert sensitivity["exact"] is True + assert sensitivity["caveat"] is None assert ( metrics["constituency"].index.tolist() == clone.frame.table("household")["household_id"].tolist() @@ -335,3 +371,299 @@ def receipt(self): provider = receipts[0]["resolution"][0]["provider"] assert provider["source_path"] == "simulation-input.h5" assert str(tmp_path) not in json.dumps(receipts[0]) + + +@pytest.mark.parametrize("blocks", [1, 2]) +def test_each_engine_block_is_collected_before_the_next_loads( + monkeypatch, tmp_path, toy_ladder, blocks +): + """A policyengine simulation is a large cyclic object graph that `del` + alone does not reclaim; the first K=25 build (2026-10-05) retained all 25 + per-clone engines until the machine killed it. The loop collects after + every block, including a single-block run.""" + frame = source_frame() + if blocks == 2: + frame = clone_uk_dataset_with_ladder_geography( + frame, + toy_ladder[0], + n_clones=2, + seed=7, + source_year=2023, + expected_constituency_vintage="2024_pcon", + ).frame + registry = TargetRegistry([], country="uk") + + class Resolver: + def __init__(self, *, frame, **kwargs): + self.frame = frame + self.simulation = object() + + def receipt(self): + return {"mode": "stub", "policyengine_uk_version": "test"} + + monkeypatch.setattr( + full_measure, + "resolve_target_measures", + lambda _factory, _registry, provider, **kwargs: SimpleNamespace( + receipt={"attached": {}, "provider": {}, "rounds": []}, + measure_inputs={}, + ), + ) + collections = [] + monkeypatch.setattr( + full_measure, "gc", SimpleNamespace(collect=lambda: collections.append(1) or 0) + ) + full_measure.resolve_uk_full_measures( + frame, + registry, + period=2025, + scratch_dir=tmp_path / "scratch", + resolver_factory=Resolver, + blocks=blocks, + local_grains=(), + ) + assert len(collections) == blocks + + +def _national_problem_fixture(monkeypatch, toy_ladder): + """A two-clone frame, one national target materialized from an injected + engine input, and the stubs the per-block path needs (the measure + inputs carry a person-level and two household-level columns, exactly as + the prepared-measures test above).""" + frame = clone_uk_dataset_with_ladder_geography( + source_frame(), + toy_ladder[0], + n_clones=2, + seed=7, + source_year=2023, + expected_constituency_vintage="2024_pcon", + ).frame + registry = TargetRegistry( + [ + TargetSpec( + name="national", + entity="household", + measure="prepared_count", + value=33.0, + period=2025, + family="fixture", + source="fixture", + ) + ], + country="uk", + ) + + class Resolver: + def __init__(self, *, frame, **kwargs): + self.frame = frame + self.simulation = object() + + def receipt(self): + return {"mode": "stub", "policyengine_uk_version": "test"} + + monkeypatch.setattr( + full_measure, + "resolve_target_measures", + lambda _factory, _registry, provider, **kwargs: SimpleNamespace( + receipt={"attached": {}, "provider": {}, "rounds": []}, + measure_inputs={ + ("person", "region"): np.zeros(provider.frame.n("person")), + ("household", "raw_engine_input"): provider.frame.table("household")[ + "household_id" + ].to_numpy(), + ("household", "ons/corporate_land_value"): np.ones( + provider.frame.n("household") + ), + }, + ), + ) + + def materialize(adapter, _registry, **kwargs): + assert "region" in adapter.tables["person"] + table = adapter.tables["household"] + table["prepared_count"] = table["raw_engine_input"].to_numpy() + return SimpleNamespace( + skipped=(), + report=lambda: {"prepared_count": 1, "skipped_count": 0, "skipped": []}, + ) + + monkeypatch.setattr(full_measure, "materialize_uk_ledger_targets", materialize) + return frame, registry, Resolver + + +def test_national_problem_is_compiled_per_block_in_the_pools_household_order( + monkeypatch, tmp_path, toy_ladder +): + """The dense kernel's path: each clone block materializes and compiles its + own columns; the stitched problem equals the whole-pool problem, column + for column, and never materializes the pool (microcosm#932 follow-up).""" + frame, registry, resolver_factory = _national_problem_fixture( + monkeypatch, toy_ladder + ) + results = {} + for blocks in (1, 2): + problem, rows, metrics, receipt = full_measure.resolve_uk_full_national_problem( + frame, + registry, + period=2025, + scratch_dir=tmp_path / f"scratch-{blocks}", + resolver_factory=resolver_factory, + blocks=blocks, + local_grains=(), + ) + results[blocks] = problem + ids = frame.table("household")["household_id"].to_numpy() + assert problem.matrix.shape == (1, len(ids)) + # Each household's contribution sits in its own column, in frame order. + np.testing.assert_array_equal(problem.matrix.toarray()[0], ids) + np.testing.assert_array_equal(problem.target_vector, [33.0]) + assert problem.weight_entity == "household" + np.testing.assert_array_equal( + problem.initial_weights.values, frame.resolve_weights("household").values + ) + assert problem.skipped == () + assert len(rows.targets) == 1 + assert metrics == {} + assert receipt["blocks"] == blocks + assert receipt["national_materialization"] == "per_engine_block" + assert receipt["target_materialization"] == { + "prepared_count": 1, + "skipped_count": 0, + "skipped": [], + } + assert receipt["national_inputs"] == 3 + if blocks == 2: + assert receipt["deviation"] == "per_clone_block_engine_resolution" + one, two = results[1], results[2] + assert (one.matrix != two.matrix).nnz == 0 + assert one.names == two.names + # The pool itself was left untouched by both paths. + assert "prepared_count" not in frame.table("household") + + +def test_national_problem_is_none_without_national_targets( + monkeypatch, tmp_path, toy_ladder +): + frame, _registry, resolver_factory = _national_problem_fixture( + monkeypatch, toy_ladder + ) + monkeypatch.setattr( + full_measure, + "compute_household_metrics", + lambda _simulation, area_type, *, household_ids, **_kwargs: pd.DataFrame( + {f"{area_type}_metric": np.ones(len(household_ids))}, + index=household_ids, + ), + ) + problem, rows, metrics, receipt = full_measure.resolve_uk_full_national_problem( + frame, + TargetRegistry([], country="uk"), + period=2025, + scratch_dir=tmp_path / "scratch", + resolver_factory=resolver_factory, + blocks=2, + local_grains=("constituency",), + ) + assert problem is None + assert len(rows.targets) == 0 + assert ( + metrics["constituency"].index.tolist() + == frame.table("household")["household_id"].tolist() + ) + assert receipt["national_materialization"] == "per_engine_block" + + +def _representation_block( + source_ids, weights, persons=3, *, identity=True, key_values=None +): + """A block whose allocation keys all equal ``key_values`` (default 1 per row).""" + from microcosm.build.uk_runtime.full_measure import ( + UK_WEIGHT_SHARE_FORMULA_INPUTS, + ) + + n = len(weights) + columns = {"household_id": np.arange(n)} + if identity: + columns["household_source_id"] = np.asarray(source_ids) + values = np.ones(n) if key_values is None else np.asarray(key_values, dtype=float) + for inputs in UK_WEIGHT_SHARE_FORMULA_INPUTS.values(): + for column in inputs: + columns[column] = values + household = pd.DataFrame(columns) + person = pd.DataFrame(index=range(persons)) + values = np.asarray(weights, dtype=float) + return SimpleNamespace( + table=lambda name: household if name == "household" else person, + weights_for=lambda entity: SimpleNamespace( + values=values, total=float(values.sum()) + ), + ) + + +def test_engine_population_representation_is_exact_only_for_identical_copies(): + pool = SimpleNamespace(weights_for=lambda entity: SimpleNamespace(total=6.0)) + copies = [ + (0, _representation_block([10, 11, 12], [0.5, 1.0, 1.5])), + (1, _representation_block([12, 10, 11], [1.5, 0.5, 1.0])), + ] + exact = full_measure._engine_population_representation(pool, copies) + assert exact["exact"] is True + assert exact["factor_by_block"] == { + "0": pytest.approx(2.0), + "1": pytest.approx(2.0), + } + assert exact["checks"]["max_abs_mass_share_deviation"] == pytest.approx(0.0) + + skewed = [ + (0, _representation_block([10, 11, 12], [0.5, 1.0, 1.5])), + (1, _representation_block([10, 11, 12], [0.5, 1.0, 2.0])), + ] + pool = SimpleNamespace(weights_for=lambda entity: SimpleNamespace(total=6.5)) + approximate = full_measure._engine_population_representation(pool, skewed) + assert approximate["exact"] is False + assert approximate["checks"]["identical_source_weight_multisets"] is False + assert approximate["factor_by_block"]["0"] == pytest.approx(6.5 / 3.0) + assert approximate["factor_by_block"]["1"] == pytest.approx(6.5 / 3.5) + + single = full_measure._engine_population_representation(pool, [(None, object())]) + assert single == {"mode": "single_block", "exact": True, "blocks": 1} + + # microcosm#1115 review: identical (source, weight) multisets are not + # enough; the allocation keys of the weight-share formulas must carry the + # same sum(x * w) in every block, and a frame without a source identity is + # never exact. + pool = SimpleNamespace(weights_for=lambda entity: SimpleNamespace(total=6.0)) + keyed = [ + (0, _representation_block([10, 11, 12], [0.5, 1.0, 1.5], key_values=[1, 2, 3])), + (1, _representation_block([12, 10, 11], [1.5, 0.5, 1.0], key_values=[3, 1, 2])), + ] + assert full_measure._engine_population_representation(pool, keyed)["exact"] is True + drawn = [ + (0, _representation_block([10, 11, 12], [0.5, 1.0, 1.5], key_values=[1, 2, 3])), + (1, _representation_block([12, 10, 11], [1.5, 0.5, 1.0], key_values=[9, 1, 2])), + ] + inexact = full_measure._engine_population_representation(pool, drawn) + assert inexact["exact"] is False + assert inexact["checks"]["identical_source_weight_multisets"] is True + assert inexact["checks"]["weight_share_inputs_match"] is False + shares = inexact["checks"]["weight_share_inputs"]["shareholding"] + assert shares["match"] is False and shares["max_rel_diff"] > 0.1 + assert set(shares["sum_by_block"]) == {"0", "1"} + nameless = [ + (0, _representation_block([10, 11, 12], [0.5, 1.0, 1.5], identity=False)), + (1, _representation_block([12, 10, 11], [1.5, 0.5, 1.0], identity=False)), + ] + anonymous = full_measure._engine_population_representation(pool, nameless) + assert anonymous["exact"] is False + assert anonymous["checks"]["identity_key"] is None + assert anonymous["checks"]["source_identity_present"] is False + partial = [ + (0, _representation_block([10, 11, 12], [0.5, 1.0, 1.5])), + (1, _representation_block([12, 10, 11], [1.5, 0.5, 1.0])), + ] + partial[1][1].table("household").drop(columns=["corporate_wealth"], inplace=True) + unverified = full_measure._engine_population_representation(pool, partial) + assert unverified["exact"] is False + assert unverified["checks"]["weight_share_inputs"]["corporate_land_value"][ + "missing_columns" + ] == ["corporate_wealth"] diff --git a/packages/microcosm-build/tests/engine_free/uk/test_uk_graph_terminal.py b/packages/microcosm-build/tests/engine_free/uk/test_uk_graph_terminal.py index 80761b106..988457606 100644 --- a/packages/microcosm-build/tests/engine_free/uk/test_uk_graph_terminal.py +++ b/packages/microcosm-build/tests/engine_free/uk/test_uk_graph_terminal.py @@ -58,6 +58,67 @@ def test_export_drops_native_aliases_and_keeps_identity_keyed_assignment(tmp_pat assert "geography_household_key" in stored and "data_zone_code" not in stored +def test_export_writes_area_codes_under_consumer_names_and_refuses_ladder_names( + tmp_path, +): + """microcosm#1114: the artifact carries ``constituency_code_oa``, + ``la_code_oa`` and ``region_code_oa`` and never the ladder names.""" + pytest.importorskip("tables") + from microcosm.build.uk_runtime.geography_ladder import ( + UK_EXPORT_AREA_CODE_COLUMNS, + export_area_code_columns, + ) + from microcosm.build.uk_runtime.graph_terminal import _tables + + frame = _frame() + ladder = frame.table("household") + exported = _tables(frame)["household"] + for ladder_name, export_name in UK_EXPORT_AREA_CODE_COLUMNS.items(): + assert ladder_name not in exported.columns + assert exported[export_name].tolist() == ladder[ladder_name].tolist() + # The in-memory frame is untouched: gates and diagnostics keep the ladder names. + assert "constituency_code" in frame.table("household").columns + descriptor = describe_uk_export(frame, bindings={"target_scope": "all"}) + assert descriptor["area_codes"]["columns"] == dict(UK_EXPORT_AREA_CODE_COLUMNS) + frames = descriptor["area_codes"]["frames"] + assert set(frames) == set(UK_EXPORT_AREA_CODE_COLUMNS.values()) + assert frames["constituency_code_oa"]["uk_ew_output_area_2021"] == "2024_pcon" + assert frames["la_code_oa"]["uk_ew_output_area_2021"] == "2023_april_lad" + assert frames["region_code_oa"]["uk_ew_output_area_2021"] == "2024_rgn" + assert "constituency_code_oa" in descriptor["tables"]["household"]["columns"] + assert "constituency_code" not in descriptor["tables"]["household"]["columns"] + path = tmp_path / "full.h5" + materialize_uk_export(frame, descriptor, path) + assert validate_uk_export(path, descriptor)["passed"] is True + with pd.HDFStore(path) as store: + stored = store["household"] + assert {"constituency_code_oa", "la_code_oa", "region_code_oa"} <= set( + stored.columns + ) + # A stored table that slipped a ladder name or an empty code through is refused. + with pd.HDFStore(path) as store: + household = store["household"].rename( + columns={"constituency_code_oa": "constituency_code"} + ) + household["la_code_oa"] = ["", "E07000008"] + store.put("household", household, format="table") + report = validate_uk_export(path, descriptor) + assert report["passed"] is False + assert any("ladder area code 'constituency_code'" in f for f in report["failures"]) + assert any( + "lacks the declared area code 'constituency_code_oa'" in f + for f in report["failures"] + ) + assert any("'la_code_oa' is empty on 1 household" in f for f in report["failures"]) + # A stale consumer column on the frame cannot ride through the boundary. + stale = ladder.assign(constituency_code_oa=["stale", "stale"]) + with pytest.raises(ValueError, match="already carries consumer area-code"): + export_area_code_columns(stale) + # A table without geography passes unchanged (the national line). + plain = pd.DataFrame({"household_id": [1], "region": ["LONDON"]}) + assert export_area_code_columns(plain) is plain + + def test_export_roundtrip_preserves_dtype_weights_lineage_period_and_gate(tmp_path): pytest.importorskip("tables") frame = _frame() @@ -772,3 +833,41 @@ def test_package_validates_materialized_evidence_against_graph_bytes(tmp_path): context.artifacts["export_readback"].payload = canonical_json(readback) with pytest.raises(ValueError, match="H5 readback failed"): UKPackageInventoryKernel().run(context) + + +def test_diagnostics_receive_the_holdout_in_the_schemas_shape(): + """The holdout artifact carries the kernel's binding and, when skipped, a + reason; the diagnostics records forbid undeclared fields and state a skipped + holdout as the bare marker. The first K=25 build (2026-10-06) reached its + terminal gates after 28 hours and the battery refused the graph's marker.""" + from pydantic import ValidationError + + from microcosm.build.uk_runtime.graph_terminal import diagnostics_rotated_holdout + from microcosm.diagnostics.schema import UKSkippedRotatedHoldout + + skipped = { + "report_only": True, + "skipped": True, + "reason": "Explicit development request; no holdout claim.", + "graph_binding": {"original_problem_artifact": "a" * 64, "artifacts": {}}, + } + assert diagnostics_rotated_holdout(skipped) == {"skipped": True} + UKSkippedRotatedHoldout.model_validate({"skipped": True}) + # The schema stays strict: the graph's artifact fields are the converter's + # to strip, and an undeclared field is refused. + with pytest.raises(ValidationError): + UKSkippedRotatedHoldout.model_validate(skipped) + with pytest.raises(ValidationError): + UKSkippedRotatedHoldout.model_validate({**skipped, "invented": 1}) + + measured = { + "report_only": True, + "method": "rotated_folds", + "n_folds": 5, + "fold_losses": [0.1, 0.2, 0.3, 0.4, 0.5], + "graph_binding": {"original_problem_artifact": "a" * 64, "artifacts": {}}, + } + shaped = diagnostics_rotated_holdout(measured) + assert "graph_binding" not in shaped + assert shaped == {k: v for k, v in measured.items() if k != "graph_binding"} + assert "graph_binding" in measured # the artifact's report is left intact diff --git a/packages/microcosm-build/tests/engine_free/uk/test_uk_measure_simulation.py b/packages/microcosm-build/tests/engine_free/uk/test_uk_measure_simulation.py index f18cfba39..0b264a3f1 100644 --- a/packages/microcosm-build/tests/engine_free/uk/test_uk_measure_simulation.py +++ b/packages/microcosm-build/tests/engine_free/uk/test_uk_measure_simulation.py @@ -1060,3 +1060,255 @@ def test_counterfactual_delta_refuses_bands_and_other_periods(monkeypatch, tmp_p ) with pytest.raises(ValueError, match="does not match target period"): resolver.counterfactual_delta(_relief_binding("income_tax"), 2024) + + +def _toy_uk_frame_with_native_aliases(): + """A validated UK national frame as the atomic geography path leaves it: + the nation-native alias codes are NA outside their own nation.""" + from microcosm.build.uk_runtime import uk_national_frame + from microcosm.frame import MassChangeRecord, WeightKind + + household = pd.DataFrame( + { + "household_id": np.asarray([100, 200, 300], dtype=np.int64), + "household_weight": [1.0, 1.0, 1.0], + "region": pd.array( + ["LONDON", "SCOTLAND", "NORTHERN_IRELAND"], dtype="string" + ), + "oa_code": pd.array( + ["E00000001", "S00090001", "N00000101"], dtype="string" + ), + "output_area_code": pd.array( + ["E00000001", "S00090001", None], dtype="string" + ), + "data_zone_code": pd.array( + [None, "S01006506", "N00000101"], dtype="string" + ), + "intermediate_zone_code": pd.array( + [None, "S02001236", None], dtype="string" + ), + "super_data_zone_code": pd.array([None, None, "N21000001"], dtype="string"), + "district_electoral_area_code": pd.array( + [None, None, "N10000101"], dtype="string" + ), + } + ) + person = pd.DataFrame( + { + "person_id": np.asarray([1, 2, 3], dtype=np.int64), + "person_benunit_id": np.asarray([10, 20, 30], dtype=np.int64), + "person_household_id": np.asarray([100, 200, 300], dtype=np.int64), + "age": [40, 41, 42], + } + ) + benunit = pd.DataFrame({"benunit_id": np.asarray([10, 20, 30], dtype=np.int64)}) + return uk_national_frame( + person=person, + benunit=benunit, + household=household, + time_period="2025", + weight_kind=WeightKind.IMPORTANCE, + mass_log=(MassChangeRecord("household", 3.0, 3.0, 1.0, "toy"),), + ) + + +def test_scratch_engine_input_leaves_the_native_alias_codes_behind( + monkeypatch, tmp_path: Path +): + """microcosm#931 derives nation-native alias codes that are NA outside + their own nation; policyengine-uk refuses a dataset with any NaN column, + so the scratch export the engine loads drops them like the release export + does. The first K=25 dense build from main (2026-10-04) failed on exactly + this at ``uk.full.measures``.""" + loaded = [] + + class FakeMicrosimulation: + def __init__(self, *, dataset): + loaded.append(dataset) + self.tax_benefit_system = SimpleNamespace(variables={}) + + monkeypatch.setitem( + sys.modules, + "policyengine_uk", + SimpleNamespace(__version__="9.9.9", Microsimulation=FakeMicrosimulation), + ) + frame = _toy_uk_frame_with_native_aliases() + before = frame.table("household").copy(deep=True) + resolver = UKMeasureResolver( + simulation_source=None, scratch_dir=tmp_path, year=2025, frame=frame + ) + written = pd.read_hdf(tmp_path / "simulation-input.h5", "household") + aliases = [ + "output_area_code", + "data_zone_code", + "intermediate_zone_code", + "super_data_zone_code", + "district_electoral_area_code", + ] + assert not set(aliases) & set(written.columns) + assert not written.isna().any().any() + assert written["oa_code"].tolist() == ["E00000001", "S00090001", "N00000101"] + assert written["household_id"].tolist() == [100, 200, 300] + # The resolver's own frame is untouched; only the engine's copy changed. + pd.testing.assert_frame_equal(frame.table("household"), before) + assert resolver.frame is frame + assert loaded == [str(tmp_path / "simulation-input.h5")] + receipt = resolver.receipt() + assert receipt["mode"] == "scratch_frame_export" + assert receipt["engine_scratch_dropped_columns"] == aliases + + +def test_scratch_engine_input_without_aliases_is_the_frame_itself( + monkeypatch, tmp_path +): + monkeypatch.setitem( + sys.modules, + "policyengine_uk", + SimpleNamespace( + __version__="9.9.9", + Microsimulation=lambda *, dataset: SimpleNamespace( + tax_benefit_system=SimpleNamespace(variables={}) + ), + ), + ) + writes = [] + monkeypatch.setattr( + measure_simulation, + "write_uk_national_frame", + lambda frame, path: writes.append(frame) or path, + ) + frame = FrameStub() + resolver = UKMeasureResolver( + simulation_source=None, scratch_dir=tmp_path, year=2025, frame=frame + ) + assert writes == [frame] + assert "engine_scratch_dropped_columns" not in resolver.receipt() + + +def _tiny_national_frame(weights): + from microcosm.build.uk_runtime.national_frame import uk_national_frame + from microcosm.frame import WeightKind + + person = pd.DataFrame( + { + "person_id": [1, 2, 3, 4], + "person_benunit_id": [10, 10, 20, 20], + "person_household_id": [100, 100, 200, 200], + "age": [35, 8, 40, 38], + "is_benunit_head": [True, False, True, False], + "is_parent": [True, False, False, False], + "is_uc_claimant": [True, False, True, True], + } + ) + benunit = pd.DataFrame( + { + "benunit_id": [10, 20], + "benunit_source_id": [10, 20], + "dependent_children": [1, 0], + "is_married": [False, False], + } + ) + household = pd.DataFrame( + { + "household_id": [100, 200], + "household_source_id": [100, 200], + "region": ["WALES", "NORTHERN_IRELAND"], + "council_tax": [1200.0, 0.0], + "rent": [6000.0, 6000.0], + "tenure_type": ["RENT_PRIVATELY", "RENT_PRIVATELY"], + } + ) + return uk_national_frame( + person=person, + benunit=benunit, + household=household, + time_period=2024, + household_weights=np.asarray(weights, dtype=float), + weight_kind=WeightKind.IMPORTANCE, + ) + + +def test_engine_scratch_frame_scales_household_weights_for_the_engine_only(): + frame = _tiny_national_frame([250.0, 125.0]) + before = frame.weights_for("household").values.copy() + + engine_frame, dropped = measure_simulation._engine_scratch_frame( + frame, engine_weight_scale=4.0 + ) + + assert dropped == () + np.testing.assert_allclose( + engine_frame.weights_for("household").values, before * 4.0 + ) + assert ( + engine_frame.weights_for("household").kind + is frame.weights_for("household").kind + ) + assert len(engine_frame.mass_log) == len(frame.mass_log) + 1 + record = engine_frame.mass_log[-1] + assert record.entity == "household" + assert record.declared_factor == 4.0 + assert record.new_total == pytest.approx(record.old_total * 4.0) + assert record.reason == measure_simulation.ENGINE_WEIGHT_SCALE_REASON + # the resolver's own frame keeps the block's true mass + np.testing.assert_array_equal(frame.weights_for("household").values, before) + # no scale (or a unit scale) on an alias-free frame: the frame itself + assert measure_simulation._engine_scratch_frame(frame)[0] is frame + assert ( + measure_simulation._engine_scratch_frame(frame, engine_weight_scale=1.0)[0] + is frame + ) + with pytest.raises(ValueError, match="positive finite"): + measure_simulation._engine_scratch_frame(frame, engine_weight_scale=0.0) + + +def test_resolver_writes_the_scaled_engine_frame_and_records_the_scale( + monkeypatch, tmp_path: Path +): + created = [] + + class FakeMicrosimulation: + def __init__(self, *, dataset): + created.append(dataset) + self.tax_benefit_system = SimpleNamespace(variables={}) + + monkeypatch.setitem( + sys.modules, + "policyengine_uk", + SimpleNamespace(__version__="9.9.9", Microsimulation=FakeMicrosimulation), + ) + writes = [] + monkeypatch.setattr( + measure_simulation, + "write_uk_national_frame", + lambda frame, path: writes.append((frame, path)) or path, + ) + monkeypatch.setattr( + measure_simulation, "validate_uc_claimant_input", lambda *args, **kwargs: None + ) + frame = _tiny_national_frame([250.0, 125.0]) + + resolver = UKMeasureResolver( + simulation_source=None, + scratch_dir=tmp_path, + year=2025, + frame=frame, + engine_weight_scale=4.0, + ) + + assert resolver.frame is frame + written, _path = writes[0] + np.testing.assert_allclose( + written.weights_for("household").values, + frame.weights_for("household").values * 4.0, + ) + assert resolver.receipt()["engine_weight_scale"] == 4.0 + assert created == [str(tmp_path / "simulation-input.h5")] + with pytest.raises(ValueError, match="scratch-mode"): + UKMeasureResolver( + simulation_source=tmp_path / "input.h5", + scratch_dir=tmp_path, + year=2025, + frame=frame, + engine_weight_scale=4.0, + ) diff --git a/packages/microcosm-build/tests/engine_free/uk/test_uk_release_input_coverage.py b/packages/microcosm-build/tests/engine_free/uk/test_uk_release_input_coverage.py index 5d0ecbaac..3775dc469 100644 --- a/packages/microcosm-build/tests/engine_free/uk/test_uk_release_input_coverage.py +++ b/packages/microcosm-build/tests/engine_free/uk/test_uk_release_input_coverage.py @@ -135,6 +135,32 @@ def test_binding_hands_the_spine_frame_to_the_build_state_half(self) -> None: assert result.passed, result.failures assert result.details["family_build_state_frame"] == "spine" + def test_binding_refuses_to_judge_the_calibrated_frame_without_build_state( + self, + ) -> None: + """microcosm#1115 review: no spine build state means evidence absent, + never a verdict read off the calibrated frame's weights.""" + from microcosm.build.gate_battery import EvidenceContext + from microcosm.build.uk_runtime.battery_bindings import ( + UK_GATE_REGISTRY, + _evaluate_release_input_coverage, + ) + + _, release = self._frames() + artifacts = { + "coverage_engine": _StubEngine({"gift_aid": 0.0}), + "coverage_manifest": self._contract(), + } + with pytest.raises(ValueError, match="spine build state"): + _evaluate_release_input_coverage( + EvidenceContext(frame=release, artifacts=artifacts), {} + ) + binding = UK_GATE_REGISTRY["release_input_coverage"] + assert "spine_frame" in binding.required_artifacts({}) + assert "spine_frame" not in binding.required_artifacts( + {"check": "manifest_current"} + ) + class TestFamilyMassSemantics: """Every family conserves household mass (microcosm#1063 item 3).""" diff --git a/packages/microcosm-build/tests/engine_free/uk/test_uk_rowwise_candidate.py b/packages/microcosm-build/tests/engine_free/uk/test_uk_rowwise_candidate.py index 54bda8f72..ed3519446 100644 --- a/packages/microcosm-build/tests/engine_free/uk/test_uk_rowwise_candidate.py +++ b/packages/microcosm-build/tests/engine_free/uk/test_uk_rowwise_candidate.py @@ -181,14 +181,15 @@ def interrupt_support(self, target): assert not output_dir.exists() -def test_release_verdict_requires_single_block_engine() -> None: +def test_release_verdict_requires_an_exact_engine_population() -> None: builder = rowwise_cli releasable, posture = builder._release_verdict( sample_fraction=1.0, engine_blocks=1, release_blocking_gates_passed=True ) assert releasable is True and all(posture.values()) - # A per-block engine resolution never writes a releasable artifact, even - # with every release-blocking gate passed on the full rung (#736 erratum). + # A per-block engine resolution whose blocks are not shown to represent the + # pool exactly never writes a releasable artifact, even with every + # release-blocking gate passed on the full rung (#736 erratum). releasable, posture = builder._release_verdict( sample_fraction=1.0, engine_blocks=15, release_blocking_gates_passed=True ) @@ -196,8 +197,20 @@ def test_release_verdict_requires_single_block_engine() -> None: assert posture == { "full_rung": True, "single_block_engine": False, + "engine_population_exact": False, "release_blocking_gates_passed": True, } + # The measures loop's record that every block is an identical copy scaled + # to the pool makes the per-block resolution exact, and releasable. + releasable, posture = builder._release_verdict( + sample_fraction=1.0, + engine_blocks=15, + release_blocking_gates_passed=True, + engine_population_exact=True, + ) + assert releasable is True + assert posture["single_block_engine"] is False + assert posture["engine_population_exact"] is True assert ( builder._release_verdict( sample_fraction=0.1, engine_blocks=1, release_blocking_gates_passed=True diff --git a/packages/microcosm-build/tests/engine_free/uk/test_uk_terminal_gates.py b/packages/microcosm-build/tests/engine_free/uk/test_uk_terminal_gates.py index 109c66366..a9637942d 100644 --- a/packages/microcosm-build/tests/engine_free/uk/test_uk_terminal_gates.py +++ b/packages/microcosm-build/tests/engine_free/uk/test_uk_terminal_gates.py @@ -282,6 +282,78 @@ def test_export_surface_allows_claimant_roles_but_not_unreviewed_columns() -> No assert any("unreviewed_extra" in failure for failure in unrelated.failures) +def test_export_candidate_columns_and_release_frame_carry_consumer_area_code_names() -> ( + None +): + """microcosm#1114: the surface the export gate compares, and the frame the + national boundary writes, name the three area codes as consumers read them.""" + import numpy as np + import pandas as pd + + from microcosm.build.uk_runtime.geography_ladder import ( + UK_EXPORT_AREA_CODE_COLUMNS, + ) + from microcosm.build.uk_runtime.national_frame import ( + uk_national_frame, + uk_release_export_frame, + ) + from microcosm.build.uk_runtime.terminal_gates import ( + UK_ALLOWED_EXTRA_EXPORT_COLUMNS, + uk_export_candidate_columns, + ) + + frame = uk_national_frame( + person=pd.DataFrame( + { + "person_id": [1, 2], + "person_benunit_id": [1, 2], + "person_household_id": [1, 2], + "age": [30, 40], + } + ), + benunit=pd.DataFrame({"benunit_id": [1, 2]}), + household=pd.DataFrame( + { + "household_id": [1, 2], + "region": ["LONDON", "SOUTH_EAST"], + "constituency_code": ["E14000001", "E14000002"], + "local_authority_code": ["E09000001", "E07000008"], + "region_code": ["E12000007", "E12000008"], + "ward_code": ["E05000001", "E05000002"], + } + ), + time_period="2024", + household_weights=np.asarray([1.0, 2.0]), + ) + surface = uk_export_candidate_columns(frame) + assert { + "household.constituency_code_oa", + "household.la_code_oa", + "household.region_code_oa", + "household.ward_code", + } <= surface + assert ( + not { + "household.constituency_code", + "household.local_authority_code", + "household.region_code", + } + & surface + ) + for export in UK_EXPORT_AREA_CODE_COLUMNS.values(): + assert f"household.{export}" in UK_ALLOWED_EXTRA_EXPORT_COLUMNS + for ladder in UK_EXPORT_AREA_CODE_COLUMNS: + assert f"household.{ladder}" not in UK_ALLOWED_EXTRA_EXPORT_COLUMNS + exported = uk_release_export_frame(frame) + household = exported.table("household") + assert household["constituency_code_oa"].tolist() == ["E14000001", "E14000002"] + assert household["la_code_oa"].tolist() == ["E09000001", "E07000008"] + assert household["region_code_oa"].tolist() == ["E12000007", "E12000008"] + assert "constituency_code" not in household.columns + assert "constituency_code" in frame.table("household").columns + assert exported.weights_for("household").values.tolist() == [1.0, 2.0] + + def test_export_candidate_columns_strip_ids_and_carry_the_weight() -> None: """microcosm#1063 c9: the certifier rehearsal listed every id column as an unreviewed extra and the frame's weights as a missing reference column.""" @@ -844,3 +916,24 @@ def test_weight_gates_evaluate_on_family_folded_weights_and_report_rows_beside() assert plain.details["evaluated_on"] == "row_weights" assert "family_folded" not in plain.details assert not plain.passed + + +def test_degenerate_gate_skips_columns_the_release_boundary_drops() -> None: + """A gate fed the pre-export frame must not report a column the exported + H5 never carries: the skipped names are recorded, nothing else changes.""" + dataset = _dataset(signal=0.0) # person.employment_income is all-zero + reported = uk_degenerate_release_surface_gate(dataset) + assert reported.passed is False + assert any("person.employment_income" in line for line in reported.failures) + + skipped = uk_degenerate_release_surface_gate( + dataset, dropped_at_export={"person": ("employment_income",)} + ) + assert skipped.passed is True + assert skipped.details["dropped_at_export"] == ["person.employment_income"] + assert skipped.details["columns_checked"] == reported.details["columns_checked"] - 1 + # a drop list naming an absent column is inert + inert = uk_degenerate_release_surface_gate( + _dataset(), dropped_at_export={"household": ("not_a_column",)} + ) + assert inert.passed is True and inert.details["dropped_at_export"] == [] diff --git a/packages/microcosm-build/tests/engine_workflow/us/test_us_multispine_pool_tool.py b/packages/microcosm-build/tests/engine_workflow/us/test_us_multispine_pool_tool.py index ca014679b..9606d75b4 100644 --- a/packages/microcosm-build/tests/engine_workflow/us/test_us_multispine_pool_tool.py +++ b/packages/microcosm-build/tests/engine_workflow/us/test_us_multispine_pool_tool.py @@ -338,7 +338,7 @@ def capture_equality(expected: object, actual: object) -> None: "country": "us", "schema_id": "country_spec", "schema_version": 1, - "spec_sha256": "d1df6b316e4937e809f5f020ea2d3a7bd38bccc1845f3ae6e104f7eed6a73323", + "spec_sha256": "f76f921d4f74392b106b8c0729ce250494e0a55455eaba7b910b19cdf7825178", }, } diff --git a/packages/microcosm-calibrate/src/microcosm/calibrate/artifacts.py b/packages/microcosm-calibrate/src/microcosm/calibrate/artifacts.py index b7cd2783b..172b396f6 100644 --- a/packages/microcosm-calibrate/src/microcosm/calibrate/artifacts.py +++ b/packages/microcosm-calibrate/src/microcosm/calibrate/artifacts.py @@ -11,7 +11,7 @@ import hashlib import json from collections.abc import Mapping, Sequence -from dataclasses import dataclass, replace +from dataclasses import dataclass, field, replace from io import BytesIO from numbers import Integral from zipfile import ZIP_DEFLATED, ZipFile, ZipInfo @@ -171,15 +171,32 @@ def _read_target(raw: Mapping) -> Target: ) +def _axis_array(entity_ids: tuple[int | str, ...]) -> np.ndarray: + """The entity axis as one array every compiled row compares against. + + Rebuilding the axis as a Python tuple from the frame on every row is + quadratic in practice: 22,053 local rows over a 1.6 million-household + pool is ~3.5e10 scalar conversions (the first K=25 UK dense build, + 2026-10-05). One shared array and a vectorised comparison keep the guard. + """ + + if all(isinstance(value, Integral) for value in entity_ids): + return np.asarray(entity_ids, dtype=np.int64) + return np.asarray(entity_ids, dtype=object) + + @dataclass(frozen=True) class _CompiledRow: entity: str entity_ids: tuple[int | str, ...] values: sparse.csr_array + axis: np.ndarray | None = field(default=None, compare=False, hash=False, repr=False) def __call__(self, frame: Frame) -> np.ndarray: column = frame.schema.entity_id_column(self.entity) - if tuple(frame.table(self.entity)[column]) != self.entity_ids: + actual = frame.table(self.entity)[column].to_numpy() + axis = _axis_array(self.entity_ids) if self.axis is None else self.axis + if actual.shape != axis.shape or not np.array_equal(actual, axis): raise ValueError( "Compiled contribution requires its exact ordered entity axis." ) @@ -204,6 +221,7 @@ def to_target_set(self) -> TargetSet: Contributions already include original filters and entity aggregation. No raw measure is recomputed; consuming a different row axis refuses. """ + axis = _axis_array(self.entity_ids) return TargetSet( [ replace( @@ -214,6 +232,7 @@ def to_target_set(self) -> TargetSet: self.problem.weight_entity, self.entity_ids, self.problem.matrix[index : index + 1], + axis, ), ) for index, target in enumerate(self.problem.targets) diff --git a/packages/microcosm-calibrate/src/microcosm/calibrate/solve.py b/packages/microcosm-calibrate/src/microcosm/calibrate/solve.py index dfc40b928..95828a026 100644 --- a/packages/microcosm-calibrate/src/microcosm/calibrate/solve.py +++ b/packages/microcosm-calibrate/src/microcosm/calibrate/solve.py @@ -205,6 +205,14 @@ #: the search's cost at ``budget_iters`` optimizations; ~10 bisection steps cut #: the ``log10`` bracket by 2^10, far finer than the count is resolvable. _DEFAULT_BUDGET_ITERS = 10 + +#: Warm start (microcosm#1115): with a user-supplied ``initial_lambda`` the +#: budget search probes it first and, when that probe misses, brackets the +#: budget within this many decades of it before interpolating, instead of +#: re-bisecting the whole ``[_L0_SEARCH_LO, _L0_SEARCH_HI]`` range. Half a +#: decade (a factor 3.16 either way) covers the drift of a re-solved pool; a +#: miss beyond it hands the search the global bracket on that side. +_L0_WARM_SPAN_DECADES = 0.5 #: What the ``target_records`` budget search measures against the budget. BUDGET_BASIS_NONZERO_COUNT = "nonzero_count" BUDGET_BASIS_OPEN_PROBABILITY_MASS = "open_probability_mass" @@ -1446,7 +1454,12 @@ def _search_l0_lambda_for_budget( prune_atol: Threshold counting a weight as non-zero (a survivor). initial_lambda: A user-supplied ``l0_lambda`` to evaluate first as a warm start (clamped into the bracket); ``None`` starts at the bracket - mid-point. + mid-point. A warm probe that misses is followed by one probe + ``_L0_WARM_SPAN_DECADES`` away on the side it steered to: if that + edge steers back the search interpolates between the two measured + ends (the measure is close to log-linear in the penalty), otherwise + it continues on the global bracket beyond the edge. The cold path is + unchanged: mid-point first, then bisection. budget_iters: Maximum number of optimizations the search may spend. Returns: @@ -1655,31 +1668,76 @@ def settled() -> bool: ) # Warm start: evaluate the user's lambda (or the bracket mid-point) first. - if initial_lambda is not None and initial_lambda > 0: + warm = initial_lambda is not None and initial_lambda > 0 + global_lo_u, global_hi_u = lo_u, hi_u + if warm: first_u = min(max(math.log10(initial_lambda), lo_u), hi_u) else: first_u = (lo_u + hi_u) / 2.0 iters_left = budget_iters n_nonzero, verdict = consider(10.0**first_u) iters_left -= 1 + # The measure at each bracket end (None while unprobed or over-pruned): + # the warm path interpolates on them, the cold path never reads them. + lo_m: int | None = None + hi_m: int | None = None + first_m = None if n_nonzero == _over_pruned else int(n_nonzero) # Seed the bracket so the side the warm start landed on is tightened. An # over-pruned (infeasible) probe groups with "too few survivors". if steer(n_nonzero, verdict) == _steer_larger: - lo_u = first_u # too many survivors -> need a larger penalty + lo_u, lo_m = first_u, first_m # too many survivors -> larger penalty else: - hi_u = first_u # too few survivors / over-pruned -> need a smaller penalty + hi_u, hi_m = first_u, first_m # too few / over-pruned -> smaller penalty + edge_probe: dict[str, object] | None = None + if warm and iters_left > 0 and not settled(): + # The warm probe missed. Probe _L0_WARM_SPAN_DECADES away on the side + # it steered to: an edge that steers back closes a narrow bracket + # around the budget; one that steers on hands the search the global + # bracket beyond the edge, as a cold start would have had. + steered_larger = lo_u == first_u + if steered_larger: + edge_u = min(first_u + _L0_WARM_SPAN_DECADES, global_hi_u) + else: + edge_u = max(first_u - _L0_WARM_SPAN_DECADES, global_lo_u) + if edge_u != first_u: + n_nonzero, verdict = consider(10.0**edge_u) + iters_left -= 1 + edge_m = None if n_nonzero == _over_pruned else int(n_nonzero) + direction = steer(n_nonzero, verdict) + edge_probe = {"l0_lambda": 10.0**edge_u, "steer": direction} + if steered_larger: + if direction == _steer_smaller: + hi_u, hi_m = edge_u, edge_m + elif direction == _steer_larger: + lo_u, lo_m = edge_u, edge_m + hi_u, hi_m = global_hi_u, None + else: + if direction == _steer_larger: + lo_u, lo_m = edge_u, edge_m + elif direction == _steer_smaller: + hi_u, hi_m = edge_u, edge_m + lo_u, lo_m = global_lo_u, None # Keep searching while no acceptable run is known yet, or the best is # outside tolerance, until the iteration budget is spent. while iters_left > 0 and not settled(): - mid_u = (lo_u + hi_u) / 2.0 + if warm and lo_m is not None and hi_m is not None and lo_m != hi_m: + # Interpolate on the measured ends (regula falsi with a progress + # safeguard): bisection would spend about six probes halving a + # decade down to the draw's window; the secant lands in one or two. + aim = target_records + _warm_aim_margin(probes) + frac = (lo_m - aim) / (lo_m - hi_m) + mid_u = lo_u + (hi_u - lo_u) * min(max(frac, 0.1), 0.9) + else: + mid_u = (lo_u + hi_u) / 2.0 n_nonzero, verdict = consider(10.0**mid_u) iters_left -= 1 + mid_m = None if n_nonzero == _over_pruned else int(n_nonzero) direction = steer(n_nonzero, verdict) if direction == _steer_smaller: - hi_u = mid_u # over-pruned / too few / short mass -> smaller penalty + hi_u, hi_m = mid_u, mid_m # over-pruned / too few / short mass elif direction == _steer_larger: - lo_u = mid_u # too many survivors or certainties -> larger penalty + lo_u, lo_m = mid_u, mid_m # too many survivors or certainties else: break @@ -1712,6 +1770,9 @@ def settled() -> bool: "feasible_draw_pi_hi": feasible_draw_pi_hi, "evaluations": int(evaluation), "probes": probes, + "initial_lambda": float(initial_lambda) if warm else None, + "warm_span_decades": _L0_WARM_SPAN_DECADES if warm else None, + "warm_edge_probe": edge_probe, "selected_l0_lambda": None if best is None else float(best[2]), "selected_measure": None if best is None else int(best[3]), "selected_feasible": ( @@ -1742,6 +1803,22 @@ def settled() -> bool: return best[:4] +def _warm_aim_margin(probes: list[dict[str, object]]) -> float: + """Where a warm search aims above the budget. + + A feasibility-aware search can only stop inside the draw's window, which + runs from the budget up to the budget plus the gates' boundary mass; aim + at its middle, taken from the latest probe that measured one. A plain + count search aims at the budget itself. + """ + + for probe in reversed(probes): + mass = probe.get("boundary_mass") + if isinstance(mass, int | float) and math.isfinite(mass) and mass > 0: + return 0.5 * float(mass) + return 0.0 + + def _project_to_total( weights: np.ndarray, total: float, diff --git a/packages/microcosm-calibrate/tests/engine_free/shared/test_informed_gates.py b/packages/microcosm-calibrate/tests/engine_free/shared/test_informed_gates.py index eede278d6..eb4e0f6c3 100644 --- a/packages/microcosm-calibrate/tests/engine_free/shared/test_informed_gates.py +++ b/packages/microcosm-calibrate/tests/engine_free/shared/test_informed_gates.py @@ -287,3 +287,93 @@ def test_feasibility_aware_search_returns_the_closest_run_when_nothing_is_drawab assert search["selected_measure"] == 90 assert abs(search["selected_measure"] - k) > search["tolerance"] assert result.gate_open_probabilities is not None + + +def test_warm_start_probes_the_given_penalty_first_and_brackets_a_miss(monkeypatch): + """microcosm#1115: a known penalty is probed first; a hit ends the search + after one probe, a miss brackets within half a decade and interpolates.""" + import numpy as np + import pandas as pd + + from microcosm.calibrate import Target, TargetSet, calibrate + from microcosm.calibrate import solve as solve_module + from microcosm.calibrate.solve import ( + _L0_WARM_SPAN_DECADES, + BUDGET_BASIS_OPEN_PROBABILITY_MASS, + ) + from microcosm.frame import EntitySchema, Frame, WeightKind, Weights + + n, k = 1000, 500 + thresholds = -7.12 + 8.0 * (np.arange(n) + 0.5) / n + monkeypatch.setattr( + solve_module, "_optimize", _stub_polarised_optimizer(thresholds, 10.0) + ) + household = pd.DataFrame({"household_id": np.arange(1, n + 1)}) + person = pd.DataFrame( + {"person_id": np.arange(1, n + 1), "person_household_id": np.arange(1, n + 1)} + ) + frame = Frame( + {"person": person, "household": household}, + EntitySchema(group_entities=("household",)), + {"household": Weights(np.ones(n), WeightKind.DESIGN)}, + ) + targets = TargetSet( + [Target("count", "household", lambda f: np.ones(f.n("household")), 500.0)] + ) + common = dict( + epochs=3, + seed=1, + budget_iters=10, + target_records=k, + budget_basis=BUDGET_BASIS_OPEN_PROBABILITY_MASS, + feasible_draw_pi_hi=0.95, + ) + + cold = calibrate(frame, targets, **common) + cold_search = cold.options["budget_search"] + # The cold path is untouched: mid-point first, no warm receipt. + assert cold_search["probes"][0]["l0_lambda"] == 1e-3 + assert cold_search["initial_lambda"] is None + assert cold_search["warm_edge_probe"] is None + assert cold_search["stopped_on"] == "acceptable_within_tolerance" + known = cold_search["selected_l0_lambda"] + assert cold_search["evaluations"] > 2 + + # A hit: the known penalty is probed first and the search stops there. + warm = calibrate(frame, targets, l0_lambda=known, **common) + search = warm.options["budget_search"] + assert search["initial_lambda"] == known + assert search["evaluations"] == 1 + assert search["probes"][0]["l0_lambda"] == known + assert search["warm_edge_probe"] is None + assert search["selected_feasible"] is True + assert warm.l0_lambda == known + np.testing.assert_array_equal(warm.weights, cold.weights) + + # A miss on the over-pruned side: the next probe sits half a decade below + # the hint, the bracket closes on the measured ends, and the search lands + # within the probes bisection would have needed from the hint alone. + stale = known * 10**0.35 + missed = calibrate(frame, targets, l0_lambda=stale, **common) + search = missed.options["budget_search"] + probes = search["probes"] + assert probes[0]["l0_lambda"] == pytest.approx(stale) + assert probes[0]["verdict"] in ("boundary_mass_short", "boundary_short_of_draw") + assert probes[1]["l0_lambda"] == pytest.approx(stale / 10**_L0_WARM_SPAN_DECADES) + assert search["warm_edge_probe"]["l0_lambda"] == pytest.approx( + probes[1]["l0_lambda"] + ) + assert search["stopped_on"] == "acceptable_within_tolerance" + assert search["selected_feasible"] is True + assert search["evaluations"] <= 4 + assert abs(search["selected_measure"] - k) <= search["tolerance"] + + # A miss beyond the span hands the search the global bracket on that side + # and it still settles within the iteration budget. + far = known * 10**2.0 + recovered = calibrate(frame, targets, l0_lambda=far, **common) + search = recovered.options["budget_search"] + assert search["probes"][0]["l0_lambda"] == pytest.approx(far) + assert search["warm_edge_probe"]["steer"] == "smaller_penalty" + assert search["stopped_on"] == "acceptable_within_tolerance" + assert search["selected_feasible"] is True diff --git a/packages/microcosm-calibrate/tests/engine_free/shared/test_ordered_artifacts.py b/packages/microcosm-calibrate/tests/engine_free/shared/test_ordered_artifacts.py index 285940abf..2ca56c138 100644 --- a/packages/microcosm-calibrate/tests/engine_free/shared/test_ordered_artifacts.py +++ b/packages/microcosm-calibrate/tests/engine_free/shared/test_ordered_artifacts.py @@ -138,3 +138,42 @@ def test_complete_result_rebuild_preserves_dense_and_search_state(monkeypatch): np.testing.assert_array_equal( restored.gate_open_probabilities, result.gate_open_probabilities ) + + +def test_compiled_rows_guard_the_axis_without_iterating_the_id_column(monkeypatch): + """Every compiled row checks the frame's entity axis against one shared + array with a vectorised comparison. Rebuilding the axis as a Python tuple + per row made 22,053 local rows over a 1.6 million-household pool cost + ~3.5e10 scalar conversions (the first K=25 UK dense build, 2026-10-05).""" + restored = decode_problem(encode_problem(problem(), entity_ids=(10, 20))) + targets = restored.to_target_set() + axes = {id(target.measure.axis) for target in targets.targets} + assert len(axes) == 1 + assert targets.targets[0].measure.axis.dtype == np.int64 + + def frame_for(ids): + return Frame( + { + "person": pd.DataFrame( + {"person_id": [1, 2], "person_household_id": list(ids)} + ), + "household": pd.DataFrame({"household_id": list(ids)}), + }, + EntitySchema(group_entities=("household",)), + {"household": problem().initial_weights}, + pd.Series(["a", "a"]), + ) + + def refuse_iteration(self): + raise AssertionError("the axis guard must not iterate the id column") + + monkeypatch.setattr(pd.Series, "__iter__", refuse_iteration) + recompiled = build_constraint_matrix( + frame_for((10, 20)), targets, weight_entity="household" + ) + np.testing.assert_array_equal( + recompiled.matrix.toarray(), problem().matrix.toarray() + ) + # A frame on another axis (same length, valid ascending ids) is refused. + with pytest.raises(ValueError, match="exact ordered entity axis"): + targets.targets[0].measure(frame_for((10, 30))) diff --git a/packages/microcosm-data/src/microcosm/data/contract.py b/packages/microcosm-data/src/microcosm/data/contract.py index 8e505e9a2..883e83e51 100644 --- a/packages/microcosm-data/src/microcosm/data/contract.py +++ b/packages/microcosm-data/src/microcosm/data/contract.py @@ -416,13 +416,13 @@ def compatibility_claim_declarer_error(declared_by: object) -> str | None: # fingerprint derives from the manifest digest. Editing the spec moves all # three here in the same reviewed change. _UK_GATE_BATTERY_POLICY_SHA256 = ( - "6e1a3a38efe1e984426d4a40a63f1671434d27c44e89d75abd7290a955af18b0" + "49bac274b87f743371a78f1bea6b19b703cb5eceb2efcfdea483c41fddcfc63d" ) _UK_GATE_BATTERY_GATES_MANIFEST_SHA256 = ( - "f0c8e1d4f3e97ff6fd6ff51dff9db6414816785807cec30a5ea9526630af6a65" + "dd9a99f66bba9dd9e4a0eee3b506d6fa8e2119db7e61178b21a034e6e4d82173" ) _UK_GATE_BATTERY_SPEC_FINGERPRINT = ( - "c361fe49d7edd364f31a1f30282fe5b9050e779d311ab438ab4aef7e3e95148c" + "f1b1fb935fde60d0d0febbda70fee1c7fb88310d1ee3cffbdb1d59410ddbb35c" ) #: Spec entry id -> the legacy gate name whose observable detail checks #: apply unchanged (the battery re-keys the report by entry id; the gate @@ -867,10 +867,10 @@ def _uk_line_cut_tag_re(release_id: str) -> re.Pattern[str] | None: }, "release_cut": { "gates_manifest_sha256": ( - "e072281905223b9346f77da11dd52c390421b22046292cef92094a3b894ab365" + "70a0add2714853b18c9e7e75efe051a17c4e5e3c2ae0351943a9b434802e00b7" ), "policy_sha256": ( - "a673ec3efe1be53567ddc12b344e44ab872e05a6ebcea2862b1762952901ce78" + "1201c699695cbd7a3d5faf54e55e6557a50a094245013aa928e5b57fdc94fc00" ), }, } diff --git a/packages/microcosm-data/tests/engine_free/shared/test_contract.py b/packages/microcosm-data/tests/engine_free/shared/test_contract.py index 288b0a72a..b1e037b65 100644 --- a/packages/microcosm-data/tests/engine_free/shared/test_contract.py +++ b/packages/microcosm-data/tests/engine_free/shared/test_contract.py @@ -127,13 +127,13 @@ def _trusted_terminal_gate_signing_key(monkeypatch) -> None: UK_GATE_BATTERY_PRODUCER = "microcosm.build.gate_battery" UK_GATE_BATTERY_SIGNING_KEY_ENV = "MICROCOSM_UK_TERMINAL_GATE_SIGNING_KEY" UK_GATE_BATTERY_POLICY_SHA256 = ( - "6e1a3a38efe1e984426d4a40a63f1671434d27c44e89d75abd7290a955af18b0" + "49bac274b87f743371a78f1bea6b19b703cb5eceb2efcfdea483c41fddcfc63d" ) UK_GATE_BATTERY_GATES_MANIFEST_SHA256 = ( - "f0c8e1d4f3e97ff6fd6ff51dff9db6414816785807cec30a5ea9526630af6a65" + "dd9a99f66bba9dd9e4a0eee3b506d6fa8e2119db7e61178b21a034e6e4d82173" ) UK_GATE_BATTERY_SPEC_FINGERPRINT = ( - "c361fe49d7edd364f31a1f30282fe5b9050e779d311ab438ab4aef7e3e95148c" + "f1b1fb935fde60d0d0febbda70fee1c7fb88310d1ee3cffbdb1d59410ddbb35c" ) UK_GATE_BATTERY_DEGENERATE_EVIDENCE_SHA256 = ( "6f0243bcda09dad26945376230c44ec3cf55d4e417c3a25e29bae8c59bc1a69d" diff --git a/packages/microcosm-graph/tests/fixtures/parity/kernels/calibrate/pins.json b/packages/microcosm-graph/tests/fixtures/parity/kernels/calibrate/pins.json index 8f7bc6a46..8740c755c 100644 --- a/packages/microcosm-graph/tests/fixtures/parity/kernels/calibrate/pins.json +++ b/packages/microcosm-graph/tests/fixtures/parity/kernels/calibrate/pins.json @@ -1 +1 @@ -{"dependencies":{"numpy":"2.4.6","pandas":"3.0.3","scipy":"1.17.1","torch":"2.12.0"},"implementation_hash":"dff5e83dd41cc95587d08b8a5c0eb06e09c61ed96cd0e32dc0fb680a6ee0842c","kernel":"calibrate.adam@1","node":"calibrate","node_key":"d942d46f44d9fc9302ee1789375296bf03c423a0597791d85d7d005c903bd59b","numeric":"bitwise","platform":"arm64/darwin/py3.14","platforms":{"arm64/darwin/py3.14":{"direct":"direct.csv","node_key":"d942d46f44d9fc9302ee1789375296bf03c423a0597791d85d7d005c903bd59b"}},"seed":0} +{"dependencies":{"numpy":"2.4.6","pandas":"3.0.3","scipy":"1.17.1","torch":"2.12.0"},"implementation_hash":"7b5f1e03cfcd5b4942c951657ac67a7b49b9ddb48cd3c62cd780a2e3a007a37a","kernel":"calibrate.adam@1","node":"calibrate","node_key":"bec0ef105306c0a25d41eaae5833ef613c15a3ff4e165eefd1d3273373d22ab1","numeric":"bitwise","platform":"arm64/darwin/py3.14","platforms":{"arm64/darwin/py3.14":{"direct":"direct.csv","node_key":"bec0ef105306c0a25d41eaae5833ef613c15a3ff4e165eefd1d3273373d22ab1"}},"seed":0} diff --git a/test_support/microcosm_build/uk_full_build_cli.py b/test_support/microcosm_build/uk_full_build_cli.py index 4434ae0c1..98465aed5 100644 --- a/test_support/microcosm_build/uk_full_build_cli.py +++ b/test_support/microcosm_build/uk_full_build_cli.py @@ -8,6 +8,7 @@ import hashlib import importlib.util import json +import sys from dataclasses import replace from pathlib import Path @@ -308,6 +309,26 @@ def gate_payload(phase, failed=None): LADDER_TARGET = "ons.census.households@E14000001" +#: The synthetic measures node's engine resolution, as the problem bindings +#: carry it: ``run_dense_main`` sets the block count from ``--engine-blocks``; +#: tests flip ``ENGINE_POPULATION_EXACT`` to stage an inexact per-block run. +ENGINE_BLOCKS = 1 +ENGINE_POPULATION_EXACT = True + + +def measure_resolution_evidence() -> dict: + evidence: dict = {"blocks": ENGINE_BLOCKS} + if ENGINE_BLOCKS > 1: + evidence["engine_population_representation"] = { + "mode": "block_weights_scaled_to_pool", + "exact": bool(ENGINE_POPULATION_EXACT), + "blocks": ENGINE_BLOCKS, + "factor_by_block": { + str(index): float(ENGINE_BLOCKS) for index in range(ENGINE_BLOCKS) + }, + } + return evidence + def problem_payload() -> bytes: """One ordered local problem over the two fixture households.""" @@ -345,7 +366,7 @@ def problem_payload() -> bytes: "stood_on": {"census_households/constituency": ["fixture"]} }, "rung_surface": {"dropped_cells": 0}, - "measure_resolution": {"blocks": 1}, + "measure_resolution": measure_resolution_evidence(), "cross_geography": { "unbound_bridges": [], "empty_legs_licensed": [], @@ -669,6 +690,7 @@ def prepare(args, *, telemetry=None, attempt=None): monkeypatch.setattr(cli, "parse_args", lambda argv: args) monkeypatch.setattr(cli, "prepare_full_build", prepare) + monkeypatch.setattr(sys.modules[__name__], "ENGINE_BLOCKS", int(args.engine_blocks)) return cli.main([]), args.out diff --git a/test_support/microcosm_build/uk_full_target_graph.py b/test_support/microcosm_build/uk_full_target_graph.py index 48ace78a6..7056bb100 100644 --- a/test_support/microcosm_build/uk_full_target_graph.py +++ b/test_support/microcosm_build/uk_full_target_graph.py @@ -32,6 +32,7 @@ from microcosm.build.uk_runtime.local_rowwise import UKRowwiseNationalRows from microcosm.calibrate import TargetRegistry, TargetSpec from microcosm.calibrate.artifacts import decode_problem +from microcosm.calibrate.matrix import build_constraint_matrix from microcosm.frame import Frame from microcosm.graph import ( ArtifactInput, @@ -170,13 +171,14 @@ def surface(): ), ) - def measures(frame, national_registry, *, local_grains, **kwargs): + def prepared_for(frame): + """The pool with the fixture's two national measures materialized.""" tables = {e: frame.table(e).copy() for e in frame.entities} tables["household"]["test_ones"] = 1.0 tables["household"]["test_london"] = ( tables["household"]["region"] == "LONDON" ).astype(float) - prepared = Frame( + return Frame( tables, frame.schema, {"household": frame.weights_for("household")}, @@ -184,12 +186,22 @@ def measures(frame, national_registry, *, local_grains, **kwargs): mass_log=frame.mass_log, metadata=frame.metadata, ) + + def measures(frame, national_registry, *, local_grains, **kwargs): + prepared = prepared_for(frame) + rows = UKRowwiseNationalRows( + national_registry.to_target_set(), national_registry, ("fixture",) + ) + # The kernel receives the compiled national problem, as the per-block + # resolver returns it; the fixture compiles its two columns here. + problem = ( + build_constraint_matrix(prepared, rows.targets, "household") + if len(rows.targets) + else None + ) return ( - prepared, - lambda _: frame, - UKRowwiseNationalRows( - national_registry.to_target_set(), national_registry, ("fixture",) - ), + problem, + rows, { g: pd.DataFrame( {"households": np.ones(frame.n("household"))}, @@ -200,13 +212,14 @@ def measures(frame, national_registry, *, local_grains, **kwargs): {"fixture": True}, ) - monkeypatch.setattr(graph_targets, "resolve_uk_full_measures", measures) + monkeypatch.setattr(graph_targets, "resolve_uk_full_national_problem", measures) return { "national": national, "local": local, "inputs": inputs, "surface": surface, "measures": measures, + "prepared": prepared_for, } diff --git a/tools/evaluate_uk_incumbent_surface.py b/tools/evaluate_uk_incumbent_surface.py index f217a6646..e8bea8efb 100644 --- a/tools/evaluate_uk_incumbent_surface.py +++ b/tools/evaluate_uk_incumbent_surface.py @@ -90,6 +90,31 @@ def _check_candidate_chronicle_identity( ) +def _area_code_columns(household: pd.DataFrame) -> dict[str, str]: + """The area-code column per ladder name, on either naming. + + Artifacts written since microcosm#1114 carry the consumers' names + (``constituency_code_oa``, ``la_code_oa``, ``region_code_oa``); earlier + graph exports carry the ladder names. Refuses a table with neither. + """ + + from microcosm.build.uk_runtime.geography_ladder import ( + UK_EXPORT_AREA_CODE_COLUMNS, + ) + + resolved = {} + for ladder, export in UK_EXPORT_AREA_CODE_COLUMNS.items(): + if export in household.columns: + resolved[ladder] = export + elif ladder in household.columns: + resolved[ladder] = ladder + else: + raise RuntimeError( + f"candidate household table carries neither {export!r} nor {ladder!r}." + ) + return resolved + + def main(argv=None) -> int: args = _parse_args(argv) artifact = load_ledger_consumer_artifact( @@ -251,10 +276,12 @@ def digest(path): household = frame.table("household") person = frame.table("person") hh_index = pd.Index(household["household_id"].to_numpy()) - assigned_region = household["region_code"].astype(str).map(GSS_REGION_CODES) + area_code_columns = _area_code_columns(household) + region_column = area_code_columns["region_code"] + assigned_region = household[region_column].astype(str).map(GSS_REGION_CODES) if assigned_region.isna().any(): unknown = sorted( - household["region_code"].astype(str)[assigned_region.isna()].unique() + household[region_column].astype(str)[assigned_region.isna()].unique() ) raise RuntimeError(f"unknown region codes on the household table: {unknown}") region_agree = ( @@ -287,8 +314,9 @@ def digest(path): rollup_receipt = { "rows": int(rollup_mask.sum()), "region_source": ( - "household.region_code (assigned area) mapped through GSS_REGION_CODES" + f"household.{region_column} (assigned area) mapped through GSS_REGION_CODES" ), + "area_code_columns": dict(area_code_columns), "frs_region_agreement_share": float(region_agree.mean()), "frs_region_agreement_weighted": float( weights[region_agree].sum() / weights.sum() @@ -300,7 +328,10 @@ def digest(path): print("evaluating the local surface ...", file=sys.stderr, flush=True) household = frame.table("household") - area_cols = {"constituency": "constituency_code", "la": "local_authority_code"} + area_cols = { + "constituency": area_code_columns["constituency_code"], + "la": area_code_columns["local_authority_code"], + } local_est: dict[tuple[str, str, str], float] = {} for grain, metrics in local_metrics.items(): codes = household[area_cols[grain]].astype(str).to_numpy()