Skip to content

Make the dense role run on a current spine: identity branch, scratch aliases, per-block national problem, compiled-row axis, holdout shape, engine blocks that represent the pool - #1115

Merged
juaristi22 merged 15 commits into
mainfrom
uk-identity-cgt-residential-branch
Oct 7, 2026
Merged

juaristi22 merged 15 commits into
mainfrom
uk-identity-cgt-residential-branch

Conversation

@juaristi22

@juaristi22 juaristi22 commented Oct 6, 2026 •

Copy link
Copy Markdown
Collaborator

Make the dense role run on a current spine. Ten commits, each answering a failure the first --release-role dense build from main's head after #1089 hit, in the order the build hit them. Measurement build (K=25, 2,000 epochs, 60,000 households) on the branch completed 2026-10-06; gate-blocked, staged for inspection only.

1. The identity kernel learns the residential-split branch

atomic_household_identity._BRANCHES gains cgt_residential_split between cgt_incidence_clone and geographic_support. The population node has keyed residential arms on that branch since microcosm#1063 item 7 landed the stage, but the kernel's roster was never extended, so household_draw_key refused every arm with "UK geography identity: unknown structural branch" 28 minutes into the build. Any current spine (2,735 residential arms in the 2026-10-04 spine) breaks main's dense role here; the national role assigns no geography, which is why #1063's final build and every #1095 arm passed.

2. Nation-native alias codes stay out of the engine's scratch dataset

Once the identity node passed, uk.full.measures refused: "Column 'data_zone_code' contains NaN values". The atomic path derives alias codes that are NA outside their own nation by design; the release export drops them at the boundary, but the scratch-mode UKMeasureResolver wrote the whole cloned frame to the engine's input H5 and policyengine-uk refuses a single-year dataset with any NaN column. One helper, without_uk_native_alias_columns, now serves both boundaries; the resolver's receipt lists what the engine copy left behind (engine_scratch_dropped_columns). The 28 September dense rungs ran --geography-assignment legacy, so the atomic dense path had never met real data.

3. Garbage collection after every engine block

With --engine-blocks 25 (one engine per clone; a single engine over the 1.6 million-household pool would need on the order of 100 GB) the measures loop retained every block's simulation: a policyengine simulation is a large cyclic object graph that del alone does not reclaim. The loop now collects after every block; a test counts one collection per block. Real, but not the cause of the kill that followed.

4. The national problem is compiled per engine block, never on the whole pool

With the engines released per block the build still died at a 140 GB footprint on a 24 GiB machine, three minutes after the last block: the measures node injected every national measure input and materialised all ~1,170 national targets as columns on the 1.6 million-household pool, copied the prepared pool to compile the national matrix and copied it again to check it was untouched. resolve_uk_full_national_problem now materialises and compiles per engine block and stitches the per-block CSR matrices into the pool's household order; resolve_uk_full_measures keeps its contract for the national seam and the evaluation tool, sharing the block loop, the metric rejoin, the provenance checks and the receipt. The measures receipt records national_materialization: per_engine_block. A test compiles the same fixture in one block and in two and compares the matrices column for column.

5. A compiled row's entity axis is one shared array

With the problem compiled, the solver sat on one core for an hour before its first epoch: _CompiledRow.__call__ rebuilt the frame's entity axis as a Python tuple on every compiled row, so compiling 22,053 rows over 1.6 million households cost ~3.5 × 10¹⁰ scalar conversions (about ten hours per compile, and the L0 search recompiles). to_target_set builds the axis array once and every row compares the id column vectorised; a different or reordered axis is refused as before, and a test pins that the guard never iterates the id column. Shared microcosm-calibrate package: the US consumers of decoded problems were run too (800 tests green).

6. Test fix for 5

The axis test's foreign frame is refused on a valid (ascending) axis; the first version of the test had been committed with its exit code masked by a pipe.

7 and 8. The terminal gates receive the holdout in the schema's shape, and the schema admits the graph's holdout fields

With the solve complete, the terminal gate kernel handed the calibration diagnostics the graph's holdout artifact as is. That artifact carries the kernel's own graph_binding and, under --skip-holdout, a reason; the diagnostics records forbid undeclared fields, so the battery refused with nineteen validation errors and the driver met a bare KeyError 28 hours in. Commit 7 has graph_terminal.diagnostics_rotated_holdout pass the schema's shape (a skipped holdout as the bare marker, a measured one without its binding); commit 8 lets UKSkippedRotatedHoldout carry optional report_only, reason and graph_binding, and UKMeasuredRotatedHoldout an optional graph_binding, so the artifact can also be stored as written. The build resumed from the content store with commit 8 alone.

9. Engine blocks represent the pool: weight-share formulas measured exactly per block

Why the release posture refused --engine-blocks > 1: a per-clone block carries a K-th of the pool, and policyengine-uk formulas that allocate a national total by weighted share (corporate_land_value = corporate_sector_wealth × weight / Σ(corporate_sector_wealth × weight) × ONS aggregate; shareholding and consumption_share have the same shape) see only the block's denominator, so every block reproduced the whole aggregate (the K-times land-value artefact of the #736 erratum; the K=25 measurement build's ONS land rows at +2023 % / +589 % are exactly this). The measures loop now hands each block's resolver an engine weight scale of pool mass / block mass: the scratch frame the engine loads carries the scaled weights with a declared mass-log record, while the resolver's own frame, the per-block constraint matrix and the local metrics keep the block's true mass. For blocks that are identical copies of one another (the clone expansion's contract; verified on the K=25 pool: 25 × 63,827 households, identical weights per spine household) every weighted sum equals the pool's and the formulas are exact. _engine_population_representation records the factor per block and the copy checks (counts, and the multiset of source household and weight) in the measures receipt, which drops its block-sensitivity caveat when they pass. release_verdict gains engine_population_exact (single block, or a per-block run with an exact representation); the dense release-candidate posture no longer refuses --engine-blocks > 1; the terminal reads the representation from the problem bindings; the size parity check accepts an exact per-block resolution. A real-engine test shows a quarter-mass block alone reporting four times the pool's corporate land value and the same block scaled by four reproducing it, with household_land_value untouched either way. Both verdicts are covered end to end through the synthetic measures evidence. Still to do on the data side: re-measure the K=25 build's land rows in a single block (or rebuild with this commit) before any ruling on them.

10. The dense battery evaluates the register gates the way the certifier does; the allow-list names the graph export

Gate triage of the K=25 measurement build found three battery defects and one stale register, none of them findings about the data. (a) uk_release_input_coverage failed all 17 required families on weight kind alone: the binding reads the build-state half from a spine_frame artifact the certifier supplies (--spine-h5) and the dense battery never did, so it fell back to the calibrated frame. The spine checkpoint now publishes a spine_build_state artifact (household weight kind, period, mass log: everything that half reads), the terminal gate node consumes it when the bound checkpoint declares it, and the binding accepts either a spine Frame or that surface. (b) uk_degenerate_release_surface reported person.incapacity_benefit_reported, a column the release boundary drops before writing and the certifier never sees: the gate now skips UK_RELEASE_EXPORT_DROPPED_COLUMNS and records them (dropped_at_export). (c) uk_cgt_projection_entrants had no evidence: the national calibration kernel computes the projection and the dense battery did not; build_full_gate_context now computes it when the caller has not. (d) uk_export_surface refused ten columns because the allow-list named the rowwise tool's export (household.clone_index, the _oa codes); it now names the graph export's per-entity clone indices, plain constituency / local-authority / region codes, and the ITL1–3 and ward codes; the three gate-battery vintage pins move with the spec. household.atomic_area_basis (constant by construction) and the threshold-hugging charitable_investment_gifts tail exclusion are left for María's register rulings. The commit also removes experiments/1063-uk-spine-followups.md.

11. The three area codes leave under the consumers' names (closes #1114)

policyengine.py's constituency and local-authority filters and impact outputs, the simulation API's geographic reports and the enhanced FRS all read constituency_code_oa, la_code_oa and region_code_oa; the graph export wrote the ladder names, so every local-area run on a Microcosm file failed on a missing column. Rather than carry both names, the single-year export boundary (graph_terminal._tables, uk_release_export_frame) renames these three and only these three; the graph keeps the ladder names everywhere in memory (gates, diagnostics, target compilation, the engine's scratch dataset), so no stored frame or node identity moves. The geography ladder gate runs on the ladder-named tables before the rename; uk_export_candidate_columns translates the names so the export-surface gate compares what the artifact carries; the allow-list names the _oa columns again (reverting that part of commit 10); the export read-back refuses a ladder name or an empty consumer column; the export descriptor and the rowwise candidate manifest record area_codes (the columns and, per support system, the code frames behind them: 2024_pcon for Westminster constituencies, 2023_april_lad / 2019_council_area / 2014_lgd for the April 2023 local-authority code set, 2024_rgn for the English regions with the FRS sentinel codes elsewhere), read from the committed support provenance rather than restated. The incumbent-surface evaluator accepts either naming so earlier exports stay measurable. The national line carries region only and is unchanged. The gate-battery vintage pins and the release-cut part digests move with the spec.

12. The size selection's budget search can start from a known penalty

--selection-initial-lambda (through select_uk_dataset_size(initial_lambda=...) and the size-search node's initial_lambda parameter, recorded in the run parameters and the search receipt) hands the shared L0 budget search a penalty to probe first. The K=25 build spent seven probes of about two hours each bisecting from the global bracket to λ = 1.15e-6; a re-selection on the same pool and targets probes that value first and stops there when the draw is feasible and within tolerance. A miss is followed by one probe half a decade away on the side it steered to, then interpolation between the measured ends (regula falsi with a progress safeguard, aiming at the middle of the draw's window) instead of re-bisecting a decade down to a window a few hundred rows wide; a miss beyond the span falls back to the global bracket on that side. The search still verifies every probe, so a stale hint costs probes, never feasibility (any penalty whose draw lands inside the budget window is a valid stop, so a warm and a cold search can settle on different penalties and select different rows), and the cold path is byte-for-byte unchanged. This is what the re-measure now running uses, and what every selection experiment on the stored checkpoint will use.

13. Review round 1 (vahid-ahmadi, at 0afabc1)

(1) A per-block engine resolution is exact only when the blocks are identical copies with a source identity and the allocation keys of policyengine-uk's weight-share formulas carry the same sum(x · w) in every block. The formulas and their frame inputs are named in UK_WEIGHT_SHARE_FORMULA_INPUTS (corporate land value and shareholding allocate by corporate_sector_wealth, which adds corporate_wealth and private_pension_wealth; the consumption share by the engine's consumption, the twelve LCFS category columns); the representation check records each block's sum per formula and a missing key column leaves the formula unverified and the resolution inexact; the measures receipt lists the formulas and states the copied-input assumption, and the release verdict's docstring names it. A formula of this shape that reads a per-clone input has to be added to the constant to be seen, never assumed exact. (2) A weights-only comparison (no source identity on the frame) is never exact; the receipt records source_identity_present. (3) The release_input_coverage binding requires the spine build state: its artifact selector marks the gate evidence_absent without it, and the evaluator refuses rather than judging the calibrated frame as a backstop; the certifier supplies --spine-h5 and a graph build the checkpoint's spine_build_state. (4) One mechanism for the holdout shape: the terminal node's converter stays (commit 7) and the schema loosening (commit 8) is withdrawn, so the shared diagnostics schema keeps its undeclared-field guard; the raw graph artifact is refused by the schema and converted before the diagnostics see it. On the note about commit 11: agreed that the next national cut's release notes should say the local-area exports carry the _oa names; one correction, the certified national release (aa31bdf6) carries region only and no area codes under either name, so policyengine.py's local-area paths need the local dataset, not a renamed national one. The branch is rebased on main (through #1129), linear, no merge commits.

14. Review round 2 and the pins the solver edit moved

Round 2 (at 6eab007) found nothing blocking and one optional nit: UK_WEIGHT_SHARE_FORMULA_INPUTS mirrors engine inputs by hand. test_uk_weight_share_formula_inputs.py (engine shard) now reads policyengine-uk's own allocation keys (corporate_sector_wealth and consumption, through their adds) and holds the constant to them, checks each key column is an engine input of the household entity, and checks every named column is a release household column (an enhanced-FRS input or a reviewed extra of the export surface), so an engine bump that changes a key fails the test. The warm-start wording now says a stale hint costs probes, never feasibility: a warm and a cold search may settle on different penalties inside the budget window and select different rows. The warm start edited microcosm.calibrate.solve, which three pinned surfaces hash: the calibrate graph parity fixture is re-pinned on every platform (tools/graph_parity_repin.py calibrate), the spec-engine loader golden vector and the US seed-protocol and seed-map digests carry the new values, and docs/evidence/spec-engine/us-f0-coverage.json is regenerated. Locally green on the spec-engine, graph-parity and US coverage suites (170), the engine tests and the UK measures tests.

What the measurement build exercised

Commits 1–4 and 8 ran in the completed build (resume tree = commit 4 + commit 8). Commits 5 and 7 are covered by tests and were not exercised in a completed build: adding them to the resume tree would have re-keyed the calibration and gate nodes and forced the 19-hour re-solve. Commits 9 and 10 (the engine-block representation; the battery and allow-list fixes) are not exercised by the completed build either. A re-measure on commit 10 ran its spine and its dense solve and is being restarted from its content store on commit 10 plus commit 12 (the warm start), which reuses the geography, measures and problem nodes and re-solves the dense problem; commit 11 and the review fixes of commit 13 touch modules the pre-solve nodes hash, so they are not on the measurement tree and will be exercised by the next build. Still expected on this PR as the gate triage proceeds: atomic_area_basis, the two stale exclusions, the tail-concentration margin, the input-mass reference.

Tests

  • test_uk_atomic_identity_graph.py, test_uk_atomic_household_lineage.py: residential arms key their declared path at K=1 and K=2; a flag/index disagreement is still refused; the arm is accepted between the incidence clone and the pool and refused before the incidence clone.
  • test_uk_measure_simulation.py: a validated toy frame with NA alias codes writes a scratch H5 without them and without NaN; the resolver's own frame is untouched; the receipt names the dropped columns.
  • test_uk_full_measure.py: one collection per block; one-block and two-block compilation agree column for column.
  • test_ordered_artifacts.py: the axis guard never iterates the id column; a foreign or reordered axis is refused.
  • test_uk_graph_terminal.py: skipped and measured holdouts reach the diagnostics in the schema's shape; the schema accepts the graph's fields.
  • test_uk_graph_terminal.py, test_uk_terminal_gates.py, test_uk_atomic_area_support.py: the export carries the _oa names with the ladder values and never the ladder names; the read-back refuses a ladder name, a missing or an empty consumer column; a stale _oa column on the frame is refused; the candidate-column surface and the national export frame translate; the recorded frames equal the committed support provenance. Touched surface: 888 build-shard and 306 data-shard tests green.
  • test_informed_gates.py: a warm hit ends the search after one probe with the cold result's weights; a miss within the span probes the edge then interpolates and settles within four probes; a miss beyond the span recovers on the global bracket; the cold path's first probe and receipt are unchanged. test_uk_full_calibration_graph.py: the hint reaches the search as its penalty, is recorded on the selection and refused when invalid.
  • test_uk_full_measure.py: identical multisets with different allocation keys are inexact (per-block sums recorded), no identity is never exact, a missing key column leaves the formula unverified; the real per-clone test carries the keys. test_uk_release_input_coverage.py, test_uk_battery_bindings.py: the binding is evidence absent without the spine build state and refuses to judge the calibrated frame. test_uk_graph_terminal.py: the raw holdout artifact is refused by the strict schema and the converter's shape validates.

Not in this PR

The K=25 / 60k candidate itself (a measurement build, not a release candidate) and its gate triage.

🤖 Generated with Claude Code

@juaristi22 juaristi22 changed the title Make the dense role run on a current spine: identity branch, scratch aliases, per-block national problem, compiled-row axis, holdout shape Make the dense role run on a current spine: identity branch, scratch aliases, per-block national problem, compiled-row axis, holdout shape, engine blocks that represent the pool Oct 6, 2026
@juaristi22
juaristi22 marked this pull request as ready for review October 7, 2026 10:17
@vahid-ahmadi

Copy link
Copy Markdown
Contributor

Automated review pass (Claude Code, high effort) — round 1 at 0afabc18

Verdict: nothing blocking in the code. Each of the eleven commits fixes the failure it names, and the targeted tests pass. By the description's own account it isn't finished: commits 9 and 10 haven't run in a completed build, a re-measure on commit 10 is queued, and more gate-triage items are still expected here. Two should-fixes and two nits below.

State: CI is green on 9 of 10 checks (engine-us 3.14 was still running). The branch is one merge behind main (#1081). I merged main in locally: no conflicts, the coverage manifest --check is current, ci_test_plan verify passes, and the spine, coverage, gate-battery, contract and FRS tests all pass on the merged tree. The touched UK test files plus test_ordered_artifacts.py all pass, including test_uk_national_graph.py.

Findings

  1. Should-fix: "exact" engine blocks are judged by source household and weight, not by the formula inputs (full_measure.py, _engine_population_representation). Scaling each block's engine weights by pool mass ÷ block mass makes a weight-share formula (x·w / Σ x·w × aggregate) exact only if every block has the same Σ x·w. The check confirms that blocks carry identical (source household, weight) multisets. It doesn't check that x is identical across blocks.
    • That holds today for corporate_land_value, shareholding and consumption_share, whose inputs are spine values copied into every clone.
    • It breaks silently if a future formula of that shape reads anything drawn per clone: the clone's geography, or a draw keyed on its identity.
    • Yet exact: true is what now lets a multi-block run be releasable (release_verdict, rowwise_cli.py:799-824).
    • Fix: record Σ x·w per block for the known weight-share formulas and require them to match, or name those formulas and the copied-input assumption in the receipt and the release verdict.
    • Commit 9 also isn't exercised by the completed measurement build. The K=25 land rows at +2023% / +589% are from before it, so the re-measure on commit 10 is the first real test.
  2. Should-fix (no fallbacks): the coverage gate still evaluates the calibrated frame when no spine build state arrives. release_input_coverage.py:1166 uses frame if build_state_frame is None. This PR supplies the spine build state when the bound checkpoint declares spine_build_state. A checkpoint that doesn't, such as an older store, still drops back to the calibrated frame. That's what failed all 17 families on weight kind. The receipt records family_build_state_frame: "release", but the gate's verdict is about the wrong frame.
    • Since the dense battery always runs from a graph checkpoint, refuse there when the spine build state is missing, instead of evaluating the release frame.
  3. Nit (fallback): identity_key = "weights_only" (_engine_population_representation). When none of household_source_id, source_household_id or source_household_key is present, blocks are compared on weights alone and can still report exact. Refuse instead of downgrading, or have exact require a source identity.
  4. Nit: two mechanisms for one holdout shape. Commit 7 converts the graph's holdout to the schema's shape. Commit 8 loosens the shared microcosm-diagnostics schema to admit graph_binding, reason and report_only as optional fields. One is enough. If the stored artifact needs to validate as written, keep 8 and say so in the schema docstring. Otherwise drop the loosening, which weakens the shared schema's undeclared-field guard for every country.

Questions and notes

Checked and correct

  • Scratch alias drop: without_uk_native_alias_columns serves both the export and the engine's scratch H5. The resolver's own frame is untouched, and the receipt lists the dropped columns.
  • Garbage collection: one gc.collect() per engine block, 25 per build, which is a negligible cost against an engine simulation.
  • Per-block national problem: the matrices compiled in one block and in two agree column for column, and the receipt records national_materialization: per_engine_block. The national role (graph_national.py) isn't touched, and test_uk_national_graph.py passes.
  • Compiled-row axis guard: a one-off int64/object array with np.array_equal. It's semantically the same as the old tuple comparison for integer and string ids, and a foreign or reordered axis is still refused.
  • Gate fixes: the degenerate-surface gate skips and records UK_RELEASE_EXPORT_DROPPED_COLUMNS, the CGT projection evidence is computed for the dense battery, and the export allow-list follows the graph export.
  • Tests: each fix has a test, and the fixture tests pin each failing case.

@vahid-ahmadi

Copy link
Copy Markdown
Contributor

Automated review pass (Claude Code, high effort) — round 2 at 6eab0077

Verdict: the code is approvable. All four round-1 items are closed in 0e60d164, and the new warm start in 07e80c47 is sound. What keeps this short of merge is your own description: the measurement tree doesn't yet include commits 9–13. The description says more gate-triage items are expected, and the re-measure on commit 10 plus 12 is still running. I'd approve once that re-measure is in and the remaining triage items have either landed here or been split out. CI was still running at review time (lint and select-countries green).

Round-1 item Status Evidence
1. "Exact" judged by source household and weight, not formula inputs Closed UK_WEIGHT_SHARE_FORMULA_INPUTS (full_measure.py) names each weight-share formula's allocation-key columns. _engine_population_representation records Σ x·w per block and requires agreement to 1e-9; a missing key column leaves the result inexact. test_uk_full_measure.py builds blocks with identical (source household, weight) multisets but differing inputs: identical_source_weight_multisets is true, weight_share_inputs_match false, exact false. The receipt and the release verdict state the copied-input assumption.
2. Coverage gate falls back to the calibrated frame Closed _coverage_required_artifacts (battery_bindings.py) requires spine_frame for evaluation, so the gate is evidence_absent without it. The evaluator also raises, as a backstop, instead of reading the calibrated frame. Tests check both the required-artifact set and the manifest_current preflight exception.
3. Weights-only fallback could report exact Closed identity_key is None without a source-identity column, source_identity_present is recorded, and the result is never exact; tested.
4. Two mechanisms for the holdout shape Closed The commit-7 converter stays and the commit-8 loosening of the shared diagnostics schema is withdrawn (schema.py: graph_binding and the undeclared-field allowance removed).
Note: commit 11 column rename Answered The description now records the _oa export names and the per-system code frames. Your correction stands: the certified national release carries region only, so local-area runs need the local dataset.

Warm start (07e80c47), checked:

  • The cold path is unchanged: with no initial_lambda, the search still starts at the mid-point and bisects. The new branches only run when warm.
  • Determinism holds per parameter. initial_lambda is a parameter of the size-search node and recorded in the receipt, so a warm and a cold run get different node keys and neither reuses the other's result.
  • It can select a different, equally valid penalty from a cold search. Any penalty whose draw lands inside the budget window is accepted, so warm and cold can stop at different λ inside that window and pick different rows. That is fine as the search contract stands, but the description's "a stale hint costs probes, never correctness" should say "never feasibility": the selected dataset can differ from a cold run's.
  • The regula falsi step is clamped to [0.1, 0.9] of the bracket, so it always makes progress, and a miss beyond half a decade falls back to the global bracket. Both new tests pass (probe the hint first, bracket a miss, refuse a bad hint).

Checks run locally at this head:

  • test_uk_full_measure.py, test_uk_release_input_coverage.py, test_uk_battery_bindings.py, test_uk_graph_terminal.py and test_uk_full_calibration_graph.py: all pass.
  • test_informed_gates.py (microcosm-calibrate): passes.

One nit, optional: UK_WEIGHT_SHARE_FORMULA_INPUTS mirrors policyengine-uk formula inputs by hand. A test that reads the engine's formula dependencies (or at least asserts the named columns exist on the spine) would catch the list going stale after an engine bump such as #1121's move to 2.122.2.

juaristi22 and others added 13 commits October 7, 2026 13:33
The population node has keyed residential arms (microcosm#1063 item 7) on a
`cgt_residential_split` branch since the stage landed, but the kernel's
structural-branch roster was never extended, so `--release-role dense`
refused every spine carrying a residential arm with "unknown structural
branch" (the main-head spine of 2026-10-04 carried 2,735 arms). The roster
gains the branch between the incidence clone and the geographic pool; arms
key their declared path and every other household's key is unchanged.

Tests: residential arms through the identity node at K=1 and K=2, the
flag/index disagreement refusal, and the projector's branch order (accepted
after the incidence clone, refused before it).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…asure engine loads

The atomic geography path (microcosm#931) derives nation-native alias codes
(`data_zone_code` and its siblings) that are NA outside their own nation by
design; the release export drops them at the boundary, but the scratch-mode
measure resolver wrote the whole cloned frame to the engine's input H5, and
policyengine-uk refuses a single-year dataset with any NaN column. Every
`--release-role dense` build under the default atomic assignment therefore
failed at `uk.full.measures` ("Column 'data_zone_code' contains NaN values");
the 28 Sept dense rungs ran the legacy assignment and never met it.

One helper, `without_uk_native_alias_columns`, now serves the export boundary
and the resolver's scratch export; the resolver's own frame is untouched and
its receipt lists the columns the engine copy left behind.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
A policyengine simulation is a large cyclic object graph that `del` alone
does not reclaim. With `--engine-blocks 25` the measures node constructs one
engine per clone block in sequence; the first K=25 dense build (2026-10-05)
retained every block's engine, reached a 140 GB memory footprint on a 24 GiB
machine and was killed by the system before the solve started. The loop now
drops the block's resolution and resolver and runs a full collection before
the next block loads; a probe on a 3,200-household block showed each
simulation leaving ~390,000 cyclic objects behind until collected.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…the whole pool

The measures node materialized every national target as a column on the
K-clone pool, then copied the prepared pool to compile the national matrix
and again to check the pool was left untouched. On a 1.6 million-household
pool (K=25 on the current spine, ~1,170 national targets) that step needs
tens of gigabytes: the first K=25 dense build (2026-10-05) reached a 140 GB
footprint on a 24 GiB machine and was killed before the solve, twice.

`resolve_uk_full_national_problem` resolves each engine block, materializes
the national targets on that block while its engine is alive, compiles the
block's columns of the national matrix, checks the block came back untouched
and releases it; the per-block matrices are stitched into the pool's
household order with the pool's weights as the start. Materialization is
row-wise (band edges come from the compiled register), so the stitched
problem is the pool's problem; a test compiles the same fixture in one block
and in two and compares the matrices column for column. The measures receipt
records `national_materialization: per_engine_block`.

`resolve_uk_full_measures` keeps its contract for the national seam and the
evaluation tools; both share the block loop, the metric rejoin, the
provenance checks and the receipt.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ow Python tuple

`_CompiledRow.__call__` rebuilt the frame's entity axis as a Python tuple on
every call to compare it with the decoded problem's axis. The solver compiles
every target through that measure, so on the first K=25 UK dense build
(2026-10-05) 22,053 local rows over a 1.6 million-household pool cost ~3.5e10
scalar conversions, roughly ten hours per compile, with the solver pinned on
one core before its first epoch. `to_target_set` now builds the axis array
once and every row compares the frame's id column against it vectorised; a
different or reordered axis is refused as before. A test pins that the guard
never iterates the id column.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The negative case handed the guard a descending household axis, which the
frame constructor refuses before the guard runs; it now presents a valid
ascending axis of the same length that differs from the compiled one.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The `uk.full.holdout` artifact carries the kernel's own binding
(`graph_binding`) and, when the holdout was skipped, the reason. The
diagnostics records forbid undeclared fields and state a skipped holdout as
the bare `{"skipped": true}` marker, so the terminal gate kernel's call
refused: the first K=25 dense build (2026-10-06, `--skip-holdout`) reached its
terminal gates after 28 hours and the battery recorded nineteen validation
errors, leaving the driver a bare KeyError for the missing diagnostics. The
kernel now passes the schema's shape and the artifact keeps the full report;
the solve stayed in the content store and resumes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The graph's `uk.full.holdout` artifact carries its report-only flag, the
reason a holdout was skipped and the kernel's binding to its inputs. The
skipped and measured holdout records now declare those as optional fields,
so the artifact validates as emitted; the bare marker stays valid and
undeclared fields are still refused.

The kernel-side coercion stays as the primary fix. This schema-side one
exists because the module hosting the gate kernels is hashed into the
preflight gate node, whose report the dense calibration node consumes, so a
change there re-keys the stored solve; the schema is hashed by no kernel,
which lets the first K=25 build resume from its content store.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…s measure the pool

A per-clone engine block carries a K-th of the pool. policyengine-uk formulas
that allocate a national total by weighted share (corporate_land_value:
corporate_sector_wealth * weight / sum(corporate_sector_wealth * weight);
shareholding and consumption_share have the same shape) see only the block's
denominator, so every block reproduced the whole aggregate: the K-times
land-value artefact behind the #736 erratum, and the +2023 % / +589 % ONS
land rows of the K=25 measurement build.

The measures loop now hands each block's resolver an engine weight scale of
pool mass / block mass. The scratch frame the engine loads carries the scaled
household weights with a declared mass-log record; the resolver's own frame,
the per-block constraint matrix and the local metrics keep the block's true
mass. For blocks that are identical copies of one another (the clone
expansion's contract) every weighted sum then equals the pool's and the
formulas are exact; _engine_population_representation records the factor per
block and the copy checks (counts, and the multiset of source household and
weight), and the receipt drops its block-sensitivity caveat when they pass.
A real-engine test shows the quarter-mass block alone reporting four times the
pool's corporate land value and the scaled block reproducing it, with
household_land_value untouched either way.

release_verdict gains engine_population_exact (true for a single block, or
for a per-block run whose receipt says the representation was exact); the
dense release-candidate posture no longer refuses --engine-blocks greater
than one; the terminal reads the representation from the problem bindings;
the size parity check accepts an exact per-block resolution. The end-to-end
CLI tests cover both verdicts through the synthetic measures evidence.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…s; allow the graph export's columns

Gate triage of the K=25 measurement build found three battery defects and one
stale register, none of them findings about the data.

uk_release_input_coverage failed all 17 required families on weight kind: the
binding reads the family build-state half from a spine_frame artifact the
certifier supplies (--spine-h5) and the dense battery never did, so it fell
back to the calibrated release frame. The spine checkpoint now publishes a
spine_build_state artifact (household weight kind, period, mass log: all that
half reads), the terminal gate node consumes it when the bound checkpoint
declares it, and the binding accepts a spine Frame or that surface.

uk_degenerate_release_surface reported person.incapacity_benefit_reported, a
column the release boundary drops before writing and the certifier never sees
on the exported H5: the gate skips UK_RELEASE_EXPORT_DROPPED_COLUMNS and
records them as dropped_at_export.

uk_cgt_projection_entrants had no evidence: the national calibration kernel
computes the projection and the dense battery did not; build_full_gate_context
computes it when the caller has not supplied one.

uk_export_surface refused ten columns because the allow-list named the rowwise
tool's export (household.clone_index, the _oa codes). It now names the graph
export's per-entity clone indices, the plain constituency, local-authority and
region codes, and the ITL1-3 and ward codes the atomic geography derives; the
runtime mirror and the gate-battery vintage pins (policy, manifest,
fingerprint, release-cut part digests) move with the spec.

Removes experiments/1063-uk-spine-followups.md.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…oundary

microcosm#1114: policyengine.py, the simulation API and the enhanced FRS read
constituency_code_oa, la_code_oa and region_code_oa. The graph keeps the
ladder names in memory; _tables and uk_release_export_frame rename the three
at the export boundary and keep no alias. The geography ladder gate runs on
the ladder-named tables; uk_export_candidate_columns translates the names so
the export-surface gate compares the artifact's surface; the allow-list names
the _oa columns; the read-back refuses a ladder name or an empty consumer
column; the export descriptor and the rowwise candidate manifest record the
area-code columns and the code frames behind them, read from the committed
support provenance. The incumbent-surface evaluator accepts either naming.
The gate-battery vintage pins and the release-cut part digests move with the
spec.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
--selection-initial-lambda (select_uk_dataset_size(initial_lambda=...), the
size-search node's initial_lambda parameter, recorded in the run parameters
and the search receipt) hands the shared search a penalty to probe first. A
hit ends the search after one probe instead of the seven a cold bisection
spent on the K=25 pool. A miss is followed by one probe half a decade away
on the side it steered to, then interpolation between the measured ends
(regula falsi with a progress safeguard) instead of re-bisecting the whole
bracket; a miss beyond the span falls back to the global bracket on that
side. The cold path is byte-for-byte unchanged.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…keys, the coverage gate never judges the calibrated frame, one holdout shape

(1) A per-block engine resolution is exact only when the blocks are
identical copies with a source identity and the allocation keys of
policyengine-uk's weight-share formulas (UK_WEIGHT_SHARE_FORMULA_INPUTS:
corporate land value and shareholding by corporate_sector_wealth,
consumption share by consumption) carry the same sum(x * w) in every block;
the receipt records the per-block sums and the formulas, and the verdict's
docstring names the assumption. (2) A weights-only comparison is never
exact. (3) The release_input_coverage binding requires the spine build
state: evidence absent without it, and a refusal rather than a verdict read
off the calibrated frame. (4) The diagnostics schema stays strict; the
terminal node's converter hands the holdout in the schema's shape and the
loosening that admitted the graph artifact's fields is withdrawn.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…inst the engine (review round 2)

The warm start edited microcosm.calibrate.solve, which the calibrate graph
parity kernel, the spec-engine loader golden vector and the US seed-protocol
and seed-map digests all hash: the calibrate parity fixture is re-pinned on
every platform, the loader golden and EXPECTED_HASHES carry the new digests,
and docs/evidence/spec-engine/us-f0-coverage.json is regenerated. Review
round 2's nit: an engine test reads policyengine-uk's own allocation keys
(corporate_sector_wealth and consumption, through their `adds`) and holds
UK_WEIGHT_SHARE_FORMULA_INPUTS to them, and checks the named columns are
release household columns, so an engine bump that changes a key fails the
test instead of leaving the per-block exactness check stale. The warm-start
docstring, help and changelog say "never feasibility", not "never
correctness": a warm and a cold search may settle on different penalties
inside the budget window.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@juaristi22
juaristi22 force-pushed the uk-identity-cgt-residential-branch branch from ede569c to 736719a Compare October 7, 2026 12:53
@juaristi22

Copy link
Copy Markdown
Collaborator Author

Thanks, all four landed in d2f8b19 (commit 13), with the warm-start commit 486ece1 (commit 12) just before it and the branch rebased on main (through #1129).

  1. Exact blocks verify their allocation keys. Agreed that identical (source, weight) multisets only prove exactness while x is copied into every clone. The representation check now names the weight-share formulas and their frame inputs in UK_WEIGHT_SHARE_FORMULA_INPUTS (corporate land value and shareholding allocate by corporate_sector_wealth = corporate_wealth + private_pension_wealth; the consumption share by consumption, the twelve LCFS category columns), records each block's sum(x · w) per formula, and requires them to agree (relative 1e-9) for exact; a key column a block lacks leaves the formula unverified and the resolution inexact. The measures receipt lists the formulas and states the copied-input assumption; release_verdict's docstring names it. A new formula of this shape that reads a per-clone input has to be added there to be seen. And yes, commit 9 is first exercised by the re-measure, which has now run its spine and dense solve and restarts from its store with the warm start.
  2. No fallback in the coverage binding. The binding's artifact selector now requires spine_frame outside the preflight check, so the battery marks the gate evidence_absent when no build state arrives, and the evaluator refuses rather than judging the calibrated frame as a backstop. The certifier supplies --spine-h5; a graph build the checkpoint's spine_build_state.
  3. Weights only is never exact. identity_key is None and source_identity_present false without a source identity, and exact requires one.
  4. One holdout mechanism. Nothing validates the raw graph artifact against the diagnostics schema (the manifest embeds it as written and the certifier checks its fields directly), so the loosening is withdrawn and the schema keeps its undeclared-field guard; the converter stays and the test now asserts the raw artifact is refused.

On commit 11: agreed the next national cut's release notes should say the local-area exports carry the _oa names. One correction: the certified national release (aa31bdf6) carries region only, no area codes under either name, so policyengine.py's local-area paths need the local dataset rather than a renamed national one; the national line is unchanged by the commit.


Round 2. The nit is in 736719a: test_uk_weight_share_formula_inputs.py (engine shard) reads the engine's own allocation keys (corporate_sector_wealth and consumption, through their adds) and holds UK_WEIGHT_SHARE_FORMULA_INPUTS to them, checks each key column is an engine input of the household entity, and checks every named column is a release household column (an enhanced-FRS input or a reviewed extra of the export surface), so an engine bump such as #1121's 2.122.2 fails the test rather than leaving the exactness check stale. Agreed on "never feasibility": the docstring, help and changelog now say a warm and a cold search may settle on different penalties inside the budget window and select different rows. The same commit re-cuts the pins the solver edit moved (the calibrate graph parity fixture on every platform, the loader golden vector, the US seed-protocol and seed-map digests and the regenerated coverage evidence), which is what CI was failing on.

The solver edit moved the US spec's seed-protocol digest, and with it the
resolved US spec_sha256 that the stacked pool tool binds into run_config.
Commit 14 re-cut the coverage evidence to f76f921d but left this test's
literal at main's d1df6b31, so engine-us failed on
test_constants_adapter_equals_live_constants_and_stays_out_of_identities.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@juaristi22
juaristi22 merged commit e3e3d88 into main Oct 7, 2026
19 of 20 checks passed
juaristi22 added a commit that referenced this pull request Oct 8, 2026
policyengine-uk defines property_wealth as residential_property_value
(main_residence_value + other_residential_property_value) plus
non_residential_property_value. The WAS stage saved the survey's own
total, which also counts owned land and the overseas and other-property
remainder; a persisted copy overrides the engine's sum, and the variable
has no uprating index, so the published aa31bdf6 year files hold it flat
from 2026 to 2030 while the sum grows 13% (the uk-data#543 defect).

- UK_RELEASE_EXPORT_DROPPED_COLUMNS gains household.property_wealth, and
  the measure resolver's scratch H5 applies the same drops in #1115's
  _engine_scratch_frame, beside the alias codes it already leaves out;
  engine_scratch_dropped_columns in the receipt names them. The spine
  keeps the column for the WAS stage's identity gate and donor checks.
- household.property_wealth is a reviewed export exclusion (constant and
  gates.json), and the coverage gate's two coverage halves read the frame
  the release writes, where the raw calibrated frame would report it as a
  stale exclusion; the build-state half keeps reading the spine (#1115).
- The coverage ledger gets a declared engine_derived_exclusions section,
  separate from the evidence-derived known gaps; the builder refuses
  drift, overlap with the gaps, and any column that is not a formula-
  owned override the reference persists. Manifest: 144 required, 1
  reviewed exclusion; the three components stay required.
- Tests: the scratch H5 and receipt, the export columns, the export-
  frame coverage binding, the household spine read for nonnegativity,
  the ledger and manifest, and an engine test that the engine's total is
  the component sum in 2024 and 2025 and that a persisted total would
  override it.

The gates.json change moves the release contract's UK gate-battery pins
(the full-manifest policy, manifest and fingerprint digests and the
release_cut part's two digests; microcosm-data contract.py and its test
mirror), re-derived from the live spec as the pin tests do.

The input-mass parity register entry (efrs-post-calibration) waits for
María's signature.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

2 participants