Repository navigation
Build the UK national release role on the graph (PR-2 of the #901 line) - #1057
Conversation
Program reviewBase repository: PolicyEngine/microcosm Source DocumentsNo source documents registered; see scope and validation. CriticalC1 — CRITICAL: published national completion hashes become stale (OPEN)Location: C2 — CRITICAL: repeated national build silently replaces a prior candidate (OPEN)Location: Should AddressA1 — SHOULD ADDRESS: operator interruption leaves no terminal Logbook disposition (OPEN)Location: SuggestionsNone. Evidence GapsNo material gaps reported. Notes
Validation SummaryLocal tests NOT RUN because the checkout has no installed project interpreter. CI at head 6954cb0 passed engine-free, engine-uk, integration-uk, lint, wheels and other jobs. Policy role compared the pinned target loader, exclusions, doctrine and seven gate scope with the prior seam. No external statutory source or target value changed. TimingNot measured. Review SeverityREQUEST_CHANGES. Open findings: 2 critical, 1 should address, 0 suggestions. |
|
Automated review pass (Claude Code, high effort) — round 1 at Verdict: the program review's C1 (build-record digest part) and C2 stand as blocking; I'd downgrade the rest of C1. My own findings add two should-fixes: a scratch path that now leaks into stored artifacts (it also hits the dense line), and a clock change that isn't listed under deviations. The rest are nits and one question. On the program review's C1 / C2 / A1 (checked at
|
…ition The certified national line (microcosm#823) runs through the calibration seam library, an in-process pipeline the graph driver dispatches to before preparing any graph. This commit composes the same computation as graph nodes over the bound spine checkpoint (uk_runtime/graph_national.py): - uk.full.national_targets compiles the national register from the pinned Chronicle artifact (exclusions, band edges, the frozen-register check, the feed pin with --allow-unpinned-feed); full_targets gains the seam's loader as load_uk_national_target_inputs. No ladder, no local surface. - uk.full.national_problem refuses non-finite inputs, resolves the engine measures on the bound frame (one block, no local grains; the seam's measure_resolver=None route stays as resolver_factory=None), materialises the register and encodes one ordered problem whose bindings carry the national doctrine's family_equal loss weights, mass reason and bounds. - uk.full.dense / uk.full.calibrated are the shared calibration nodes under the national doctrine's epochs, learning rate and seed. The dense kernel's admission is declared by the node: the dense line keeps its preflight battery (parameter absent, node keys unchanged); the national line admits the bound spine's provenance artifact, because it runs no pre-solve battery (its source and reference gates are the release-cut certifier's). - uk.full.gates.calibrated evaluates the seven UK_CALIBRATION_GATE_SCOPE gates on the calibrated population with the seam's evidence (the calibration manifest block, the parity rows, the admin anchors, the CGT projection) and stores the phase report, the manifest block, the diagnostic rows and the register census as artifacts. - A readback node over the written national H5 compares its content identity with the calibrated population. Two shared-UK adjustments the composition needed: the bound spine checkpoint keeps its H5's own time period instead of the CREATE normaliser's 2024 default (the projection fence declares the frame's period), and resolve_uk_full_measures records the materialisation report in its receipt (the seam manifest carries it). rowwise_cli's Logbook envelope takes the role's pipeline. Tests (test_uk_national_graph.py, 8) run the composed graph on the seam suite's synthetic frame, sidecar and register: the ordered problem is the seam's matrix row for row with the doctrine bound, the graph's weights equal an in-process calibrate call under the national doctrine, the calibration mass record is the seam's, and the gate battery evaluates the seam scope. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
microcosm-build-uk --release-role national no longer dispatches to the calibration seam library before preparing a graph: the driver prepares the national graph (uk_runtime/graph_national.py) over the bound checkpoint and materialises the seam-shaped evidence from the stored artifacts. - full_build_cli: prepare_national_build (the bound checkpoint, the doctrine with its receipted overrides, the register node with the request's Ledger pins), execute_national_build (the staged temporary directory and the published bundle, as the dense line), _execute_national_build (the checkpoint, the register frozen beside the outputs, the solve and the seam battery in the graph, the seam diagnostics writer with the build block and the digest measured from the file, the target-support sidecars, the battery replayed from the stored phase report through GateBatteryRun.record_phase and signed, write_uk_national_frame, the graph's readback of the written H5, the schema-1 build_record.json the release-cut certifier reads, the schema-4 national manifest), the national envelope (the seam's attempt id and Logbook pipeline, staging telemetry, the failure sidecar and the failed row) and _close_national_attempt (the staged bundle, the incumbent evaluation after staging, the record and manifest rewritten with the delivery summary, the sums). A blocked seam battery leaves the report on disk, writes no H5 and records a failed row. - national_role keeps what the role adds around the graph: the dry-run plan (now with the compiled graph's inventory), the doctrine overrides, the national manifest (over reported paths and the graph's artifact keys), the incumbent evaluation, the pin and output-path checks, the sums; its seam dispatch is gone and its private spellings stay aliases. - graph_national gains the driver-side projections: the decoded result rebound to the calibrated population, the battery replay, the seam run_config and the build record. Tests: test_uk_rowwise_national_role.py (21) runs the driver end to end on the seam suite's synthetic frame bound as a checkpoint and checks the seam shapes: the build record's bindings and provenance, the signed report, the diagnostics build block, the frozen registries, the manifest, the local and remote staging, the Logbook row on uk-frs-calibration, the doctrine overrides, the unpinned-feed refusal and override, the incumbent evaluation after staging, a blocked battery, and that the H5's weights are an in-process calibrate call under the national doctrine. The driver's dispatch tests follow the graph path. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
With the national release role on the graph, calibration_run's run_uk_calibration, its attempt runner, its gate battery runner and the two run dataclasses go, and national_calibration's UKNationalCalibrationStage (with its checkpoint resume and post-solve fence) goes. What both lines and the release-cut certifier still consume stays: the Logbook pipeline and attempt-id prefix, the four gate scopes and their closed-world exclusions, the bound spine sidecar and checkpoint authentication, the spine provenance blocks, the scoped gate manifest, the CGT projection and admin-anchor artifacts, the Ledger and runtime provenance blocks, the register census and doctrine payloads, the scoped-report signing, the frame adapter with its prepared/restore lifecycle, the measure-input injection pair, the canonical mass reason and the scoring route. The problem encoding is factored out of the national problem kernel as encode_uk_national_problem, so the seam suites re-anchor on the graph's national route as a library: the national-calibration suite solves the encoded problem under the doctrine's bindings and checks the evidence block (the fact-moving solve, the nested frame, the resolver injection and restore, the band-edge threading, the packaged binding classes, the materialisation skip refusal, the id preservation, the writer-clean tables, the evidence block's refusals, the doctrine's loss weights in the bindings); the seam suite keeps the checkpoint, scope, admin-anchor, provenance and Logbook-scope contracts and gains the band-edge reconciliation and attempt-id tests the runner tests carried; the CGT observation-period test's national route materialises through the library. The stage lifecycle, checkpoint-resume and runner-envelope tests retire with the code they exercised (the driver's national suite covers the envelope). Docs: the graph guide's national section and owner rows, the national calibration runbook, the changelog. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…national line The first licensed A/B (the seam on main 6a70cd4 vs the graph, the 10 % smoke spine) matched the solve exactly (the same 1,062 x 53,806 matrix, identical final weights, losses and diagnostics numbers, all seven gate outcomes and details equal, the same Logbook identity digest) and left two shape differences in calibration_diagnostics.json: the target descriptors (the graph solves the compiled rows, so each target read as a callable on the household entity where the seam's read as the spec's column measure on its own entity) and the measure_resolution block (the graph recorded the engine receipt where the seam recorded the resolution loop's own receipt: the attached routes, the rounds, the provider's receipt). The driver now hands the frozen register to national_result_from_manifest, which re-describes the result's targets from the specs (row names checked against the solved problem), and resolve_uk_full_measures carries the resolution loop's receipt per block in its receipt, from which the national evidence block takes the seam-shaped measure_resolution. The receipts file opens with the hermetic parts. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
R5: on the licensed 10 % smoke spine, the seam (main at the #901 merge) and the graph (this branch) produce the same 1,062 x 53,806 matrix, target vector, design and final weights to the bit, the same diagnostics outside the build block, the same seven gate outcomes and details (both blocked by the projection fence, this spine's property) and the same Logbook identity digest; 638 s vs 653 s wall. The one residual is the resolver's own description of where it read its input. R4: the sweeps on the final tree. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The retired seam validated the national role's output paths before any work (`_validate_output_paths` in `_run_national_attempt`) and raised FileExistsError on an existing candidate artifact. The graph driver stages into a temporary directory and `publish_staged_bundle` replaces whatever is at `--out`, so without that check a second run into the same directory silently replaced the first candidate's bytes while the first Logbook row still named them. `prepare_national_build` now runs the retained `national_role.validate_output_paths` on the role's output paths (dry runs excepted: they write nothing), so the refusal lands in `failure.json` and a failed row before the graph is compiled. Test: a second run into an occupied `--out`, chained on the first row, exits 1 with the FileExistsError, leaves the H5, build record and completion marker byte-identical and records a failed row. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… (review C1) `publish_staged_bundle` places `build.json` last, binding the staged bytes. The close step then appends the staging receipts to the rowwise manifest on both lines and, on the national line, the delivery summary to `build_record.json` (the national release assembler reads it there), so the marker's `build_record` digest was stale on every successful national run, and the manifest digest carried the note that receipts follow. `_refresh_completion` re-issues the marker after the close step with those files' final digests and drops the note, on both lines: the marker is the bundle's binding of what is on disk. Tests: the national end-to-end run with the incumbent evaluation checks the marker's build-record, manifest and dataset digests against the files; the dense driver run checks the manifest digest after the close step. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… (review A1) The seam wrapped the national attempt in `except BaseException`, recorded a KeyboardInterrupt as `disposition="discarded"` (every terminal disposition records a row) and re-raised. The graph driver caught only `Exception`, so a Ctrl-C during the solve left the attempt without a terminal row and the staging run open; the dense `main` had the same gap. Both envelopes now record the discarded row through `_record_failure` (which takes the disposition and, on the dense line, the prepared build for its output check), fail the staging telemetry and re-raise; `record_candidate_error` takes the disposition. Tests: an interrupt raised from the execute step on either line writes the failure sidecar naming it, spools one `discarded` row and propagates. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ew should-fix 1) A scratch-mode `UKMeasureResolver` records the file it simulated from (`source_path`), which lies under the per-run temporary directory the measure node creates. Since 2182268 the resolution loop's receipt, which carries that provider receipt, rides the measure receipt into the national problem's bindings and the dense `uk.full.measures` artifact, so an absolute scratch path changed a deterministic node's output between identical runs. `resolve_uk_full_measures` now records every path under the scratch root relative to it (`simulation-input.h5`, `clone-0/simulation-input.h5`), in both the resolution receipts and the resolver receipts. Test: two resolutions in different scratch directories with a resolver that records its scratch path yield equal receipts with no temporary directory in them. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…l tidy-ups (review nits 4, 5)
The national problem binds `mass_rule`, `scale_rule` and `l0_lambda`
as the doctrine the solve ran under, while `UKDenseSolveKernel` solves
under free mass, the default target-loss scales and no L0 penalty. The
values agree today (the doctrine dataclass allows nothing else), but a
problem binding another value would have been recorded without being
honoured; the node now refuses it (`refuse_unhonoured_solve_doctrine`).
In `graph_national`: the no-op `if …: pass` on the override rule is
gone; the gate node carries `gate_scope` from the posture and the kernel
evaluates that scope (the manifest digest still pins its compilation)
instead of re-reading the seam constant; `national_run_config` refuses
an empty national register instead of recording `{}`.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…932 review 1) `sources.yaml` and the provenance register pin each atomic-area support's sha256 and byte size, but nothing compared the supports a build used with them: `_prepare_geography` checked the bytes only against the digest pinned on the command line, and no release check read the manifest's `geography.assignment`. A rebuilt support therefore passed as a release candidate as long as the caller pinned its own digest. `country_adapter.uk_atomic_support_register` reads the register from the UK spec by system. Under `--release-candidate` each supplied support pin must equal its register row, and `tools/preflight_uk_local_release_candidate.py` refuses a manifest whose `geography.assignment` is not the atomic law or whose `support_pins` differ from the register (a pre-#932 manifest with no binding is refused like a pre-role candidate). The synthetic dense fixtures bind synthetic pins (`SYNTHETIC_SUPPORT_PINS`, the `RELEASE_PINS` digests) and the preflight and assembler suites register those; one preflight test reads the committed register and shows the synthetic pins refused on all three systems. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…view 2) `atomic-assignment-cells.json` suppressed `rows < 3` as `<3` but published `expected_rows` and `z` for the same cells, from which the review recovered every suppressed count exactly. The harness now blanks `z` wherever it suppresses `rows` (`_suppressed_area`), the purpose text says so, and the committed file is re-suppressed in place under that rule (985 z-scores blanked, no cell regenerated; sha256 32f494df… → 7d120d52…, recorded in the receipts). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…review 5) `_normalise_ew_lad_lookup` and `load_ew_oa_lad_region_lookup` accepted any `LAD??CD` column, so a `LAD25CD` frame (the Barnsley/Sheffield recode, a separate follow-up) was silently relabelled `lad23_code`, and a frame with two such columns failed with an AttributeError. Both now accept exactly `LAD23CD` or `LAD24CD` (the December 2024 OA lookup carries the April 2023 set under the latter), ignore other names and refuse a frame carrying both accepted names. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…review 4, 6, 8) The pool checkpoint's docstring said the platform scope "stops here"; it only avoids the typed-artifact refusal, since the shared assignment kernels' columns already reach the UK chain at platform scope through population slices (reworded; the question to Max stays open). The five nation-native alias names in the export allowlist are there because the full-build export-surface gate reads the graph frame's in-memory columns before `graph_terminal._tables` drops them (the F8 gap on #932); the allowlist now says so, and the names stay so the gate-register digests do not move again. The CLI test's "dispatched to the seam" comment is replaced, and the sample-admission test gains its sibling under the default atomic law. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…parison as evidence Records the review-date clock as a deviation from the seam (the dense convention, a node parameter), the occupied-output refusal, the interrupt disposition, the re-issued completion marker, the scratch-relative resolver path and the registered-support check in the graph doc and the #623 runbook; extends the changelog fragment; corrects the receipts' test counts (R2, R3), notes the committed comparison in R5 and adds R6 (round 1 and the #932 folds); commits R5's seam-vs-graph comparison under docs/evidence/uk-901-national-ab (aggregates and field-level differences only, machine paths omitted). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
6954cb0 to
63771af
Compare
|
Posting on María's call. Thanks for the round. The fixes are on the branch at
The #932 round-1 items are folded here on María's decision, linked from #932: a4308eb (finding 1), 11f4014 (2), c7057e9 (5), edb478f (4, 6, 8). Verification at this head is in the body (345 targeted; the whole engine-free uk directory plus the shared-spec lane 3,055 with the one environmental failure). |
Automated review pass (Claude Code, high effort) — round 2 at
|
| Item | Status | Evidence |
|---|---|---|
| C1 completion marker stale after close | Closed | _refresh_completion (full_build_cli.py:1080) re-issues build.json inside the finally of both close steps (:1133 dense, :2250 national), after the manifest, build-record and sha256sums rewrites. It is now the last write to the bundle; only the spool row follows, and that lives outside it. build.json is not a manifest output, so neither sha256sums.txt nor the staged bundle lists it, and nothing else goes stale. tools/assemble_uk_release_dir.py reads build_record.json, not the marker, so it is unaffected. |
| C2 second national build replaces the first | Closed | prepare_national_build calls national_role.validate_output_paths before the graph compiles (:1622), skipping dry runs, which write nothing. The new test leaves the first H5, build record and marker byte-identical and chains a failed row. One side effect is item 9 below. |
| A1 Ctrl-C leaves no terminal row | Closed, both lines | except KeyboardInterrupt in _national_main and in the dense main records a discarded row through _record_failure(..., disposition=...), fails the telemetry and re-raises. The dense prepared = None is set before the try, so an early interrupt can't hit a NameError. state.spool_path prevents a second row if the interrupt lands after the close step has written one. |
| 1 scratch path in stored artifacts | Closed | _without_scratch_paths (full_measure.py:36) rewrites any absolute path under the scratch root, recursively, in both the resolution and resolver receipts. I reran my probe on the national call shape (blocks=1, local_grains=()), adding a nested scratch path to the stub receipt. The evidence is now identical across two scratch directories, and the temp root no longer appears in it. That evidence is the only run-dependent input to the problem bindings, so problem.sha256 is stable too. The dense uk.full.measures artifact goes through the same function. |
2 exclusion clock follows --review-date |
Closed | It is listed under Deviations, in docs/uk-full-build-graph.md and in the #623 runbook. |
| 3 "node keys unchanged" | Closed | The body now says the dense node's implementation hash moves. |
| 4 unhonoured doctrine fields | Closed | refuse_unhonoured_solve_doctrine (graph_calibration.py) runs on the problem's bindings before the solve. The new test refuses non-default mass_rule, scale_rule and l0_lambda. |
5 graph_national tidy-ups |
Closed | The no-op if/pass is gone. gate_scope is now a node parameter read from the posture. run_config refuses an empty register (with a test). |
| 6 test counts | Closed | Collected counts are 23 / 18 / 2, plus 11 for test_uk_national_graph. |
| 7 R5 comparison in the tree | Closed | docs/evidence/uk-901-national-ab/compare_a_vs_b2.json holds aggregates and field diffs only. It contains no unit records, and a grep for /Users, /private, /tmp, /home and .h5 finds nothing. |
| #932-1 registered supports | Closed | Under --release-candidate, _prepare_geography now requires each pin to equal the sources.yaml row (uk_atomic_support_register, which is tested against the provenance resource). tools/preflight_uk_local_release_candidate.py refuses a manifest whose geography.assignment is not atomic or whose support_pins differ from the register; that path matches what graph_terminal.py:1452 writes. build_uk_country_graph(release_candidate=True) with a legacy config still compiles, but the preflight now catches the manifest it produces, which is the guard I asked for. |
| #932-2 z-scores beside suppressed counts | Partly | z is null on every suppressed row, in all four cells (474 / 495 / 10 / 6). The subtraction route below remains. |
| #932-5 LAD column | Closed | Only LAD23CD and LAD24CD are accepted. LAD25CD is refused as a missing column, and two accepted names are refused. |
| #932-4/6/8 docstring, allowlist comment, atomic-law admission test | Closed | The docstring now says only the typed-artifact edge stops, and the open question points at #932. The alias comment explains the F8 in-memory read. test_sample_admission_holds_under_the_default_atomic_law is added. |
New
8. should-fix (low risk): the per-level totals still reveal the suppressed K=15 counts.
- The file publishes zeros, so a suppressed count is 1 or 2.
- It also publishes each cell's total
rows, and each household sits in exactly one constituency and one LA. Within a level, the total minus the published counts is therefore the sum of the suppressed counts. - At K=15 that sum pins every suppressed value:
- legacy constituency: 5 suppressed, 10 rows;
- keyed constituency: 3 suppressed, 6 rows;
- keyed LA: 3 suppressed, 6 rows.
- So 11 of the 14 suppressed K=15 constituency and LA counts are exactly 2. Only legacy LA (5 cells, 8 rows) stays ambiguous.
- The keyed constituency codes whose count is recovered this way are
E14001175,E14001495andE14001518. levels.*.min_rows = 2.0andbreaches.constituency.rows.min = 2.0say the same thing independently. At K=1 the sums are too large to single any cell out.
The counts are synthetic geography draws, so the practical risk is small. But the file claims "counts below the minimum are suppressed together with their z-scores", and at K=15 that is not true. The cheapest fix is secondary suppression: drop rows at the cell level and the min_rows/min_sources/rows.min summaries wherever a level has suppressed cells. Alternatively, drop the claim and state why these counts don't need protecting. #1059 has the same pattern (a count under 10 recovered by subtraction), so a shared helper may be worth it.
9. nit: a refused second run writes into the candidate's own directory. The C2 refusal is raised inside the try, so _record_failure writes failure.json (with "release_authorized": false) and a logbook-spool row into the occupied --out. create_staging_telemetry has already opened a run there too: a local staging/runs/<id> marked failed, or, with remote staging, a failed run on the Hub. Nothing reads failure.json today (only full_build_cli.py and one US tool mention it), but a valid candidate directory now carries a failure sidecar from a different attempt. Running validate_output_paths in _national_main before telemetry opens would make the refusal leave no trace, though the failed Logbook row would then need a home. At minimum, the test could assert that the first run's files are the only candidate artifacts and that failure.json names the refusing attempt.
10. nit: the comparison script isn't in the tree. compare_a_vs_b2.json's purpose cites data/ukds/acceptance/901-national-ab/compare_national_ab.py, which lives in the licensed data tree. The script contains no data, so committing it beside the JSON, for example under tools/ or experiments/, would let R5 be regenerated by anyone with the licensed inputs.
…ppressed one (#1057 review round 2, item 8) Vahid's round 2 on #1057: every household row lies in exactly one area per level, so a level total less its published counts is the sum of the suppressed ones; with zeros published each suppressed count is 1 or 2, and at K=15 the residual pinned eleven of fourteen (all 2). Dropping the cell `rows` would not close it: the expected rows sum to the same total (and the integer pool counts with the public census shares fix it anyway). The harness's new `resuppress_evidence` pass: - withholds one published count per level that has suppressed counts (`withheld`, with its z-score, ESS and source count), chosen in a value-independent order (sha256 of level and area code) among counts above the minimum, so the residual carries an unknown of at least minimum + 1; - withholds the level minima of rows, ESS and sources there, the breach-table minima, and any breach quantile inside the suppressed range; - recomputes max |z| and the share above 3 over published areas only (at K=1 the reported LA maximum belonged to a suppressed area). `main` applies the pass before writing; the committed file is re-suppressed in place (no cell regenerated, sha 7d120d52... -> 99013cc1...). Tests: the committed file is a fixed point, each level with suppressed counts carries one complement and a non-degenerate feasible range for their sum, and two invented-level cases. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…iew round 2, item 9) The C2 refusal (#1057 round 1) ran inside prepare_national_build, after the attempt id was minted and telemetry opened, so a refused second run wrote failure.json (`release_authorized: false`) and a failed Logbook row into the first candidate's directory, and opened a failed staging run. The check moves to `refuse_occupied_national_output`, called in `_national_main` with the other argument refusals, before the staged- dataset pre-flight and before any attempt exists: a refused run writes nothing into the occupied directory. No Logbook row is owed, as for every refusal that precedes the attempt (the driver's standing rule: argument refusals cost nothing). The input's existence is still checked inside the attempt; the collision check needs only its path. Test: the second run raises FileExistsError and leaves every file under the first candidate's --out byte-identical, with no failure.json and one Logbook row. The runbook and the graph doc say so. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…nd 2, item 10) `compare_a_vs_b2.json` cited a script that lived only in the licensed data tree. The script (no data, no machine paths) is committed as docs/evidence/uk-901-national-ab/scripts/compare_national_ab.py, formatted under the evidence-scripts ruff waiver, so R5 can be regenerated by anyone with the licensed inputs; the JSON's purpose and the #1057 receipts point at it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Round 2 items 8–10 are folded into #1059 on María's decision, one commit each:
|
What this does
microcosm-build-uk --release-role nationalnow builds the certified national line on the shared executable graph instead of dispatching, before any graph is prepared, to the in-process calibration seam library. This is the second change on the #901 line, planned with #901's re-base (her ruling 2026-09-25: the national role through the retained seam engine in #901, the posture-driven graph national path as the next PR).The national posture composes (
uk_runtime/graph_national.py):uk.full.spine_checkpoint, the bound checkpoint as Register one UK full-build graph with all-geography calibration #901 admits it (sidecar, content identity, gate report), now keeping its H5's own time period.uk.full.national_targets, the national register compiled from the pinned Chronicle artifact with the measure exclusions, the band edges, the frozen-register check and the feed pin (--allow-unpinned-feedas a node parameter). No ladder, no local surface.uk.full.national_problem: refuse non-finite inputs, resolve the engine measures on the bound frame (one block, no local grains; the seam'smeasure_resolver=Noneroute stays asresolver_factory=None), materialise the register, compile the constraint matrix row for row, restore the pristine tables, and encode one ordered problem whose bindings carry the national doctrine'sfamily_equalloss weights, mass reason and bounds.uk.full.denseanduk.full.calibrated, the shared calibration nodes under the doctrine's epochs, learning rate and seed. The dense node's admission is declared by the node: the dense line keeps its source preflight battery (parameter absent; the dense node's implementation hash moves with its_admitsource and the bound checkpoint's withcalibration_runand itstime_periodparameter, souk.full.dense,uk.full.spine_checkpointand everything downstream get new keys; nothing committed pins them and no Register one UK full-build graph with all-geography calibration #901 graph store resumes across this change); the national line admits the bound spine's provenance artifact, because it runs no pre-solve battery (its source and reference gates are the release-cut certifier's).uk.full.gates.calibrated, the sevenUK_CALIBRATION_GATE_SCOPEgates evaluated on the calibrated population with the seam's evidence (the calibration manifest block, the parity rows, the admin anchors, the CGT projection), storing the phase report, the manifest block, the diagnostic rows and the register census.uk.full.national.readback, the written H5 read back and its content identity compared with the calibrated population.The driver (
full_build_cli) materialises the seam-shaped evidence from the stored artifacts: the diagnostics through the seam writer with the build block (digest measured from the file), the target-support sidecars, the calibration-seam battery replayed from the stored phase report throughGateBatteryRun.record_phase(#901's replay) and signed,microcosm_uk_2024_25.h5, the schema-1build_record.jsonthe release-cut certifier reads, the frozennational_target_registry.jsonandnational_contract_registry.json, the schema-4rowwise_candidate_manifest.json, the Logbook row on the seam'suk-frs-calibrationpipeline with the seam's attempt id, staging telemetry and the staged bundle, and the end-of-build incumbent evaluation after staging. A blocked battery leaves the report on disk (blocked_at_phase: "terminal"), writes no H5 and records a failed row.Retired:
calibration_run.run_uk_calibration, its attempt runner and battery runner,national_calibration.UKNationalCalibrationStage. Kept: the seam's contracts (gate scopes and closed-world exclusions, sidecar and checkpoint authentication, provenance blocks, scoped manifest, projection and admin-anchor artifacts, register census, doctrine payload, scoped-report signing), the frame adapter, the measure-input injection pair, the mass reason and the scoring route.Shared touches (country-agnostic, none)
None. Every change is under
uk_runtimeand its tests, plus the docs and changelog.New UK-side pieces (flagged)
uk_runtime/graph_national.py(the four kernels, the composition, the driver-side projections).full_build_cli.py(prepare_national_build,execute_national_build,_execute_national_build,_close_national_attempt,_national_main).full_targets.load_uk_national_target_inputs(the seam's loader as the target node runs it).time_periodparameter, the materialisation report in the measure receipt, and apipelineargument on the rowwise Logbook envelope.Deviations from the plan, flagged
uniformoverride the graph binds an explicit all-ones loss-weight vector where the seam passedNone; the solver computes a weighted sum over a mean, so the last bit may differ under that override only. The doctrine'sfamily_equalis exact by construction.--review-date(default today), the dense graph's convention and a node parameter, where the seam evaluated both at the run clock; a back-dated review date keeps expired exclusions and deferrals in force, which the seam never allowed. Documented indocs/uk-full-build-graph.mdand the Calibrate the UK national build from Ledger-backed targets #623 runbook (review round 1).Review round 1 (2026-09-30) and the #932 folds
Rebased over #932 as merged. One commit per finding: eb1704e C2 (an occupied
--outrefused before any work, the seam'sFileExistsError), cb0d390 C1 (build.jsonre-issued after the close step on both lines with the final build-record and manifest digests; the national assembler reads the delivery summary from the build record, so the record could not be left alone), 1e8e252 A1 (KeyboardInterruptrecords adiscardedrow on both lines and re-raises), 15bff5b should-fix 1 (scratch-relativesource_pathin the measure receipt, so the measure and national problem nodes are deterministic between runs), e61dd97 nits 4 and 5 (the dense solve refuses a boundmass_rule,scale_ruleorl0_lambdait does not honour; the no-op removed;gate_scopefrom the posture as a node parameter; an empty register refused inrun_config), 63771af docs, changelog, receipts R6 and nit 7 (R5's comparison committed underdocs/evidence/uk-901-national-ab/). Folded from #932's round 1 on María's decision: a4308eb finding 1 (country_adapter.uk_atomic_support_register;--release-candidaterequires thesources.yamlsupport rows, and the dense pre-flight refuses a manifest whosegeography.assignmentis not the atomic law on those pins), 11f4014 finding 2 (z-scores blanked beside suppressed counts indocs/evidence/uk-931/atomic-assignment-cells.json, sha 7d120d52…), c7057e9 finding 5 (LAD23CD/LAD24CDonly), edb478f findings 4, 6 and 8 (pool docstring, the alias-name comment with the names kept so the gate-register digests do not move, the CLI test comment, the sample-admission sibling under the atomic law).Verification (receipts R1–R4)
test_uk_national_graph.py(11): the composed graph on the seam suite's synthetic frame bound as a checkpoint; the ordered problem is the seam's matrix row for row with the doctrine bound; the graph's weights equal an in-processcalibratecall under the national doctrine; the calibration mass record is the seam's; the seven seam gates evaluate.test_uk_rowwise_national_role.py(23): the driver end to end on the same fixture, local and remote staging (fake Hub), the build record's bindings and provenance, the signed report, the diagnostics build block, the frozen registries, the manifest, the Logbook row onuk-frs-calibration, the doctrine overrides, the unpinned-feed refusal and override, the incumbent evaluation after staging, a blocked battery, the H5's weights equal to an in-process solve.test_uk_national_calibration.py26,test_uk_calibration_run.py18,test_uk_cgt_observation_period.py2; the driver's dispatch tests follow the graph path.test_uk_cgt_projection::test_engine_is_reported_unavailable_without_the_uk_extra) that asserts the uk extra is absent and fails identically on main in a venv that has it; the shared-spec lane (test_country_spec,test_spec_engine_country_bundles, the gate-register pins, the data contract,test_stage_evidence,test_gate_battery_replay) 432 passed;tools/ci_test_groups.py --verifyok;ruff checkclean; every edited file formatted.tests/engine_free/ukdirectory plus the shared-spec lane (test_country_spec, the gate-register pins, the data contract) 3,055 tests, 25 skipped, one failure (the environmental engine-absence test above);ruff check .clean;tools/ci_test_plan.py verifyok; gate-register digests unchanged.Parity (receipts R5)
Licensed A/B on the 10 % smoke spine of the #901 pre-merge run (53,806 households), the pinned Chronicle artifact, the doctrine with no solve flags: A = the seam on main at the #901 merge, B = this branch. Both sides blocked at the terminal battery on the same gate,
uk_cgt_projection_entrants, with the same numbers (a property of this spine under the engine's uprating, not of either path), so neither wrote the H5 or the build record. Before the battery, identical on both sides: the 1,062 × 53,806 constraint matrix, the target vector, the design weights, the final weights to the bit (max abs diff 0.0), the diagnostics file minus the build block (every target row, the UK block), all seven gate outcomes and details, and the Logbook row's identity digest (the graph'srun_configis the seam's byte for byte). Wall time 638 s seam, 653 s graph. The one remaining difference is deliberate:build.measure_resolution.provider.modereadsdirect_h5on the seam andscratch_frame_exporton the graph (the problem node exports the bound frame it holds to scratch for the engine); the measures are the same matrix. An f100 A/B on a spine that passes the projection fence is owed when one exists. The comparison is committed asdocs/evidence/uk-901-national-ab/compare_a_vs_b2.json(aggregates and field-level differences, machine paths omitted); since round 1 the graph records the scratchsource_pathrelative to its scratch root, so only the mode difference would remain on a re-run.🤖 Generated with Claude Code