Skip to content

Build the UK national release role on the graph (PR-2 of the #901 line) - #1057

Merged
juaristi22 merged 15 commits into
mainfrom
uk-national-graph-path
Sep 30, 2026
Merged

juaristi22 merged 15 commits into
mainfrom
uk-national-graph-path

Conversation

@juaristi22

@juaristi22 juaristi22 commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

What this does

microcosm-build-uk --release-role national now builds the certified national line on the shared executable graph instead of dispatching, before any graph is prepared, to the in-process calibration seam library. This is the second change on the #901 line, planned with #901's re-base (her ruling 2026-09-25: the national role through the retained seam engine in #901, the posture-driven graph national path as the next PR).

The national posture composes (uk_runtime/graph_national.py):

  • uk.full.spine_checkpoint, the bound checkpoint as Register one UK full-build graph with all-geography calibration #901 admits it (sidecar, content identity, gate report), now keeping its H5's own time period.
  • uk.full.national_targets, the national register compiled from the pinned Chronicle artifact with the measure exclusions, the band edges, the frozen-register check and the feed pin (--allow-unpinned-feed as a node parameter). No ladder, no local surface.
  • uk.full.national_problem: refuse non-finite inputs, resolve the engine measures on the bound frame (one block, no local grains; the seam's measure_resolver=None route stays as resolver_factory=None), materialise the register, compile the constraint matrix row for row, restore the pristine tables, and encode one ordered problem whose bindings carry the national doctrine's family_equal loss weights, mass reason and bounds.
  • uk.full.dense and uk.full.calibrated, the shared calibration nodes under the doctrine's epochs, learning rate and seed. The dense node's admission is declared by the node: the dense line keeps its source preflight battery (parameter absent; the dense node's implementation hash moves with its _admit source and the bound checkpoint's with calibration_run and its time_period parameter, so uk.full.dense, uk.full.spine_checkpoint and everything downstream get new keys; nothing committed pins them and no Register one UK full-build graph with all-geography calibration #901 graph store resumes across this change); the national line admits the bound spine's provenance artifact, because it runs no pre-solve battery (its source and reference gates are the release-cut certifier's).
  • uk.full.gates.calibrated, the seven UK_CALIBRATION_GATE_SCOPE gates evaluated on the calibrated population with the seam's evidence (the calibration manifest block, the parity rows, the admin anchors, the CGT projection), storing the phase report, the manifest block, the diagnostic rows and the register census.
  • uk.full.national.readback, the written H5 read back and its content identity compared with the calibrated population.

The driver (full_build_cli) materialises the seam-shaped evidence from the stored artifacts: the diagnostics through the seam writer with the build block (digest measured from the file), the target-support sidecars, the calibration-seam battery replayed from the stored phase report through GateBatteryRun.record_phase (#901's replay) and signed, microcosm_uk_2024_25.h5, the schema-1 build_record.json the release-cut certifier reads, the frozen national_target_registry.json and national_contract_registry.json, the schema-4 rowwise_candidate_manifest.json, the Logbook row on the seam's uk-frs-calibration pipeline with the seam's attempt id, staging telemetry and the staged bundle, and the end-of-build incumbent evaluation after staging. A blocked battery leaves the report on disk (blocked_at_phase: "terminal"), writes no H5 and records a failed row.

Retired: calibration_run.run_uk_calibration, its attempt runner and battery runner, national_calibration.UKNationalCalibrationStage. Kept: the seam's contracts (gate scopes and closed-world exclusions, sidecar and checkpoint authentication, provenance blocks, scoped manifest, projection and admin-anchor artifacts, register census, doctrine payload, scoped-report signing), the frame adapter, the measure-input injection pair, the mass reason and the scoring route.

Shared touches (country-agnostic, none)

None. Every change is under uk_runtime and its tests, plus the docs and changelog.

New UK-side pieces (flagged)

  • uk_runtime/graph_national.py (the four kernels, the composition, the driver-side projections).
  • The national path in full_build_cli.py (prepare_national_build, execute_national_build, _execute_national_build, _close_national_attempt, _national_main).
  • full_targets.load_uk_national_target_inputs (the seam's loader as the target node runs it).
  • A dense-solve admission parameter on the shared UK dense kernel (default unchanged), the bound checkpoint's time_period parameter, the materialisation report in the measure receipt, and a pipeline argument on the rowwise Logbook envelope.

Deviations from the plan, flagged

  • The incumbent evaluation stays a driver step reusing the seam's function verbatim (it never fails the build), not a graph node as the plan's wording said.
  • Under the uniform override the graph binds an explicit all-ones loss-weight vector where the seam passed None; the solver computes a weighted sum over a mean, so the last bit may differ under that override only. The doctrine's family_equal is exact by construction.
  • The measure-exclusion windows and the target-fit deferral register are evaluated at --review-date (default today), the dense graph's convention and a node parameter, where the seam evaluated both at the run clock; a back-dated review date keeps expired exclusions and deferrals in force, which the seam never allowed. Documented in docs/uk-full-build-graph.md and the Calibrate the UK national build from Ledger-backed targets #623 runbook (review round 1).

Review round 1 (2026-09-30) and the #932 folds

Rebased over #932 as merged. One commit per finding: eb1704e C2 (an occupied --out refused before any work, the seam's FileExistsError), cb0d390 C1 (build.json re-issued after the close step on both lines with the final build-record and manifest digests; the national assembler reads the delivery summary from the build record, so the record could not be left alone), 1e8e252 A1 (KeyboardInterrupt records a discarded row on both lines and re-raises), 15bff5b should-fix 1 (scratch-relative source_path in the measure receipt, so the measure and national problem nodes are deterministic between runs), e61dd97 nits 4 and 5 (the dense solve refuses a bound mass_rule, scale_rule or l0_lambda it does not honour; the no-op removed; gate_scope from the posture as a node parameter; an empty register refused in run_config), 63771af docs, changelog, receipts R6 and nit 7 (R5's comparison committed under docs/evidence/uk-901-national-ab/). Folded from #932's round 1 on María's decision: a4308eb finding 1 (country_adapter.uk_atomic_support_register; --release-candidate requires the sources.yaml support rows, and the dense pre-flight refuses a manifest whose geography.assignment is not the atomic law on those pins), 11f4014 finding 2 (z-scores blanked beside suppressed counts in docs/evidence/uk-931/atomic-assignment-cells.json, sha 7d120d52…), c7057e9 finding 5 (LAD23CD/LAD24CD only), edb478f findings 4, 6 and 8 (pool docstring, the alias-name comment with the names kept so the gate-register digests do not move, the CLI test comment, the sample-admission sibling under the atomic law).

Verification (receipts R1–R4)

  • test_uk_national_graph.py (11): the composed graph on the seam suite's synthetic frame bound as a checkpoint; the ordered problem is the seam's matrix row for row with the doctrine bound; the graph's weights equal an in-process calibrate call under the national doctrine; the calibration mass record is the seam's; the seven seam gates evaluate.
  • test_uk_rowwise_national_role.py (23): the driver end to end on the same fixture, local and remote staging (fake Hub), the build record's bindings and provenance, the signed report, the diagnostics build block, the frozen registries, the manifest, the Logbook row on uk-frs-calibration, the doctrine overrides, the unpinned-feed refusal and override, the incumbent evaluation after staging, a blocked battery, the H5's weights equal to an in-process solve.
  • The seam suites re-anchored on the route as a library: test_uk_national_calibration.py 26, test_uk_calibration_run.py 18, test_uk_cgt_observation_period.py 2; the driver's dispatch tests follow the graph path.
  • Sweeps on the final tree: the whole engine-free suite including the uk directory, 11,980 passed, 119 skipped, one failure (test_uk_cgt_projection::test_engine_is_reported_unavailable_without_the_uk_extra) that asserts the uk extra is absent and fails identically on main in a venv that has it; the shared-spec lane (test_country_spec, test_spec_engine_country_bundles, the gate-register pins, the data contract, test_stage_evidence, test_gate_battery_replay) 432 passed; tools/ci_test_groups.py --verify ok; ruff check clean; every edited file formatted.
  • Round 1 (receipts R6): the 21 touched UK files 345 passed; the whole tests/engine_free/uk directory plus the shared-spec lane (test_country_spec, the gate-register pins, the data contract) 3,055 tests, 25 skipped, one failure (the environmental engine-absence test above); ruff check . clean; tools/ci_test_plan.py verify ok; gate-register digests unchanged.

Parity (receipts R5)

Licensed A/B on the 10 % smoke spine of the #901 pre-merge run (53,806 households), the pinned Chronicle artifact, the doctrine with no solve flags: A = the seam on main at the #901 merge, B = this branch. Both sides blocked at the terminal battery on the same gate, uk_cgt_projection_entrants, with the same numbers (a property of this spine under the engine's uprating, not of either path), so neither wrote the H5 or the build record. Before the battery, identical on both sides: the 1,062 × 53,806 constraint matrix, the target vector, the design weights, the final weights to the bit (max abs diff 0.0), the diagnostics file minus the build block (every target row, the UK block), all seven gate outcomes and details, and the Logbook row's identity digest (the graph's run_config is the seam's byte for byte). Wall time 638 s seam, 653 s graph. The one remaining difference is deliberate: build.measure_resolution.provider.mode reads direct_h5 on the seam and scratch_frame_export on the graph (the problem node exports the bound frame it holds to scratch for the engine); the measures are the same matrix. An f100 A/B on a spine that passes the projection fence is owed when one exists. The comparison is committed as docs/evidence/uk-901-national-ab/compare_a_vs_b2.json (aggregates and field-level differences, machine paths omitted); since round 1 the graph records the scratch source_path relative to its scratch root, so only the mode difference would remain on a re-run.

🤖 Generated with Claude Code

@juaristi22

Copy link
Copy Markdown
Collaborator Author

Program review

Base repository: PolicyEngine/microcosm
PR number: 1057
Reviewed head SHA: 6954cb0
Merge base SHA: 6a70cd4
Mode: full
Scope: changed behavior and affected dependencies
Source manifest: /private/tmp/policyengine-command-runs/7d60d9d077c4/pr-1057-review-sources.json
Review status: COMPLETE

Source Documents

No source documents registered; see scope and validation.

Critical

C1 — CRITICAL: published national completion hashes become stale (OPEN)

Location: packages/microcosm-build/src/microcosm/build/uk_runtime/full_build_cli.py:2035-2054 and :2104-2124. Trigger: any successful national run, including --no-staging. _execute_national_build stores the current rowwise_candidate_manifest.json SHA-256 in build.json, and execute_national_build publishes that completion marker last. _close_national_attempt then adds staged_dataset and evaluation to the manifest and rewrites it, so the marker's recorded digest differs from the final file on every success. With active staging telemetry, _close_national_attempt also changes build_record.json from the initial delivery summary to the completed one; its digest in build.json then becomes stale too. Expected: a published completion marker binds the final bytes and remains the last mutation of the bundle, as publish_staged_bundle documents. Observed by direct code trace: the final mutation occurs after the marker is present, with no refresh of build.json. The old seam also appended staging evidence after upload, but did not publish this new digest-bearing completion marker. The added end-to-end test checks final manifest-to-build-record identity yet never compares either against build.json.

C2 — CRITICAL: repeated national build silently replaces a prior candidate (OPEN)

Location: packages/microcosm-build/src/microcosm/build/uk_runtime/full_build_cli.py:1616-1660 and :1495-1614. Trigger: run microcosm-build-uk --release-role national twice with the same --out and a valid pinned spine, for example with a different receipted doctrine override on the second run. Expected: the previous national seam's national_role.validate_output_paths raises FileExistsError when a candidate artifact already exists, preserving the first H5 and its Logbook artifact location (national_role.py at the merge base, called from _run_national_attempt). Actual: the new preparation does not call that retained validation; _output_locations checks only source overlap, and publish_staged_bundle moves existing destination files to temporary backup then replaces them, deleting the backup on success. The first Logbook row still names the now-replaced output path, so its artifact reference no longer retrieves the recorded candidate. The new tests do not exercise reuse of an occupied national output directory.

Should Address

A1 — SHOULD ADDRESS: operator interruption leaves no terminal Logbook disposition (OPEN)

Location: packages/microcosm-build/src/microcosm/build/uk_runtime/full_build_cli.py:2186-2194. Trigger: KeyboardInterrupt during national graph execution before a gate verdict (for example Ctrl-C during the solve). Expected: the prior seam's calibration_run.run_uk_calibration caught BaseException and _record_failed_attempt wrote a terminal Logbook row with disposition='discarded' for KeyboardInterrupt, fulfilling its documented every-terminal-disposition contract. Actual: _national_main catches only Exception; KeyboardInterrupt bypasses _record_failure and fail_staging_telemetry, leaving the attempt without a terminal row and the staging run open. Ordinary Exception does close with a failed row, so this is specific to interruption.

Suggestions

None.

Evidence Gaps

No material gaps reported.

Notes

  • No target values or policy rules changed. The PR's licensed A/B compared the solve and gates on a sampled spine; both paths blocked before an H5 or build record, so that receipt does not test a successful publication.

Validation Summary

Local tests NOT RUN because the checkout has no installed project interpreter. CI at head 6954cb0 passed engine-free, engine-uk, integration-uk, lint, wheels and other jobs. Policy role compared the pinned target loader, exclusions, doctrine and seven gate scope with the prior seam. No external statutory source or target value changed.

Timing

Not measured.

Review Severity

REQUEST_CHANGES. Open findings: 2 critical, 1 should address, 0 suggestions.

@juaristi22
juaristi22 marked this pull request as ready for review September 30, 2026 10:06
@vahid-ahmadi

Copy link
Copy Markdown
Contributor

Automated review pass (Claude Code, high effort) — round 1 at 6954cb0e

Verdict: the program review's C1 (build-record digest part) and C2 stand as blocking; I'd downgrade the rest of C1. My own findings add two should-fixes: a scratch path that now leaks into stored artifacts (it also hits the dense line), and a clock change that isn't listed under deviations. The rest are nits and one question.

On the program review's C1 / C2 / A1 (checked at 6954cb0e)

  • C1, partly confirmed.
    • The manifest half is the dense line's existing convention, not new. build.json carries rowwise_candidate_manifest with the note "staging receipts are appended after publication" (full_build_cli.py:2045-2048), exactly as main's dense completion does (:1446-1449; also on 6a70cd4e). A stale manifest digest there is declared, so I'd call that part should-fix, and it applies to both lines.
    • The build-record half is new and real. build.json binds build_record by file_artifact with no note (:2041-2044). _close_national_attempt then rewrites build_record.json with the final staging_delivery after publish_staged_bundle has placed the marker (:2104-2107, called from :1669). The dense _close_attempt rewrites only the manifest. With active telemetry the delivery summary changes, so the marker's build_record.sha256 no longer matches the file the certifier reads.
    • Fix: either leave the build record alone after publication (the manifest already carries staging_delivery), or re-issue build.json last.
  • C2, confirmed. On 6a70cd4e the seam called _validate_output_paths(...) (national_role.py:341), which raises FileExistsError on any existing candidate artifact. At this head national_role.validate_output_paths (:582) has no caller in src. execute_national_build (:1616-1662) stages into a temp dir, and publish_staged_bundle replaces whatever is at --out. A second national run into the same directory therefore silently replaces the first candidate, and the first Logbook row's artifact_location then points at different bytes. The dense line has no such refusal either, so this is a national-line regression rather than a shared bug. Restoring the call in prepare_national_build, with a test on an occupied --out, closes it.
  • A1, confirmed. The seam wrapped the attempt in except BaseException and recorded KeyboardInterrupt as disposition="discarded" (calibration_run.py:400,503-505 on 6a70cd4e). _national_main catches only Exception (full_build_cli.py:2190), so a Ctrl-C during the solve leaves no terminal row and the staging run open. The dense _main (:2311) has the same gap, so a shared fix would cover both lines.

None of my findings below duplicates these three.

Checked and fine:

  • Merges cleanly onto current main 5187fce2. The six commits since the Register one UK full-build graph with all-geography calibration #901 merge touch nothing under uk_runtime.
  • Nothing still calls the retired symbols. From calibration_run the PR removes run_uk_calibration, _run_uk_calibration_attempt, _run_calibration_gate_battery, _record_failed_attempt, _notify_run_event and UKCalibrationRunPaths/Result. From national_calibration it removes UKNationalCalibrationStage and uk_national_calibration_stage. None of these is referenced in src, tools/ or tests. tools/diagnose_uk_uc_matrix.py imports only the kept adapter and injection pair.
  • test_uk_national_graph, test_uk_rowwise_national_role, test_uk_national_calibration, test_uk_calibration_run and test_uk_cgt_observation_period: 76 passed, engine-free.
  • Also green: test_uk_full_build_cli, test_uk_full_calibration_graph, test_uk_full_measure, test_uk_full_build_preparation, test_uk_rowwise_candidate, test_uk_release_certification and test_gate_battery_contract_pins. ci_test_groups.py --verify is ok and ruff check is clean.
  • The uniform override claim holds. solve.py:582 computes (loss*w).sum()/w.sum(), so an all-ones vector gives the mean up to rounding.
  • UK_ROWWISE_NATIONAL_POSTURE.gate_scope == UK_CALIBRATION_GATE_SCOPE (7 entries), with gate_posture=calibration_seam and pipeline=uk-frs-calibration.
  • The build record has the seam's top-level keys plus graph. The certifier reads it with .get and doesn't reject the extra key.
  • On a blocked battery, record_phase → enforce raises before finalize/write. That is the seam's order, and the driver returns before write_uk_national_frame.

1. should-fix: the resolver's scratch path now reaches stored artifacts, so the problem node's output changes on every run (national and dense lines)

full_measure.py:137 now adds each block's resolution.receipt to the measure receipt. That receipt carries the provider's own receipt (target_materialization.py:174). In scratch mode, UKMeasureResolver records source_path = scratch_dir / "simulation-input.h5" (measure_simulation.py:291,333). scratch_dir is a fresh TemporaryDirectory in both callers:

  • National: graph_national.py:587. It lands in bindings["measure_resolution"] of the encoded problem (:699), so problem.sha256 changes between identical runs. That sha then flows into the solution and result bindings, the diagnostics build block and the build record's graph.artifacts. R5's "scratch_frame_export with the matching source_path" is this field, and the path it records was deleted once the node finished.
  • Dense: graph_targets.py:381. The same receipt goes into the uk.full.measures artifact's metadata["evidence"] (graph_targets.py:428), with one path per engine block. This contradicts "Shared touches: none" and "default unchanged".

Probe: I called resolve_uk_full_measures twice with a stub resolver shaped like the scratch-mode receipt, each call in its own temp dir. On this branch evidence["resolution"][0]["provider"]["source_path"] came back as …/microcosm-uk-full-measures-ii_xerw7/simulation-input.h5 and then …-co6l2xoe/…, which differ. On 6a70cd4e no scratch path appears in the evidence. Both kernels are declared Determinism.DETERMINISTIC, which kernel.py:82 defines as "a function of its inputs and seed".

Suggested fix: drop source_path from the recorded provider receipt in scratch mode, or replace it with the scratch H5's sha256. Add a test that runs the problem kernel twice and asserts byte-identical output, with a stub resolver whose receipt includes its scratch path.

2. should-fix: the exclusion and deferral clocks now follow --review-date, not the run clock; list this as a deviation

On main the seam evaluated the measure-exclusion windows at datetime.now(UTC) (the loader was called with no exclusions_evaluated_on, which falls back through exclusion_evaluation_date(None)). It did the same for the terminal target-fit deferral register (calibration_run.py:806 on main).

The graph passes review_date to both: to the target node (graph_national.py:483) and to exclusions_evaluated_on in the gate kernel (:934). --review-date defaults to date.today() and the operator can set it. A back-dated review date therefore brings expired exclusions and deferrals back into force on the national line, which the seam never allowed.

This is the dense graph's existing convention (graph_targets.py:237, graph_terminal.py:567), and it keeps node keys reproducible, so it may well be the right call. But it's a change in behaviour: R5's parity holds only because both runs used today's date. Either add it to "Deviations from the plan", or have the national role refuse a --review-date other than today unless the override is recorded.

3. nit: "node keys unchanged" is not literally true for the dense line

The dense node's params are unchanged, but UKDenseSolveKernel's class source changed (_admit), and source_hash(type(self), …) feeds the implementation hash. Measured with UKDenseSolveKernel().implementation_hash(): f0b42806… on 6a70cd4e, 903b781b… here. UKBoundSpineKernel also moves, from d5991eb1… to 8d94cb4b…, through calibration_run in its hash plus the new time_period param that R1 already notes.

So uk.full.dense and everything downstream get new keys. R4's "no pinned node key moved" is still accurate because nothing pins them. Please reword the body so nobody expects to resume a #901 graph store.

4. nit (latent): the dense kernel ignores three of the doctrine fields the national problem binds

encode_uk_national_problem binds mass_rule, scale_rule and l0_lambda (graph_national.py:685-695). UKDenseSolveKernel.run hard-codes mass="free", passes no l0_lambda and uses the default scales (graph_calibration.py:406-430). The seam passed all three to calibrate.

The results are identical today: _ALLOWED_MASS_RULES = ("free",), _ALLOWED_SCALE_RULES = ("default_target_loss_scales",), and l0_lambda = 0.0 is not overridable. But the problem and manifest record these fields as the doctrine the solve ran under. A one-line check in the national problem encoder (or in _admit) that refuses any value the dense node cannot honour would stop this drifting silently.

5. nit: small tidy-ups in graph_national.py

  • :281-284 is an if …: pass with no effect.
  • :896 has the gate kernel rebuild its manifest from the UK_CALIBRATION_GATE_SCOPE constant, while the composition uses posture.gate_scope (:256). The two are equal (checked above), and the manifest comparison at :901 would catch a divergence, but reading the scope from the node's params would be simpler.
  • :1185: registry_document["national_registry"] and registry_from_payload(...).version records {} rather than failing when the register is empty. The problem node refuses an empty register later, but run_config is built before that.

6. nit: test counts in the body and receipts

The body and R2/R3 give 21 / 12 / 1 tests for test_uk_rowwise_national_role / test_uk_calibration_run / test_uk_cgt_observation_period. At this head, collection gives 22 / 18 / 2.

7. question: can R5's comparison be committed?

R5 rests on data/ukds/acceptance/901-national-ab/compare_a_vs_b2.json, which is outside the tree. If that file holds only aggregates and field-level diffs (no unit records), committing it under docs/evidence/ would let the parity claim be checked from the repository, as other UK receipts are.

juaristi22 and others added 15 commits September 30, 2026 12:59
…ition

The certified national line (microcosm#823) runs through the calibration
seam library, an in-process pipeline the graph driver dispatches to before
preparing any graph. This commit composes the same computation as graph
nodes over the bound spine checkpoint (uk_runtime/graph_national.py):

- uk.full.national_targets compiles the national register from the pinned
  Chronicle artifact (exclusions, band edges, the frozen-register check, the
  feed pin with --allow-unpinned-feed); full_targets gains the seam's loader
  as load_uk_national_target_inputs. No ladder, no local surface.
- uk.full.national_problem refuses non-finite inputs, resolves the engine
  measures on the bound frame (one block, no local grains; the seam's
  measure_resolver=None route stays as resolver_factory=None), materialises
  the register and encodes one ordered problem whose bindings carry the
  national doctrine's family_equal loss weights, mass reason and bounds.
- uk.full.dense / uk.full.calibrated are the shared calibration nodes under
  the national doctrine's epochs, learning rate and seed. The dense kernel's
  admission is declared by the node: the dense line keeps its preflight
  battery (parameter absent, node keys unchanged); the national line admits
  the bound spine's provenance artifact, because it runs no pre-solve
  battery (its source and reference gates are the release-cut certifier's).
- uk.full.gates.calibrated evaluates the seven UK_CALIBRATION_GATE_SCOPE
  gates on the calibrated population with the seam's evidence (the
  calibration manifest block, the parity rows, the admin anchors, the CGT
  projection) and stores the phase report, the manifest block, the
  diagnostic rows and the register census as artifacts.
- A readback node over the written national H5 compares its content
  identity with the calibrated population.

Two shared-UK adjustments the composition needed: the bound spine checkpoint
keeps its H5's own time period instead of the CREATE normaliser's 2024
default (the projection fence declares the frame's period), and
resolve_uk_full_measures records the materialisation report in its receipt
(the seam manifest carries it). rowwise_cli's Logbook envelope takes the
role's pipeline.

Tests (test_uk_national_graph.py, 8) run the composed graph on the seam
suite's synthetic frame, sidecar and register: the ordered problem is the
seam's matrix row for row with the doctrine bound, the graph's weights equal
an in-process calibrate call under the national doctrine, the calibration
mass record is the seam's, and the gate battery evaluates the seam scope.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
microcosm-build-uk --release-role national no longer dispatches to the
calibration seam library before preparing a graph: the driver prepares the
national graph (uk_runtime/graph_national.py) over the bound checkpoint and
materialises the seam-shaped evidence from the stored artifacts.

- full_build_cli: prepare_national_build (the bound checkpoint, the doctrine
  with its receipted overrides, the register node with the request's Ledger
  pins), execute_national_build (the staged temporary directory and the
  published bundle, as the dense line), _execute_national_build (the
  checkpoint, the register frozen beside the outputs, the solve and the seam
  battery in the graph, the seam diagnostics writer with the build block and
  the digest measured from the file, the target-support sidecars, the
  battery replayed from the stored phase report through
  GateBatteryRun.record_phase and signed, write_uk_national_frame, the
  graph's readback of the written H5, the schema-1 build_record.json the
  release-cut certifier reads, the schema-4 national manifest), the national
  envelope (the seam's attempt id and Logbook pipeline, staging telemetry,
  the failure sidecar and the failed row) and _close_national_attempt (the
  staged bundle, the incumbent evaluation after staging, the record and
  manifest rewritten with the delivery summary, the sums). A blocked seam
  battery leaves the report on disk, writes no H5 and records a failed row.
- national_role keeps what the role adds around the graph: the dry-run
  plan (now with the compiled graph's inventory), the doctrine overrides,
  the national manifest (over reported paths and the graph's artifact
  keys), the incumbent evaluation, the pin and output-path checks, the
  sums; its seam dispatch is gone and its private spellings stay aliases.
- graph_national gains the driver-side projections: the decoded result
  rebound to the calibrated population, the battery replay, the seam
  run_config and the build record.

Tests: test_uk_rowwise_national_role.py (21) runs the driver end to end on
the seam suite's synthetic frame bound as a checkpoint and checks the seam
shapes: the build record's bindings and provenance, the signed report, the
diagnostics build block, the frozen registries, the manifest, the local
and remote staging, the Logbook row on uk-frs-calibration, the doctrine
overrides, the unpinned-feed refusal and override, the incumbent
evaluation after staging, a blocked battery, and that the H5's weights are
an in-process calibrate call under the national doctrine. The driver's
dispatch tests follow the graph path.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
With the national release role on the graph, calibration_run's
run_uk_calibration, its attempt runner, its gate battery runner and the
two run dataclasses go, and national_calibration's UKNationalCalibrationStage
(with its checkpoint resume and post-solve fence) goes. What both lines and
the release-cut certifier still consume stays: the Logbook pipeline and
attempt-id prefix, the four gate scopes and their closed-world exclusions,
the bound spine sidecar and checkpoint authentication, the spine provenance
blocks, the scoped gate manifest, the CGT projection and admin-anchor
artifacts, the Ledger and runtime provenance blocks, the register census
and doctrine payloads, the scoped-report signing, the frame adapter with
its prepared/restore lifecycle, the measure-input injection pair, the
canonical mass reason and the scoring route.

The problem encoding is factored out of the national problem kernel as
encode_uk_national_problem, so the seam suites re-anchor on the graph's
national route as a library: the national-calibration suite solves the
encoded problem under the doctrine's bindings and checks the evidence block
(the fact-moving solve, the nested frame, the resolver injection and
restore, the band-edge threading, the packaged binding classes, the
materialisation skip refusal, the id preservation, the writer-clean tables,
the evidence block's refusals, the doctrine's loss weights in the
bindings); the seam suite keeps the checkpoint, scope, admin-anchor,
provenance and Logbook-scope contracts and gains the band-edge
reconciliation and attempt-id tests the runner tests carried; the CGT
observation-period test's national route materialises through the library.
The stage lifecycle, checkpoint-resume and runner-envelope tests retire
with the code they exercised (the driver's national suite covers the
envelope). Docs: the graph guide's national section and owner rows, the
national calibration runbook, the changelog.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…national line

The first licensed A/B (the seam on main 6a70cd4 vs the graph, the 10 %
smoke spine) matched the solve exactly (the same 1,062 x 53,806 matrix,
identical final weights, losses and diagnostics numbers, all seven gate
outcomes and details equal, the same Logbook identity digest) and left two
shape differences in calibration_diagnostics.json: the target descriptors
(the graph solves the compiled rows, so each target read as a callable on
the household entity where the seam's read as the spec's column measure on
its own entity) and the measure_resolution block (the graph recorded the
engine receipt where the seam recorded the resolution loop's own receipt:
the attached routes, the rounds, the provider's receipt).

The driver now hands the frozen register to national_result_from_manifest,
which re-describes the result's targets from the specs (row names checked
against the solved problem), and resolve_uk_full_measures carries the
resolution loop's receipt per block in its receipt, from which the national
evidence block takes the seam-shaped measure_resolution. The receipts file
opens with the hermetic parts.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
R5: on the licensed 10 % smoke spine, the seam (main at the #901 merge) and
the graph (this branch) produce the same 1,062 x 53,806 matrix, target
vector, design and final weights to the bit, the same diagnostics outside
the build block, the same seven gate outcomes and details (both blocked by
the projection fence, this spine's property) and the same Logbook identity
digest; 638 s vs 653 s wall. The one residual is the resolver's own
description of where it read its input. R4: the sweeps on the final tree.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
The retired seam validated the national role's output paths before any
work (`_validate_output_paths` in `_run_national_attempt`) and raised
FileExistsError on an existing candidate artifact. The graph driver
stages into a temporary directory and `publish_staged_bundle` replaces
whatever is at `--out`, so without that check a second run into the same
directory silently replaced the first candidate's bytes while the first
Logbook row still named them. `prepare_national_build` now runs the
retained `national_role.validate_output_paths` on the role's output
paths (dry runs excepted: they write nothing), so the refusal lands in
`failure.json` and a failed row before the graph is compiled.

Test: a second run into an occupied `--out`, chained on the first row,
exits 1 with the FileExistsError, leaves the H5, build record and
completion marker byte-identical and records a failed row.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… (review C1)

`publish_staged_bundle` places `build.json` last, binding the staged
bytes. The close step then appends the staging receipts to the rowwise
manifest on both lines and, on the national line, the delivery summary
to `build_record.json` (the national release assembler reads it there),
so the marker's `build_record` digest was stale on every successful
national run, and the manifest digest carried the note that receipts
follow. `_refresh_completion` re-issues the marker after the close step
with those files' final digests and drops the note, on both lines: the
marker is the bundle's binding of what is on disk.

Tests: the national end-to-end run with the incumbent evaluation checks
the marker's build-record, manifest and dataset digests against the
files; the dense driver run checks the manifest digest after the close
step.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… (review A1)

The seam wrapped the national attempt in `except BaseException`,
recorded a KeyboardInterrupt as `disposition="discarded"` (every terminal
disposition records a row) and re-raised. The graph driver caught only
`Exception`, so a Ctrl-C during the solve left the attempt without a
terminal row and the staging run open; the dense `main` had the same
gap. Both envelopes now record the discarded row through
`_record_failure` (which takes the disposition and, on the dense line,
the prepared build for its output check), fail the staging telemetry and
re-raise; `record_candidate_error` takes the disposition.

Tests: an interrupt raised from the execute step on either line writes
the failure sidecar naming it, spools one `discarded` row and propagates.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ew should-fix 1)

A scratch-mode `UKMeasureResolver` records the file it simulated from
(`source_path`), which lies under the per-run temporary directory the
measure node creates. Since 2182268 the resolution loop's receipt,
which carries that provider receipt, rides the measure receipt into the
national problem's bindings and the dense `uk.full.measures` artifact,
so an absolute scratch path changed a deterministic node's output
between identical runs. `resolve_uk_full_measures` now records every
path under the scratch root relative to it (`simulation-input.h5`,
`clone-0/simulation-input.h5`), in both the resolution receipts and the
resolver receipts.

Test: two resolutions in different scratch directories with a resolver
that records its scratch path yield equal receipts with no temporary
directory in them.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…l tidy-ups (review nits 4, 5)

The national problem binds `mass_rule`, `scale_rule` and `l0_lambda`
as the doctrine the solve ran under, while `UKDenseSolveKernel` solves
under free mass, the default target-loss scales and no L0 penalty. The
values agree today (the doctrine dataclass allows nothing else), but a
problem binding another value would have been recorded without being
honoured; the node now refuses it (`refuse_unhonoured_solve_doctrine`).

In `graph_national`: the no-op `if …: pass` on the override rule is
gone; the gate node carries `gate_scope` from the posture and the kernel
evaluates that scope (the manifest digest still pins its compilation)
instead of re-reading the seam constant; `national_run_config` refuses
an empty national register instead of recording `{}`.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…932 review 1)

`sources.yaml` and the provenance register pin each atomic-area
support's sha256 and byte size, but nothing compared the supports a
build used with them: `_prepare_geography` checked the bytes only
against the digest pinned on the command line, and no release check
read the manifest's `geography.assignment`. A rebuilt support therefore
passed as a release candidate as long as the caller pinned its own
digest.

`country_adapter.uk_atomic_support_register` reads the register from
the UK spec by system. Under `--release-candidate` each supplied support
pin must equal its register row, and
`tools/preflight_uk_local_release_candidate.py` refuses a manifest whose
`geography.assignment` is not the atomic law or whose `support_pins`
differ from the register (a pre-#932 manifest with no binding is
refused like a pre-role candidate). The synthetic dense fixtures bind
synthetic pins (`SYNTHETIC_SUPPORT_PINS`, the `RELEASE_PINS` digests)
and the preflight and assembler suites register those; one preflight
test reads the committed register and shows the synthetic pins refused
on all three systems.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…view 2)

`atomic-assignment-cells.json` suppressed `rows < 3` as `<3` but
published `expected_rows` and `z` for the same cells, from which the
review recovered every suppressed count exactly. The harness now blanks
`z` wherever it suppresses `rows` (`_suppressed_area`), the purpose
text says so, and the committed file is re-suppressed in place under
that rule (985 z-scores blanked, no cell regenerated; sha256
32f494df… → 7d120d52…, recorded in the receipts).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…review 5)

`_normalise_ew_lad_lookup` and `load_ew_oa_lad_region_lookup` accepted
any `LAD??CD` column, so a `LAD25CD` frame (the Barnsley/Sheffield
recode, a separate follow-up) was silently relabelled `lad23_code`, and
a frame with two such columns failed with an AttributeError. Both now
accept exactly `LAD23CD` or `LAD24CD` (the December 2024 OA lookup
carries the April 2023 set under the latter), ignore other names and
refuse a frame carrying both accepted names.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…review 4, 6, 8)

The pool checkpoint's docstring said the platform scope "stops here";
it only avoids the typed-artifact refusal, since the shared assignment
kernels' columns already reach the UK chain at platform scope through
population slices (reworded; the question to Max stays open). The five
nation-native alias names in the export allowlist are there because the
full-build export-surface gate reads the graph frame's in-memory columns
before `graph_terminal._tables` drops them (the F8 gap on #932); the
allowlist now says so, and the names stay so the gate-register digests
do not move again. The CLI test's "dispatched to the seam" comment is
replaced, and the sample-admission test gains its sibling under the
default atomic law.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…parison as evidence

Records the review-date clock as a deviation from the seam (the dense
convention, a node parameter), the occupied-output refusal, the interrupt
disposition, the re-issued completion marker, the scratch-relative
resolver path and the registered-support check in the graph doc and the
#623 runbook; extends the changelog fragment; corrects the receipts'
test counts (R2, R3), notes the committed comparison in R5 and adds R6
(round 1 and the #932 folds); commits R5's seam-vs-graph comparison
under docs/evidence/uk-901-national-ab (aggregates and field-level
differences only, machine paths omitted).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@juaristi22
juaristi22 force-pushed the uk-national-graph-path branch from 6954cb0 to 63771af Compare September 30, 2026 12:03
@juaristi22

Copy link
Copy Markdown
Collaborator Author

Posting on María's call. Thanks for the round. The fixes are on the branch at 63771aff6, rebased over #932 as merged, one commit per finding; the body carries the same map under "Review round 1".

  • C1 (cb0d390): confirmed for the build record, and taken for the manifest on both lines. build.json is re-issued after the close step with the final digests and the note dropped. Leaving the record alone was not an option: tools/assemble_uk_release_dir.py reads the delivery summary from build_record.json. Tests compare the marker with the files on both lines.
  • C2 (eb1704e): confirmed. prepare_national_build runs the seam's validate_output_paths before the graph is compiled (dry runs excepted). Test: two runs into one directory, the first candidate's H5, build record and marker byte-identical, a failed second row chained on the first.
  • A1 (1e8e252): confirmed. KeyboardInterrupt records a discarded row on both lines through _record_failure (which now takes the disposition), fails the telemetry and re-raises, as the seam did.
  • 1, scratch path (15bff5b): resolve_uk_full_measures records every path under the scratch root relative to it, in both the resolution and the resolver receipts, so the provider block reads simulation-input.h5. Test: two resolutions in different scratch directories with a resolver that records its scratch path yield equal receipts. This covers the dense uk.full.measures artifact as well.
  • 2, review date: recorded as a deviation in the body, docs/uk-full-build-graph.md and the Calibrate the UK national build from Ledger-backed targets #623 runbook; the dense convention stands and the date is a node parameter.
  • 3: reworded in the body; uk.full.dense and uk.full.spine_checkpoint move, nothing pins them.
  • 4 (e61dd97): UKDenseSolveKernel refuses a bound mass_rule, scale_rule or l0_lambda it does not honour; the national problem's bindings pass the check in the existing problem test.
  • 5 (e61dd97): the three tidy-ups as suggested (gate_scope is a node parameter read from the posture; the manifest digest still pins its compilation).
  • 6: the counts were 22 / 18 / 2 at your head, miscounted in the body; 23 / 18 / 2 now with the C2 test, and 11 for test_uk_national_graph.py.
  • 7 (63771af): committed as docs/evidence/uk-901-national-ab/compare_a_vs_b2.json, aggregates and field diffs only, machine paths omitted.

The #932 round-1 items are folded here on María's decision, linked from #932: a4308eb (finding 1), 11f4014 (2), c7057e9 (5), edb478f (4, 6, 8). Verification at this head is in the body (345 targeted; the whole engine-free uk directory plus the shared-spec lane 3,055 with the one environmental failure).

@vahid-ahmadi

Copy link
Copy Markdown
Contributor

Automated review pass (Claude Code, high effort) — round 2 at 63771aff

Verdict: nothing blocking. Everything from round 1 is closed, including the program review's C1, C2 and A1. So are the four #932 items folded in here, except the evidence suppression, which is only partly closed (item 8 below; low risk, since the counts are synthetic draws). I'd approve once the four CI jobs still running (engine-free and engine-us, started 12:03Z) come back green.

I verified this at this head:

  • 245 targeted engine-free tests pass locally: the five national-line files, test_uk_full_build_cli, test_uk_full_measure, test_uk_full_calibration_graph, the six test files the UK geography step 2: identity-keyed atomic-area assignment on rebuilt supports (#931) #932 folds touch (country adapter, dense release assembler, full-build preparation, local release preflight, atomic-assignment measurement, geography sources), test_uk_graph_evidence and test_uk_full_population_graph.
  • ruff check is clean and tools/ci_test_plan.py verify is ok.
  • Collected counts are 23 / 18 / 2 / 11, which match the body.

Status of earlier items

Item Status Evidence
C1 completion marker stale after close Closed _refresh_completion (full_build_cli.py:1080) re-issues build.json inside the finally of both close steps (:1133 dense, :2250 national), after the manifest, build-record and sha256sums rewrites. It is now the last write to the bundle; only the spool row follows, and that lives outside it. build.json is not a manifest output, so neither sha256sums.txt nor the staged bundle lists it, and nothing else goes stale. tools/assemble_uk_release_dir.py reads build_record.json, not the marker, so it is unaffected.
C2 second national build replaces the first Closed prepare_national_build calls national_role.validate_output_paths before the graph compiles (:1622), skipping dry runs, which write nothing. The new test leaves the first H5, build record and marker byte-identical and chains a failed row. One side effect is item 9 below.
A1 Ctrl-C leaves no terminal row Closed, both lines except KeyboardInterrupt in _national_main and in the dense main records a discarded row through _record_failure(..., disposition=...), fails the telemetry and re-raises. The dense prepared = None is set before the try, so an early interrupt can't hit a NameError. state.spool_path prevents a second row if the interrupt lands after the close step has written one.
1 scratch path in stored artifacts Closed _without_scratch_paths (full_measure.py:36) rewrites any absolute path under the scratch root, recursively, in both the resolution and resolver receipts. I reran my probe on the national call shape (blocks=1, local_grains=()), adding a nested scratch path to the stub receipt. The evidence is now identical across two scratch directories, and the temp root no longer appears in it. That evidence is the only run-dependent input to the problem bindings, so problem.sha256 is stable too. The dense uk.full.measures artifact goes through the same function.
2 exclusion clock follows --review-date Closed It is listed under Deviations, in docs/uk-full-build-graph.md and in the #623 runbook.
3 "node keys unchanged" Closed The body now says the dense node's implementation hash moves.
4 unhonoured doctrine fields Closed refuse_unhonoured_solve_doctrine (graph_calibration.py) runs on the problem's bindings before the solve. The new test refuses non-default mass_rule, scale_rule and l0_lambda.
5 graph_national tidy-ups Closed The no-op if/pass is gone. gate_scope is now a node parameter read from the posture. run_config refuses an empty register (with a test).
6 test counts Closed Collected counts are 23 / 18 / 2, plus 11 for test_uk_national_graph.
7 R5 comparison in the tree Closed docs/evidence/uk-901-national-ab/compare_a_vs_b2.json holds aggregates and field diffs only. It contains no unit records, and a grep for /Users, /private, /tmp, /home and .h5 finds nothing.
#932-1 registered supports Closed Under --release-candidate, _prepare_geography now requires each pin to equal the sources.yaml row (uk_atomic_support_register, which is tested against the provenance resource). tools/preflight_uk_local_release_candidate.py refuses a manifest whose geography.assignment is not atomic or whose support_pins differ from the register; that path matches what graph_terminal.py:1452 writes. build_uk_country_graph(release_candidate=True) with a legacy config still compiles, but the preflight now catches the manifest it produces, which is the guard I asked for.
#932-2 z-scores beside suppressed counts Partly z is null on every suppressed row, in all four cells (474 / 495 / 10 / 6). The subtraction route below remains.
#932-5 LAD column Closed Only LAD23CD and LAD24CD are accepted. LAD25CD is refused as a missing column, and two accepted names are refused.
#932-4/6/8 docstring, allowlist comment, atomic-law admission test Closed The docstring now says only the typed-artifact edge stops, and the open question points at #932. The alias comment explains the F8 in-memory read. test_sample_admission_holds_under_the_default_atomic_law is added.

New

8. should-fix (low risk): the per-level totals still reveal the suppressed K=15 counts.

  • The file publishes zeros, so a suppressed count is 1 or 2.
  • It also publishes each cell's total rows, and each household sits in exactly one constituency and one LA. Within a level, the total minus the published counts is therefore the sum of the suppressed counts.
  • At K=15 that sum pins every suppressed value:
    • legacy constituency: 5 suppressed, 10 rows;
    • keyed constituency: 3 suppressed, 6 rows;
    • keyed LA: 3 suppressed, 6 rows.
  • So 11 of the 14 suppressed K=15 constituency and LA counts are exactly 2. Only legacy LA (5 cells, 8 rows) stays ambiguous.
  • The keyed constituency codes whose count is recovered this way are E14001175, E14001495 and E14001518.
  • levels.*.min_rows = 2.0 and breaches.constituency.rows.min = 2.0 say the same thing independently. At K=1 the sums are too large to single any cell out.

The counts are synthetic geography draws, so the practical risk is small. But the file claims "counts below the minimum are suppressed together with their z-scores", and at K=15 that is not true. The cheapest fix is secondary suppression: drop rows at the cell level and the min_rows/min_sources/rows.min summaries wherever a level has suppressed cells. Alternatively, drop the claim and state why these counts don't need protecting. #1059 has the same pattern (a count under 10 recovered by subtraction), so a shared helper may be worth it.

9. nit: a refused second run writes into the candidate's own directory. The C2 refusal is raised inside the try, so _record_failure writes failure.json (with "release_authorized": false) and a logbook-spool row into the occupied --out. create_staging_telemetry has already opened a run there too: a local staging/runs/<id> marked failed, or, with remote staging, a failed run on the Hub. Nothing reads failure.json today (only full_build_cli.py and one US tool mention it), but a valid candidate directory now carries a failure sidecar from a different attempt. Running validate_output_paths in _national_main before telemetry opens would make the refusal leave no trace, though the failed Logbook row would then need a home. At minimum, the test could assert that the first run's files are the only candidate artifacts and that failure.json names the refusing attempt.

10. nit: the comparison script isn't in the tree. compare_a_vs_b2.json's purpose cites data/ukds/acceptance/901-national-ab/compare_national_ab.py, which lives in the licensed data tree. The script contains no data, so committing it beside the JSON, for example under tools/ or experiments/, would let R5 be regenerated by anyone with the licensed inputs.

@juaristi22
juaristi22 merged commit 12cc970 into main Sep 30, 2026
10 checks passed
juaristi22 added a commit that referenced this pull request Sep 30, 2026
…ppressed one (#1057 review round 2, item 8)

Vahid's round 2 on #1057: every household row lies in exactly one area per
level, so a level total less its published counts is the sum of the
suppressed ones; with zeros published each suppressed count is 1 or 2, and
at K=15 the residual pinned eleven of fourteen (all 2). Dropping the cell
`rows` would not close it: the expected rows sum to the same total (and the
integer pool counts with the public census shares fix it anyway). The
harness's new `resuppress_evidence` pass:

- withholds one published count per level that has suppressed counts
  (`withheld`, with its z-score, ESS and source count), chosen in a
  value-independent order (sha256 of level and area code) among counts
  above the minimum, so the residual carries an unknown of at least
  minimum + 1;
- withholds the level minima of rows, ESS and sources there, the
  breach-table minima, and any breach quantile inside the suppressed range;
- recomputes max |z| and the share above 3 over published areas only (at
  K=1 the reported LA maximum belonged to a suppressed area).

`main` applies the pass before writing; the committed file is re-suppressed
in place (no cell regenerated, sha 7d120d52... -> 99013cc1...). Tests: the
committed file is a fixed point, each level with suppressed counts carries
one complement and a non-degenerate feasible range for their sum, and two
invented-level cases.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
juaristi22 added a commit that referenced this pull request Sep 30, 2026
…iew round 2, item 9)

The C2 refusal (#1057 round 1) ran inside prepare_national_build, after
the attempt id was minted and telemetry opened, so a refused second run
wrote failure.json (`release_authorized: false`) and a failed Logbook row
into the first candidate's directory, and opened a failed staging run. The
check moves to `refuse_occupied_national_output`, called in
`_national_main` with the other argument refusals, before the staged-
dataset pre-flight and before any attempt exists: a refused run writes
nothing into the occupied directory. No Logbook row is owed, as for every
refusal that precedes the attempt (the driver's standing rule: argument
refusals cost nothing). The input's existence is still checked inside the
attempt; the collision check needs only its path.

Test: the second run raises FileExistsError and leaves every file under
the first candidate's --out byte-identical, with no failure.json and one
Logbook row. The runbook and the graph doc say so.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
juaristi22 added a commit that referenced this pull request Sep 30, 2026
…nd 2, item 10)

`compare_a_vs_b2.json` cited a script that lived only in the licensed data
tree. The script (no data, no machine paths) is committed as
docs/evidence/uk-901-national-ab/scripts/compare_national_ab.py, formatted
under the evidence-scripts ruff waiver, so R5 can be regenerated by anyone
with the licensed inputs; the JSON's purpose and the #1057 receipts point
at it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
juaristi22 added a commit that referenced this pull request Sep 30, 2026
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
juaristi22 added a commit that referenced this pull request Sep 30, 2026
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@juaristi22

Copy link
Copy Markdown
Collaborator Author

Round 2 items 8–10 are folded into #1059 on María's decision, one commit each:

  • 8, suppressed counts recoverable at K=15 (73d1736).
    • Dropping the cell rows alone wouldn't have closed it. The per-area expected rows sum to the same total, and the integer pool counts together with the public census shares fix it anyway.
    • So the harness's new resuppress_evidence pass withholds one more count per level that has suppressed counts, marked withheld with its z-score, ESS and source count. It is chosen in a value-independent order (sha256 of level and area code) among counts above the minimum, so the residual carries an unknown of at least four.
    • The pass also withholds the level and breach-table minima there and any breach quantile inside the suppressed range.
    • It computes max_abs_z and share_abs_z_gt_3 over published areas only. At K=1 the reported LA maximum (1.96 legacy, 1.94 keyed) belonged to a suppressed area.
    • The committed file is re-suppressed in place with no cell regenerated (sha 7d120d52… → 99013cc1…).
    • The test checks that the file is a fixed point of the pass, and that at every level with suppressed counts the reader's feasible range for their sum spans more than one value.
  • 9, a refused second run writing into the candidate's directory (9c13b32). The refusal now runs in _national_main with the other argument refusals, before the attempt id is minted and telemetry opens.
    • A refused run writes nothing into the occupied directory: no failure.json, no Logbook row, no staging run.
    • No row is owed, as for every refusal that precedes the attempt.
    • The input's existence is still checked inside the attempt.
    • The test compares every file under the first candidate's --out byte for byte and finds one Logbook row. The runbook and the graph doc are updated.
  • 10, the comparison script (0073c5b): committed as docs/evidence/uk-901-national-ab/scripts/compare_national_ab.py. The JSON's purpose and the receipts point at it.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants