UK certification is started by hand today (scripts/bundle.py certify-data, provenance/certification.py). Its CLI checks manifests, compatibility claims and whether the default artifact exists. It runs no simulation and does not check core compatibility.
The last three certifications took 10h37m (#555), about 20h43m (#536) and about 8d6h (#491) from model release to merge. Manual review caught real defects:
With no model freeze through the UK Autumn Budget (28 October 2026), every new pair needs to be qualified automatically and promoted only if it passes.
Scope
-
Trigger. .github/workflows/certify-uk-release.yaml runs on:
- a policyengine-uk wheel publish, by repository_dispatch;
- a uk-data release, by repository_dispatch;
- a scheduled catch-up.
Candidates are deduplicated by pair identity.
-
Candidate commit. A candidate commit regenerates the manifest, pins, uv.lock, TRO and test constants together. That commit is what qualification runs, what the replay corpus replays, and what promotion merges.
-
Mandatory gates (scripts/check_uk_candidate.py), none of which can be skipped:
- installed wrapper, model and core versions;
- wheel hashes, by RECORD verification;
- model/core compatibility;
- authenticated data access;
- dataset bytes;
- supported years;
- required Budget inputs;
- W1's reform contract;
- W3's metrics under review bounds against the previous certificate;
- W6's replay corpus, via an immutable pre-promotion request and an artifact-bound receipt.
-
Ledger. Qualification receipts are recorded in a ledger with a verdict that is a pure function of the gates, and failure evidence is kept.
-
Promotion. Promotion is serialized and ordered: a stale completion never displaces a newer accepted candidate. Qualification, consumer adoption and deployment are tracked separately.
-
Compatibility claims. Exact tested compatibility claims (==version) are backed by the qualification receipt and never broadened to get a green check.
-
Consumer dispatch after release goes to sim-api, the dashboard and the scorecard, with a richer payload.
Automatic promotion and the review bounds are methodology calls under cos d1344. Until that is ruled, promotion is a reviewed PR merge.
Refs #462, #570.
UK certification is started by hand today (
scripts/bundle.py certify-data,provenance/certification.py). Its CLI checks manifests, compatibility claims and whether the default artifact exists. It runs no simulation and does not check core compatibility.The last three certifications took 10h37m (#555), about 20h43m (#536) and about 8d6h (#491) from model release to merge. Manual review caught real defects:
With no model freeze through the UK Autumn Budget (28 October 2026), every new pair needs to be qualified automatically and promoted only if it passes.
Scope
Trigger.
.github/workflows/certify-uk-release.yamlruns on:Candidates are deduplicated by pair identity.
Candidate commit. A candidate commit regenerates the manifest, pins,
uv.lock, TRO and test constants together. That commit is what qualification runs, what the replay corpus replays, and what promotion merges.Mandatory gates (
scripts/check_uk_candidate.py), none of which can be skipped:Ledger. Qualification receipts are recorded in a ledger with a verdict that is a pure function of the gates, and failure evidence is kept.
Promotion. Promotion is serialized and ordered: a stale completion never displaces a newer accepted candidate. Qualification, consumer adoption and deployment are tracked separately.
Compatibility claims. Exact tested compatibility claims (
==version) are backed by the qualification receipt and never broadened to get a green check.Consumer dispatch after release goes to sim-api, the dashboard and the scorecard, with a richer payload.
Automatic promotion and the review bounds are methodology calls under cos d1344. Until that is ruled, promotion is a reviewed PR merge.
Refs #462, #570.