Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
17 commits
Select commit Hold shift + click to select a range
5907a72
UK was_lisa runtime: Lifetime ISA holdings from the WAS round-8 perso…
juaristi22 Sep 29, 2026
014f4bc
UK was_lisa joins the spine after was_wealth (#1003)
juaristi22 Sep 29, 2026
0964936
UK was_lisa gates, export cells and support bounds (#1003)
juaristi22 Sep 29, 2026
9627c6c
UK was_lisa receipts, evidence, topic doc and signed differences (#1003)
juaristi22 Sep 29, 2026
13f293a
Receipts: label the Part F table and cite HMRC's 2024-25 LISA subscri…
juaristi22 Sep 29, 2026
1e6448c
Withhold the released exact/banded holder split (review round 1, item 1)
juaristi22 Sep 30, 2026
575b23b
Label the donor's youngest age group 16-24 wherever it is published (…
juaristi22 Sep 30, 2026
58a3b95
Stop reading the self-employment income no LISA model uses (review ro…
juaristi22 Sep 30, 2026
73d1736
Withhold a complementary count so the #931 level totals recover no su…
juaristi22 Sep 30, 2026
9c13b32
Refuse an occupied national --out before the attempt opens (#1057 rev…
juaristi22 Sep 30, 2026
0073c5b
Commit the R5 comparison script beside its evidence (#1057 review rou…
juaristi22 Sep 30, 2026
fd9b713
Changelog: the #1057 review round 2 folds
juaristi22 Sep 30, 2026
fb87946
Receipts: the measurement base after the rebase onto #1057 and #1045
juaristi22 Sep 30, 2026
3c5999d
Receipts: the WAS donor against HMRC's LISA subscription counts (revi…
juaristi22 Sep 30, 2026
61fe7da
Receipts: option (d), the donor's income predictors uprated to the sp…
juaristi22 Sep 30, 2026
23f7a91
Render the #931 receipts' summary cells from the re-suppressed eviden…
juaristi22 Sep 30, 2026
84a3169
Receipts: arm D's coefficients move slightly, not at all (review roun…
juaristi22 Sep 30, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
20 changes: 20 additions & 0 deletions changelog.d/1003-uk-was-lisa-stage.added.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,20 @@
A person-grain `was_lisa` stage joins the UK spine right after `was_wealth` (microcosm#1003; María's rulings of 2026-09-23 and 2026-09-29). It gives every adult a Lifetime ISA holding, imputed from the Wealth and Assets Survey round-8 person tab (UKDS SN 7215, End User Licence), and exports three new cells: `person.has_lifetime_isa`, `person.lifetime_isa_balance` and `household.household_lifetime_isa_balance`.

The donor. The person tab is pinned by size and sha256 beside the `was_wealth` household tab and read selectively (`read_pinned_tab(..., columns=)`). `clean_was_lisa_donor` joins each person to its household on `CASER8` and refuses the tab unless the released values `DVFLISAvR8` sum to the household's `DVFLISAVR8_aggr`. A missing or negative value is refused, never read as zero. The response class is kept for the receipt and never used as a predictor: observed, banded, ONS-imputed or rule-impossible. A declared credibility rule recodes holders in age bands from 45 up to non-holders (LISAs open only to adults under 40 since April 2017). Holders above £40,000 stay holders but leave the balance fit.

The model. Ownership is a weighted ridge logistic with an unpenalised intercept (`impute_lifetime_isa_ownership`). Its predictors are the age group, log earnings, household net income, gross financial wealth, cash ISA, stocks-and-shares ISA and savings, private renting, sex and the number of children. An adult holds when a person-keyed uniform falls below their probability. The balance is the regime-gated QRF fitted on the credible holders and drawn at a second person-keyed quantile (`impute_lifetime_isa_balance`), so it is positive exactly when the person holds.

Coherence. A household total above its `gross_financial_wealth` draw is scaled down pro rata (`cap_lifetime_isa_to_financial_wealth`, receipted), because WAS counts LISAs inside gross financial wealth. The household cell is the sum of its persons (`aggregate_person_to_household`). No `was_wealth` column is rewritten. Balances are 2020-22 pounds from a Great Britain donor, not uprated. No contribution or bonus column is created.

Gates. `uk_stage_was_lisa_support` (`stage_health`, support clip, release-blocking) joins the spine gate scope. `uk_support` gains `was_lisa_support_bounds.json`, generated with `--check` by `tools/build_uk_was_lisa_support_bounds.py`.

Surfaces that move in lockstep:

- the sources manifest and its projection;
- five schema branches and the operation-kind allow-list;
- the graph roster, kernel registry and spine tool (`--was-person-tab`);
- the H2 fixture (35 stages, a synthetic `was_person.csv`) and the coverage manifest;
- the export allow-lists, the country package and the gate-policy digests;
- two net-new-column signed differences.

Receipts are in `experiments/1003-uk-lisa-receipts.md`, aggregates in `docs/evidence/uk-lisa-1003/`, and the topic doc is `docs/uk-lisa-1003.md`.
1 change: 1 addition & 0 deletions changelog.d/1057-review-round-2-folds.changed.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
Folded from microcosm#1057's review round 2 into microcosm#1059. The national line refuses an occupied `--out` with the other argument refusals, before the attempt opens, so a refused second run writes nothing into the first candidate's directory (no `failure.json`, no Logbook row, no staging run). The #931 atomic-assignment evidence withholds one more count per level that has suppressed counts, because every row lies in one area per level and the level total less the published counts recovered them. It also withholds the level and breach-table minima there, and computes its z summaries over published areas only. The harness applies the same pass (`resuppress_evidence`) and the committed file is its fixed point. The R5 national comparison script is committed beside its evidence.
2 changes: 1 addition & 1 deletion docs/evidence/uk-901-national-ab/compare_a_vs_b2.json
Original file line number Diff line number Diff line change
@@ -1,5 +1,5 @@
{
"purpose": "R5 of experiments/901-uk-national-graph-path-receipts.md: the licensed seam-vs-graph national A/B on the 10 % smoke spine of the #901 pre-merge run (53,806 households, 1,062 targets), compared field by field by data/ukds/acceptance/901-national-ab/compare_national_ab.py. Aggregates and field-level differences only; no unit records.",
"purpose": "R5 of experiments/901-uk-national-graph-path-receipts.md: the licensed seam-vs-graph national A/B on the 10 % smoke spine of the #901 pre-merge run (53,806 households, 1,062 targets), compared field by field by docs/evidence/uk-901-national-ab/scripts/compare_national_ab.py (run in the licensed lane as data/ukds/acceptance/901-national-ab/compare_national_ab.py; committed by microcosm#1059 on review round 2 of #1057, item 10). Aggregates and field-level differences only; no unit records.",
"date": "2026-09-29",
"a": "a-seam (main at the #901 merge, 6a70cd4ee, run_uk_calibration)",
"b": "b-graph-2 (uk-national-graph-path at 218226898, microcosm-build-uk --release-role national)",
Expand Down
235 changes: 235 additions & 0 deletions docs/evidence/uk-901-national-ab/scripts/compare_national_ab.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,235 @@
"""Compare the seam (A) and graph (B) national builds on one spine.

Usage (from a tree with the uk extra): uv run --no-sync python compare_national_ab.py <a-dir> <b-dir>

Reports as JSON: the solve through the target-support sidecars both sides
write before the battery (row names, target vector, design and final weights,
the CSR matrix), the diagnostics minus the build block and the build block
minus its operational fields, the gate report outcomes and details, the
Logbook rows' identity and pin digests, the H5 payload when both sides wrote
one, and the wall time.
"""

from __future__ import annotations

import hashlib
import json
import sys
from pathlib import Path

import numpy as np
import scipy.sparse as sp

from microcosm.build.logbook import load_spool_rows

A, B = Path(sys.argv[1]), Path(sys.argv[2])
DATASET = "microcosm_uk_2024_25.h5"
GATES = "microcosm_uk_2024_25.terminal_gates.json"
OPERATIONAL = {
"build_id",
"created_at",
"code_pin",
"runtime",
"git_commit",
"git_dirty",
"staging_delivery",
"path",
"graph",
"rowwise_driver_parameters",
"certification",
"sha256",
"size_bytes",
"bytes",
"elapsed_seconds",
"seconds",
"evaluated_at",
"timestamp",
}


def sha(path: Path) -> str:
return hashlib.sha256(path.read_bytes()).hexdigest()


def strip(payload, drop=OPERATIONAL):
if isinstance(payload, dict):
return {k: strip(v, drop) for k, v in payload.items() if k not in drop}
if isinstance(payload, list):
return [strip(v, drop) for v in payload]
return payload


def diff_paths(a, b, prefix=""):
out = []
if isinstance(a, dict) and isinstance(b, dict):
for key in sorted(set(a) | set(b)):
if key not in a or key not in b:
out.append(
f"{prefix}/{key}: {'only in A' if key in a else 'only in B'}"
)
else:
out.extend(diff_paths(a[key], b[key], f"{prefix}/{key}"))
elif isinstance(a, list) and isinstance(b, list):
if len(a) != len(b):
out.append(f"{prefix}: list length {len(a)} vs {len(b)}")
else:
for i, (x, y) in enumerate(zip(a, b, strict=True)):
out.extend(diff_paths(x, y, f"{prefix}[{i}]"))
elif a != b:
out.append(f"{prefix}: {str(a)[:80]!r} vs {str(b)[:80]!r}")
return out


report = {"a": str(A), "b": str(B)}

# 1. The solve, from the target-support sidecars.
va = np.load(A / "target_support_vectors.npz", allow_pickle=True)
vb = np.load(B / "target_support_vectors.npz", allow_pickle=True)
ma, mb = (
sp.load_npz(A / "target_support_matrix.npz"),
sp.load_npz(B / "target_support_matrix.npz"),
)
fa, fb = va["final_weights"], vb["final_weights"]
report["solve"] = {
"targets": (int(ma.shape[0]), int(mb.shape[0])),
"households": (int(ma.shape[1]), int(mb.shape[1])),
"row_names_equal": bool(np.array_equal(va["names"], vb["names"])),
"target_vector_equal": bool(
np.array_equal(va["target_vector"], vb["target_vector"])
),
"matrix_identical": bool(ma.shape == mb.shape and (ma != mb).nnz == 0),
"design_weights_equal": bool(
np.array_equal(va["design_weights"], vb["design_weights"])
),
"household_ids_equal": bool(
np.array_equal(va["household_ids"], vb["household_ids"])
),
"final_weights_identical": bool(np.array_equal(fa, fb)),
"final_weights_max_abs_diff": float(np.max(np.abs(fa - fb)))
if fa.shape == fb.shape
else None,
"final_weight_totals": (float(fa.sum()), float(fb.sum())),
}

# 2. The diagnostics.
da = json.loads((A / "calibration_diagnostics.json").read_text())
db = json.loads((B / "calibration_diagnostics.json").read_text())
report["diagnostics"] = {
"scalars": {
k: (da.get(k), db.get(k), da.get(k) == db.get(k))
for k in (
"schema_version",
"initial_loss",
"final_loss",
"n_nonzero",
"n_records",
"realized_max_weight_ratio",
"fraction_within_10pct",
)
},
"diff_minus_build": diff_paths(
strip(da, OPERATIONAL | {"build"}), strip(db, OPERATIONAL | {"build"})
),
"build_diff": diff_paths(strip(da.get("build", {})), strip(db.get("build", {}))),
"bytes_identical": sha(A / "calibration_diagnostics.json")
== sha(B / "calibration_diagnostics.json"),
}

# 3. The gate report.
ga, gb = json.loads((A / GATES).read_text()), json.loads((B / GATES).read_text())
report["gate_report"] = {
"blocked_at_phase": (ga.get("blocked_at_phase"), gb.get("blocked_at_phase")),
"statuses": {
k: (ga["gates"][k]["status"], gb["gates"][k]["status"]) for k in ga["gates"]
},
"diff_minus_attestation": diff_paths(
strip(ga, OPERATIONAL | {"attestation", "release_evidence"}),
strip(gb, OPERATIONAL | {"attestation", "release_evidence"}),
),
"signed": (
bool(ga.get("attestation", {}).get("signature")),
bool(gb.get("attestation", {}).get("signature")),
),
}

# 4. The Logbook rows.
ra, rb = (
load_spool_rows(A / "logbook-spool")[0].to_mapping(),
load_spool_rows(B / "logbook-spool")[0].to_mapping(),
)
report["logbook"] = {
k: (ra[k], rb[k], ra[k] == rb[k])
for k in (
"pipeline",
"identity_digest",
"input_pins_digest",
"disposition",
"rung",
"seed",
)
}
report["logbook"]["phases"] = (ra["phases_reached"], rb["phases_reached"])

# 5. The H5 and the build record, when both sides wrote them.
if (A / DATASET).is_file() and (B / DATASET).is_file():
from microcosm.build.uk_runtime.national_frame import load_uk_national_frame

fra, _ = load_uk_national_frame(A / DATASET)
frb, _ = load_uk_national_frame(B / DATASET)
h5 = {"bytes_identical": sha(A / DATASET) == sha(B / DATASET), "tables": {}}
for entity in fra.entities:
ta, tb = fra.table(entity), frb.table(entity)
entry = {
"rows": (len(ta), len(tb)),
"columns_equal": list(ta.columns) == list(tb.columns),
}
unequal, dtype_only = [], []
for col in [c for c in ta.columns if c in tb.columns]:
sa, sb = ta[col], tb[col]
if str(sa.dtype) != str(sb.dtype):
dtype_only.append(f"{col}: {sa.dtype} vs {sb.dtype}")
try:
same = np.array_equal(sa.to_numpy(), sb.to_numpy()) or sa.astype(
object
).equals(sb.astype(object))
except Exception:
same = sa.astype(str).equals(sb.astype(str))
if not same:
unequal.append(col)
entry["value_unequal_columns"] = unequal
entry["dtype_only_differences"] = dtype_only
h5["tables"][entity] = entry
h5["weights_identical"] = bool(
np.array_equal(
fra.weights_for("household").values, frb.weights_for("household").values
)
)
h5["mass_log_identical"] = fra.mass_log == frb.mass_log
report["h5"] = h5
else:
report["h5"] = {"written": ((A / DATASET).is_file(), (B / DATASET).is_file())}
if (A / "build_record.json").is_file() and (B / "build_record.json").is_file():
bra, brb = (
json.loads((A / "build_record.json").read_text()),
json.loads((B / "build_record.json").read_text()),
)
report["build_record"] = {"diff": diff_paths(strip(bra), strip(brb))}
else:
report["build_record"] = {
"written": (
(A / "build_record.json").is_file(),
(B / "build_record.json").is_file(),
)
}


# 6. Wall time from the run logs.
def wall(path: Path):
for line in path.read_text().splitlines():
if line.strip().endswith("real") or " real " in line:
return line.strip().split()[0]
return None


report["wall_seconds"] = {"a": wall(A / "run.log"), "b": wall(B / "run.log")}
print(json.dumps(report, indent=2, default=str))
Loading
Loading