Skip to content

model-v2: certified bundle, engine port and guarded rehearsal - #24

Draft
MaxGhenis wants to merge 261 commits into
mainfrom
model-v2
Draft

MaxGhenis wants to merge 261 commits into
mainfrom
model-v2

Conversation

@MaxGhenis

@MaxGhenis MaxGhenis commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Moves the engine to the released certified policyengine[uk]==6.2.6 bundle: UK 2.125.1, core 3.33.0 and Enhanced FRS 1.58.0. The released wheel, immutable HF manifest and downloaded dataset SHA-256 are independently verified. The upstream audit covers fiscal lists, GB masks, resolved household heads, pension types, additional pension and the remaining horizon extensions. Publication and validation entry points enforce certification before model jobs.

The historical 2.120.0 / 1.56.16 pilot remains uncertified. Frozen source/formula archives and independently pinned privacy proofs authenticate retained fiscal content. Whole single-record diagnostic families are omitted while surviving values and original provenance stay unchanged. New publishers and dashboard readers omit unsupported concentration diagnostics.

OPEN draft; pre-Budget numerical execution is incomplete. Host load exceeded the documented admission limit of 36 on 18 CPUs; a fresh vm_stat/load check at 11:51 ET found load 148.79. The certified central/40-path bridge, GB-versus-DWP coverage, Microcosm load check and timed 20 Enhanced FRS + 2 Microcosm sample did not complete. No certified savings or current measured runtime estimate is available to quote.

The ruling-(c) rehearsal uses an unauthorized dry handoff and private output directories. It cannot authorize a binding build or write results.json. Receipts retain safe timing/invariant metadata; completed cold timings exclude cancellation races. Runtime planning accounts for serial stages and leaves estimates unavailable without measurements. Bridge publication checks linked gross/net/offset, UK/GB and HB complements under the ten-record rule. Provenance is frozen across stages and checkpoints, and final receipt replacement rolls back on drift. Ageing reports stay private pending a separate joint contrast support audit.

Independent HARD review: APPROVE at 590ddac292bc9ec54d8f513bb99e288ad9dfb612; review record. This verdict covers implementation and explicitly excludes the missing numerical closeout.

Validation at final head b31137ea055ab7e2b2ef7dd8b63df1302f5affdf: Dashboard passed lint, tests and build. Pipeline is running the full suite; exact-head Python validation remains pending. Both workflows passed at the preceding documentation head 89a1309: 1,544 Python tests passed, one skipped and exactly three permitted stale-results failures (615.61 seconds; completion gate confirmed 1,548 collected tests). Local dashboard: 95 passed, lint zero errors/two existing warnings, production build passed. No unexpected test failures are waived.

Execution report, certification, upstream audit and REBUILD record the evidence and remaining work. After committed 28 October inputs: binding C2/from 29 October uk-triple-lock-c2, certified rebuild, review, then paper/dashboard update. The launchd host schedule still needs verification. Rehearsal and historical pilot outputs supply no binding replacement figures.

vahid-ahmadi and others added 30 commits October 1, 2026 16:07
…peline constant

The engine records how a run treats the population in fixed_inputs.population
(weights, ages, pension types), read from what it pins and the model's ages,
and the pipeline words the item from it, with #14 §3's vocabulary for ONS
reweighting and cohort types; a missing record, a value without wording or a
missing figure fails the build. The dashboard computes the strip only for a
file with no block; a damaged block shows its valid items or nothing. The
paths item's calibration, floor and shock-model wording come from the results.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…est reads the figure, not a literal

With the backtests on the fiscal 0.1-point rounding, the realised 2012-start gap sits at the 55.8th percentile of
the untilted model (was 55.2), so the method text and docs/METHOD.md say 56th; the tilted model stays 93rd. The test
now checks the text quotes whatever percentile the file records, so it fails until the rebuild, like test_not_stale,
and can't drift again. Also notes that a descendant starting its own session (setsid) is not reached by run_child.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…n and data/scenarios

pipeline.build runs every registered scenario after the paths (the same cached
engine jobs as --scenario) inside the build's start/end check and returns them
with the results; triple-lock-build writes each to data/scenarios/NAME.json.
--scenario NAME still reruns one alone. A test checks that a scenario run holds
the central run's fixed inputs (a key only later engines record, such as
population, is compared once both files have it); stubbed tests check the
build writes every scenario and what the scenario jobs and records hold.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…el run; figures unchanged)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… by the dataset load

The dataset load (_managed) declares what it did to the survey's weights and
ages (engine.LOADED_POPULATION), since a change made inside the dataset leaves
nothing to compare against; population_treatment fails without a declaration
and checks it against pinned weights and the model's ages. Pension types are
recorded from the rule pinned_inputs follows (PENSION_TYPE_RULE), replacing a
comparison that could not fail. The block fails unless the central run and
every trajectory record the same treatment. Sensitivities outside the
reweightings to past dynamics, and models without a short name, are worded
generically; scripts/check_assumptions.py checks a pilot's results in seconds.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…what a finished job left behind

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…the OBR-wedge scenario

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ions block from the results

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

# Conflicts:
#	docs/RESULTS_SCHEMA.md
#	tests/test_results.py
… review

How jobs are started, isolated and stopped never changes what they compute, so it now
lives in triple_lock/jobs.py, outside engine.ENGINE_FILES: a runner fix no longer forces a
full rebuild. engine.py keeps the jobs themselves, the cache key and the cache. The job
process is `python -m triple_lock.jobs --job`, which watches its parent and runs
engine._job. Tests: no engine file imports jobs.py, and the cache stores each job's output
exactly as the job wrote it.

Fixes from the independent review of b4607a2:
- run_jobs keeps its signal handlers while it stops the running jobs; kill_children's
  SIGKILL pass runs even if a second signal cuts the grace period short; run_child kills
  a child started just before an interruption (no gap before its try).
- A failed job is named in the log as soon as run_jobs sees it, not only when the running
  jobs finish; failures seen before a Ctrl-C are logged before the stop propagates.
- rules.floor_binds counts a tie with 2.5% as the floor, as triple_lock_source does
  (unused in the pipeline).

Changes engine.py's and rules.py's code, so it reruns every job: it is meant to ride on the
next full rebuild.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 4b15274)
- A failing worker sets `stop` itself, so the job queued behind it never starts (on one
  worker it sometimes did: 16 of 100 trials).
- run_child kills the rest of a job's process group when the job exits, so nothing it
  left running outlives it or holds its output pipes open.
- The cache key parses source bytes, so it decodes a file as Python does (a coding cookie
  moves it, a byte-order mark does not); a u prefix no longer moves it.
- pyproject.toml counts only its dependencies, optional dependencies, Python version and
  build system: rewording the description or bumping the version reruns nothing.
- __init__.py, which runs in every job, is in the key. The docstring guard also catches
  "__doc__" as a string and getdoc imports; the import walk sees absolute imports.
- held_pension_types reports shapes, not sizes; model_triple_lock_rate takes float()s,
  since round() of a numpy float rounds as numpy does.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 07c7bb2)
….py): one docstring

#16 kills a finished job's process group before the worker directory is released;
07c7bb2 kills it as soon as the job exits. jobs.run_child keeps both; #16's test now
runs against jobs.run_child.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…zon to policyengine-uk 2.102.5

policyengine-uk 2.102.5 (PolicyEngine/policyengine-uk#1901) lets families whose
adults are all over State Pension age make new Housing Benefit claims (SI
2014/1230 reg 6A(4)). The example renters therefore no longer need a reported
claim (housing_benefit_reported), the claims_all_entitled_benefits override, or
would_claim_uc False. No other example input is one of the seven reported
benefits that switch claims_all_entitled_benefits off, so its formula gives True
and the override goes too.

model_horizon copies policyengine-uk's parameter builders and hash-checks them.
Under 2.102.5 three of them changed:
- lag_cpi and lag_average_earnings now build through
  lagged_series.add_lagged_parameter and derive their end from the source;
- convert_to_fiscal_year_parameters skips parameters flagged
  preserve_calendar_dates (the alcohol and tobacco duties).
The copies now mirror those builders, still extended to 2042.

Verified on all four published paths, both rules, 2027-28 to 2039-40 (2,912
cells). policyengine-uk 2.90.2 with the three inputs and 2.102.5 without them
give bit-identical State Pension, income tax, Pension Credit, Housing Benefit,
council tax reduction, Winter Fuel Payment and net income for every example.
The rerun at 2.90.2 matches data/results.json exactly.

New tests:
- no example sets a take-up input, and every take-up flag computes True;
- Hypothesis differential: the old inputs change none of those outputs for a
  wholly pension-age family (single or couple, any tenure). It fails on
  2.90.2 and passes on 2.102.5;
- differential: every processed parameter under model_horizon equals
  upstream's on every date to 2034. It fails if the calendar-dates skip is
  removed.

Not mergeable yet: the pin stays at policyengine[uk]==5.3.0 (policyengine-uk
2.90.2), and under that pin check_upstream refuses these hashes. It needs a
policyengine.py release bundling policyengine-uk >= 2.102.5, then a full
rebuild.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit ccf4996)
…t policyengine.py; record the pair as uncertified

policyengine.py 6.2.1 pins policyengine-uk 2.102.3 and raises on import beside 2.118.0
(its data certification finds no release certified for the runtime model), and no bundle
yet carries the pensioner fixes (d778). So:

- pyproject.toml pins policyengine-uk==2.118.0 (with huggingface-hub and openpyxl, which
  policyengine.py used to bring), with a comment saying why; requirements-lock.txt is a
  clean install of .[uk,dev] constrained to the previous versions, so only policyengine.py's
  own packages leave and policyengine-uk/core move (core 3.32.16).
- triple_lock/datasets.py: the Enhanced FRS 1.56.16 (the build 5.3.0 and 6.2.1 certified),
  1.57.4 and Microcosm populace_uk_2023, each pinned to a revision and a SHA-256, fetched once
  into .cache/datasets and hashed before every use, as the managed loader did.
- engine._managed loads the file by path (each simulation reads it afresh and extends it with
  its own parameters, as before) and attaches the provenance; run_path reads the unreformed
  parameters without loading a dataset (one load fewer per path).
- Every run records model: the installed policyengine-uk and core, the dataset's pin, and
  certified: false with the reason. The results and scenario provenance carry it in place of
  release_bundle; the dashboard's replication line reads either.
- datasets.py is an engine file (in the job key); TRACKED_PACKAGES drops policyengine and adds
  tables.

Tests: the pins agree across pyproject and the lock and with the environment; the recorded
model version equals the installed policyengine-uk (and __version__ where a release has one;
2.118.0 has none); materialize fetches once, refuses a file off its pin. Both Enhanced FRS
releases load and compute on 2.118.0 with no warnings and no NaN (pilot load check).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
MaxGhenis and others added 30 commits October 11, 2026 06:36
The bridge and rehearsal runners refused whenever one-minute load exceeded
twice the CPU count (36 on the 18-CPU host). On a shared host load measures
contention, not danger; Fleet ops measured 150-350 for much of 11 October,
which stalled Part J with 57-65 GiB free.

One shared helper, triple_lock.host, now decides admission for both: refuse
iff available memory is below the runner's floor or load exceeds
TRIPLE_LOCK_MAX_LOAD_PER_CPU x CPUs (default 6, Subfleet's background guard;
--max-load-per-cpu overrides). Non-finite readings refuse; a malformed limit
stops the runner at startup. Memory floors are unchanged (bridge 40 GiB;
rehearsal 40, 40 x workers for Microcosm, 8 during a job). Every host check
records minimum_available_gib, max_load_per_cpu, load_limit and admitted;
the configured limit and its source go in provenance.host_admission (bridge)
and host_admission (rehearsal).

host.py is outside engine.ENGINE_FILES: no job key, recipe hash, C2 handoff
fingerprint or committed receipt changes.

REBUILD documents the cutoff and the Fleet ops reservation
triple-lock-rebuild-2026-10 (ready/release steps), and corrects step 2: the
binding C2 re-score is the uk-triple-lock-c2 step of the loaded launchd job
com.maxghenis.triple-lock-2027, not a separate label.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch was successfully deployed

1 active deployment
Preview — 0e02f268 Deployed Oct 11, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants