Skip to content

Calibrate the ACS local release on a sparse target matrix and bind district SOI targets (state_cd) - #1053

Merged
MaxGhenis merged 14 commits into
mainfrom
us-acs-local-sparse-cd-soi
Sep 30, 2026
Merged

MaxGhenis merged 14 commits into
mainfrom
us-acs-local-sparse-cd-soi

Conversation

@MaxGhenis

@MaxGhenis MaxGhenis commented Sep 28, 2026 •

Copy link
Copy Markdown
Contributor

Summary

d487 (Max, 2026-09-27, amended): congressional-district fidelity comes from ACS-based content, not a CPS-clone bridge. The ACS local release already exists (tools/build_us_acs_local_release.py, 1.59M households). It left district SOI targets out because its materializer built a dense float32 households × targets matrix: 184 GiB for the full surface. This PR removes that wall and adds the district surface.

  1. Sparse target matrix, end to end.
    • Materialize writes a (targets × households) float32 CSR (target_matrix.npz), row-aligned with Unify calibration diagnostics schema and publication #1007's target_registry.json, plus a row-aligned target_roles.json beside a structure-only lean H5. No dense matrix exists at any point.
    • District SOI rows are carriers. The engine pass materializes one geography-free carrier column per SOI concept; each district row is its carrier restricted to the district's households. The first chunk of every run also materializes one district row per carrier directly and compares it with the row actually stored in the CSR, refusing any difference.
    • Calibrate builds each training target from its registry spec, with its value, metadata and hierarchy unchanged, and a callable measure that reads its CSR row (the UK rowwise precedent). The calibrate kernel is unchanged and compiles it with its own build_constraint_matrix. Unify calibration diagnostics schema and publication #1007's calibration_summary.json and schema-8 diagnostics are kept.
    • The dense construction survives only in the differential test.
  2. --soi-mode state_cd: the TY2022 SOI district file's district rows on top of state, with one vintage per state concept. See "Choices" below.
  3. CD holdout and pro-rata baseline.
    • A hash-assigned 10% of (state × SOI concept family) district blocks never reach the calibrator (microcosm.build.holdout.hash_holdout_unit).
    • Held blocks are scored under design, calibrated and state-only weights against a population pro-rata baseline.
    • calibration_summary.json also records Kish ESS nationally, per state and per district, ESS over distinct households (household_spine, household_source_id), the top-1% weight share and household-weight share by spine, at design and calibrated weights.
  4. Development rungs. --sample-fraction accepts f001/f004/f010/f025 (the stacked pool's rungs), stratified by spine × district. Drawn households are weighted by their stratum's inverse sampling rate and each spine is scaled to its full mass; zero-draw strata are recorded. Package refuses any rung but f100.
  5. Calibration evidence and the calibrated H5 are bound to their materialization.
    • run_identity.json hashes the registry, roles, matrix and lean H5.
    • weights_latest.npz and calibration_summary.json carry its digest and the solver settings. --resume and the "already complete" shortcut refuse other materializations or settings.
    • consumer_export.json records the calibrated H5's sha256 and the digest, and spine_qa.json the digest. Finalize and package refuse an H5 or QA evidence from another materialization before any release directory exists.
    • A new materialize deletes the previous calibration outputs, consumer export, spine QA and gate report.
  6. A factor band on the district file. A (state, concept) whose Historic Table 2 / district-file ratio sits more than 1.25× from its measure's median across states loses its district rows (or its bridged state row) and is recorded. See "Choices".

The default state surface calibrates to the same problem as before (checked below). No release is published or packaged here.

Evidence

All numbers below come from runs in this session; paths are under ~/PolicyEngine/_build_artifacts/acs-local-state-cd-20260927/.

Real-data dense vs sparse (the 2026-09-23 state checkpoint, 1,588,854 households × 4,459 targets).

  • Streaming the dense 28.7 GB checkpoint into the CSR gave 24,773,532 nonzeros, exactly the matrix_nnz that run published. The CSR checkpoint is 198 MB.
  • Re-solving it through the new path with the 09-23 settings gave final loss 0.01553 (09-23: 0.015491) and ESS 13,633 (09-23: 13,631).
  • At torch's default thread count, which 09-23 used, the sparse path reproduces the 09-23 dense weights bit for bit: max absolute difference 0.0 over all 1,588,854 weights, final loss 0.015491, ESS 13,631.3 (state/dense_comparison.json). A first rerun at 4 threads differed by a median 0.58%, which is thread-order float nondeterminism, not the sparse path. On the fixture, the differential test shows the two paths give bitwise-identical problems and weights.

Full-scale state_cd (engine-free, same checkpoint).

  • How the rows were built. Each district row whose state parent is a Historic Table 2 row is exactly that parent's CSR row restricted to the district, because parent and child materialize identically apart from geography, and state_cd_soi_surface refuses otherwise. The three district-file-only families (charitable, interest paid, QBI deduction) have no engine column in that checkpoint and are left out.
  • Surface (v2/). North Carolina included, before the band. There are 3,972 state admin targets and 487 population targets, plus 19,215 district rows whose parent is a Historic Table 2 row. Matrix nnz is 24.8M state rows plus 12.9M district rows = 37.7M, and the whole CSR is about 0.45 GB as float32. 2,394 district targets are held out (103 of 860 state × family units).
  • Held-out district targets (v2/, 2,394 targets in 103 units; lower is better):
Weights Mean abs rel error Median Within 10% Capped loss
State-only calibration (today's method) 55.7% 34.6% 17.9% 0.404
Design weights 48.3% 32.7% 16.2% 0.382
state_cd calibration 33.9% 19.7% 35.1% 0.280
Pro-rata baseline 34.2% 17.2% 32.9% 0.268

The first run (state_cd/, NC excluded, 2,310 held targets) gave the same picture: 34.2% against pro-rata's 34.6%, and 56.3% for state-only. v2's train fit is loss 0.0298, 94.0% within 10%, ESS 19,664.

The band barely touches this evidence. v2 predates the band. On the banded surface, its held-out set loses one unit (Utah tax-exempt interest, 4 targets). Rescored without it, the result is state_cd 33.85% against pro-rata's 34.22% (medians 19.7% and 17.2%), better on 48.9% of 2,390 targets. The band's other change is dropping 30 Hawaii and New York rows from training, out of about 21,000. That would need a re-solve to measure, and I did not run one: it is a heavy engine-free calibration, and the host guidance keeps those off the shared Mac.

  • District targets cut held-out district error by 39% against today's state-only calibration, which is worse than the design weights. Against pro-rata they only tie (mean 33.9% vs 34.2%, median 19.7% vs 17.2%; better on 48.9% of held targets). They win on income-linked families (AGI 6% vs 34%, income tax 3–4% vs 19–21%, ordinary and qualified dividends 26–40% vs 73–92%, itemized deductions) and lose on benefit-linked ones (EITC 40% vs 23%, unemployment, Social Security, medical, IRA, Schedule C, rental). Train fit: loss 0.030, 94.0% within 10% (state-only: 0.016, 97.6%).

  • Weight by origin. Adding district targets moves household weight off the ASEC donor rows (random geography) and onto ACS rows (real PUMA geography): donor share 68.0% → 48.0%. The design weights split 50/50.

    Concentration, as the Route A session asked (concentration_v2.json). The state-only arm is the 09-23 weights, reproduced bit for bit; state_cd is v2/.

Design State-only state_cd
Kish ESS, rows 36,288 13,631 19,664
ESS, distinct households 30,098 10,995 16,876
Top-1% weight share 50.7% 69.3% 51.3%
Median state ESS 736 304 323
Median district ESS 79.5 33.5 49.8
Massachusetts ESS (its districts) 1,020 (108–134) 446 (32–68) 479 (34–65)

Flag: state_cd lowers ESS in some places. Relative to state-only it lowers ESS in 17 of 51 states (worst WV −25.5%, NM −17.3%, HI −13.6%) and 99 of 436 districts (worst MN-06 −44%, WA-05 −41%, AR-04 −40%). The median effect is +10.7% per state and +33% per district, and +44% nationally. The cap and l2_lambda are unchanged: that is the open methodology question (low_effective_sample_size_lambda_zero), and this PR only reports it. The first run (concentration_v1.json) gave the same picture: 18 states and 111 districts lower.

  • Memory. Peak RSS 8.1 GB for the whole solve.

Real-engine smoke on Modal, final code (smoke4-modal/; run acs-state-cd-smoke4c-f004-20260930, commit 7d536ab, --soi-mode state_cd --sample-fraction 0.04 --hh-chunk 700, capped pilot staging, policyengine-us):

  • Stages. materialize, calibrate, qa and finalize all COMPLETED. Package refused before any release directory, because the staging is a capped smoke. The fetched state verifies against the final receipt: 23 files, no problems.
  • Cost. About $0.57 at list price. That includes a materialize attempt that Modal preempted and restarted.
  • Surface. Counts match the pinned contract exactly: 25,864 SOI rows, 21,743 district rows, 8 band drops. The district file's internal gap is 4.0e-16 over 2,193 blocks. The matrix is 26,504 × 2,082.
  • Carrier checks. The first-chunk row check passed on all 51 carriers. The per-chunk parent rebuild checked 6,567 blocks, which is every one of the 2,189 blocks in each of the 3 chunks. It covered 22,494 nonzero households, with 0 rows unchecked and no difference.
  • Bindings. The run-identity digest is the same in the summary, consumer_export.json and spine_qa.json. The weights digest is the same in the summary, the export and the saved weights. The H5 sha is the same in the export, the file and QA's certified bytes. The solver settings match, and the lean-H5 digest matches its file.
  • Gates. All 10 finalize gates are green. The calibration fit is poor at this size (loss 0.87 on 2,082 households); the smoke tests the pipeline, not the fit.
  • Staging checks. Every one of the staging's 62,240 households has a district inside its own state on the 119th plan, so the strict per-chunk check cannot false-fail on a stray district.
  • Lessons. f001 of this staging draws no ACS household (each spine × district stratum floors to zero), and the sampler refuses it by design. The first run found that Unify calibration diagnostics schema and publication #1007's hierarchy-bearing specs broke carrier renaming; that is fixed, with a regression test using production-shaped hierarchies.

Choices and why

  • One vintage per state concept: Historic Table 2.
    • It is the level state already calibrates to, so a state_cd build stays comparable to Build O, Build P and 09-23 state by state.
    • The district file is TY2022 data stamped tax year 2023. It is therefore aged one year less, and its taxable interest is not rebased to Table 4.3. A within-state share is immune to both; a level is not.
    • Two vintages of one concept are contradictory constraints.
    • District rows are rebased to their Historic Table 2 parent as shares.
    • Rebase factors, by measure across the 43 states with district rows, are mostly 1.00–1.10. Counts sit near 1.017 and are never aged, so that is the district file's ~1.7% coverage gap. Amounts sit near 1.05, which is that gap times the ~3% aging gap. Taxable interest is 2.35–2.70 (the Table 4.3 rebase).
    • The six measures Historic Table 2 lacks (charitable, interest paid, QBI deduction; 302 state rows and 2,562 district rows) keep the district file's own state row. It is lifted onto the Historic Table 2 basis by a sibling's two state levels (STATE_CD_LEVEL_BRIDGES: return_count for counts, itemized_deductions_amount or adjusted_gross_income for amounts), so every state level shares one basis. Bridge factors: counts median 1.016, itemized amounts median 1.062, QBI amounts median 1.049. The bridge is an estimate: it assumes the measure shares its sibling's ratio in that state. It is stamp-invariant.
  • Factor band: 1.25× from the measure's median across states (STATE_CD_FACTOR_BAND).
    • The median absorbs the part of a factor common to every state: the coverage and aging gap, or a measure-wide rebase such as taxable interest's Table 4.3 factor or capital gains. That is why no separate register of known measure-wide rebases is needed.
    • A state beyond the band is one where the two publications disagree about that state. Neither the district shares nor a bridge through that sibling is trusted there.
    • On the pinned feed it drops Hawaii's ordinary and qualified dividends (2.03× and 2.37× the median), New York's rental income (0.68×) and Utah's tax-exempt interest (1.32×). That is 34 district rows; the Historic Table 2 state rows stay.
    • It also drops the Wyoming and South Dakota itemized bridges (1.65× and 1.36×). Those are 4 state rows; both states are at-large.
    • The largest gaps it keeps are 1.22× (Wisconsin partnership income, West Virginia capital gains), so any tolerance from 1.22 to 1.31 drops the same eight. The surface is 25,864 SOI rows (21,743 district) in 2,189 blocks.
  • Off the surface:
  • North Carolina is bound: Build the 117th->119th CD crosswalk from the plan registry (fixes North Carolina) #1043's registry-built crosswalk fixed its 117th→119th mapping (37% of NC's population had mapped wrong before). A test pins the crosswalk digest the exclusion register was reviewed against.
  • 117th plan.
    • District targets reach the households' 119th districts through the packaged block-population crosswalk, which assumes returns spread with population inside each 117th/119th intersection.
    • The registry (Add the US block -> congressional-district plan registry (117th-120th) #1041) plus a household congressional_district_geoid__117th_congress column (location v1, coordinated with that session) would make them exact block sums.
    • The carrier split masks on one named household column, so that is a parameter change.
  • Holdout unit: (state, concept family).
    • A single held district is pinned by adding up whenever its state total and sibling districts are trained, so it would score perfectly for free.
    • Families group measures that add up; every EITC row is one family.
    • The US-holdout port session builds its train/holdout/sealed roles on the same hash_holdout_unit.
  • Sigma. No fact in the pinned feed carries uncertainty (IRS SOI is administrative). target_roles.json records sigma: null with its basis, and the loss is unchanged. A σ-weighted objective is a separate ruling (REPORT recommendation 7).

Invariants

Each holds for every input and is tested; P marks a Hypothesis property, D a differential test, F a pinned-feed test.

  1. Sparse = dense on every target (D, P).
    • D: the kernel compiles the identical constraint row from a CSR row as from the same values stored as a dense float32 household column.
    • D: the carrier-split CSR through the checkpoint gives the kernel a bitwise-identical CalibrationProblem and bitwise-identical calibrated weights.
    • P: random matrices (negatives, tiny and huge float32 values) compile to identical rows.
  2. Carrier split = direct materialization for every district SOI row (D over AGI bands, filing status, itemizer, positive-EITC and specs without state_fips). At runtime, the first chunk checks one stored row per carrier against its direct row. Every chunk rebuilds each state parent's direct column from its block's stored district rows, bit for bit; a test corrupts only a later chunk and is refused.
  3. District rows add up to their state parent within 1e-9 relative (P; F on all 2,189 blocks; the builder re-checks its own output and refuses a failure).
    • Each state concept is bound once.
    • A rebase only rescales a block: district shares are kept.
    • States with one source-plan district bind no district rows.
  4. Held-out targets never reach the calibrator (P; a spy on calibrate checks it end to end).
    • The hash holdout is deterministic, order- and insertion-free, and nested in its fraction (P).
    • Its rate matches the fraction.
    • A held unit is held whole.
  5. A pro-rata block sums to its state parent (P).
  6. Distinct-household ESS never exceeds row ESS and equals it when no household repeats (P).
  7. A bridged district-file-only state level = its district-file value × the sibling's Historic Table 2 / district-file ratio in that state (P). F: the six measures are bridged in all 51 states, except the band's two itemized outliers.
  8. The band keeps a block iff its factor is within 1.25× of its measure's median (P). A dropped block is recorded and its state row stays (P). F: exactly the eight drops above.
  9. Integrity.
    • Calibrate and finalize refuse a registry, roles file, matrix or lean H5 whose bytes changed.
    • --resume, the complete shortcut, finalize and package refuse weights or a summary that are unstamped or carry another run identity. The digest changes with the staging file, surface, holdout and sample; resume and the shortcut also refuse other solver settings.
    • Finalize and package refuse a calibrated H5 whose sha or run identity differ from consumer_export.json's, an export and summary that record different weights, and QA evidence from another materialization or of other bytes.
    • The district file must agree with itself (state row = Σ its districts) to 1e-6 wherever both exist (F: max gap 4e-16 over 2,193 blocks), and a bridge sibling with a sign conflict is refused.
    • Materialize refuses a district row the pro-rata baseline cannot score.
    • The attach step re-draws the recorded rung and refuses a different selection.
    • Package refuses dev rungs.
    • The refresh recipe round-trips the recorded holdout fraction, 0 included.
  10. Rungs (D). With one weight per stratum, every drawn stratum keeps its exact mass and each spine its total. Zero-draw strata are recorded, and their spine's mass is spread over its drawn strata.
  11. The CD holdout hash is pinned by golden values, so a formula change cannot silently redraw the holdout.

Review

A six-lens adversarial review ran on the rebased branch, with every finding independently verified: 24 confirmed, 2 rejected. All are addressed in 698a343, 74a5fd8 and 65db50b:

  • the carrier-hierarchy crash, which the real-engine smoke found too;
  • the stale North Carolina and holdout text in the release register;
  • the district-file-only level basis (now bridged);
  • resume, shortcut, finalize and package accepting stale evidence (now bound to the run identity);
  • the recipe dropping --cd-holdout-fraction 0;
  • the carrier check comparing a recomputation rather than the stored rows;
  • the rebase-factor mechanism claim (counts are not aged);
  • rung weighting;
  • stale targets.json references.

A second independent review (subfleet run --task review --tier standard) returned REQUEST_CHANGES. Its code findings are addressed in d4c66f8 and 6596998:

  • unbounded rebase and bridge factors, now the factor band, with the doc's "measures exactly" claim corrected;
  • the calibrated H5 and QA not bound to the run identity;
  • pro-rata populations checked only after the solve;
  • a carrier check that could be vacuous, now every chunk;
  • solver settings and the lean H5 missing from the identity;
  • nits.

A third independent review approved (conditional on CI). It fixed-checked every earlier finding and raised seven lower-severity ones, all addressed in e190723:

  • the H5 and summary are now bound by a calibrated-weights digest, so an interrupted recalibration cannot pair them, and a solve deletes previous outputs first;
  • a bridge is distrusted if either verdict on its sibling fails, and a sign conflict is refused;
  • the district file must agree with itself to 1e-6;
  • the per-chunk check covers every mode and cannot be vacuous for state_cd;
  • QA must certify the H5's bytes;
  • package's evidence checks run before the release directory exists.

A fourth review, of that delta, also approved (conditional on CI and the feed-gated contract test). Its remaining points are addressed in the next commit:

  • the "already complete" shortcut also requires the summary's weights digest, so a checkpoint summarized before the digest cannot loop;
  • a household outside its state's district block gets a message naming it;
  • the district file's largest internal gap is recorded in the receipt and pinned;
  • the wording is corrected.

The 1e-6 internal-consistency refusal holds on the pinned feed: all 2,193 district-file blocks match their own state rows, with a largest relative gap of 4.0e-16.

A fifth, confirmation review of c00b096 approved with no regressions. Its two wording nits on the new refusal message are applied in 225fafd.

Coordination

Decisions for Max (queued as d619, d620 and d621; not needed to merge this opt-in mode)

state_cd is opt-in; the default surface is unchanged. Using it in any build or release is a methodology call:

Not in this PR

  • A full-scale state_cd run through the engine. No existing staging H5 passes the current reviewed-null contract (is_spm_independent_minor_role, chip filed), and staging has to be rebuilt first. Materialize RSS is now the staging frame plus one engine chunk, and the dense-checkpoint jump from 55 GB to 78 GB is gone.
  • Publishing anything. That needs Max's word.

axiom: n/a: calibration infrastructure and target-surface selection; no policy rule changes.

🤖 Generated with Claude Code

MaxGhenis and others added 14 commits September 28, 2026 15:34
… state_cd SOI surface (WIP)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…s; add the engine-free evaluation

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… evaluation to the target-registry checkpoint

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…r-district ESS and the top-1% weight share

The real-engine smoke on the rebased branch found that #1007's specs carry
calibration hierarchies whose target id must equal the spec name, so a renamed
carrier cannot keep one. The fixture's district specs now carry
production-shaped hierarchies.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…d calibration outputs to the run identity, check stored carrier rows

- District-file-only concepts (charitable, interest paid, QBI deduction) are
  lifted onto the Historic Table 2 basis by a sibling's two state levels
  (STATE_CD_LEVEL_BRIDGES), so every state level shares one basis.
- weights_latest.npz and calibration_summary.json carry the run-identity
  digest; resume, the complete shortcut, finalize and package refuse stale or
  unstamped evidence, finalize re-verifies the checkpoint digests, and
  materialize removes the previous calibration outputs.
- The first-chunk carrier check compares the CSR rows the assembler stored.
- The refresh recipe names --cd-holdout-fraction for every state_cd build,
  0 included, at full precision.
- Development rungs weight drawn households by their stratum's inverse
  sampling rate and record zero-draw strata.
- The reviewed-limitations register is built from the surface receipt (no
  stale North Carolina text) and names where the holdout is scored.
- The CD holdout hash is pinned by golden values.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…t the level bridge, rungs, binding and ESS

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…gh the Modal stage plan

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…QA and solver settings to the run identity

Second adversarial review (REQUEST_CHANGES) fixes:

- A (state, concept) rebase factor or level-bridge sibling ratio more than
  1.25x from its measure's median across states is dropped and recorded
  (receipt factor_band.out_of_band). On the pinned feed that is Hawaii's
  ordinary and qualified dividends, New York's rental income, Utah's
  tax-exempt interest, and the Wyoming and South Dakota itemized bridges:
  34 district rows and 4 bridged state rows. Contract counts updated
  (25,864 SOI rows, 2,189 blocks); the doc no longer says the bridge
  "measures exactly" the gap.
- Every chunk rebuilds each state parent from its block's stored district
  rows and refuses any difference, so the carrier check cannot be vacuous.
- consumer_export.json records the calibrated H5's sha256 and the
  run-identity digest, spine_qa.json the digest; finalize and package refuse
  an H5 or QA evidence from another materialization before any release
  directory exists. Materialize deletes them and the gate report.
- The run identity binds the lean H5; weights and summary carry the solver
  settings, and resume or reuse under other settings is refused.
- Materialize refuses a district row without positive ladder populations
  for the pro-rata baseline, before the solve.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…_cd surface is built

The reconciliation ran only in the feed-gated test; the builder now refuses
its own output if any block fails to add up to a bound parent.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…link the band's verdicts, check district rows in every mode

Second independent review (APPROVE, conditional on CI) findings, all fixed:

- A recalibration that stops between writing the H5 and the summary can no
  longer pair new weights with an old summary: consumer_export.json and
  calibration_summary.json both record a digest of the calibrated weights
  (and the export the solver settings), finalize and package require them
  to match, and a solve deletes every previous output but the resume
  weights before it starts.
- A bridge sibling is distrusted if either band verdict fails: its two state
  levels, or the rebase of its own district block (the medians differ, since
  at-large states count only in the first). A sign conflict in the sibling
  is refused before the band. Wherever the district file has a state row
  beside its district rows, the two must agree to 1e-6, or the surface is
  refused as a defect inside the file.
- The per-chunk check now covers every mode: state_cd blocks rebuild their
  named parent bit for bit (a state_cd row without a compiled parent is
  refused, so the check cannot be vacuous); other modes check each district
  row against a same-concept state row on its own households.
- Finalize and package refuse QA evidence of other bytes; package's evidence
  checks run before the release directory exists; the H5 path is no longer
  compared (the sha binds the bytes); solver literals are shared constants;
  QA reads the run identity instead of rehashing the staging file.
- Docs: the bridge comment no longer says "exactly"; the tolerance range is
  stated against the largest kept gap (1.220) and Utah's (1.316).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…seholds outside their state's block; pin the district file's internal consistency

Fourth review (APPROVE, conditional on CI and the feed-gated test):

- The "already complete" shortcut now also requires the summary to record
  the saved weights' digest; a checkpoint summarized before the digest
  existed is refused with "run without --resume" instead of looping.
- A state_cd block that fails to rebuild its parent because households sit
  in no district of the block now names them (their district is not one of
  their state's current-plan districts) rather than blaming the carrier split.
- The receipt records the district file's internal consistency (blocks
  checked, largest relative gap); on the pinned feed all 2,193 blocks agree
  to 4.0e-16, which the feed-gated contract test pins.
- Dropped a per-chunk count check that could never fire; docs and changelog
  scoped (fallback check's shared key, what a solve deletes).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…, and allow for a missing district row

The confirmation review's two wording nits on the new refusal message; no
behavior change.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…checkpoint through finalize and package

Closes the test gap two reviews noted (every finalize/package fixture used a
pre-sparse identity, so the binding returned early). A consistent checkpoint
round-trips; an interrupted recalibration is refused by finalize before any
gate report exists, and QA from another materialization by package before
any release directory exists.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@MaxGhenis
MaxGhenis marked this pull request as ready for review September 30, 2026 15:04
@MaxGhenis

Copy link
Copy Markdown
Contributor Author

Merging after every gate held on head 7d536ab:

  • CI: gh pr checks exits 0; all 10 checks pass. The PR is MERGEABLE and CLEAN, with no CHANGES_REQUESTED review.
  • Independent reviews (subfleet run --task review --tier standard, Opus 5.5):
    • 20260928-170111-pr1053-review asked for changes, all fixed in d4c66f8 and 6596998.
    • 20260928-183338-pr1053-review2 approved; its findings are fixed in e190723.
    • 20260928-224246-pr1053-review3 approved, conditional on CI and the feed-gated test; findings fixed in c00b096.
    • 20260928-233550-pr1053-review4 approved c00b096 with no regressions.
    • Later commits contain only review4's two wording nits (225fafd), an end-to-end binding test (67a36f5) and a clean merge of main (7d536ab).
  • Feed-gated contract tests pass locally on the pinned feed.
  • Real-engine smoke on Modal passed at 7d536ab (materialize through finalize), with every binding and gate checked. Details are in the description.
  • Not needed for this merge: state_cd is opt-in and the default state surface calibrates to the same problem. Using state_cd in any build is queued for Max as d619 (factor band), d620 (bar against pro-rata) and d621 (holdout families). Nothing is published.

@MaxGhenis
MaxGhenis merged commit 482efbb into main Sep 30, 2026
10 checks passed
@MaxGhenis
MaxGhenis deleted the us-acs-local-sparse-cd-soi branch September 30, 2026 15:07
MaxGhenis added a commit that referenced this pull request Sep 30, 2026
…options into its solver stamp

l2_basis and mass_parametrization join _solver_settings, so main's resume and
already-complete checks cover them; a stamp written before they existed reads
as the historical solve. calibrate_surface takes both with historical
defaults. My separate resume guard is dropped in favor of the stamp.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant