Skip to content

Chi-square design-weight penalty and softmax mass parametrization for calibration; ACS local ESS frontier - #1078

Merged
MaxGhenis merged 24 commits into
mainfrom
l2-design-basis
Oct 4, 2026
Merged

MaxGhenis merged 24 commits into
mainfrom
l2-design-basis

Conversation

@MaxGhenis

@MaxGhenis MaxGhenis commented Oct 1, 2026 •

Copy link
Copy Markdown
Contributor

Why

On 2026-09-28 David Trimmer measured concentrated weights in the ACS local release:

  • national Kish ESS about 14k of 1.59M;
  • Massachusetts 477, with half its weight on 152 records;
  • district ESS 41–64.

He asked whether to anchor calibration to the design weights. This PR adds that anchor and tests it at full scale. The full write-up is experiments/us-acs-local-l2-basis-20260928/README.md; the method is in docs/calibration-l2-basis.md.

What the evidence shows:

  1. Most of the concentration is in the starting weights. The staging's --acs-share 0.5 gives the 57,240 donor rows half the mass. Before calibration, national ESS is 36,288 and Massachusetts 1,020. Calibration takes ESS to 13,646.
  2. The old l2_lambda knob backfires. Its record-weighted form pulls toward w ∝ d², and at λ = 0.1 it lowers ESS to 8,689.
  3. The new chi-square basis works. It is GREG's distance to the design weights, which it pulls toward directly. At λ = 0.03 with the default projection parametrization (on today's equal-weight loss), against the reproduced release:
    • national ESS goes from 13,646 to 20,690;
    • Massachusetts from 448 to 589;
    • districts below ESS 50 from 352 to 278, and the smallest district from 12 to 22;
    • training targets within 10% from 97.6% to 97.3%;
    • on two rotated holdout folds, held-out fit is ahead on both measures: capped error 0.1247/0.1191 against 0.1296/0.1236, and share within 10% 64.2%/67.0% against 63.3%/65.8%. Held-out run-to-run variation was not measured and λ was picked on those folds, so read this as "no worse, probably slightly better".
  4. The ACS share is the bigger lever, at an out-of-sample cost. At share 0.9: ESS 93,524, no district below 50, and about 40% more held-out error, mostly on SOI targets. That and the default λ are queued for Max, not changed here.
  5. A release gate. No district below a quarter of its starting ESS separates cleanly at the current seeding: the release fails it in 15-24 districts across its three solves, and every penalized solve passes. A state floor near 200 sits inside the solves' spread across folds (177-216), so none is proposed.
  6. The softmax mass parametrization is opt-in, and projection stays recommended. Under mass="conserve", the historical projection solve can stall on small problems where every target pulls the same way: it is up to 0.34 off the exact optimum, against softmax's 0.0034. At the release's scale, gradients are near Adam's eps with mixed signs, so the stall does not occur. The two parametrizations agree for λ ≤ 0.03, and projection falls inside softmax's training frontier at λ ≥ 0.1. Softmax's per-step cap loop runs out of its 32 rounds on most epochs at that scale (counted in the receipt; the closing projection makes the returned weights exact). The doc states this as a known limitation.

What changes

microcosm-calibrate gets two opt-in options. The defaults are byte-identical to bda72cb.

  • l2_basis:
    • "record" (default) is the historical mean((w/d)**2).
    • "chi_square" is sum(d*(w/d-1)**2)/sum(d).
    • Threaded through calibrate, the L0 budget search, refit_l0_selection, calibrate_l0_refit (refit_l2_basis) and static_aging.
  • mass_parametrization:
    • "projection" (default) is the historical scheme.
    • "softmax" is w = total*softmax(log_w), for Adam with mass="conserve" and no L0 gates. calibrate_l0_refit takes refit_mass_parametrization.
    • Softmax epochs whose cap rounds run out are counted in the iterate-selection receipt.
  • CalibrationResult.chi_square_distance (also on L0RefitResult) and chi_square_distance(weights, anchor).

tools/build_us_acs_local_release.py. The changes are flags on top of #1053; other open PRs edit this tool too.

  • New flags --l2-basis and --mass-parametrization, with the defaults unchanged.
  • Both join _solver_settings, so resume and the already-complete shortcut cover them. A stamp written before they existed reads as the historical solve.
  • The summary and the build manifest record them with the realized chi-square distance; a legacy summary reads as the historical solve.
  • The ESS limitation keeps the "certified default" text only for the historical solve, and fills the historical basis for a legacy summary.
  • The manifest's refresh recipe always names --l2-lambda, --l2-basis and --mass-parametrization at the recorded values, so a later default change cannot move it.
  • The summary and manifest record softmax_cap_rounds_exhausted_epochs.

Evidence lives in experiments/us-acs-local-l2-basis-20260928/:

  • the harness, and the CSR checkpoint verified against the dense one;
  • the Modal grid of 57 solves, gated on reproducing the release (loss 0.015485 vs 0.015491, ESS 13,646 vs 13,631), run on three kernel heads whose solve.py differs only in docstrings, a receipt counter and guards;
  • two reruns measuring run-to-run variation (training metrics only). They are not byte-identical to their originals: trajectories differ from epoch 0 or 1, though solve.py's module hash is the same at 9ef71ed, bc763cb and d82ff85 (it differs from 35ad665, where the originals ran, only by docstrings and the counter). Over 800 epochs the release's final loss moved 2.7% (0.3% at λ = 0.03), while every concentration measure moved under 1%;
  • a receipt for the published weights' own concentration (results/published_weights.json);
  • the seeding options, the gradient scale, and the CLARABEL references;
  • the frontier tables and chart.

CI-pinned identities this PR moves

The US spec-engine seed protocol attests the source bytes of microcosm.calibrate.solve, so editing solve.py re-pins EXPECTED_HASHES["seed_protocol"] and ["seed_map"], the US spec_sha256 in test_us_multispine_pool_tool.py, and docs/evidence/spec-engine/us-f0-coverage.json (regenerated; --check passes 41/41). Each new value was computed on this tree (final values: seed_protocol be39eb6b…, seed_map 085d8d39…, US spec_sha256 e0b757ce…).

Invariants

These are property-tested in packages/microcosm-calibrate/tests/engine_free/shared/test_l2_basis.py.

  1. The chi-square penalty is exactly 0 at w = anchor and nonnegative everywhere (Hypothesis).
  2. Differential: the float32 torch penalty equals the float64 chi_square_distance, and the record penalty equals mean((w/d)**2) (Hypothesis).
  3. When sum(w) = sum(d), the distance equals sum(w**2/d)/sum(d) - 1. With uniform d it equals n/ESS - 1 (Hypothesis).
  4. Under a mass constraint, the reduced gradient vanishes at w = d for chi-square and at w ∝ d**2 for the record basis (Hypothesis).
  5. With no targets and mass="conserve", the chi-square solve returns the design weights within 1% from any warm start, under both parametrizations (Hypothesis).
  6. Softmax conserves the total to 1e-12 relative, respects the cap and keeps weights positive (Hypothesis).
  7. Differential against an exact solver, on the tests' own problems (fixture plus generator):
    • each softmax-conserve and free-mass solve lands within 0.006 of CLARABEL's optimum;
    • the bound is twice the measured maximum of 0.0026;
    • the regularization-path identity holds with the slack 2·eps/Δλ derived from that bound, on informative λ pairs only;
    • with uniform design weights, ESS = n/(1+P) holds exactly.
  8. Defaults reproduce the pre-change optimizer byte for byte across 7 configurations, with the defaults implicit and spelled out. This was mutation-tested: perturbing the record penalty by 1e-6, or reusing the loss path's exp node, fails it.

Intended violation, pinned as such: under "projection", invariant 7 does not hold in general. test_projection_parametrization_stalls_under_uniform_mass_pressure shows the stall when every target sits above its design total.

Review

Independent Opus 5.5 reviews, each through subfleet run --task review:

  • Round 1 requested changes (one high, four medium). Fixed: truthful limitation provenance for an unpenalized softmax solve; the missing README; the stall claim qualified by measured gradient scale; test bounds derived from an exact reference on the tests' own problems; the cap-loop docs corrected and exhausted rounds counted.
  • Round 2 died on a lane error after partial findings; those were fixed (no dominance claim, the prior as a limit not a bound, the λ = 1 district count, the resume text).
  • Round 5 approved the round-4 fixes (delta review; six cosmetic nits, fixed).
  • Round 4 approved (no blocker, high or medium). Its lows are fixed: the recipe names every penalty flag; the held-out claim no longer leans on training-loss noise; noise-sized differences are no longer used as arguments; a receipt for the published Massachusetts figure; floor ranges over every solve; the softmax limitation surfaced in calibrate()'s docstring, the CLI help and the summary.
  • Round 3 requested changes (four medium, seven low), all in the write-up or the tool's evidence. Fixed: the cap-loop exhaustion stated as a known limitation, with projection recommended; the frontier claim limited to λ ≤ 0.03; both held-out metrics reported, with fold selection disclosed; the gate led by the relative floor, the state floor dropped; the kernel heads listed and checked; the legacy basis in the limitation text; the refresh recipe's penalty flags; a receipt for the Massachusetts split; labels and nits.

axiom: n/a: calibration kernel and ACS local tool; no policy rules change.

🤖 Generated with Claude Code

MaxGhenis and others added 24 commits September 28, 2026 16:27
…m calibration

l2_basis selects the L2 concentration penalty's form: the historical
record-weighted mean(r**2) (default, bit-identical) or GREG's
anchor-weighted chi-square distance sum(d*(r-1)**2)/sum(d), whose
target-free optimum is the anchor itself. mass_parametrization selects
how mass='conserve' holds the total: the historical per-step uniform
log shift (default, bit-identical) or w = total*softmax(log_w), which
hands Adam the constraint-reduced gradient. Both are recorded in the
result options; chi_square_distance reports the realized distance.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Hypothesis properties for the penalty algebra (zero at the anchor,
nonnegative, float32 torch form equal to the float64 reference, the
weighting-effect identity, each basis's mass-constrained stationary
point), target-free convergence to the design weights, softmax mass
and cap invariants, the regularization-path monotonicity with a slack
derived from the measured optimizer error, the projection stall as a
pinned intended violation, option provenance and validation, and a
byte-for-byte pin of the defaults against the verbatim bda72cb
optimizer.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Both default to the historical solve and take the kernel's names. The
calibration summary and build manifest record them with the realized
chi-square distance; a resumed run refuses weights solved under other
penalty settings (a legacy checkpoint resumes only under the defaults it
was solved with); and the concentration limitation no longer claims
l2_lambda=0 for a penalized run.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… draft

The harness recalibrates the 2026-09-23 ACS local release's 4,459-target
surface from a sparse copy of its checkpoint, verified row-for-row against
3,000 sampled dense rows, with the release's settings plus the run's L2 and
mass options. optimizer_reference.py compares the kernel under each mass
setting with CLARABEL's exact optimum. The frontier sections of the doc
follow the sweep.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…s and the Modal grid

seeding_options.py measures the starting weights' Kish ESS (rows and distinct
source households) for ACS shares 0.5/0.7/0.9/proportional and 1/4/16 donor
location clones. The harness takes an acs_share prior, which becomes the
frame weights, the chi-square anchor and the cap base. modal_sweep.py runs
the grid, gated on reproducing the published release.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e frontier analysis

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…tmax cap receipt

- The ACS tool keeps the certified-default limitation only for the historical
  solve (no penalty, projection); an unpenalized softmax run is recorded as
  its own. Legacy summaries read as the historical solve in the manifest.
  CLI choices are spelled in the tool (pinned equal to the kernel's) so
  parsing no longer imports torch. The resume guard ignores the basis when
  neither side has a penalty and says which settings it checks.
- The path tests now check each solve against CLARABEL's exact optimum of
  the same program on their own problems (committed fixture and generator),
  and derive the monotonicity slack from that bound on a lambda grid dense
  where the slack is informative.
- The softmax cap loop's docstrings describe what it does; epochs whose
  rounds run out are counted in the iterate-selection receipt. The
  projection stall claim now states its conditions (gradients of one sign,
  well above Adam's eps). Numerics are unchanged from 35ad665.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e at full size

At the release's scale the median record's log-weight gradient is 1.2e-8 to
3.8e-8 with mixed signs, so the sharp same-sign stall does not occur there;
the doc now says so and leaves the full-scale comparison to the sweep.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…options into its solver stamp

l2_basis and mass_parametrization join _solver_settings, so main's resume and
already-complete checks cover them; a stamp written before they existed reads
as the historical solve. calibrate_surface takes both with historical
defaults. My separate resume guard is dropped in favor of the stamp.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…he sweep grid

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…floors

Recalibrations of the published ACS local release on Modal (reproduction
gated): the chi-square penalty at lambda 0.03 lifts national ESS 13.6k to
21.8k and beats the release on two held-out folds; the ACS share of the
mass is the larger lever and costs held-out SOI fit. Receipts are compacted
copies of each run's metrics; the README's key table is written by
analyze.py.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The Primary-QRF worker identity pins uv.lock's sha256. Hypothesis already
reaches every test environment through microcosm-graph's dev group, so the
calibrate dev-group entry was redundant; drop it rather than re-pin.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…imit not a bound, and the merged resume stamp

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…oop limit, lead the gate with the relative floor

- README and doc: projection at lambda 0.03 is the recommendation, with both
  held-out metrics on both folds and the fold selection disclosed; the
  frontier claim is limited to lambda <= 0.03; softmax's cap-loop exhaustion
  at production scale is a stated limitation; the gate leads with "no
  district below 25% of its starting ESS", the state floor of 200 is dropped,
  and the relative gate's behaviour at share 0.9 is shown; the release row is
  labelled as reproduced and reconciled with the published and reported
  Massachusetts figures; kernel heads are listed with the solve.py diff.
- analyze.py: parametrization in every key-table label, a held-out
  within-10% column, kernel_head in frontier.csv, and a byte-level
  head-equivalence check for the dup_* reruns.
- Tool: the limitation text fills the historical basis for a legacy summary;
  the refresh recipe carries the penalty flags when they differ from the
  historical solve.
- seeding_options.py emits the Massachusetts split by spine.
- Rename test_path_reference.py to path_reference.py; drop the fixture's
  unchecked module hash; say eps is empirical; assert the unreachable
  softmax best-iterate branch.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…by fingerprint

engine-us CI failed test_us_coverage_is_exact_complete_and_honest: the US
seed protocol attests the installed source bytes of microcosm.calibrate.solve
(spec_engine/seeds.py), so this PR's solve.py edits move three identities:
- EXPECTED_HASHES["seed_protocol"] 91989ed5... -> 0e7ab5fc...
- EXPECTED_HASHES["seed_map"] 3cf10a52... -> f84f4fab...
- the US spec_sha256 pinned in test_us_multispine_pool_tool.py
  ffbb93ed... -> a7eee025...
The new values come from build_inventory_coverage on this tree, where that
was the only failing item. docs/evidence/spec-engine/us-f0-coverage.json was
regenerated with tools/spec_engine_coverage.py; --check passes with 41/41
inventory checks.

engine-free CI failed the path test's byte digest of the generated problem
for one of twelve cases while the other eleven matched. The fixture now
stores per-problem moments (verified locally against the old digests before
they were dropped), compared with rel 1e-9.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The head-equivalence reruns are not byte-identical to their originals: the
trajectories already differ at epoch 0 (softmax, the loss at the starting
weights) or epoch 1 (the release), so the solves are not bit-reproducible
across Modal containers. analyze.py now writes results/rerun_variation.json
(bytes, first divergence, every frontier metric's relative change) instead
of a byte-identity check. The release's final loss moved 2.7% between two
identical runs (0.3% at lambda 0.03); concentration measures moved under 1%.

The held-out margins of projection lambda 0.03 over the release (3.6-3.8% in
capped error, 0.9-1.2 points within 10%) are of that order, so the README and
the doc now say "no worse out of sample" rather than "better", and list one
run per configuration as a limit.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… the published weights, keep claims inside the noise

- release_refresh_recipe always names --l2-lambda, --l2-basis and
  --mass-parametrization, so a later default change cannot move a recorded
  release's recipe; the test parses the historical recipe back.
- The calibration summary and build manifest record
  softmax_cap_rounds_exhausted_epochs; calibrate()'s docstring and the
  --mass-parametrization help say the rounds run out at the ACS release's
  scale.
- published_weights.py writes results/published_weights.json, the receipt for
  the published weights' Massachusetts 446 / 149 and national 13,631.
- README and doc: the held-out gain is "ahead on both measures and folds,
  held-out variation unmeasured, optimistic", not justified by the training
  loss noise; softmax versus projection rests on training fit and the cap
  loop; lambda 0 is no longer "neither dominates" on a 3% loss gap; the
  share >= 0.9 floor ranges cover every solve; the 0.3% loss-gap bound is
  shown to say nothing about the overshoot (projection has the same gap);
  module hashes back the head account; the fingerprint docstring names the
  failing CI run without inventing a cause.
- Re-pin the spec-engine digests the docstring edit moves (seed_protocol
  be39eb6b..., seed_map 085d8d39..., US spec_sha256 e0b757ce...); coverage
  --check passes 41/41.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…, changelog

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
engine-free CI failed test_every_pinned_platform_key_is_the_one_that_platform_would_compute[calibrate]:
the calibrate kernel's implementation hash covers solve.py's bytes, so this
PR's docstring, counter and guard edits move every platform's node key while
leaving outputs alone. Re-pinned with `uv run python tools/graph_parity_repin.py
calibrate`, which checks the pinned dependency versions match this machine,
refuses if the local direct call's bytes moved (they did not: the defaults are
byte-identical), and derives the other platforms' keys. The parity pin tests
pass locally (11).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The minimal-bundle golden in test_spec_engine_loader.py hashes the
resolved legacy-v1 seed protocol, whose legacy_v1_direct_draws kernel
attests microcosm.calibrate.solve's source bytes. A differential run
that substitutes only the merge base's solve.py bytes reproduces the
old golden (6af478ff...) exactly; on this tree the only envelope fields
that differ are that kernel's source_sha256 and the derived
implementation_sha256 (be39eb6b..., the value inventory_coverage.py
already pins for US).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
MaxGhenis added a commit that referenced this pull request Oct 2, 2026
…ests solve.py moves

docs/calibration-l2-basis.md replaces the cap-rounds limitation with the
projection, its tested invariants and the full-scale measurement; the
experiment README adds the re-measurement; a changelog fragment records the
change and the #1078 fragment drops its exhaustion-count claims.

solve.py's bytes are attested by the US spec-engine seed protocol and the H1
calibrate parity case. Re-pinned on this tree:
- EXPECTED_HASHES seed_protocol be39eb6b -> 5795dc07, seed_map 085d8d39 -> 292ad19c
  (build_inventory_coverage; tools/spec_engine_coverage.py --check: 41/41)
- US spec_sha256 e0b757ce -> edcd6475 (docs/evidence/spec-engine/us-f0-coverage.json,
  test_us_multispine_pool_tool.py)
- loader golden fee5893a -> 5135dd79 (test_spec_engine_loader.py)
- calibrate parity pins (tools/graph_parity_repin.py calibrate; direct-call
  bytes unchanged)

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@MaxGhenis

Copy link
Copy Markdown
Contributor Author

Merge audit (head b7b4c29):

  • gh pr checks 1078 exits 0 on run 37001587847 (all 10 jobs pass); mergeable MERGEABLE / CLEAN; no CHANGES_REQUESTED.
  • Independent review: Opus 5.5 rounds 4 and 5 via subfleet run --task review --tier standard approved (reports in _build_artifacts/acs-local-l2-basis-20260928/review/review-r4.md, review-r5.md). Every finding and nit from rounds 1-5 is fixed in the branch.
  • Changes after round 5's approval are mechanical re-pins only: the H1 calibrate parity pin (b115745, tools/graph_parity_repin.py calibrate, which refuses if outputs move) and the spec-engine loader golden vector (b7b4c29). Both move because solve.py's source bytes are attested; numerics are byte-identical to bda72cb under the defaults (pinned by the frozen-oracle test).
  • Main has moved since the merge base, but not in any seed-kernel module or pinned spec-engine file, so the pins hold on the merge ref.
  • No default changes. Adopting the chi-square penalty (λ 0.03, projection) and the 25%-of-starting-ESS district gate is queued for Max as d797 (supersedes d792); the seeding question is d793.

@MaxGhenis
MaxGhenis merged commit 45d1ae3 into main Oct 4, 2026
10 checks passed
@MaxGhenis
MaxGhenis deleted the l2-design-basis branch October 4, 2026 19:07
MaxGhenis added a commit that referenced this pull request Oct 4, 2026
Resolve the calibrate-stage conflicts: keep both the penalty settings and the
target_loss stamp in _solver_settings, keep _stamped_settings (it does not fill
target_loss, so a stamp from before the weights still never matches), name any
recorded --target-family-loss-multiplier in the refresh recipe, and give
#1078's two do_calibrate tests a surface the loss-weight mapping can classify.
Regenerate results.json: its weight digests predated the name@period row
names (every other field is byte-identical).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant