Repository navigation
Chi-square design-weight penalty and softmax mass parametrization for calibration; ACS local ESS frontier - #1078
Merged
Merged
Conversation
…m calibration l2_basis selects the L2 concentration penalty's form: the historical record-weighted mean(r**2) (default, bit-identical) or GREG's anchor-weighted chi-square distance sum(d*(r-1)**2)/sum(d), whose target-free optimum is the anchor itself. mass_parametrization selects how mass='conserve' holds the total: the historical per-step uniform log shift (default, bit-identical) or w = total*softmax(log_w), which hands Adam the constraint-reduced gradient. Both are recorded in the result options; chi_square_distance reports the realized distance. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Hypothesis properties for the penalty algebra (zero at the anchor, nonnegative, float32 torch form equal to the float64 reference, the weighting-effect identity, each basis's mass-constrained stationary point), target-free convergence to the design weights, softmax mass and cap invariants, the regularization-path monotonicity with a slack derived from the measured optimizer error, the projection stall as a pinned intended violation, option provenance and validation, and a byte-for-byte pin of the defaults against the verbatim bda72cb optimizer. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Both default to the historical solve and take the kernel's names. The calibration summary and build manifest record them with the realized chi-square distance; a resumed run refuses weights solved under other penalty settings (a legacy checkpoint resumes only under the defaults it was solved with); and the concentration limitation no longer claims l2_lambda=0 for a penalized run. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… draft The harness recalibrates the 2026-09-23 ACS local release's 4,459-target surface from a sparse copy of its checkpoint, verified row-for-row against 3,000 sampled dense rows, with the release's settings plus the run's L2 and mass options. optimizer_reference.py compares the kernel under each mass setting with CLARABEL's exact optimum. The frontier sections of the doc follow the sweep. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…s and the Modal grid seeding_options.py measures the starting weights' Kish ESS (rows and distinct source households) for ACS shares 0.5/0.7/0.9/proportional and 1/4/16 donor location clones. The harness takes an acs_share prior, which becomes the frame weights, the chi-square anchor and the cap base. modal_sweep.py runs the grid, gated on reproducing the published release. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e frontier analysis Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…tmax cap receipt - The ACS tool keeps the certified-default limitation only for the historical solve (no penalty, projection); an unpenalized softmax run is recorded as its own. Legacy summaries read as the historical solve in the manifest. CLI choices are spelled in the tool (pinned equal to the kernel's) so parsing no longer imports torch. The resume guard ignores the basis when neither side has a penalty and says which settings it checks. - The path tests now check each solve against CLARABEL's exact optimum of the same program on their own problems (committed fixture and generator), and derive the monotonicity slack from that bound on a lambda grid dense where the slack is informative. - The softmax cap loop's docstrings describe what it does; epochs whose rounds run out are counted in the iterate-selection receipt. The projection stall claim now states its conditions (gradients of one sign, well above Adam's eps). Numerics are unchanged from 35ad665. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…e at full size At the release's scale the median record's log-weight gradient is 1.2e-8 to 3.8e-8 with mixed signs, so the sharp same-sign stall does not occur there; the doc now says so and leaves the full-scale comparison to the sweep. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…options into its solver stamp l2_basis and mass_parametrization join _solver_settings, so main's resume and already-complete checks cover them; a stamp written before they existed reads as the historical solve. calibrate_surface takes both with historical defaults. My separate resume guard is dropped in favor of the stamp. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…he sweep grid Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…floors Recalibrations of the published ACS local release on Modal (reproduction gated): the chi-square penalty at lambda 0.03 lifts national ESS 13.6k to 21.8k and beats the release on two held-out folds; the ACS share of the mass is the larger lever and costs held-out SOI fit. Receipts are compacted copies of each run's metrics; the README's key table is written by analyze.py. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The Primary-QRF worker identity pins uv.lock's sha256. Hypothesis already reaches every test environment through microcosm-graph's dev group, so the calibrate dev-group entry was redundant; drop it rather than re-pin. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…imit not a bound, and the merged resume stamp Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…oop limit, lead the gate with the relative floor - README and doc: projection at lambda 0.03 is the recommendation, with both held-out metrics on both folds and the fold selection disclosed; the frontier claim is limited to lambda <= 0.03; softmax's cap-loop exhaustion at production scale is a stated limitation; the gate leads with "no district below 25% of its starting ESS", the state floor of 200 is dropped, and the relative gate's behaviour at share 0.9 is shown; the release row is labelled as reproduced and reconciled with the published and reported Massachusetts figures; kernel heads are listed with the solve.py diff. - analyze.py: parametrization in every key-table label, a held-out within-10% column, kernel_head in frontier.csv, and a byte-level head-equivalence check for the dup_* reruns. - Tool: the limitation text fills the historical basis for a legacy summary; the refresh recipe carries the penalty flags when they differ from the historical solve. - seeding_options.py emits the Massachusetts split by spine. - Rename test_path_reference.py to path_reference.py; drop the fixture's unchecked module hash; say eps is empirical; assert the unreachable softmax best-iterate branch. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…by fingerprint engine-us CI failed test_us_coverage_is_exact_complete_and_honest: the US seed protocol attests the installed source bytes of microcosm.calibrate.solve (spec_engine/seeds.py), so this PR's solve.py edits move three identities: - EXPECTED_HASHES["seed_protocol"] 91989ed5... -> 0e7ab5fc... - EXPECTED_HASHES["seed_map"] 3cf10a52... -> f84f4fab... - the US spec_sha256 pinned in test_us_multispine_pool_tool.py ffbb93ed... -> a7eee025... The new values come from build_inventory_coverage on this tree, where that was the only failing item. docs/evidence/spec-engine/us-f0-coverage.json was regenerated with tools/spec_engine_coverage.py; --check passes with 41/41 inventory checks. engine-free CI failed the path test's byte digest of the generated problem for one of twelve cases while the other eleven matched. The fixture now stores per-problem moments (verified locally against the old digests before they were dropped), compared with rel 1e-9. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The head-equivalence reruns are not byte-identical to their originals: the trajectories already differ at epoch 0 (softmax, the loss at the starting weights) or epoch 1 (the release), so the solves are not bit-reproducible across Modal containers. analyze.py now writes results/rerun_variation.json (bytes, first divergence, every frontier metric's relative change) instead of a byte-identity check. The release's final loss moved 2.7% between two identical runs (0.3% at lambda 0.03); concentration measures moved under 1%. The held-out margins of projection lambda 0.03 over the release (3.6-3.8% in capped error, 0.9-1.2 points within 10%) are of that order, so the README and the doc now say "no worse out of sample" rather than "better", and list one run per configuration as a limit. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… the published weights, keep claims inside the noise - release_refresh_recipe always names --l2-lambda, --l2-basis and --mass-parametrization, so a later default change cannot move a recorded release's recipe; the test parses the historical recipe back. - The calibration summary and build manifest record softmax_cap_rounds_exhausted_epochs; calibrate()'s docstring and the --mass-parametrization help say the rounds run out at the ACS release's scale. - published_weights.py writes results/published_weights.json, the receipt for the published weights' Massachusetts 446 / 149 and national 13,631. - README and doc: the held-out gain is "ahead on both measures and folds, held-out variation unmeasured, optimistic", not justified by the training loss noise; softmax versus projection rests on training fit and the cap loop; lambda 0 is no longer "neither dominates" on a 3% loss gap; the share >= 0.9 floor ranges cover every solve; the 0.3% loss-gap bound is shown to say nothing about the overshoot (projection has the same gap); module hashes back the head account; the fingerprint docstring names the failing CI run without inventing a cause. - Re-pin the spec-engine digests the docstring edit moves (seed_protocol be39eb6b..., seed_map 085d8d39..., US spec_sha256 e0b757ce...); coverage --check passes 41/41. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…, changelog Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
engine-free CI failed test_every_pinned_platform_key_is_the_one_that_platform_would_compute[calibrate]: the calibrate kernel's implementation hash covers solve.py's bytes, so this PR's docstring, counter and guard edits move every platform's node key while leaving outputs alone. Re-pinned with `uv run python tools/graph_parity_repin.py calibrate`, which checks the pinned dependency versions match this machine, refuses if the local direct call's bytes moved (they did not: the defaults are byte-identical), and derives the other platforms' keys. The parity pin tests pass locally (11). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The minimal-bundle golden in test_spec_engine_loader.py hashes the resolved legacy-v1 seed protocol, whose legacy_v1_direct_draws kernel attests microcosm.calibrate.solve's source bytes. A differential run that substitutes only the merge base's solve.py bytes reproduces the old golden (6af478ff...) exactly; on this tree the only envelope fields that differ are that kernel's source_sha256 and the derived implementation_sha256 (be39eb6b..., the value inventory_coverage.py already pins for US). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
MaxGhenis
added a commit
that referenced
this pull request
Oct 2, 2026
…ests solve.py moves docs/calibration-l2-basis.md replaces the cap-rounds limitation with the projection, its tested invariants and the full-scale measurement; the experiment README adds the re-measurement; a changelog fragment records the change and the #1078 fragment drops its exhaustion-count claims. solve.py's bytes are attested by the US spec-engine seed protocol and the H1 calibrate parity case. Re-pinned on this tree: - EXPECTED_HASHES seed_protocol be39eb6b -> 5795dc07, seed_map 085d8d39 -> 292ad19c (build_inventory_coverage; tools/spec_engine_coverage.py --check: 41/41) - US spec_sha256 e0b757ce -> edcd6475 (docs/evidence/spec-engine/us-f0-coverage.json, test_us_multispine_pool_tool.py) - loader golden fee5893a -> 5135dd79 (test_spec_engine_loader.py) - calibrate parity pins (tools/graph_parity_repin.py calibrate; direct-call bytes unchanged) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Contributor
Author
|
Merge audit (head b7b4c29):
|
MaxGhenis
added a commit
that referenced
this pull request
Oct 4, 2026
Resolve the calibrate-stage conflicts: keep both the penalty settings and the target_loss stamp in _solver_settings, keep _stamped_settings (it does not fill target_loss, so a stamp from before the weights still never matches), name any recorded --target-family-loss-multiplier in the refresh recipe, and give #1078's two do_calibrate tests a surface the loss-weight mapping can classify. Regenerate results.json: its weight digests predated the name@period row names (every other field is byte-identical). Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
On 2026-09-28 David Trimmer measured concentrated weights in the ACS local release:
He asked whether to anchor calibration to the design weights. This PR adds that anchor and tests it at full scale. The full write-up is
experiments/us-acs-local-l2-basis-20260928/README.md; the method is indocs/calibration-l2-basis.md.What the evidence shows:
--acs-share 0.5gives the 57,240 donor rows half the mass. Before calibration, national ESS is 36,288 and Massachusetts 1,020. Calibration takes ESS to 13,646.l2_lambdaknob backfires. Its record-weighted form pulls towardw ∝ d², and at λ = 0.1 it lowers ESS to 8,689.mass="conserve", the historical projection solve can stall on small problems where every target pulls the same way: it is up to 0.34 off the exact optimum, against softmax's 0.0034. At the release's scale, gradients are near Adam'sepswith mixed signs, so the stall does not occur. The two parametrizations agree for λ ≤ 0.03, and projection falls inside softmax's training frontier at λ ≥ 0.1. Softmax's per-step cap loop runs out of its 32 rounds on most epochs at that scale (counted in the receipt; the closing projection makes the returned weights exact). The doc states this as a known limitation.What changes
microcosm-calibrategets two opt-in options. The defaults are byte-identical to bda72cb.l2_basis:"record"(default) is the historicalmean((w/d)**2)."chi_square"issum(d*(w/d-1)**2)/sum(d).calibrate, the L0 budget search,refit_l0_selection,calibrate_l0_refit(refit_l2_basis) andstatic_aging.mass_parametrization:"projection"(default) is the historical scheme."softmax"isw = total*softmax(log_w), for Adam withmass="conserve"and no L0 gates.calibrate_l0_refittakesrefit_mass_parametrization.CalibrationResult.chi_square_distance(also onL0RefitResult) andchi_square_distance(weights, anchor).tools/build_us_acs_local_release.py. The changes are flags on top of #1053; other open PRs edit this tool too.--l2-basisand--mass-parametrization, with the defaults unchanged._solver_settings, so resume and the already-complete shortcut cover them. A stamp written before they existed reads as the historical solve.--l2-lambda,--l2-basisand--mass-parametrizationat the recorded values, so a later default change cannot move it.softmax_cap_rounds_exhausted_epochs.Evidence lives in
experiments/us-acs-local-l2-basis-20260928/:solve.pydiffers only in docstrings, a receipt counter and guards;solve.py's module hash is the same at 9ef71ed, bc763cb and d82ff85 (it differs from 35ad665, where the originals ran, only by docstrings and the counter). Over 800 epochs the release's final loss moved 2.7% (0.3% at λ = 0.03), while every concentration measure moved under 1%;results/published_weights.json);CI-pinned identities this PR moves
The US spec-engine seed protocol attests the source bytes of
microcosm.calibrate.solve, so editingsolve.pyre-pinsEXPECTED_HASHES["seed_protocol"]and["seed_map"], the USspec_sha256intest_us_multispine_pool_tool.py, anddocs/evidence/spec-engine/us-f0-coverage.json(regenerated;--checkpasses 41/41). Each new value was computed on this tree (final values:seed_protocolbe39eb6b…,seed_map085d8d39…, USspec_sha256e0b757ce…).Invariants
These are property-tested in
packages/microcosm-calibrate/tests/engine_free/shared/test_l2_basis.py.w = anchorand nonnegative everywhere (Hypothesis).chi_square_distance, and the record penalty equalsmean((w/d)**2)(Hypothesis).sum(w) = sum(d), the distance equalssum(w**2/d)/sum(d) - 1. With uniformdit equalsn/ESS - 1(Hypothesis).w = dfor chi-square and atw ∝ d**2for the record basis (Hypothesis).mass="conserve", the chi-square solve returns the design weights within 1% from any warm start, under both parametrizations (Hypothesis).2·eps/Δλderived from that bound, on informative λ pairs only;ESS = n/(1+P)holds exactly.expnode, fails it.Intended violation, pinned as such: under
"projection", invariant 7 does not hold in general.test_projection_parametrization_stalls_under_uniform_mass_pressureshows the stall when every target sits above its design total.Review
Independent Opus 5.5 reviews, each through
subfleet run --task review:calibrate()'s docstring, the CLI help and the summary.axiom: n/a: calibration kernel and ACS local tool; no policy rules change.🤖 Generated with Claude Code