Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
24 commits
Select commit Hold shift + click to select a range
f8a4bed
Add the chi-square L2 basis and a softmax mass parametrization to Ada…
MaxGhenis Sep 28, 2026
7599cd1
Test the chi-square L2 basis and softmax mass parametrization
MaxGhenis Sep 28, 2026
fa8df5d
Plumb --l2-basis and --mass-parametrization through the ACS local tool
MaxGhenis Sep 28, 2026
949599a
Add the ACS local L2-basis sweep harness, optimizer reference and doc…
MaxGhenis Sep 30, 2026
da7f465
Seed the ACS local sweep's prior from --acs-share; add seeding option…
MaxGhenis Sep 30, 2026
35ad665
Make the sweep launcher importable inside its Modal container; add th…
MaxGhenis Sep 30, 2026
5e7203f
Address review round 1: truthful provenance, derived test bounds, sof…
MaxGhenis Sep 30, 2026
1ba923c
State the projection stall's conditions and measure the gradient scal…
MaxGhenis Sep 30, 2026
9ef71ed
Extend the sweep grid with holdout at the knee and intermediate priors
MaxGhenis Sep 30, 2026
30706e3
Merge origin/main (#1053 sparse ACS local tool) and fold the penalty …
MaxGhenis Sep 30, 2026
bc763cb
Add a second holdout fold and projection at the held-out optimum to t…
MaxGhenis Sep 30, 2026
5888b65
Write up the ACS local L2-basis frontier: seeding, penalty, holdout, …
MaxGhenis Oct 1, 2026
a471113
Add projection at lambda 0.03 to the ACS local frontier
MaxGhenis Oct 1, 2026
d9ea709
Keep uv.lock at main's approved digest
MaxGhenis Oct 1, 2026
2e4d443
Address review round 2: state the training-fit cost, the prior as a l…
MaxGhenis Oct 1, 2026
d82ff85
Add head-equivalence reruns to the sweep grid
MaxGhenis Oct 1, 2026
8ebf901
Address review round 3: recommend projection, state the softmax cap-l…
MaxGhenis Oct 1, 2026
51cef49
Merge remote-tracking branch 'origin/main' into l2-design-basis
MaxGhenis Oct 1, 2026
412c8f7
Re-pin the spec-engine digests solve.py moves; compare path problems …
MaxGhenis Oct 1, 2026
ab65596
Measure run-to-run variation; hold the held-out claim to "no worse"
MaxGhenis Oct 1, 2026
3dfe38d
Address review round 4: pin every penalty flag in the recipe, receipt…
MaxGhenis Oct 1, 2026
f8e48b2
Address review round 5 nits: exact district minimum, frontier wording…
MaxGhenis Oct 1, 2026
b115745
Re-pin the H1 calibrate parity case solve.py's edits move
MaxGhenis Oct 1, 2026
b7b4c29
Re-pin the loader golden vector solve.py's edits move
MaxGhenis Oct 2, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 5 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -334,7 +334,11 @@ calibrates to the SOI `state` surface by default, the 4,459-target contract of
Build O and Build P; `--soi-mode totals` and `--soi-mode full` are explicit
opt-ins. See
[the ACS local-area SOI target surface](docs/us-acs-local-soi-target-surface.md)
for what each mode contains and where the build records it.
for what each mode contains and where the build records it. Its
`--l2-basis chi_square` and `--mass-parametrization softmax` options (defaults:
the historical `record` and `projection`) penalize distance from the design
weights; see [penalized calibration toward the design weights](docs/calibration-l2-basis.md)
for the algebra, the evidence and the measured frontier.

National and ACS local-area builds now use the same typed schema-8 calibration
diagnostics writer. The local builder adds its Census population marginals to a
Expand Down
1 change: 1 addition & 0 deletions changelog.d/calibrate-chi-square-l2-basis.added.md
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
Add two opt-in options to Adam calibration and the ACS local-area tool. `l2_basis="chi_square"` (`--l2-basis chi_square`) makes the `l2_lambda` penalty GREG's design-weighted chi-square distance `sum(d * (w / d - 1) ** 2) / sum(d)`, which pulls toward the design weights themselves; the historical record-weighted `mean((w / d) ** 2)`, whose mass-conserved optimum is `w ∝ d ** 2`, stays the default. `mass_parametrization="softmax"` (`--mass-parametrization softmax`) conserves mass through `w = total * softmax(log_w)`; the historical per-step uniform shift, still the default, can stall away from the constrained optimum when every record's gradient shares a sign (on the ACS local release it does not, and softmax's per-step cap rounds run out on most epochs there, which the solve counts). Defaults reproduce the previous solver's bytes. Results record both settings and `CalibrationResult.chi_square_distance`; the tool records them, with the softmax cap-exhaustion count, in its calibration summary and build manifest, names all three penalty flags in the manifest's refresh recipe, and refuses to resume a checkpoint solved under other settings. See docs/calibration-l2-basis.md.
237 changes: 237 additions & 0 deletions docs/calibration-l2-basis.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,237 @@
# Penalized calibration toward the design weights

`microcosm.calibrate.calibrate` has two opt-in options for spreading calibrated
weights: `l2_basis`, the functional form of the `l2_lambda` penalty, and
`mass_parametrization`, how the Adam solve holds the total under
`mass="conserve"`. Their defaults (`"record"`, `"projection"`) are the historical
solve, byte for byte. This page gives the algebra, the evidence for each claim,
and what the options can and cannot do for the ACS local-area release. The
full-scale sweep and the recommendation are in
[`experiments/us-acs-local-l2-basis-20260928/README.md`](../experiments/us-acs-local-l2-basis-20260928/README.md).

## The two bases

Write `d` for the anchor (the initial, design weights unless `l2_anchor` says
otherwise), `w` for the calibrated weights, `r = w / d`, and `D = sum(d)`.

| | `l2_basis="record"` (default) | `l2_basis="chi_square"` |
|---|---|---|
| penalty | `mean(r ** 2)` | `sum(d * (r - 1) ** 2) / D` = `sum((w - d) ** 2 / d) / D` |
| value at `w = d` | 1 | 0 |
| target-free optimum, `mass="conserve"` | `w ∝ d ** 2` | `w = d` (an explicit anchor whose total differs is rescaled to the input total) |
| target-free optimum, `mass="free"` | `w → 0` | `w = d` |
| cost of a record collapsing to 0 | none (0 is its cheapest ratio) | its anchor share `d_i / D` |
| `options["l2_penalty"]` (initial anchor) | `mean_initial_pre_gate_weight_ratio_squared` | `chi_square_initial_pre_gate_weight_distance` |

The chi-square form is GREG's distance function. Minimizing target loss plus
`l2_lambda` times it is penalized ("ridge") calibration, the soft form of GREG:
it works where exact GREG cannot, with thousands of partly inconsistent soft
targets, nonnegativity, a ratio cap and 1.6 million records.

Identities (tested in
`packages/microcosm-calibrate/tests/engine_free/shared/test_l2_basis.py`):

- When `sum(w) = D`, the distance equals `sum(w ** 2 / d) / D - 1`: the
weighting effect of `w` relative to `d`, minus one.
- With uniform `d` it equals `n / ESS(w) - 1`, so it is a direct Kish ESS
control.
- The float32 penalty the solver differentiates agrees with the float64
reference `microcosm.calibrate.chi_square_distance`.
- Under a mass constraint the reduced gradient of the chi-square penalty
vanishes at `w = d`, and that of the record penalty at `w ∝ d ** 2`.

`CalibrationResult.chi_square_distance` reports the realized distance from the
input weights for every solve, whatever its penalty.

Kish ESS is not in general monotone in `l2_lambda`. What is monotone, for exact
minimizers, is the penalty itself: adding the optimality inequalities of two
solves at `l2_lambda` values `a < b` gives
`(b - a) * (P(w_b) - P(w_a)) <= 0`. With unequal design weights, calibration can
raise Kish ESS above the design's, and pulling back toward `d` then lowers it.
In the ACS local release, calibration lowers ESS from the design's 36,288 to
13,631 in the published weights (13,646 in the sweep's reproduction), and as `l2_lambda → ∞` the solve returns to the design's figure; the
sweep measures the path in between.

The path tests check each solve against the exact optimum of the same convex
program, computed by CLARABEL on the tests' own problems (fixture
`tests/fixtures/l2_basis_path_reference.json`, generator
`experiments/us-acs-local-l2-basis-20260928/path_reference.py`). The
largest excess measured there is 0.0026. The excess bound `eps` = 0.006 is
empirical (twice that maximum, rounded up to 1e-3); the monotonicity slack
`2·eps/Δλ` is derived from it, since two `eps`-optimal solves of the same
feasible set satisfy `(λb − λa)(Pb − Pa) ≤ 2·eps`. The exact optima also show the textbook structure: between
`l2_lambda = 0.01` and `0.3` the optimal distance is flat, because the solve
meets every feasible target exactly and picks the closest such point to `d`,
which is GREG's solution.

## The mass parametrization

Under `mass="conserve"` the historical solve takes an Adam step on the
log-weights, then shifts every log-weight by the same amount so the total
returns to the input total (`mass_parametrization="projection"`). Adam divides
each coordinate's gradient by its own running RMS plus `eps` (1e-8) before
stepping. When a gradient is well above `eps`, the coordinate moves about `lr`
in the direction of the gradient's sign, whatever its size. If every record's
gradient then has the same sign, all records step by about `lr` and the shift
takes them straight back. Every target missing on the same side produces
this; a net demand for mass with targets on both sides does not. Such a point
is a fixed point of the projected iteration whether or not it is the
constrained optimum. At the constrained optimum the log-weight gradient is
proportional to `w` (the constraint's normal), not zero, so the iteration has
no systematic pull toward it; it gets there only through records whose
gradients change sign. Records whose gradient is zero or near `eps` take
smaller steps, which the shift does not cancel, and that weakens the
mechanism. At the ACS release's scale it is weak in exactly this way
(`experiments/us-acs-local-l2-basis-20260928/gradient_scale.py`). At the
design weights the median record's log-weight gradient is 3.8e-8 and 23% are
at or below `eps`; at the release's weights the median is 1.2e-8 and 47% are
at or below `eps`. The signs are mixed, 33% positive at the design and 52% at
the release. So the sharp same-sign stall does not occur there, and whether
the projection solve still falls short is an empirical question the
full-scale sweep answers by solving the release's own surface both ways.

`mass_parametrization="softmax"` optimizes `w = D * softmax(log_w)` instead.
The total holds by construction, and autograd hands Adam the gradient with its
component along the constraint already removed. For a smooth objective that
reduced gradient vanishes at the constrained optimum; the capped-MAPE loss has
kinks, so Adam still oscillates there on the scale of `lr`, as it does under
`mass="free"`. The ratio cap remains a per-step clamp, alternated with the
softmax-invariant renormalization for up to 32 rounds, plus the closing
float64 projection that both parametrizations share. It requires
`mass="conserve"` and no L0 gates.

**Known limitation: the cap rounds run out at production scale.** Each
renormalization lifts the clamped records back over the cap by a shrinking
amount, so the rounds converge geometrically but need not reach float32
exactness in 32. A step that runs out ends on a clamp, and its next forward
pass optimizes `total·softmax(log_w)` past the cap, by an amount that is not
recorded.
`options["iterate_selection_receipt"]["softmax_cap_rounds_exhausted_epochs"]`
counts those steps. On small problems it is rare. On the ACS local release it
is the normal state: 165-400 of the last 400 epochs in every run that records
it, including 400 at the share-0.5, `l2_lambda = 0.03` configuration on the
full surface and on one fold, and 400 in every share-0.9 run. Only the
closing projection makes the returned vector exact. The returned loss stays
within 0.3% of the last trajectory loss in every run that records the count,
but projection runs, which have no cap loop, show the same gap (0.01-0.33%), so
that says nothing about the overshoot; its size is not measured. An
exact capped-softmax step (water-filling the excess onto the uncapped records)
would remove the limitation. Until then, prefer `"projection"` at that scale,
where the stall does not bite (below).

Evidence:

- `test_projection_parametrization_stalls_under_uniform_mass_pressure` pins
the stall in its sharpest form. When every target sits above its design
total, the projection
solve returns the design weights unchanged after 200 epochs, while the
softmax solve halves the loss. The same mechanism means the record penalty
alone never moves a projection solve: its log-space gradient is positive for
every record.
- `experiments/us-acs-local-l2-basis-20260928/optimizer_reference.py` runs the
kernel under each setting and compares it with CLARABEL's exact optimum of
the same convex program:

| Setting | Problems | Mean excess objective at λ = 0 / 0.001 / 0.01 / 0.1 / 1 | Max excess | Problems with non-monotone distance |
|---|---|---|---:|---:|
| `mass="conserve"`, `"projection"` (default) | 16 small (120 records, 6 targets; half press on the total) | 0.0285 / 0.0268 / 0.0304 / 0.1608 / 0.0697 | 0.3384 | 94% |
| `mass="conserve"`, `"softmax"` | 16 small (120 records, 6 targets; half press on the total) | 0.0013 / 0.0016 / 0.0013 / 0.0017 / 0.0008 | 0.0034 | 0% |
| `mass="free"` | 16 small (120 records, 6 targets; half press on the total) | 0.0004 / 0.0015 / 0.0020 / 0.0019 / 0.0014 | 0.0042 | 0% |
| `mass="conserve"`, `"projection"` (default) | 3 large (3,000 records, 150 targets in both directions) | 0.0023 / 0.0037 / 0.0062 / 0.0081 / 0.0061 | 0.0139 | 0% |
| `mass="conserve"`, `"softmax"` | 3 large (3,000 records, 150 targets in both directions) | 0.0012 / 0.0016 / 0.0018 / 0.0012 / 0.0010 | 0.0023 | 0% |
| `mass="free"` | 3 large (3,000 records, 150 targets in both directions) | 0.0009 / 0.0016 / 0.0018 / 0.0013 / 0.0010 | 0.0023 | 0% |

Excess objective is the kernel solve's loss plus `l2_lambda` times its
distance, minus CLARABEL's optimum of the same program; the loss cap never
binds on these surfaces. Receipts: `results/optimizer_reference.csv` and
`results/optimizer_reference_summary.json` in that experiment directory.
The small problems where every target presses on the total show the stall.
On the larger problems with targets in both directions, projection is never
non-monotone but still lands 1.9 to 6.7 times further from the optimum than
softmax.

## At the ACS release's scale

On the published ACS local release (1,588,854 households, 4,459 targets),
recalibrated from its own checkpoint (`experiments/us-acs-local-l2-basis-20260928/`):

- **Parametrization.** The per-record gradients there are near Adam's `eps`
with mixed signs, so the projection stall does not occur. The two
parametrizations trace the same training frontier for `l2_lambda ≤ 0.03`
(at `0` softmax reaches 18% more ESS, 16,132 against 13,646, at a 3% higher
loss, about the size of the unpenalized solve's 2.7% run-to-run variation). At `0.1` and above, projection falls inside softmax's
frontier: at `1` it reaches ESS 30,582 at loss 0.0785, where softmax's path,
interpolated between its `0.3` and `1` solves, gives about 0.055. Equal
`l2_lambda` is not an equal-ESS comparison. At `0.03` the held-out
comparison is a toss-up: softmax has slightly lower capped error, projection
more held-out targets within 10% on both folds. Projection's training loss is
6% lower there, against 0.3% run-to-run variation. With that and the
cap-loop limitation above, projection is the recommended parametrization.
- **Record basis.** The record basis at `0.1` lowers national ESS from 13,646
to 8,689, as its `w ∝ d ** 2` optimum predicts.
- **Chi-square basis.** The chi-square basis raises ESS smoothly with
`l2_lambda`. At `0.03` under projection it costs 0.3 points of training fit
and no measurable held-out fit: it is ahead of the release on both measures
(capped error and share within 10%) on both rotated folds. Held-out
run-to-run variation was not measured and those folds also chose
`l2_lambda`, so the gain is optimistic.
- **What limits ESS is the starting weights.** The experiment's README has the
frontier, the holdout, candidate ESS floors and the recommendation.

## Using it

```python
from microcosm.calibrate import calibrate

result = calibrate(
frame,
targets,
mass="conserve",
max_weight_ratio=5.0,
l2_lambda=...,
l2_basis="chi_square",
# mass_parametrization="softmax" for problems where every target presses
# the same way; see the cap-loop limitation above.
)
result.options["l2_basis"], result.options["mass_parametrization"]
result.chi_square_distance, result.effective_sample_size
```

The ACS local-area tool takes the same settings as
`--l2-lambda`, `--l2-basis {record,chi_square}` and
`--mass-parametrization {projection,softmax}`. It records them, with the
realized chi-square distance, in `calibration_summary.json` and in the build
manifest's `calibration` block. Both join the tool's solver-settings stamp,
so `--resume` and the already-complete shortcut refuse weights solved under
any other solver setting (cap, loss cap, `l2_lambda`, `l2_basis`,
`mass_parametrization`, seed, epoch batch); a stamp written before these two
options existed reads as the historical solve. The manifest's refresh recipe
carries the three flags whenever they differ from the historical solve.
`calibrate_l0_refit` takes `l2_basis` / `refit_l2_basis`
and `refit_mass_parametrization`; `refit_l0_selection` and `static_aging` take
`l2_basis`.

## What the penalty cannot fix in the ACS local release

The chi-square penalty pulls toward the design weights, so the most it can
recover is the design weights' own concentration. In the ACS local staging,
that concentration is already high by construction. The 57,240 donor-spine
records (the Build O sparse release) carry half the design mass with a within-
spine ESS of 9,174, while the 1,531,614 ACS records carry the other half with a
within-spine ESS of 812,434. National design ESS is therefore 36,288 (2.3% of
records). Massachusetts shows the same split
(`massachusetts_by_spine_at_release_design` in that experiment's
`results/seeding_options.json`). Its 35,056 ACS records have a
design ESS of 20,238, while its 1,522 donor records carry 49.6% of the state's
design mass with an ESS of 255, the national release's own Massachusetts figure,
since the donor spine is that release's records at half weight. Together the
state's design ESS is 1,020, and each of its nine districts is 108-134 (the ACS
records alone give 2,020-2,444 per district). As `l2_lambda → ∞` the solve
returns to those combined figures. The penalty bounds the distance from them,
not ESS itself, but no tested setting lifted national ESS above them. Reaching the ACS-only figures needs a larger ACS
share of the mass in the staging (`--acs-share`), which is a construction
decision outside calibration; `experiments/us-acs-local-l2-basis-20260928/
seeding_options.py` measures the starting ESS under other shares and under
donor location clones.
(Measured on the calibration checkpoint of the 2026-09-23 release,
`populace-us-2024-buildo-acs-local-767312d60-20260923T074941Z`.)
16 changes: 8 additions & 8 deletions docs/evidence/spec-engine/us-f0-coverage.json
Original file line number Diff line number Diff line change
Expand Up @@ -1656,13 +1656,13 @@
"compiler_ir.node_slices"
],
"expected": {
"map_sha256": "3cf10a524310ea0ada18ddc65cfa585d1d0652d23e7d7a180c050ac686224e28",
"protocol_sha256": "91989ed5e366e19598f86335b2e66723dcb2b6481a78d0dbeb83183a1aa09bd4"
"map_sha256": "085d8d39ddde7827926a6e33fbba1c7f6e85e193fcd6ae5b4d1c91efeb73a176",
"protocol_sha256": "be39eb6b08b4fbee55f2911a3e8ddef753b24c51dd2a25655e49879e12f622cb"
},
"failures": [],
"observed": {
"map_sha256": "3cf10a524310ea0ada18ddc65cfa585d1d0652d23e7d7a180c050ac686224e28",
"protocol_sha256": "91989ed5e366e19598f86335b2e66723dcb2b6481a78d0dbeb83183a1aa09bd4"
"map_sha256": "085d8d39ddde7827926a6e33fbba1c7f6e85e193fcd6ae5b4d1c91efeb73a176",
"protocol_sha256": "be39eb6b08b4fbee55f2911a3e8ddef753b24c51dd2a25655e49879e12f622cb"
},
"status": "covered"
},
Expand All @@ -1677,7 +1677,7 @@
"compiler_ir.seed_stream_map"
],
"expected": {
"implementation_sha256": "91989ed5e366e19598f86335b2e66723dcb2b6481a78d0dbeb83183a1aa09bd4",
"implementation_sha256": "be39eb6b08b4fbee55f2911a3e8ddef753b24c51dd2a25655e49879e12f622cb",
"protocol": "legacy-v1",
"streams": [
"build_model",
Expand All @@ -1698,7 +1698,7 @@
},
"failures": [],
"observed": {
"implementation_sha256": "91989ed5e366e19598f86335b2e66723dcb2b6481a78d0dbeb83183a1aa09bd4",
"implementation_sha256": "be39eb6b08b4fbee55f2911a3e8ddef753b24c51dd2a25655e49879e12f622cb",
"protocol": "legacy-v1",
"streams": [
"build_model",
Expand Down Expand Up @@ -2599,7 +2599,7 @@
"country": "us",
"schema_id": "country_spec",
"schema_version": 1,
"spec_sha256": "ffbb93ed5ae22a95536ae475d6912a42628568609b442360dc1cd3bcf26772cd"
"spec_sha256": "e0b757cee1b860b304e099282ad0f9684c10f5e04fddbce128e6bd88b3480fe8"
}
},
"report_schema_version": 3,
Expand All @@ -2609,7 +2609,7 @@
"country": "us",
"schema_id": "country_spec",
"schema_version": 1,
"spec_sha256": "ffbb93ed5ae22a95536ae475d6912a42628568609b442360dc1cd3bcf26772cd"
"spec_sha256": "e0b757cee1b860b304e099282ad0f9684c10f5e04fddbce128e6bd88b3480fe8"
},
"status": "pass"
}
Loading
Loading