Skip to content

ACS local λ frontier on the weighted loss: re-pick λ and measure the district population cost - #1105

Open
MaxGhenis wants to merge 14 commits into
mainfrom
acs-local-weighted-lambda-frontier
Open

MaxGhenis wants to merge 14 commits into
mainfrom
acs-local-weighted-lambda-frontier

Conversation

@MaxGhenis

@MaxGhenis MaxGhenis commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Why

#1078's ESS frontier for the ACS local release was measured on the release's loss, with every target weighted equally (3,819 of the 4,459 targets are IRS SOI cells). #1104 makes the ACS local build weight its targets like the national release, through one shared implementation. λ is in units of the loss, so decision d797 (which supersedes d792) asks for λ to be re-picked on the weighted loss before a default ships.

This PR is research evidence only. It publishes nothing and changes no default or package code: every change is under experiments/us-acs-local-l2-basis-20260928/, plus one bullet in docs/calibration-l2-basis.md.

What the evidence shows

46 Modal solves of the published release's 4,459-target surface, 800 epochs each. The write-up is the README's "On the weighted loss" section.

  1. The penalty still buys ESS and held-out fit, now with an interior optimum. Held-out weighted capped error (the objective's out-of-sample form, folds 0 and 1) at projection λ 0 / 0.03 / 0.1 / 0.2 / 0.3 is 0.0784 / 0.0721 / 0.0708 / 0.0742 / 0.0787. Over the same range, ESS rises from 13,707 to 24,816. Projection is ahead of softmax at every λ held out.

  2. The weighting costs district population fit, and the penalty multiplies it.

    • Under equal weights, every trained district population is within 0.75% at share 0.5 and λ 0 and 0.03 on the full surface (2% on the holdout folds' trained districts).
    • Under the shared weighting, 11 of 436 districts miss by more than 10% at λ 0 (worst TX-14 −33%), 73 at λ 0.03 and 135 at λ 0.1 (worst 50%).
    • The misses concentrate in many-district states. In the shared module, each state's pop_cd rows form one concept group scaled to its largest member's weight, so California's 52 districts carry 1.51 between them, about what one at-large district carries.
    • The held-out measure barely sees this: district rows carry 1.1-1.9% of the held-out weighted loss at share 0.5, and they are predicted poorly in every configuration.
  3. census_population ×8 with λ 0.03 is the best configuration held out, by a small margin.

    • Held-out weighted error 0.0704, 10.2% / 10.4% below the release settings, and lower than every other configuration on each fold, weighted or equal-weight. Its lead over the next best, default weights at λ 0.1, is 0.9% on fold 0 and 0.3% on fold 1 (inside the rerun noise).
    • ESS 19,874 (from 13,707); Massachusetts 574 (from 478); the smallest district 22 (from 11).
    • It passes d797's gate on all three of its solves, and it misfits 3 district populations (worst 15%).
    • Against the rule's pick (λ 0.1, no multiplier) it gives up about 9% of national ESS (19,874 against 21,751; Massachusetts 574 against 609) for 132 fewer district misses. Against the equal-weight pick (ESS 20,690, no misses) it gives up 4% of ESS for held-out error of 0.0704 against 0.0724.
  4. Whether training on the weighted loss helps out of sample depends on λ. On the weighted held-out yardstick, against the same settings trained on the equal loss:

    • λ 0: worse on both folds.
    • λ 0.03: the folds split.
    • λ 0.1, the weighted optimum: better than every equal-weight configuration held out, on both folds (by 1.6-2.9%).

    At λ 0 and 0.03 the equal-weight solves keep district populations; the weighted ones need the multiplier to.

The λ re-pick. The rule was fixed and committed before any weighted held-out result existed (43b3d0f). It picks projection λ 0.1. The README states the district cost the rule did not weigh, and recommends λ 0.03 (d797's λ) with --target-family-loss-multiplier census_population=8. The multiplier was tried after the first pass, so it is labelled post hoc. Both calls are queued for Max: the result is noted on d797 (λ and gate; d792 is superseded by it and noted too), and the multiplier is d952.

What changed

  • registry.py rebuilds the 09-23 release's TargetSpecs from its own feed and settings (snap/medicaid/soi, soi_mode=state). All 4,459 names and values match the checkpoint bit for bit, and a receipt records it.
  • sweep.py:
    • target_weighting is equal (the old harness, unchanged) or shared. shared passes us_acs_local_target_loss_weights(train_specs) as target_loss_weights.
    • family_loss_multipliers scales the training weights only; held-out scoring keeps the unmultiplied full-surface weights, one yardstick for every run.
    • The metrics record the weights with the module's own digest and distribution. Every fit block gains weighted measures, there is a yardstick training block, and the consistency block checks both the epoch-0 and the final weighted loss.
    • In the container, the shared module is loaded from its single file, because importing it through microcosm.build.us_runtime needs the whole build stack.
  • modal_sweep.py:
    • --grid weighted and --grid weighted2 (w_* run ids).
    • A harness gate for the weighted grid.
    • The checkpoint upload skips files the shared volume already holds.
  • cross_score.py scores the equal-weight solves on the weighted loss.
  • analyze.py writes the weighted outputs to their own files (results/weighted_*, weighting_shift.md, target_loss_weights.csv). The equal-weight outputs regenerate byte for byte.

Dependency on #1104

The runs used #1104's target_loss_weights.py at 4d87854 (sha256 849fbede…); the full-surface loss-vector digest is ba36f887407181dc…, which #1104's results.json now records too. #1104's review fixes later changed the file (now 0c2a6a9d…) but not these weights. check_weights.py rebuilds every w_ run's training and yardstick weights with the current module and compares both digests with the receipts: 46 runs, 7 distinct weightings, no mismatch, against #1104's head c4fd2e8 and again against main's module (sha256 5cb42663…, imported normally) after #1104 merged as e34712c. This branch has main merged in.

Invariants (executed, not property-tested; experiments have no test directory in tools/ci_test_plan.py)

  1. The solve optimizes the weighted loss the harness computes. Every w_ run records the solver's epoch-0 loss against relative_error_loss of the starting weights with the same weights, and its final loss against the recomputed weighted loss of its returned weights. The first pass's launch refused to fan out until its first run passed both. The second pass reused that cached run, so check_weights.py asserts both on all 46 receipts: epoch-0 within 1e-5 (at most 9.3e-6), final differences exactly 0.
  2. Differential: harness digest = module digest for every weighted run's training vector (a mismatch aborts the run). The module's results.json and the harness agree on the full-surface digest.
  3. Differential: the runs' weights = the current module's weights (check_weights.py), for every run's training rows and multipliers and its yardstick.
  4. Differential: file load = package import. The weights are identical either way for the full surface and folds 0 and 1, and stable across PYTHONHASHSEED.
  5. The yardstick is independent of the multipliers. The full-surface weights are equal with and without census_population=8, while the population rows' training weights scale by exactly 8 relative to the rest (smoke test).
  6. Equal mode is unchanged. An equal spec passes no target_loss_weights, and analyze.py regenerates every pre-existing output byte for byte.
  7. Registry round-trip. Specs reload in CSR row order with values equal to the checkpoint's. The receipt binds the registry's sha256 to the checkpoint's MANIFEST and to targets_meta.
  8. Resume binds the weights. A weighted run's resume.npz records its training-weight digest, and a resume under other weights, or from a file with no digest, is refused (smoke-tested).

Review

The Subfleet review (subfleet run --task review --tier standard -m opus, job 20261004-190337-review-1105-r1) could not start: every Claude lane was auth-dead or limited, the earliest until 2026-10-06 04:00Z. In its place, in-session Opus reviewers ran as workflow wf_a1e85ecd-9f1: four dimensions (harness, fidelity, analysis, claims and re-pick), with every medium-or-higher finding adversarially verified.

  • Harness and fidelity approved. Analysis and claims requested changes.
  • Verified medium: finding 5 generalized from one λ. Restated by λ.
  • Downgraded to low or nit, fixed in round 1 (64ca105):
    • the recommendation's margin over λ 0.1, and its ESS trade against the alternatives;
    • "lowest of everything tested" scoped to configurations with held-out folds;
    • the dropped "How it was run" heading;
    • two range claims (held-out district error, the 0.7% qualifier) and SOI's held-out share (92-96%);
    • weighted_frontier.md labels: the multiplied training objective is now separate from the default-weights loss, and the yardstick column has one definition;
    • the second pass's launch did not gate, so check_weights.py asserts the consistency checks on every receipt;
    • resume now binds the weights' digest;
    • module_loaded_from_file reports the actual load;
    • limits for the one surface (--soi-mode state), for registry metadata compiled at this head, and for what the multiplier left untested.

Round 2 (wf_f3233be0-63c, two reviewers on the round-1 delta, claims and code) approved at 64ca105; every round-1 finding was fixed or acceptably scoped. Its residuals, all fixed in the next commit:

  • one low: the untested list wrongly said no λ below 0.03 ran with ×8 (λ 0 did);
  • nits: the district ranges' scope (1.1-1.9% at share 0.5; fold means over two-fold configurations), two roundings (0.110; within 0.75%), "loss weight" where weight shares are meant, item 4's λ qualifier, check_weights.py now flags a missing, None or NaN consistency value instead of crashing (tested on doctored receipts) and reports its actual load, a table label, and registry.py's value comment.

Round 3 (wf_c0224711-ecb, one reviewer on that delta) approved at f81493e with four nits, fixed in the final commit: the 0.75% figure scoped to share 0.5, the 1.1-1.9% sentence's scope, two more "loss weight" labels, and check_weights.py no longer accepting a JSON boolean as a number.

Round 4 (wf_3fb8ea48-ef1) approved at 6a5d762 with one low finding: finding 5's "within 0.75%" holds on the full surface but the fold solves' trained districts reach 2% (1.98% at λ 0.03, fold 0). The final commit (4fc78f4) applies its suggested scoping, in finding 3 verbatim and in finding 5 with the two figures the other way round.

Round 5 (wf_c1126d14-6c0) approved at 4fc78f4, the PR's head, with one PR-body nit (this sentence), now fixed. Review journals: _build_artifacts/acs-local-l2-basis-20260928/review/weighted/review-r*-journal.jsonl.

axiom: n/a: calibration experiment; no policy rules change.

🤖 Generated with Claude Code

MaxGhenis and others added 7 commits October 4, 2026 15:08
registry.py rebuilds the 09-23 release's TargetSpecs from its feed (the
metadata target-loss weighting reads); all 4,459 names and values match the
checkpoint exactly. sweep.py gains target_weighting ("equal" keeps the old
harness byte for byte; "shared" passes the shared module's weights as
target_loss_weights and records them, with weighted fit blocks).
modal_sweep.py gains the weighted grid (w_* ids) and a harness gate;
cross_score.py scores the equal-weight solves on the weighted loss;
analyze.py writes the weighted outputs to their own files.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ness

sweep.py now weights with microcosm.build.us_runtime.target_loss_weights.
us_acs_local_target_loss_weights (the us_acs_local.v1 row mapping) and
records the module's own digest and distribution. The Modal image ships
that one module file and sweep.py loads it on its own, since importing it
through the us_runtime package needs the whole build stack.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Written while the w_ grid's first run was still calibrating, before any
weighted held-out result existed. Also commits cross_scores.json: the
equal-weight solves scored on the weighted loss.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ation multiplier

The first pass's held-out weighted error was still falling at projection
λ 0.1, and its weighted solves miss 11-135 trained district populations by
more than 10%. The second pass extends projection λ and tests the shared
weighting's family lever (census_population ×4, ×8). Runs gain
family_loss_multipliers (training weights only; held-out scoring keeps the
unmultiplied yardstick) and a yardstick training-fit block.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Receipts in results/runs/w_*.json (kernel c028b6a, shared weights from
#1104's target_loss_weights.py 849fbede…, full-surface digest
ba36f887…). Every run's epoch-0 loss matches the recomputed weighted
loss within 1e-5 and its final loss exactly.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…d loss

Second pass (18 Modal solves): projection λ 0.2 and 0.3 locate the
held-out optimum at λ 0.1 on the default weights; a census_population
multiplier of 8 recovers the district population fit the weighting loses.
The README's "On the weighted loss" section reports both passes, applies
the pre-registered rule (projection λ 0.1), states the district population
cost it did not weigh, and recommends λ 0.03 with census_population=8.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
MaxGhenis and others added 7 commits October 4, 2026 19:12
…le hash

#1104's review fixes changed target_loss_weights.py (849fbede… to
0c2a6a9d…) without changing these weights. check_weights.py rebuilds every
w_ run's training and yardstick weights with the current module and
compares the digests with the receipts: 46 runs, 7 weightings, no mismatch
at #1104's head c4fd2e8. Rerun against main before merging.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…fs, honest labels

- Restate finding 5 by λ: the weighted solve is worse held out at λ 0, split
  at 0.03 and better than every equal-weight run at its own optimum, 0.1.
- State the recommendation's margin over λ 0.1 and its ESS trade against the
  rule's pick and the equal-weight pick; scope 'lowest' to configurations
  with held-out folds; restore the dropped 'How it was run' heading.
- Correct two ranges and SOI's held-out share; add limits for the one
  surface, the registry metadata and the multiplier's untested range.
- weighted_frontier.md separates the default-weights loss from the
  multiplied training objective; yardstick_loss has one definition.
- check_weights.py asserts every receipt's consistency checks (the second
  pass's launch did not gate); resume binds the weights' digest;
  module_loaded_from_file reports the actual load.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The untested list no longer says ×8 lacks a λ below 0.03 (λ 0 ran); the
district ranges are scoped to what they cover; 0.110 and within 0.75%
round correctly; weight shares say 'loss weight'. check_weights.py flags
missing, None or NaN consistency values instead of crashing and reports
its actual load; the holdout table labels its training objective.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Scope the 0.75% equal-weight district figure to share 0.5 and the 1.1-1.9%
held-out share to share 0.5; label two more weight shares as loss weight;
check_weights.py no longer accepts a JSON boolean as a number.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Within 0.75% on the full surface at share 0.5 and λ 0 and 0.03; the holdout
folds' trained districts reach 2% (1.98% at λ 0.03, fold 0). Review round 4.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant