Repository navigation
Conversation
registry.py rebuilds the 09-23 release's TargetSpecs from its feed (the
metadata target-loss weighting reads); all 4,459 names and values match the
checkpoint exactly. sweep.py gains target_weighting ("equal" keeps the old
harness byte for byte; "shared" passes the shared module's weights as
target_loss_weights and records them, with weighted fit blocks).
modal_sweep.py gains the weighted grid (w_* ids) and a harness gate;
cross_score.py scores the equal-weight solves on the weighted loss;
analyze.py writes the weighted outputs to their own files.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ness sweep.py now weights with microcosm.build.us_runtime.target_loss_weights. us_acs_local_target_loss_weights (the us_acs_local.v1 row mapping) and records the module's own digest and distribution. The Modal image ships that one module file and sweep.py loads it on its own, since importing it through the us_runtime package needs the whole build stack. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Written while the w_ grid's first run was still calibrating, before any weighted held-out result existed. Also commits cross_scores.json: the equal-weight solves scored on the weighted loss. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ation multiplier The first pass's held-out weighted error was still falling at projection λ 0.1, and its weighted solves miss 11-135 trained district populations by more than 10%. The second pass extends projection λ and tests the shared weighting's family lever (census_population ×4, ×8). Runs gain family_loss_multipliers (training weights only; held-out scoring keeps the unmultiplied yardstick) and a yardstick training-fit block. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Receipts in results/runs/w_*.json (kernel c028b6a, shared weights from #1104's target_loss_weights.py 849fbede…, full-surface digest ba36f887…). Every run's epoch-0 loss matches the recomputed weighted loss within 1e-5 and its final loss exactly. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…d loss Second pass (18 Modal solves): projection λ 0.2 and 0.3 locate the held-out optimum at λ 0.1 on the default weights; a census_population multiplier of 8 recovers the district population fit the weighting loses. The README's "On the weighted loss" section reports both passes, applies the pre-registered rule (projection λ 0.1), states the district population cost it did not weigh, and recommends λ 0.03 with census_population=8. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…le hash #1104's review fixes changed target_loss_weights.py (849fbede… to 0c2a6a9d…) without changing these weights. check_weights.py rebuilds every w_ run's training and yardstick weights with the current module and compares the digests with the receipts: 46 runs, 7 weightings, no mismatch at #1104's head c4fd2e8. Rerun against main before merging. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…fs, honest labels - Restate finding 5 by λ: the weighted solve is worse held out at λ 0, split at 0.03 and better than every equal-weight run at its own optimum, 0.1. - State the recommendation's margin over λ 0.1 and its ESS trade against the rule's pick and the equal-weight pick; scope 'lowest' to configurations with held-out folds; restore the dropped 'How it was run' heading. - Correct two ranges and SOI's held-out share; add limits for the one surface, the registry metadata and the multiplier's untested range. - weighted_frontier.md separates the default-weights loss from the multiplied training objective; yardstick_loss has one definition. - check_weights.py asserts every receipt's consistency checks (the second pass's launch did not gate); resume binds the weights' digest; module_loaded_from_file reports the actual load. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The untested list no longer says ×8 lacks a λ below 0.03 (λ 0 ran); the district ranges are scoped to what they cover; 0.110 and within 0.75% round correctly; weight shares say 'loss weight'. check_weights.py flags missing, None or NaN consistency values instead of crashing and reports its actual load; the holdout table labels its training objective. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Scope the 0.75% equal-weight district figure to share 0.5 and the 1.1-1.9% held-out share to share 0.5; label two more weight shares as loss weight; check_weights.py no longer accepts a JSON boolean as a number. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Within 0.75% on the full surface at share 0.5 and λ 0 and 0.03; the holdout folds' trained districts reach 2% (1.98% at λ 0.03, fold 0). Review round 4. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
#1078's ESS frontier for the ACS local release was measured on the release's loss, with every target weighted equally (3,819 of the 4,459 targets are IRS SOI cells). #1104 makes the ACS local build weight its targets like the national release, through one shared implementation. λ is in units of the loss, so decision d797 (which supersedes d792) asks for λ to be re-picked on the weighted loss before a default ships.
This PR is research evidence only. It publishes nothing and changes no default or package code: every change is under
experiments/us-acs-local-l2-basis-20260928/, plus one bullet indocs/calibration-l2-basis.md.What the evidence shows
46 Modal solves of the published release's 4,459-target surface, 800 epochs each. The write-up is the README's "On the weighted loss" section.
The penalty still buys ESS and held-out fit, now with an interior optimum. Held-out weighted capped error (the objective's out-of-sample form, folds 0 and 1) at projection λ 0 / 0.03 / 0.1 / 0.2 / 0.3 is 0.0784 / 0.0721 / 0.0708 / 0.0742 / 0.0787. Over the same range, ESS rises from 13,707 to 24,816. Projection is ahead of softmax at every λ held out.
The weighting costs district population fit, and the penalty multiplies it.
pop_cdrows form one concept group scaled to its largest member's weight, so California's 52 districts carry 1.51 between them, about what one at-large district carries.census_population×8 with λ 0.03 is the best configuration held out, by a small margin.Whether training on the weighted loss helps out of sample depends on λ. On the weighted held-out yardstick, against the same settings trained on the equal loss:
At λ 0 and 0.03 the equal-weight solves keep district populations; the weighted ones need the multiplier to.
The λ re-pick. The rule was fixed and committed before any weighted held-out result existed (43b3d0f). It picks projection λ 0.1. The README states the district cost the rule did not weigh, and recommends λ 0.03 (d797's λ) with
--target-family-loss-multiplier census_population=8. The multiplier was tried after the first pass, so it is labelled post hoc. Both calls are queued for Max: the result is noted on d797 (λ and gate; d792 is superseded by it and noted too), and the multiplier is d952.What changed
registry.pyrebuilds the 09-23 release'sTargetSpecs from its own feed and settings (snap/medicaid/soi,soi_mode=state). All 4,459 names and values match the checkpoint bit for bit, and a receipt records it.sweep.py:target_weightingisequal(the old harness, unchanged) orshared.sharedpassesus_acs_local_target_loss_weights(train_specs)astarget_loss_weights.family_loss_multipliersscales the training weights only; held-out scoring keeps the unmultiplied full-surface weights, one yardstick for every run.microcosm.build.us_runtimeneeds the whole build stack.modal_sweep.py:--grid weightedand--grid weighted2(w_*run ids).cross_score.pyscores the equal-weight solves on the weighted loss.analyze.pywrites the weighted outputs to their own files (results/weighted_*,weighting_shift.md,target_loss_weights.csv). The equal-weight outputs regenerate byte for byte.Dependency on #1104
The runs used #1104's
target_loss_weights.pyat 4d87854 (sha256849fbede…); the full-surface loss-vector digest isba36f887407181dc…, which #1104'sresults.jsonnow records too. #1104's review fixes later changed the file (now0c2a6a9d…) but not these weights.check_weights.pyrebuilds everyw_run's training and yardstick weights with the current module and compares both digests with the receipts: 46 runs, 7 distinct weightings, no mismatch, against #1104's head c4fd2e8 and again against main's module (sha2565cb42663…, imported normally) after #1104 merged as e34712c. This branch has main merged in.Invariants (executed, not property-tested; experiments have no test directory in
tools/ci_test_plan.py)w_run records the solver's epoch-0 loss againstrelative_error_lossof the starting weights with the same weights, and its final loss against the recomputed weighted loss of its returned weights. The first pass's launch refused to fan out until its first run passed both. The second pass reused that cached run, socheck_weights.pyasserts both on all 46 receipts: epoch-0 within 1e-5 (at most 9.3e-6), final differences exactly 0.check_weights.py), for every run's training rows and multipliers and its yardstick.PYTHONHASHSEED.census_population=8, while the population rows' training weights scale by exactly 8 relative to the rest (smoke test).equalspec passes notarget_loss_weights, andanalyze.pyregenerates every pre-existing output byte for byte.targets_meta.resume.npzrecords its training-weight digest, and a resume under other weights, or from a file with no digest, is refused (smoke-tested).Review
The Subfleet review (
subfleet run --task review --tier standard -m opus, job20261004-190337-review-1105-r1) could not start: every Claude lane was auth-dead or limited, the earliest until 2026-10-06 04:00Z. In its place, in-session Opus reviewers ran as workflowwf_a1e85ecd-9f1: four dimensions (harness, fidelity, analysis, claims and re-pick), with every medium-or-higher finding adversarially verified.weighted_frontier.mdlabels: the multiplied training objective is now separate from the default-weights loss, and the yardstick column has one definition;check_weights.pyasserts the consistency checks on every receipt;module_loaded_from_filereports the actual load;--soi-mode state), for registry metadata compiled at this head, and for what the multiplier left untested.Round 2 (
wf_f3233be0-63c, two reviewers on the round-1 delta, claims and code) approved at 64ca105; every round-1 finding was fixed or acceptably scoped. Its residuals, all fixed in the next commit:check_weights.pynow flags a missing, None or NaN consistency value instead of crashing (tested on doctored receipts) and reports its actual load, a table label, andregistry.py's value comment.Round 3 (
wf_c0224711-ecb, one reviewer on that delta) approved at f81493e with four nits, fixed in the final commit: the 0.75% figure scoped to share 0.5, the 1.1-1.9% sentence's scope, two more "loss weight" labels, andcheck_weights.pyno longer accepting a JSON boolean as a number.Round 4 (
wf_3fb8ea48-ef1) approved at 6a5d762 with one low finding: finding 5's "within 0.75%" holds on the full surface but the fold solves' trained districts reach 2% (1.98% at λ 0.03, fold 0). The final commit (4fc78f4) applies its suggested scoping, in finding 3 verbatim and in finding 5 with the two figures the other way round.Round 5 (
wf_c1126d14-6c0) approved at 4fc78f4, the PR's head, with one PR-body nit (this sentence), now fixed. Review journals:_build_artifacts/acs-local-l2-basis-20260928/review/weighted/review-r*-journal.jsonl.axiom: n/a: calibration experiment; no policy rules change.🤖 Generated with Claude Code