You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
UK sparse selection: keep household mass, composition, nation shares and per-area ESS at a tractable size — test the target-weight rule and the L2 penalty first, dataset size only if they fall short #1124
The K=25 measurement build (uk-local-candidate-f100-s42-20261006T094003Z-835d98d9, #1115) shows the dense solve is sound and the 60,000-household selection is where sub-national quality is lost. Two builds now give the same shape (K=15 at 55,000 in #877's P50/P95b; K=25 at 60,000 here), so the clone count is not the lever. This issue sets the order in which the solve's own levers are tested: the target-weight rule (A) and the L2 penalty (E) first, each alone and together; dataset size (F) only if A and E leave the acceptance criteria unmet. Changes to the selection algorithm itself (allocation, carrier protection, refit mass) are recorded at the end as options that are not tested in this round. Data-quality defects of the surface and spine are #1123.
What the chain does today
Dense solve: capped relative-error loss, grain_equal weights (one third each to national, constituency and local-authority rows, uniform inside a grain; local_doctrine.uk_local_target_loss_weights), stretch bound 10 on the pool design weights, free mass, 2,000 epochs, l2_lambda = 0.
Selection (dataset_size.select_uk_dataset_size): contribution_initialization protects one row per target (the largest weighted carrier, 10,558 rows) and seeds gate probabilities from contribution shares; hard-concrete L0 gates with the same loss, weights and bound; the budget search tunes λ_L0 on open-probability mass until a draw at π_hi is feasible. Landed at λ 1.15e-6 with 59,806 certainties and 194 boundary draws, so the draw is a threshold on the learned gates. l2_lambda = 0.
Refit (refit_l0_selection via refit_uk_dataset_size): starts from the kept rows' design weights divided by inclusion probability, normalised to the full pool mass (29.0m), free mass, stretch bound 10 on that baseline, same weights and cap, 2,000 epochs, l2_lambda = 0. The dense weights are not used.
The loss weights are the same vector in all three stages, so a weighting rule changes which rows the search keeps as well as how the refit spreads weight over them.
What it costs (S60 against D25, same pool, same targets)
Weighted households 29.02m → 28.37m (−2.2 %) while weighted persons hold (69.3m → 69.2m).
Lone-person households 28.6 % → 24.1 % of weight (31 % → 18 % of rows); national rows: lone households 65+ −35 %, lone parents +43 %, couples with non-dependent children +34 %. Same mechanism as UK dataset sizes: informed L0 and exact-count candidates (#355) #877's P50 (−7.7 % households then).
Northern Ireland: 9.2 % of pool rows → 5.3 % of kept rows; weight share 2.68 % → 1.97 %; census households −25 % to −39 % in all 18 constituencies and 11 districts; tenure −45 % to −72 %.
Per-area support: constituency ESS median 59 (min 19), 212 of 650 under the ESS-50 floor (Port UK local target surface and credibility gates from uk-data #147 ruling); 28 districts. D25: median 175, none under 50. ESS/rows median 0.68. The incumbent's weight matrices reach a median constituency ESS of 456, but with about 38,000 non-zero households per area.
105 local rows past 25 % that D25 fits: 32 constituency employment-income amounts all over-estimated (+25 % to +75 %, stretch on the few high earners kept), 24 NI census cells, 22 NI tenure cells, 13 private-rent districts, 4 age 20–30 cells.
National rows past 25 %: 16 vs 5. SDLT −58 % (dense 0 %), CGT residential gains −42 %, salary sacrifice −23 %, savings interest 150k+ −41 %.
Kept rows carry 11,167 of 16,288 FRS source households; distinct sources per constituency median 83.
K: 211 constituencies under the floor at K=15/55,000, 212 at K=25/60,000; dense per-constituency ESS median 219 at K=15 (R13) against 175 at K=25.
A. The target-weight function, as named doctrine rules
The doctrine (#762 A2, local_doctrine closed vocabulary) admits named rules derived from row metadata, never per-target vectors. No open issue covers a weighting function; related: #492 (loss shape and the past-cap census), #724 (closed; the US count/amount 50/50 rule), #104 (zero-valued targets), #458 (per-target tolerances).
Candidate rules, each a declared function of row metadata:
grain_family_equal: equal grain shares, then equal family shares inside each grain. Today census households are 4.8 % of the local loss and the eight composition rows 0.7 % of the national third; the rule roughly triples the first and quintuples the second. The national build already solves under family_equal, which is why rail (0 % there, −66 % in D25), SDLT and ESA contributory fit there and not here.
A nation grain: nation-level rows (where bound) take their own share so the solve cannot trade Northern Ireland away.
A magnitude-aware count rule: the capped relative-error loss treats a 3-household cell like a 50,000-household cell; a declared function of target magnitude for count rows keeps the micro-cells from absorbing stretch.
A reliability-tier multiplier from Chronicle provenance (administrative outturn, forecast, research estimate), declared once per tier.
Mechanism: because the same weights feed the L0 search, a heavier household-count and composition share makes it cheaper for the search to keep the carriers of those rows (small households, NI households) than to close them; in the refit the same weights pull the kept rows' weights back toward the household totals.
E. The L2 penalty
Supported by the solver since #309 (l2_lambda, l2_anchor initial/uniform/design, l2_basis record/chi_square) in both calibrate and refit_l0_selection, but never threaded through dataset_size._solver_common, so every UK stage runs at 0. #285 planned the sweep for the US and never ran.
Refit: the chi-square basis to the design anchor pulls the shipped weights toward the kept rows' design spread; it should raise ESS/rows above 0.68, tame the +75 % over-stretched cells and the within-area max/median of 33.5, at a fit cost to measure. It cannot restore dropped rows, and ESS cannot exceed rows, so the thin constituencies (57 to 67 rows) need ESS/rows near 0.85 to pass 50.
Selection: an L2 term on the pre-gate weights penalises concentrating mass on few rows, so at fixed k the search should keep more, smaller-weight rows; that is the route by which L2 can help composition. Hypothesis to measure.
Implementation: thread l2_lambda, l2_basis and l2_anchor through _solver_common for the search and the refit separately (the solver already distinguishes refit_l2_lambda), record them in the size receipt and the doctrine identity.
F. Dataset size, only if A and E fall short
The dense solution's own ESS is 105,000, so a file well below that will keep losing composition; #877's S4 (110,000) was deferred. Size is tested only after A and E, alone and combined, have been measured and the acceptance criteria below are still unmet. Then: the ESS-floor count and the composition loss against k at 60,000, 80,000 and 110,000 with the best A/E configuration.
G. Measurement
Every experiment runs on the stored checkpoint (size_selection_checkpoint.npz and the stored problem and dense solution of the run), so the dense reference is fixed and every delta is selection-only.
Report per experiment: household total and nation weight shares against D25; lone-person and composition shares; per-area ESS distribution and floor count at both grains; ESS/rows; distinct sources per area; fit by family and grain; national rows past 25 %; loss; wall and peak memory.
Refit-only on the fixed 60,000 support (minutes each, no dense solve, no new search): E alone (chi-square/design at λ ∈ {1e-3, 1e-2, 1e-1}); A alone (grain_family_equal, then the nation grain and the count rule if the first moves the composition rows); A and E together.
Selection re-runs with the winning A and E settings inside the L0 search (each a full-pool search, hours, one at a time on the 24 GiB machine): A in the search; E in the search; both.
F only if step 2 leaves the criteria unmet: size curve with the best A/E configuration.
Proposed acceptance for a configuration to replace the current one: household total within 1 % of D25; each nation's weight share within 5 % relative; lone-person share within one point; no local family loses more than one point of within-10 share against S60; national rows past 25 % no more than D25 + 2; the ESS-floor count at the ruled floor (or the re-scoped floor) below a declared number.
Mass-share carrier protection in contribution_initialization (protect carriers up to a declared share of each target's dense mass instead of the single largest).
These are the direct fixes for the nation and small-carrier losses if A and E do not reach them; they change the selection algorithm and are deferred until the two solver levers have been measured.
Rulings needed
Which named rules enter the doctrine vocabulary (A), and whether grain_equal stays the default meanwhile.
Why
The K=25 measurement build (
uk-local-candidate-f100-s42-20261006T094003Z-835d98d9, #1115) shows the dense solve is sound and the 60,000-household selection is where sub-national quality is lost. Two builds now give the same shape (K=15 at 55,000 in #877's P50/P95b; K=25 at 60,000 here), so the clone count is not the lever. This issue sets the order in which the solve's own levers are tested: the target-weight rule (A) and the L2 penalty (E) first, each alone and together; dataset size (F) only if A and E leave the acceptance criteria unmet. Changes to the selection algorithm itself (allocation, carrier protection, refit mass) are recorded at the end as options that are not tested in this round. Data-quality defects of the surface and spine are #1123.What the chain does today
grain_equalweights (one third each to national, constituency and local-authority rows, uniform inside a grain;local_doctrine.uk_local_target_loss_weights), stretch bound 10 on the pool design weights, free mass, 2,000 epochs,l2_lambda = 0.dataset_size.select_uk_dataset_size):contribution_initializationprotects one row per target (the largest weighted carrier, 10,558 rows) and seeds gate probabilities from contribution shares; hard-concrete L0 gates with the same loss, weights and bound; the budget search tunes λ_L0 on open-probability mass until a draw at π_hi is feasible. Landed at λ 1.15e-6 with 59,806 certainties and 194 boundary draws, so the draw is a threshold on the learned gates.l2_lambda = 0.refit_l0_selectionviarefit_uk_dataset_size): starts from the kept rows' design weights divided by inclusion probability, normalised to the full pool mass (29.0m), free mass, stretch bound 10 on that baseline, same weights and cap, 2,000 epochs,l2_lambda = 0. The dense weights are not used.What it costs (S60 against D25, same pool, same targets)
A. The target-weight function, as named doctrine rules
The doctrine (#762 A2,
local_doctrineclosed vocabulary) admits named rules derived from row metadata, never per-target vectors. No open issue covers a weighting function; related: #492 (loss shape and the past-cap census), #724 (closed; the US count/amount 50/50 rule), #104 (zero-valued targets), #458 (per-target tolerances).Candidate rules, each a declared function of row metadata:
grain_family_equal: equal grain shares, then equal family shares inside each grain. Today census households are 4.8 % of the local loss and the eight composition rows 0.7 % of the national third; the rule roughly triples the first and quintuples the second. The national build already solves underfamily_equal, which is why rail (0 % there, −66 % in D25), SDLT and ESA contributory fit there and not here.nationgrain: nation-level rows (where bound) take their own share so the solve cannot trade Northern Ireland away.Mechanism: because the same weights feed the L0 search, a heavier household-count and composition share makes it cheaper for the search to keep the carriers of those rows (small households, NI households) than to close them; in the refit the same weights pull the kept rows' weights back toward the household totals.
E. The L2 penalty
Supported by the solver since #309 (
l2_lambda,l2_anchorinitial/uniform/design,l2_basisrecord/chi_square) in bothcalibrateandrefit_l0_selection, but never threaded throughdataset_size._solver_common, so every UK stage runs at 0. #285 planned the sweep for the US and never ran.l2_lambda,l2_basisandl2_anchorthrough_solver_commonfor the search and the refit separately (the solver already distinguishesrefit_l2_lambda), record them in the size receipt and the doctrine identity.F. Dataset size, only if A and E fall short
The dense solution's own ESS is 105,000, so a file well below that will keep losing composition; #877's S4 (110,000) was deferred. Size is tested only after A and E, alone and combined, have been measured and the acceptance criteria below are still unmet. Then: the ESS-floor count and the composition loss against k at 60,000, 80,000 and 110,000 with the best A/E configuration.
G. Measurement
size_selection_checkpoint.npzand the stored problem and dense solution of the run), so the dense reference is fixed and every delta is selection-only.Experiment ladder
grain_family_equal, then the nation grain and the count rule if the first moves the composition rows); A and E together.Proposed acceptance for a configuration to replace the current one: household total within 1 % of D25; each nation's weight share within 5 % relative; lone-person share within one point; no local family loses more than one point of within-10 share against S60; national rows past 25 % no more than D25 + 2; the ESS-floor count at the ruled floor (or the re-scoped floor) below a declared number.
Not tested in this round (recorded options)
contribution_initialization(protect carriers up to a declared share of each target's dense mass instead of the single largest).CONSERVE_MASSand the one-stretch contract of max_weight_ratio anchors differently per arm: 5x vs design (dense) but ~25x effective (sparse refit re-anchors) — declare one stretch contract #493.These are the direct fixes for the nation and small-carrier losses if A and E do not reach them; they change the selection algorithm and are deferred until the two solver levers have been measured.
Rulings needed
grain_equalstays the default meanwhile.Links
#898, #877 (merged machinery), #355 (closed), #762 (doctrine constants), #147 (floors), #493, #285, #458, #309, #302, #104, #492, #724, #346 (not needed while the pool is fixed), #931.