Skip to content

Register gate_epuf against SSA's Earnings Public-Use File (unlocked, report-only) - #509

Merged
MaxGhenis merged 12 commits into
masterfrom
epuf-gate-registration-20261001
Oct 4, 2026
Merged

MaxGhenis merged 12 commits into
masterfrom
epuf-gate-registration-20261001

Conversation

@MaxGhenis

@MaxGhenis MaxGhenis commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

What this does

Registers gate_epuf: the rules for scoring the gate-1 generator's earnings against SSA's 2006 Earnings Public-Use File (EPUF), a public 1 percent sample of Social Security numbers with capped taxable earnings for 1951-2006. A gate is a pass-or-fail test whose rules and thresholds are fixed and published before the model is scored against it.

Outcome: the gate does not lock. As registered it has not been shown to catch anything gate 1 does not already catch, and it could not meet its own check on its bite. It gates nothing. Nothing here edits gates.yaml, a locked gate or a committed runs/*.json; the registration record is docs/design/gate_epuf_block_draft.yaml (locked: false, status: unlocked_report_only). No model has been scored.

Full record: docs/amendments/gate_epuf_registration_proposal.md. Review rounds, each by an independent Opus 5.5 lane: reviews/gate_epuf_round1_referee_20261002.md (AMEND: do not lock), reviews/gate_epuf_round2_verification_20261002.md (merge after listed fixes), reviews/gate_epuf_round3_rereview_20261002.md (APPROVE, five minor points), and reviews/gate_epuf_round4_confirmation_20261002.md, a confirmation of 95291e75 (APPROVE, three optional points; also posted below as a comment).

Why it does not lock

  • The generator redraws earnings on held-out persons' observed PSID periods (even years 1998-2022, ages 25-59), so its overlap with EPUF's reliable years is 1998, 2000, 2002 and 2004. Nothing generates a career's earnings before 1998.
  • On that overlap only two cells met the rules for gating: the 1998-2004 rank persistence of capped earnings, for men and for women.
  • Those cells fail a persistence shortfall from the PSID of about 0.10 four times in five and 0.11 nine times in ten. The registration's own check asked them to catch a smaller shortfall (0.07 for men, 0.06 for women) nine times in ten, which the rules could not deliver. The pause that triggered was a flaw in the registration, not a surprise in the data.
  • A perturbation drawing each person's early years from donors of both sexes raised men's persistence to about EPUF's level, which the cells accept. No perturbation was shown to pass gate 1 and fail these cells.
  • With the PSID's distance from EPUF now public, candidate 11's verdict on these cells is largely predictable.

The round-1 referee (independent Opus 5.5) rejected loosening the check after the result. This PR takes that ruling. Section 12 of the proposal lists what a gate with bite would need: cells whose bridges no one has computed (which needs a generator of earlier years or careers), a bite dosed above the cell's own detection point, and a demonstrated catch beyond gate 1.

What EPUF does measure

On the generator's support (5,769 PSID persons; 1,311,282 EPUF persons born 1947-1973). Each figure is the unweighted mean of three birth-cohort bands.

EPUF PSID
Rank persistence 1998-2004, men 0.711 0.668
Rank persistence 1998-2004, women 0.672 0.640
Men's positive person-years at the taxable maximum 11.6% 17.8%
Men with a zero year in 1998-2002, given earnings in 2004 12.8% 9.7%

The persistence gap holds in five of the six sex-by-cohort cells. The artifact measures these gaps; it does not separate how much comes from reporting, from who the PSID samples, or from noncovered work.

On EPUF alone, rewriting careers by two of the career assembler's rules (nothing before 1968; odd years from 1997 filled from neighbours) lowers the median AIME of men born 1940-1944 by 10 percent and of men born 1930-1934 by 36 percent. A cohort born in 1946 or later loses no year to the first rule.

EPUF has 4,384,254 persons; the 4,348,254 in parts of SSA's article is a digit transposition.

What a run reports

populace_dynamics.harness.epuf_run.report_candidate reports every window cell for a candidate's 20 generated panels: the 20-seed estimate, its distance from EPUF, and that distance split into the candidate's distance from the PSID and the PSID's distance from EPUF. It returns no pass or fail. No run script calls it yet.

Invariants, and where they are tested

  • The reader refuses bytes other than the pinned SHA-256, and reproduces SSA's published Table A1, Table 4, Charts 3-4 and the research note's Table 8 from the staged file (tests/data/test_epuf.py).
  • The measurement operator keeps every value in [0, wage base], preserves positive and at-maximum status, and is weakly increasing in all 56 years (Hypothesis, tests/harness/test_epuf_operator.py).
  • Unit-weight Spearman equals scipy's; integer weights equal repeated rows; cells ignore weight scale and row order; the career AIME equals ss.statutory_aime.aime (tests/harness/test_epuf_cells.py).
  • The acceptance interval always contains EPUF and the PSID's position and never extends past EPUF by more than the tolerance; an eligible cell's interval lies inside the cap (tests/harness/test_epuf_gate.py).
  • The two-term floor matches the 20-seed spread of a generator that draws from the true law, in simulation, and a holdout-only floor is too narrow (test_floor_prices_a_faithful_generator).
  • A report's gap equals its source term plus its model term in every cell, and a report carries no pass field (tests/test_epuf_gate_floor_builder.py).
  • The floor artifact, the block and the supplement recompute from stored replicates; the derivation files hash to what the artifact recorded; the first build's window results equal the rebuild's; the reporting path and supplement are pinned by literal SHA-256 (tests/test_gate_epuf_block_draft.py).

Commits

  • eec910d6 rules first, pushed before any real-PSID value in EPUF units existed.
  • 970a9db7 fixes the report-only career AIME (it now ranks every year after 1950, as the statute does).
  • ce8d5000 floor artifact, block, results.
  • 246eab02 excludes the new modules from the birth-evidence reducer's identity seal (the earlier CI failure).
  • d606ddea referee round 1: ruling, supplement, frozen first build, corrected wording, paper sentence.
  • 93848324 verification round 2: a reporting path with no verdict; wording.
  • 95291e75 the approving re-review's five minor points.
  • 56938042 the confirmation review's three optional points: an independent check of the pooled estimate (on seed values now skewed, since a median passed the old symmetric ones), verdicts kept apart from fixes in the block's review records, and one rewrapped line. The block now records round 4.

The five derivation files are unchanged since ce8d5000, so the floor artifact still binds to the code that built it.

Paper

paper/paper.qmd, section "The PSID's limits": the last sentences of the EPUF paragraph no longer say the comparison is undesigned and unregistered. They say its rules are registered for the generated earnings, that only two cells met the rules for gating and nothing was found that they would catch beyond gate 1, so the comparison is reported without a pass or fail, and that the model has not yet been scored. The site's paper bundle needs a rebuild after merge.

Tests

150 new tests (120 unit, 30 artifact); the EPUF reader's reproduction tests skip unless the pinned bytes are staged. tests/tier_counts.json and tests/README-tiers.md are updated.

🤖 Generated with Claude Code

Registers the rules of a proposed gate that scores the gate-1 generator's
earnings against SSA's 2006 Earnings Public-Use File, before any real-PSID
value in EPUF units exists. The floor artifact and the draft gates.yaml
block follow in a later commit, built by the code committed here.

- data/epuf.py: byte-pinned EPUF reader (POPULACE_DYNAMICS_EPUF_DIR)
  that reproduces SSA's published Table A1, Table 4, Charts 3-4 and the
  research note's Table 8 from the staged bytes.
- harness/epuf_operator.py: EPUF's cap and disclosure operator for
  survey-side and model-side earnings, with per-year constants read off
  the pinned bytes.
- harness/epuf_cells.py: the window cells (r6, zint, d_anyzero, q_atmax,
  mpers, q_sexratio by sex and cohort band, 1998-2004) and the
  report-only career cells.
- harness/epuf_gate.py: floor replicate, tolerance, acceptance interval,
  eligibility, ladder and scoring.
- scripts/build_epuf_gate_floors.py: the candidate-blind floor builder.
- docs/amendments/gate_epuf_registration_proposal.md: the proposal's
  rules; its results section is filled by the floor build.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@vercel

vercel Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
social-security-model Ready Ready Preview Oct 4, 2026 3:24pm UTC

Request Review

MaxGhenis and others added 2 commits October 2, 2026 09:50
The report-only career AIME ranked only ages 22-61. The statute ranks every
year after 1950 through the cutoff, including years before age 22, and its
oracle test had been built the same way, so it matched by construction. The
test now passes the oracle every year through age 61.

The career-assembler mask now starts each career at max(1968, birth year +
22), as build_career does, rather than at 1968 for everyone. For the
cohorts it is applied to (born 1930-1944) the two agree.

Neither cell enters the gate: both belong to tranche R, which is
report-only. Found after the first floor build, whose outputs are preserved
in the evidence folder (epuf-20261001/first-floor-build).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Adds the floor artifact the rules-first commit's builder produced, the
draft gates.yaml block it renders (locked: false, not in gates.yaml), the
test that binds the two, and the proposal's results.

- Support: 5,769 PSID persons present at 1998-2004 with a last period of
  2006 or later; EPUF side 1,311,282 persons born 1947-1973.
- Gated: 1998-2004 rank persistence by sex. PSID 0.668 vs EPUF 0.711 (men)
  and 0.640 vs 0.672 (women); faithful-candidate pass probability 0.9986.
- Reported: zero-year and maximum cells, unpowered at PSID scale or with a
  bridge past the cap (PSID men at the maximum 17.8% vs EPUF 11.6%).
- Pause: the registered persistence bite fails the gate 44% of the time,
  against 90% required. Options are listed for the referee round; no rule
  was changed to clear it.
- Tranche R, EPUF only: the career assembler's rules cut the median AIME
  of men born 1940-1944 by 10% and 1930-1934 by 36%.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
MaxGhenis and others added 3 commits October 2, 2026 16:53
…ion-20261001

# Conflicts:
#	tests/README-tiers.md
#	tests/tier_counts.json
CI shard 1 failed test_reducer_input_identity_matches_reviewed_branch: the
five new opt-in modules (data/epuf.py and harness/epuf_{operator,cells,
gate,run}.py) were not in POST_REVIEW_SOURCE_EXCLUSIONS or the test's
matching tuple. Nothing historical imports them; the reachability test now
asserts that for the five, as it does for the bridge modules.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The independent referee (reviews/gate_epuf_round1_referee_20261002.md,
verdict AMEND) found that the registered gate cannot fail the generator
for anything gate 1 does not already catch, and rejected relaxing the
bite requirement after the result. This commit takes that ruling.

- The block's status is unlocked_report_only and it gates nothing. The
  two cells the registered rules selected are recorded beside the ruling.
- runs/epuf_gate_supplement_v1.json stores each bite's per-seed estimates,
  mean shift and power under the gate's own noise model. The registered
  persistence bite shifts the correlation by 0.067 (men) while the cell's
  90 percent detection point is 0.111, so the pause was built into the
  registration. It also stores the support's birth-year mix beside EPUF's.
- runs/epuf_gate_floors_v1_first_build.json is the first floor build,
  frozen; a test checks its window results equal the rebuild's.
- The block names the rules commit and the build commit separately and
  pins the scoring path's SHA-256.
- The proposal states the outcome in plain words, corrects the wording the
  referee flagged, and lists what a gate with bite would need.
- The paper's section on the PSID's limits no longer says the comparison
  is undesigned and unregistered.

No derivation file changed: the floor artifact still binds to the code
that built it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@MaxGhenis MaxGhenis changed the title Register an EPUF covered-earnings gate (proposal, unlocked) Register gate_epuf against SSA's Earnings Public-Use File (unlocked, report-only) Oct 2, 2026
The round-2 verifier (reviews/gate_epuf_round2_verification_20261002.md,
merge after listed fixes, no blockers) found that the record claimed a
report-only path the code did not provide: the scoring function split the
gap for the two persistence cells only and still returned a pass field.

- harness/epuf_run.py: report_candidate replaces score_candidate. It
  reports every window cell (estimate, distance from EPUF, and that
  distance split into model and source terms) and returns no pass or
  fail. No run script calls it yet, and the record now says so.
- The proposal, block and paper say the gate "has not been shown to catch
  anything gate 1 does not already catch" where they said it "cannot fail"
  the generator for it; no perturbation was ever scored on gate 1.
- Section 7: sex-level figures are labelled as means over the three
  cohort bands; two bridges corrected in the third decimal (+0.424,
  +0.086); one cohort cell (r6.women.c1) is eligible and reported; the
  both-sexes donor bite lands slightly past EPUF, not short of it.
- The birth-year comparison is described as a comparison of labels, not a
  measure of the shift.
- Tests pin the reporting path and supplement by literal SHA-256 and
  recompute the 80 percent detection points, the per-cell bite fail
  shares and the birth-year mix.
- The paper sentence follows the verifier's wording.

No derivation file changed (git diff ce8d500 HEAD on the five is empty).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The re-review of the round-2 fixes (reviews/gate_epuf_round3_rereview_
20261002.md) approved head 9384832 and listed five minor points.

- The paper and the proposal's outcome say only two cells "met the rules
  for gating". A third cell had the power but the ladder did not select
  it, so "had enough power" contradicted section 7.2.
- report_candidate is now tested against the committed artifact, where
  the bridges are not zero: the source term equals the stored bridge, the
  model term is the rest, and undefined values propagate.
- The supplement tests assert their cell sets and band years, so none can
  pass on an empty loop.
- Section 10 lists the round-2 commit; the checklist says "reporting
  path"; the block's ruling gives the bite shift for both sexes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@MaxGhenis

Copy link
Copy Markdown
Contributor Author

Confirmation review of the final head 95291e75 (independent Opus 5.5 lane, subfleet job 20261002-173551-epuf-gate-r4; read-only, no shell). Posted verbatim. Rounds 1-3 are committed under reviews/.

APPROVE

All five of round 3's points are fixed at 95291e75. The delta adds no false claim and nothing that reads as a lock or a certification. gates.yaml still has no EPUF entry, and the block still has locked: false and gated_cells: {}. I had no shell, so this rests on reading the files; I didn't run the tests.

# Round-3 point Status Evidence
1 "only two cells had enough power" in the paper and the proposal's outcome RESOLVED Both now say "only two cells met the rules for gating" (paper/paper.qmd:366, proposal.md:36), the wording round 3 suggested. The sentence is now true and matches §7.2–7.3. §7.2 says r6.women.c1 is eligible but not gated, because the ladder gates cohort cells only when all six are eligible (proposal.md:374-377). §7.3 says "The ladder fell back to sex-level persistence and selected two cells" (:381). The rest of the paper sentence ("we found nothing they would catch that gate 1 does not… without a pass or fail") matches §7.4 (:430-431). No "enough power" or "had the power" is left in the paper, proposal, block or renderer.
2 report_candidate tests: source and model terms on nonzero bridges, and the NaN paths RESOLVED See the notes below the table.
3 Supplement tests could pass vacuously RESOLVED See the notes below the table.
4 §10 commit record and the checklist wording RESOLVED Item 4 now names d606ddea. A new item 5 names 93848324 (report_candidate, the round-2 report and its corrections) and mentions the final commit. The checklist says "a pinned reporting path" and adds the re-review line. "Scoring path" survives only in the block's round-2 record (block.yaml:206, renderer :73), where it correctly describes what was replaced.
5 Four wording fixes RESOLVED • Block ruling (rendered from the script, block.yaml:190-199): "0.067 for men and 0.057 for women". This matches §7.4 and the supplement's −0.0665/−0.0568.
• §10 quotes "the house floor formula", the docstring's exact text (epuf_gate.py:179).
• §7.4 says bd2c moves the cells "about 0.01", matching the table's +0.007/+0.010 (proposal.md:402).
• §9 says tranche R "would report… no code computes its PSID side yet". career_cells is called only on EPUF data (build_epuf_gate_floors.py:295-296), so this is accurate.

Point 2 in detail. The new tests cannot pass vacuously:

  • Real bridges, not the synthetic fixture. test_report_candidate_terms_against_the_committed_bridges feeds the committed artifact's cells into report_candidate. candidate_window_cells is stubbed out (monkeypatched); report_candidate calls it through the module, so the stub takes effect.
  • Nonzero bridges. I counted the artifact's 41 bridges: 38 have |b| > 0.01, against the test's floor of 25. Only three are below 0.01 (0.0045, −0.0097, −0.0095).
  • Source term. The test compares the code's psid − epuf on the cell's scale with the stored bridge_psid_minus_epuf. A flipped sign or a wrong scale would now fail.
  • Model term. It is checked as estimate − transform(psid), which is computed separately from the code's gap − source.
  • Coverage. The test asserts that the report's cell ids equal cell_ids(), so all 41 cells are checked. It also asserts that neither the report nor any cell carries a pass key.
  • Defined values only. pytest.approx never treats NaN as equal to NaN, so the test cannot pass by producing NaN.
  • NaN paths. test_report_candidate_propagates_undefined_values sets one seed of r6.men to NaN, which makes its estimate and gap NaN. It sets r6.women's psid_value to None, which leaves the gap defined but makes both split terms NaN (via _on_scale, epuf_run.py:102-103).

Point 3 in detail.

  • set(bites[name]["cells"]) == set(artifact["registered"]) is added (test:161). It can't hold for empty sets: the detection-point loop indexes bites["bd1_persistence_loss_0.10"]["cells"]["r6.men"] and ["r6.women"] (:209), which would fail if those cells were missing.
  • set(bites["detection_points"]) == {"r6.men", "r6.women"} is a literal (:192).
  • The band years are now tied to COHORT_BANDS (:350-351), which defines c0 = 1947–1955, c1 = 1956–1964 and c2 = 1965–1973.

New findings. Nothing blocks merge. Three optional nits:

  • The pooled estimate is checked against itself. Test 1 compares estimate with gate.pooled_estimate, the same function the code calls, so an error in that function would go unnoticed. Fix (optional): also assert row["estimate"] == pytest.approx(transform(cell_id, statistics.fmean(per_seed))).
  • The round-3 verdict mixes in the drafting session's claim. The block records verdict: APPROVE (five minor points, applied) (block.yaml rereview_round_3), but "applied" is the drafting session's claim, not the reviewer's verdict. Round 2's record uses the same pattern. Fix (optional): verdict: APPROVE (five minor points) and fixes: applied in 95291e75.
  • One long line. proposal.md:574 ("the house floor formula" line) now runs past the wrap width.

Lock and certification check.

  • gates.yaml has 0 EPUF matches.
  • The block has status: unlocked_report_only (:18), locked: false (:19), gated_cells: {} (:109) and lock_ceremony.exists: false (:216).
  • The header says it "gates nothing" and "edits no gates.yaml byte".
  • The diff adds no wording about passing, locking or certifying. The phrase "met the rules for gating" sits beside "report the comparison without a pass or fail".

What I could not check:

  • Tests. I ran no tests. So I couldn't confirm that the two new tests pass, that the tier count of +2 (3,359 → 3,361) is right, or that the block equals a fresh render.
  • Git. I ran no git commands. I took the delta from .diag/round3-fix-delta.diff, and the claim that the derivation-core diff since ce8d5000 is empty from the drafting session.
  • Hashes. I computed none. The supplement and artifact hash literals are unchanged in the diff.
  • PR body. I didn't read it.
  • PSID data. I read no PSID microdata.

Resolves the only conflicts, in tests/tier_counts.json and
tests/README-tiers.md, by recounting the merged suite with
pytest --collect-only -m <tier>: unit 5,948, artifact 3,361,
integration_psid 1,341, reproduction_legacy 520, oracle_policyengine 220
(total 11,390). No registration file, derivation file or result changes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…l points

- reviews/gate_epuf_round4_confirmation_20261002.md commits round 4
  (job 20261002-173551-epuf-gate-r4, head 95291e7) verbatim; the block
  records it as confirmation_round_4.
- Rounds 3 and 4 keep the reviewer's verdict apart from the fixes.
- The committed-bridges test checks each estimate against the seeds'
  mean computed without pooled_estimate. The invented seed values are
  now skewed: with the old symmetric tilt, a median passed. A median
  and a mean of logs each fail the test now.
- One proposal line rewrapped; section 10 names 95291e7.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The history list stopped at round 2, and its round-2 line mixed the
verdict with the fixes. It now has one entry per review round, each
keeping the verdict apart from the commit that applied the fixes, and
the rounds test checks every round appears in the history.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ion-20261001

# Conflicts:
#	tests/README-tiers.md
#	tests/tier_counts.json
@MaxGhenis
MaxGhenis marked this pull request as ready for review October 4, 2026 15:23
@MaxGhenis
MaxGhenis merged commit 492d5a6 into master Oct 4, 2026
12 checks passed
@MaxGhenis

Copy link
Copy Markdown
Contributor Author

Merge record:

MaxGhenis added a commit that referenced this pull request Oct 4, 2026
#509 merged as 492d5a6, so this branch now targets master. Merged rather
than rebased so the registered floor build's commit (cb76ad1) stays on
the branch's history, which the bound-file test diffs against.

Only the tier-count files conflicted; recounted against the full
collection: unit 6,017, artifact 3,388 (total 11,486).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
MaxGhenis added a commit that referenced this pull request Oct 4, 2026
…ates branch

Brings in master through #509 (492d5a6). Merged, not rebased, so the
manifest's code_commit (b722382) stays on this branch's history; the
manifest, the staged fills and every pinned file are unchanged.

Only the tier-count files conflicted; recounted against the full
collection: unit 6,037, artifact 3,398 (total 11,516).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
MaxGhenis added a commit that referenced this pull request Oct 4, 2026
…eakdowns PR

Keeps both sides' reducer exclusions. Tier manifest from a full
collection, reconciled with master plus this PR's delta: unit 6,401,
artifact 3,904, integration_psid 1,346.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch was successfully deployed

1 active deployment
Preview — 358e15b5 Deployed Oct 4, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant