Repository navigation
Track C step 1: real Axiom AIME vs the Python oracle on the Track A PSID cohort - #455
Merged
Merged
Conversation
scripts/track_c_aime_agreement.py builds the real 2011-wave Track A cohort (load_psid2010_inputs, build_psid2010_cohort, prepare_track_a_cohort with registered_real provenance), selects every member whose year of attaining 62 lies in 1979-2030, and for each one records the oracle AIME exactly as Track A computes it (ss.benefits.aime with the TR2008-AWI parameters of tr2008_ssa_parameters) next to the AIME the pinned Axiom rules engine returns through axiom_benefit_bridge. Every declaration the bridge needs is explicit and recorded: career amounts limited to the contribution and benefit base and declared USD creditable earnings, shortest round-trip binary64 encoding, synthetic zero rows for elapsed-window years before 1968 and after 2010, indexing year birth + 60, and the entitlement year per pass (primary: max(birth + 62, 2011); diagnostic: Track A's recorded opening entitlement year). A strict pass without synthetic rows records the bridge's refusals. An attribution check reruns the oracle's arithmetic with the candidate's computation-year count. The 2026-cohort PIA is a formula check only. It writes result.json, RESULTS.md and executions/ and computes no COLA statistic or age-profile value; the Axiom candidates are labeled unaccepted. Tests use invented data only (one runs the pinned engine and skips when it is absent). Unit tier count 2,847 -> 2,861. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The script now checks, before any engine run, that the rebuilt cohort's diagnostics and source provenance and the parameter revision equal those recorded in runs/replication_urban2010_cola_v1.json (only those fields are used), and records the artifact's SHA-256 and the checks. RESULTS.md gains a result paragraph computed from the summaries, a binned difference distribution (per-dollar counts stay in result.json), a table by the candidate's computation-year count, and the rows sent by kind (observed, gap-imputed, zero-filled, synthetic zeros). Execution files are capped per pass for unexplained persons and recorded in result.json. The bend-point wording now says the candidate derives its 2026 bend points from SSA's published 2024 AWI. Unit tier count 2,861 -> 2,863. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…re used The not-compared list said the age profile was "not read"; the script parses the committed artifact and uses only its parameter revision and 2011-wave cohort diagnostics and source provenance, and now says so. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
MaxGhenis
added a commit
that referenced
this pull request
Sep 24, 2026
…he tier manifest Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds
scripts/track_c_aime_agreement.pyand its tests. The script runs every Track A 2011-wave person the oracle computes an AIME for (7,486 real PSID careers, cohort rebuilt and checked against the committed Track A artifact) through the real pinned Axiom rules engine viaaxiom_benefit_bridge(#450). It then compares the engine's AIME with Track A's Python-oracle AIME, person by person. It computes no COLA statistic.Result
The Axiom rule candidates remain unaccepted. Evidence:
track-c-aime-agreement-20260923/in the lane evidence directory, with engine and module hashes.Review
An independent reviewer recomputed all 7,486 primary and 1,334 diagnostic records with separate exact-fraction code, with zero mismatches. It also re-ran a 54-person sample on the pinned engine, with identical digests. It approved without changes.
🤖 Generated with Claude Code