Skip to content

Add IRS SOI TY2023 state-data US totals and IRA Tables 5/6 - #296

Draft
MaxGhenis wants to merge 3 commits into
artifact-year-restamp-guardfrom
soi-ty2023-state-ira
Draft

MaxGhenis wants to merge 3 commits into
artifact-year-restamp-guardfrom
soi-ty2023-state-ira

Conversation

@MaxGhenis

@MaxGhenis MaxGhenis commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Adds the TY2023 IRS SOI state-data United States totals and IRA Tables 5/6. They mirror the TY2022 packages that main currently builds at 2023 and stamps as ty2023 (#117). After #292 pins those TY2022 packages to ty2022, these packages supply the real ty2023 facts:

Package File Facts Examples
soi-state-2023 23in54us.xlsx (state-data page, 21 Aug 2026) 4 returns 159,949,000; AGI $15,234,086,106k; EITC 3+ children 3,155,130 returns, $15,299,042k
soi-ira-traditional-contributions-2023 23in05ira.xlsx (IRA study, June 2026) 2 5,873,053 taxpayers; $26,901,749k
soi-ira-roth-contributions-2023 23in06ira.xlsx 2 9,488,414 taxpayers; $32,809,393k

Stacked on #292. Merge after it. Before #292, the TY2022 IRA packages built at 2023 emit the same irs_soi.ty2023.*_ira_contributions… ids with TY2022 values.

What changes

  • Parser: chronicle/sources/cells.py (commit 1). 23in54us.xlsx merges its footnote row A179:XFD179. openpyxl fills every merged position with a MergedCell, so sheet.max_column became 16,384 and xlsx_used_range would have built about 70 million empty cells. A probe reached 17.9 GB before I killed it. Now the placeholders of a merged range that runs to the worksheet edge (column XFD or the last row) do not widen the bounds.
    • The anchor cell and every other cell still count, and a sheet with no such range keeps (max_row, max_column) exactly.
    • Corpus check: across every committed workbook (212 files, 2,014 sheets, zip members included), only 23in54us.xlsx changes. It goes from 4,264 x 16,384 to 4,264 x 143, which is the file's own <dimension ref="A1:EM4264">.
    • A broader rule, ignoring every merged placeholder, would have moved 309 sheets in existing DESNZ, HHS, ICI, W-2, OBR, ScotGov, SLC and Statbel packages, so I rejected it.
    • The state build now parses 609,752 cells at 0.9 GB peak.
  • Packages (commit 2). Each is its TY2022 twin with literal TY2023 labels and the workbook's moved layout:
    • 23in54us.xlsx: the geography label moved from A8 to B3 (merged B3:L3), the "All returns" header from row 3 to row 4, and the EITC block two rows down (EITC 3+ rows 146/147, guards on rows 138 and 146).
    • IRA tables: the sheet is Estimates (was Sheet1), "All taxpayers" is row 7 (was 8), and the header reads Number of taxpayers without the line break.
    • record_set_spec_id is kept, as in the Historic Table 2 and Table 2.5 vintage pairs, so each TY2023 package continues its twin's source series.
  • Raw custody. I registered the files with chronicle fetch-artifact --upload-r2, which writes storage.r2 only after a successful upload. I then mirrored them to chronicle-raw and read all three back from both buckets, hash-verified. The IRA manifests sit beside the TY2022 ones as manifest_{traditional,roth}_2023_source_package.yaml.
  • Registered in SOURCE_PACKAGE_ALIASES; soi-state-2023 is also in the SOI filer coverage family and docs/pe-calibration-targets.md.

Chronicle governance

  • Approved agent role: ledger-source-ingestor.
  • Deterministic checks run:
  • LLM judge verdicts:
    • ledger-source-fidelity: PASS. An independent Opus 5.5 review via Subfleet, posted as a comment below, re-read every declaration against its TY2022 twin, the layout moves, the manifests, the parser change and the corpus scan. It also hand-checked the cross-file identity (4 of 4).
    • ledger-contract: no schema or consumer-contract change.
    • ledger-boundary: PASS (same review).

Invariants (tested)

  1. Each fact is its publisher cell x scale, with period tax_year 2023, vintage tax_year_2023 and the registered sha.
  2. The state totals equal the HT2 CSV's US row.
  3. Building at 2022 or 2024 yields the same TY2023 facts. The labels are literal, so Refuse pinned-artifact builds that only relabel another year (#117) #292's guard allows the build.
  4. Mirror (differential): apart from year labels, file locations and the moved cell positions, the declarations equal TY2022's.
  5. Parser: without an edge-to-edge merge, the bounds equal openpyxl's (enumerated layouts). With one, only its placeholders stop counting.

Default-bundle snapshot

The re-pin is computed from in-memory builds of these packages plus every package declaring one of their concepts. Relative to #292's tree it adds +8 facts (US, tax_year:2023, tax_unit), +3 packages and 0 new semantic-duplicate keys. The absolute values currently sit on #292's un-re-pinned numbers, so this PR stays draft until #292 re-pins, then gets rebased and re-pinned.

Note: TY2022 IRA restatement

The June 2026 IRA study that published TY2023 also restated TY2022. The live 22in05ira.xlsx reports 5,900,726 traditional contributors, and Chronicle pins the February 2025 release's 5,101,648. For Roth the figures are 9,170,792 live against 10,036,960 pinned. TY2023 compares cleanly only with the restated TY2022. That re-registration (#225 item 2) is a separate follow-up.

🤖 Generated with Claude Code

MaxGhenis and others added 2 commits September 27, 2026 02:00
openpyxl fills every position of a merged range with a MergedCell
placeholder. The IRS's TY2023 23in54us.xlsx merges its footnote row
A179:XFD179, so sheet.max_column became 16,384 and xlsx_used_range would
have built about 70 million empty cells (a probe hit 17.9 GB). Placeholders
of a merged range that runs to the worksheet edge (column XFD or the last
row) no longer widen the bounds; the anchor cell and every other cell still
do, and sheets without such a range keep openpyxl's bounds exactly.

Scanned every committed workbook (212 files, 2,014 sheets): only
23in54us.xlsx changes, from 4264 x 16384 to 4264 x 143, which is the
sheet's own <dimension> (A1:EM4264). A broader rule that ignored every
merged placeholder would have moved 309 sheets and was rejected.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
soi-state-2023 (23in54us.xlsx, 4 facts) and
soi-ira-{traditional,roth}-contributions-2023 (23in05ira.xlsx and
23in06ira.xlsx from the June 2026 IRA study, 2 facts each) mirror their
TY2022 twins with literal TY2023 labels, so the ty2023 ids now carry
TY2023 data. They follow each workbook's moved layout: the state workbook's
geography label (A8 -> B3), column header (row 3 -> 4) and EITC block
(two rows down), and the IRA tables' Estimates sheet with All taxpayers
on row 7. The raw files were registered with fetch-artifact --upload-r2 and
hash-verified in ledger-raw and chronicle-raw.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…match

Addresses the #296 judge review: the max_row edge branch had no test, and
a spot-value check could have passed without checking anything.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@MaxGhenis

Copy link
Copy Markdown
Contributor Author

ledger-source-fidelity: PASS
ledger-boundary: PASS

Recommendation: APPROVE, on three conditions: #292 merges first, the default-bundle snapshot is re-pinned, and a maintainer runs the two checks under "Not verified" before the PR leaves draft. I found no fidelity or boundary defects.

How I checked: I only had file-read, search and web tools here. I could not run git, uv run python or pytest, and I could not open the .xlsx files. So I have not checked the workbook cells, hashes or sizes directly. I say below exactly what I checked and what I didn't.

The review prompt mixes in PR-A items. Several verdict checks describe PR-A, not #296: CSV column letters, the two columns inserted at DP, "interest measures", "four packages" and gen/bundle_delta.py → bundle-delta-pr-a.json. #296 adds 3 packages and 8 facts, all read from xlsx cells, and its snapshot evidence is gen/bundle_delta_pr_b.py → evidence/bundle-delta-pr-b.out. I applied each check where it has a PR-B equivalent. Because the PR has only 8 facts, the requested 15-fact spot check isn't possible.

Findings, by severity

1. Medium, blocks merge (disclosed in the PR body): the default-bundle snapshot is not re-pinned.

2. Low: the only check that doesn't depend on the declared cell addresses is skipped in CI.

  • tests/test_chronicle_soi_state_ira_2023.py:169-171 skips when 23in55cmcsv.csv is missing, which is always the case on this branch.
  • The path is right: db/data/irs_soi/historic_table_2/23in55cmcsv.csv exists in the PR-A worktree.
  • I ran the comparison by hand against that CSV (US row, AGI_STUB 0):
CSV variable CSV value Value in the PR
N1 159,949,000 159,949,000
A00100 15,234,086,106 $15,234,086,106k
N59664 3,155,130 3,155,130
A59664 15,299,042 $15,299,042k
  • 4 of 4 match, and I aligned the positions from the end of the header, so they are exact.
  • The test's scales are right (×1 for counts, ×1000 for amounts), and split(".")[4] picks the right part of the record id.
  • Suggestion: once Add IRS SOI Historic Table 2 TY2023 state AGI-band facts #291 is on main, turn the skip into a hard failure.

3. Low: parser test coverage.

  • The max_row >= _XLSX_MAX_ROW branch (chronicle/sources/cells.py:225) has no test; only the column edge is exercised (tests/test_chronicle_source_cells.py:494).
  • The function reads openpyxl's private sheet._cells (cells.py:238).
  • Neither is a correctness bug.
  • The logic is correct and minimal. When no range reaches the sheet edge, it returns openpyxl's own (max_row, max_column) unchanged. When one does, it takes the maximum over the loaded cells and skips only the MergedCell placeholders inside those edge ranges. The anchor cell is a real cell, so it still counts. That is the same calculation openpyxl uses for max_row/max_column (default 1), minus the placeholders.
  • Corpus scan. evidence/xlsx-bounds-scan.json has exactly one "same": false entry and no load errors. It is 23in54us.xlsx, going from 4264×16384 to 4264×143. Note that the scan read the copy under _reviews/irs/ (passed as an extra file), not the committed fixture.

4. Low: a test assertion could pass without checking anything.

  • The spot-value check at test_chronicle_soi_state_ira_2023.py:160-162 is wrapped in if record_id in facts, so it would pass silently if the ids drifted.
  • For now, the id set-equality check at :149 prevents that.
  • The amount facts aren't in PUBLISHED.

5. Low / informational: nothing ties the file's own contents to the year 2023.

  • Neither the state package nor the IRA packages guard a title cell that says 2023. The year rests on the filename, the manifest and the hash.
  • Their TY2022 twins have the same gap, so this isn't a regression. A title-cell guard would be cheap to add.

6. Informational (disclosed): the series now mixes release vintages.

7. Informational: the snapshot script's candidate filter is a text match.

  • gen/bundle_delta_pr_b.py:23 picks candidate packages by searching for the substring concept: X\n, so it would miss a quoted concept declaration. It also builds candidates at 2023 only.
  • The method for counting new duplicates is correct: dups(base+new) - dups(base).
  • build_semantic_fact_key hashes the canonical measure.concept, which is what the filter matches on.
  • The reported +8 facts, 0 new semantic duplicate keys and 688 MB peak are plausible. I did not re-run the script.

What I checked directly

TY2023 packages against their TY2022 twins. I compared each file line by line and ran gen_state_ira_2023.py's edits in my head. The only differences are:

  • Year text: literal 2023 labels, ids, period, vintage, legal_vintage and artifact_year. There is no {year} anywhere.
  • Kept on purpose: record_set_spec_id.
  • Layout moves:
    • State: the geography guard moved from A8 to B3 ("UNITED STATES "), the header row from 3 to 4, EITC 3+ from rows 144/145 to 146/147, and the guards from rows 136/144 to 138/146.
    • IRA: the sheet is Estimates, the row moved from 8 to 7, and the header reads Number of taxpayers without the line break.
  • Unchanged: scales (1000 on the amounts, none on the counts), concepts, filters, the EITC >= 3 constraint, the geography 0100000US, and the entity and domain.

The EITC two-row shift has independent support. Comparing irs/h22.txt with h23.txt, TY2023 replaces N07260/A07260 with N07262/A07262 and N07265/A07265. That adds one net credit pair before the EITC block, which explains a two-row move. N1, A00100, N59664 and A59664 exist in both years.

Manifests. Paths, source_url, bucket ledger-raw and the key pattern raw/irs_soi/<pkg>/2023/<sha>/<file> match the TY2022 twin's convention. The hashes and sizes match the constants in the test.

Registration. The aliases (chronicle/source_package.py:361, :378-382), the coverage family (chronicle/targets/us_poverty.py:88), the docs row and the fixture counts (tests/test_chronicle_source_package.py:1660-1673) are all consistent.

bundle-state-ira-2023.out. It shows 8 facts, 0 errors and warnings, lineage coverage 1.0, and every fact at tax_year:2023.

Not verified

  • Workbook cells: B3 and the B3:L3 merge, header rows 4 and 5, the header text, A138/A146, B4 = Total, and the "Estimates" sheet name.
  • The four IRA values (5,873,053; $26,901,749k; 9,488,414; $32,809,393k). The only independent sign is that the live 23in05ira.xlsx is about 11.1 KB, which fits the 11,345 bytes in the manifest.
  • File hashes, and the R2 and chronicle-raw read-back.
  • Variable meanings from the doc guides. The .doc files are binary and I couldn't search them.
  • Any test run.

The build still gives good protection here: it checks the row headers, column headers and guards against the bytes, and it fails if any of them doesn't match. Before the PR leaves draft, a maintainer should run:

Boundary

The PR adds three source-package declarations, their manifests and fixture bytes, a parser bounds fix, alias and coverage registrations, and tests. It adds no reconciliation, aging, imputation, support-aware activation or solver construction. us_poverty.py only gains an alias in a coverage list. ledger-boundary: PASS.

@MaxGhenis

Copy link
Copy Markdown
Contributor Author

Main session's response to the judge review above (independent Opus 5.5 via Subfleet; that lane had no shell):

  • Verdicts: ledger-source-fidelity PASS, ledger-boundary PASS.
  • Finding 1 (snapshot): the +8 facts / +3 packages deltas are already in f4a0ad8. The period count moves through expected_period_counts["tax_year:2023"] += 8. What remains is Refuse pinned-artifact builds that only relabel another year (#117) #292's own re-pin; once it lands I'll rebase and re-apply these on top.
  • Finding 2: the cross-file identity was executed against Add IRS SOI Historic Table 2 TY2023 state AGI-band facts #291's CSV and matches 4/4. The skip becomes a hard failure once Add IRS SOI Historic Table 2 TY2023 state AGI-band facts #291 is on main.
  • Findings 3 and 4: fixed in the latest commit.
    • The row-edge branch is now tested with a stub sheet, because a real D4:D1048576 merge would materialize about 1M MergedCells.
    • Spot values must now all match a package id.
  • Finding 5 (no 2023 title guard): the same gap exists in the TY2022 twins, so it is not a regression. It is left as is so the TY2023 packages mirror them.
  • Items the reviewer could not verify: I ran these locally.
    • The workbook cells are checked by test_each_fact_is_its_publisher_cell (openpyxl read-only), the layout guards by the source-suite builds, and the hashes by the manifest tests. The R2 read-backs from both buckets were hash-verified. 94 of 95 targeted tests pass; the one skip is the cross-file test.

🤖 Generated with Claude Code

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant