Skip to content

Add IRS SOI Historic Table 2 TY2023 national, state broad and state EITC facts - #295

Open
MaxGhenis wants to merge 3 commits into
soi-ty2023-state-agi-bandsfrom
soi-ty2023-ht2-national-broad-eitc
Open

MaxGhenis wants to merge 3 commits into
soi-ty2023-state-agi-bandsfrom
soi-ty2023-ht2-national-broad-eitc

Conversation

@MaxGhenis

@MaxGhenis MaxGhenis commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Adds IRS SOI Historic Table 2 for tax year 2023 at the national and state level, mirroring the TY2022 packages:

Package Facts TY2022 twin
soi-historic-table-2-2023 (US, 11 AGI stubs x 55 measures) 605 soi-historic-table-2
soi-historic-table-2-state-broad-2023 (51 states x 53 measures) 2,703 soi-historic-table-2-state-broad-2022
soi-historic-table-2-state-eitc-2023 (51 states x 10 measures) 510 soi-historic-table-2-state-eitc-2022
soi-historic-table-2-state-agi-2023: adds taxable interest (N00300, A00300) +1,020 the TY2022 state AGI package's interest measures

Every fact is a single published cell of 23in55cmcsv.csv (sha d1f7c890…, registered by #291). Every label is literal TY2023: period, record ids, vintage and legal_vintage. The packages are therefore already clean under the artifact-year restamp guard in #292. This addresses #117 item 3 ("Ingest HT2 TY2023"), and it adds the state interest measures #291 deferred until the TY2023 national table was packaged.

Stacked on #291 (base soi-ty2023-state-agi-bands). Merge after #291, and preferably after #292 (see "Semantic duplicates").

What changes

  • Column shift. The TY2023 CSV has 165 columns, two more than TY2022. N07262/A07262/N07265/A07265 replace N07260/A07260 at DP, so every column from DP on sits two places right. This moves 8 national/broad measures (premium tax credit, EITC, ACTC and income tax) and all 10 EITC measures. Each measure is remapped by its CSV variable name. Every measure now guards its header (source_column_id plus expected_column_header on row 1), so a future shift fails the build instead of reading a neighbouring variable. The TY2022 national package guarded only some measures.
  • Variable semantics. 23incmdocguide.doc and 22incmdocguide.doc describe every variable these packages read with the same text, except N11070/A11070. That pair was renamed from "Refundable child tax credit or additional child tax credit" to "Additional child tax credit", which matches the existing irs_soi.additional_child_tax_credit concept.
  • Generation. The YAML is a deterministic line-level transform of each TY2022 file, so git diff --no-index <2022> <2023> shows only year labels, column letters and added guards. An independent dict-level derivation of the same mapping must agree with it, and it does. The same equivalence is tested in test_packages_mirror_their_2022_twins.
  • QBI correction (from the judge review). The TY2022 national and state broad packages read the QBI measures from N03270/A03270. Both IRS guides define that pair as the self-employment health insurance deduction. The TY2023 packages read N04475/A04475 (CT/CU), the qualified business income deduction: US all returns are 26,391,030 returns and $213.7B, against the 3.58M and $31.1B the inherited binding would give. DOCUMENTED_VARIABLE pins all 63 measures to their documented variables. The TY2022 packages get a separate fix.
  • Series continuity. record_set_spec_id, source_table, concepts, units, scales, rows, filters and constraints are the TY2022 values, so the two vintages form one source series.
  • Registered in SOURCE_PACKAGE_ALIASES, in the irs_soi_filer_income_tax_credits coverage family and in docs/pe-calibration-targets.md.

Chronicle governance

  • Approved agent role: ledger-source-ingestor.
  • Deterministic checks run:
    • chronicle validate-package <id> --year 2023: all four packages are valid. Counts are 605, 2,703, 510 and 2,040 source records.
    • chronicle build-bundle --year 2023 over the four packages is valid: 5,858 facts, 0 errors, 0 warnings, 0 duplicate keys. Lineage coverage is 1.0 for each package, with 0 agent-acceptance errors. The only acceptance warnings are the standard concept_alignment_validation_skipped / no_concept_alignments.
    • Independent oracle: I parsed every exported consumer fact against the CSV using the standard library only. 5,858 of 5,858 equal cell x scale, with geography, stub, variable, period, vintage and sha all checked. Every YAML column letter heads its declared variable.
    • The targeted tests pass: the new module, Add IRS SOI Historic Table 2 TY2023 state AGI-band facts #291's module, fixture counts, alias drift, us_poverty coverage and governance (45 passed). ruff check passes.
  • LLM judge verdicts:
    • ledger-source-fidelity: PASS after re-review. The first independent Opus 5.5 review (Subfleet) caught the QBI mis-binding, which is fixed in 73cb3b9. The re-review checked 38 documented variables and the CT/CU binding and approved. Both reviews are posted below.
    • ledger-contract: no schema or consumer-contract change.
    • ledger-boundary: PASS (independent Opus 5.5 review via Subfleet).

Invariants (tested in tests/test_chronicle_soi_ht2_2023.py and tests/test_chronicle_soi_state_agi_2023.py)

  1. One cell per fact. value == the named CSV cell x scale, with one source row, and the fact's geography, stub and variable name that cell.
  2. Header binding. Every measure's column letter heads its source_column_id in the TY2023 header, and the package guards it.
  3. Mirror (differential). With year labels removed and each column replaced by the variable it heads in its own year's file, the TY2023 declarations equal the TY2022 ones.
  4. Pinned vintage. Building at 2022 or 2024 yields the same TY2023 facts.
  5. Accounting identities on the facts.
    • National AGI bands sum to the all-returns row, dollars exactly.
    • The 50 states and DC plus the file's OA and PR rows sum to the national total, for every shared broad and EITC measure; dollars match exactly.
    • EITC by 0/1/2/3+ children partitions each state's EITC total.
    • State interest bands sum to each state's total.
    • Counts, which the IRS rounds to tens, meet each identity within their rounding.

Default-bundle snapshot and semantic duplicates

A full build-bundle --year 2023 needs more disk than the build machine had free. The re-pin in tests/test_chronicle_bundle.py instead comes from in-memory builds of the new packages plus every package that declares one of their concepts:

CI confirms the re-pin. Whichever of this PR and #292 merges second re-pins the snapshot.

Consumers

Facts are inert until Microcosm re-pins its Chronicle feed. That re-pin is queued separately as a decision for Max.

🤖 Generated with Claude Code

…ITC facts

Three packages mirror their TY2022 twins on 23in55cmcsv.csv, with literal
TY2023 labels: soi-historic-table-2-2023 (605 facts),
soi-historic-table-2-state-broad-2023 (2,703) and
soi-historic-table-2-state-eitc-2023 (510). The TY2023 file inserts two
columns at DP, so every column from DP on moves two places right; each
measure is remapped by its CSV variable and guards that header.

The TY2023 state AGI package also gains the TY2022 package's taxable
interest measures (N00300, A00300; 1,020 facts), which chronicle#291 left
for when the TY2023 national table was packaged.

Tests read the publisher CSV with the standard library: every fact is one
cell, the declarations equal the TY2022 twins' up to year labels and column
moves, and the published accounting identities hold on the facts (AGI bands
and state totals add up, EITC child counts partition the total).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…023 HT2

The TY2022 national and state broad packages bind qbi_claims/qbi_amount to
N03270/A03270, which both IRS documentation guides define as the
self-employment health insurance deduction. The TY2023 packages now read
N04475/A04475 (columns CT/CU), the qualified business income deduction:
US all returns 26,391,030 returns and $213,733,168,000 instead of
3,578,530 and $31,114,668,000.

Tests now also check each measure against the variable the documentation
names (not only against the variable its package names), and the mirror
test states the QBI correction explicitly. Found by the #295 judge review.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@MaxGhenis

Copy link
Copy Markdown
Contributor Author

ledger-source-fidelity: FAIL
ledger-boundary: PASS

  1. P1 — 124 new QBI facts report the wrong deduction. National package:865 and state-broad package:695 bind qbi_claims/qbi_amount to N03270/A03270 at BF/BG. Both TY2022 documentation and TY2023 documentation, section G, identify these as self-employment health insurance deductions. QBI uses N04475/A04475, at CT/CU.

    In CSV row 2, the PR emits 3,578,530 returns and $31,114,668,000 as national QBI; the published QBI cells contain 26,391,030 returns and $213,733,168,000. This affects 22 national and 102 state facts. Although inherited from TY2022, these are newly published, incorrectly labeled TY2023 facts. Correct the variables, columns and header guards.

  2. P2 — The tests preserve the inherited semantic error. test_chronicle_soi_ht2_2023.py:164 obtains the expected variable from the fact itself, so it verifies extraction without checking whether the variable represents the named measure. The mirror test:199 additionally requires the erroneous TY2022 mapping. All 16 selected tests in this module passed with the QBI error present. Add documentation-backed measure-to-variable expectations and account explicitly for the corrected QBI mapping in the mirror test.

Overall recommendation: REQUEST_CHANGES.

I independently rebuilt all four packages and checked 5,858 facts, including the 4,838 added by this PR, against the local publisher CSV. Every value equals its declared cell times the correct scale; that mechanical agreement does not resolve the QBI semantic error. I verified the CSV’s SHA-256 (d1f7c8901fcefb2c46f5dd14f22715b7c8513794ce7b9ea6a6a947f1c548668f) and 757,078-byte size.

Direct checks covered all column/header bindings, selected STATE/AGI_STUB rows, state FIPS, source-row and cell lineage, filters, AGI bounds, EITC child constraints, and literal TY2023 identifiers and vintages. The column shift is handled correctly. The documentation comparison covered all 63 selected variables, including the ACTC rename; the workbook confirms the EITC “three or more” category. All 102 added interest measure declarations match TY2022 apart from the added numeric/header guards.

I made 19 explicit spot-checks; 16 are shown below. Amount variables (A…) are multiplied by 1,000; count variables (N…) are unscaled. Addresses refer to the original CSV.

Package STATE / stub Variable / cell Published cell Emitted value
National US / 0 N1 / C2 159,949,000 159,949,000
National US / 0 A00100 / T2 15,234,086,106 $15,234,086,106,000
National US / 0 A59660 / EC2 65,006,308 $65,006,308,000
National US / 0 A85770 / DW2 59,881,173 $59,881,173,000
State broad CA / 0 A00200 / X57 1,414,115,351 $1,414,115,351,000
State broad NY / 0 A00300 / Z365 33,449,796 $33,449,796,000
State broad TX / 0 A11070 / EO486 4,187,936 $4,187,936,000
State broad WY / 0 A06500 / EU563 4,496,701 $4,496,701,000
State EITC AL / 0 N59661 / ED13 89,230 89,230
State EITC CA / 0 A59662 / EG57 2,326,336 $2,326,336,000
State EITC TX / 0 N59663 / EH486 677,300 677,300
State EITC DC / 0 A59664 / EK101 25,543 $25,543,000
State AGI CA / 1 A00300 / Z58 751,857 $751,857,000
State AGI CO / 10 A00300 / Z78 2,574,931 $2,574,931,000
State AGI NY / 9 N00300 / Y374 119,160 119,160
State AGI TX / 4 A00300 / Z490 623,280 $623,280,000

The boundary verdict is PASS: the diff adds publisher-cell declarations and registrations, without reconciliation, aging, imputation, support-aware activation or solver construction. Builds at alternative years retain TY2023 facts.

I found no arithmetic or candidate-selection problem in the bundle delta. Independently scanning registered packages and sequentially rebuilding all 17 relevant packages reproduced +4,838 facts, +605 US facts, +83 California facts and +1,650 semantic-duplicate keys. The latter comprise 1,553 congressional-district overlaps and 97 Publication 1304 overlaps. These support the snapshot:139 totals of 356,751 facts and 2,117 duplicate keys, conditional on #291’s baseline.

Validation limits: 24 selected tests passed, and all four package and consumer-contract validations passed. I excluded four artifact-writing tests and did not run the full default bundle, independently confirm its unchanged warning count, verify remote CI, or fetch fresh IRS/R2 copies. Commands ran sequentially; peak measured RSS was 693 MiB. The checkout remains clean at 553ab64.

@MaxGhenis

Copy link
Copy Markdown
Contributor Author

Main session's response to the judge review above (independent Opus 5.5 via Subfleet):

  • P1 (QBI mis-binding): confirmed against both IRS documentation guides. N03270/A03270 is the self-employment health insurance deduction; the QBI deduction is N04475/A04475. Fixed in 73cb3b9, where the TY2023 national and state broad QBI measures read CT/CU with header guards.
    • US all returns are now 26,391,030 returns and $213,733,168,000, where the inherited binding gave 3,578,530 and $31,114,668,000.
    • The TY2022 packages carry the same bug and get a separate fix; CD 2022 already reads N04475.
  • P2 (tests blind to semantics): DOCUMENTED_VARIABLE now pins all 63 measures to the variable the documentation names, via test_every_measure_reads_its_documented_variable. The mirror test states the QBI correction explicitly.
  • Re-checked after the fix: 46 tests pass, the four-package bundle is valid, and the oracle finds 0 errors in 5,858 facts. The snapshot is unchanged, because semantic keys depend on the concept, not the column.
  • A focused re-review is running.

🤖 Generated with Claude Code

…exactly

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@MaxGhenis

Copy link
Copy Markdown
Contributor Author

ledger-source-fidelity: PASS
ledger-boundary: PASS

Recommendation: APPROVE. The P1 and P2 findings are fixed. I could not run the tests or git show myself, so please run step 3 before merging.

What I couldn't do: this session had only file-read and search tools, with no shell. So I did not run git show 73cb3b9, and I did not run step 3 (uv run pytest -q tests/test_chronicle_soi_ht2_2023.py tests/test_chronicle_soi_state_agi_2023.py). Everything below comes from reading the checkout at its current state, the TY2023 CSV and the documentation guide. I made no changes.

1. P1 (QBI reading the wrong variable): fixed

  • National package: packages/irs_soi/historic_table_2_2023/source_package.yaml:862-884 now reads qbi_claims from column CT (N04475) and qbi_amount from column CU (A04475, scale 1000). Both have expected_column_header_row: 1 and expected_column_header set to N04475 / A04475.
  • State-broad package: packages/irs_soi/historic_table_2_state_broad_2023/source_package.yaml:692-714 has the same bindings and header guards.
  • No leftovers: no TY2023 package still refers to 03270. It survives only in the TY2022 twins, which are unchanged.
  • CSV header (irs/h23.txt): field 98 is N04475 and field 99 is A04475, which are columns CT and CU. Field 58 is N03270, column BF, the old wrong binding. A regex that skips 97 fields finds N04475,A04475, at that position in the CSV header.
  • Documentation guide (23incmdocguide.doc):
    • N04475 is "Number of returns with Qualified business income deduction" (1040:13), and A04475 is "Qualified business income deduction amount".
    • N03270 is "…returns with self-employment health insurance deduction" (Schedule 1:17).
  • US all-returns values: "26,391,030","213,733,168" sit at fields 98–99 of CSV line 2, and both numbers appear only on that line. Scaled, that is 26,391,030 returns and $213,733,168,000. The test pins both at tests/test_chronicle_soi_ht2_2023.py:232-233.

2. P2 (tests couldn't catch a wrong binding): fixed

  • DOCUMENTED_VARIABLE (lines 72–136) has 63 entries. That equals 55 national measures plus the 8 EITC-by-children measures. The state-broad package has 53 measures (2,703 facts ÷ 51) and the state EITC package has 10 (510 ÷ 51), both inside that set.

  • Spot-check: I checked 38 variables against the guide and all matched their measure:

    • N1, N2, A00100, N02650
    • N/A00200, N00300, N00400, N00600, N00650, N/A00900, A01000, A01400, A01700
    • N/A02300, N/A02500, N/A26270, N/A25870
    • N/A04470, N/A17000, N/A18500, N/A18460, N/A04475, N/A04800, N/A05800, N/A06500
    • N/A07225, N/A11070, N/A85770, N/A59660, N/A59661–N/A59664

    None is bound to the wrong variable.

  • test_every_measure_reads_its_documented_variable (line 279) is sound. The expected variable now comes from a list written from the guide, not from the fact itself. The test at line 261 ties each column's header to its variable name, and the test at line 237 ties each fact's value to that CSV cell. Together they check every step from measure to CSV value.

  • TY2022_QBI_CORRECTION (lines 140–143, used at 308–311) is sound. The mirror test asserts that TY2022's column holds the wrong variable (N03270/A03270) before substituting the correct one. If the TY2022 twin is fixed later, this test will fail and the correction must be removed. That is good behaviour.

Two non-blocking notes (P3), both inherited from TY2022. Neither is a wrong binding:

  • N/A07225: labelled "child tax credit", but the guide defines it as the nonrefundable child and other dependent tax credit.
  • N/A06500: labelled "Total income tax", but the guide calls it "income tax after credits" (1040:22). The same file has A10300 as "Total tax liability amount", so the label could be confused with that column.

A small optional hardening: assert that the union of measure IDs across the three packages equals DOCUMENTED_VARIABLE's keys, so stale entries can't linger.

3. Tests: not run

See above: step 3 is the one outstanding check.

4. Snapshot counts: unaffected

tests/test_chronicle_bundle.py:139,142,153,185,1231 still pins 356,751 facts and 2,117 semantic duplicate keys. The fix swaps one column for another, so the number of facts is the same. The concepts are unchanged (irs_soi.returns_with_qualified_business_income_deduction and irs_soi.qualified_business_income_deduction), and semantic keys depend on concept rather than column. So both counts are unaffected, which matches the reasoning in the request.

Boundary

The diff still only declares publisher cells and their header guards. It adds no reconciliation, aging, imputation or solver logic.

APPROVE, once step 3 passes.

@MaxGhenis

Copy link
Copy Markdown
Contributor Author

Heads-up: this PR's two TY2023 capital-gains measures need a relabel before it merges (chronicle#304).

historic_table_2_2023 and historic_table_2_state_broad_2023 clone their TY2022 twins line by line, so N01000/A01000 carry "Returns with taxable net capital gains" / "Taxable net capital gains". They also carry Table 1.4's Schedule D concept, irs_soi.returns_with_taxable_net_capital_gains / irs_soi.taxable_net_capital_gains.

The TY2023 IRS guide (23incmdocguide.doc, lines 271–278) defines them as "Number of returns with net capital gain (less loss)" / "Net capital gain (less loss) amount" at Form 1040 line 7. That line covers Schedule D gains, limited Schedule D losses, and distributions-only returns. For TY2023, US N01000 is 29,481,840, against Table 1.4's Schedule D gain count of 12,392,020.

#304 relabels TY2022 Historic Table 2 and the congressional-district file to irs_soi.returns_with_form_1040_capital_gain_or_loss / irs_soi.form_1040_capital_gain_or_loss. It also adds a guard test that covers every vintage. Whichever of #304 and this PR merges second, CI blocks the old label. I checked this on this PR's head f55ae34 with #304 cherry-picked:

  • Without the patch: Label IRS SOI N01000/A01000 as Form 1040 line 7 capital gain or loss #304's test_every_line_7_column_carries_the_line_7_concept fails on historic_table_2_2023:net_capital_gains_returns, and this PR's test_packages_mirror_their_2022_twins fails for the national and state-broad packages (3 failed; the EITC mirror, which has no capital-gains measures, passes).
  • With the patch: tests/test_chronicle_soi_ht2_2023.py and tests/test_chronicle_soi_capital_gain_concepts.py pass in full (29 passed), mirror tests included.

The patch (8 lines in 2 files; git apply it after rebasing onto a main that has #304):

diff --git a/packages/irs_soi/historic_table_2_2023/source_package.yaml b/packages/irs_soi/historic_table_2_2023/source_package.yaml
index 2d4d1ca..c3f8ddb 100644
--- a/packages/irs_soi/historic_table_2_2023/source_package.yaml
+++ b/packages/irs_soi/historic_table_2_2023/source_package.yaml
@@ -611,11 +611,11 @@ record_sets:
         value_scale: 1000
 
       - measure_id: net_capital_gains_returns
-        label: Returns with taxable net capital gains
+        label: Returns with net capital gain (less loss)
         ordinal: 27
         column: AK
         source_column_id: N01000
-        concept: irs_soi.returns_with_taxable_net_capital_gains
+        concept: irs_soi.returns_with_form_1040_capital_gain_or_loss
         unit: count
         aggregation: sum
         expected_cell_type: number
@@ -623,11 +623,11 @@ record_sets:
         expected_column_header: N01000
 
       - measure_id: net_capital_gains_amount
-        label: Taxable net capital gains
+        label: Net capital gain (less loss)
         ordinal: 28
         column: AL
         source_column_id: A01000
-        concept: irs_soi.taxable_net_capital_gains
+        concept: irs_soi.form_1040_capital_gain_or_loss
         unit: usd
         aggregation: sum
         expected_cell_type: number
diff --git a/packages/irs_soi/historic_table_2_state_broad_2023/source_package.yaml b/packages/irs_soi/historic_table_2_state_broad_2023/source_package.yaml
index 912cb84..f67577a 100644
--- a/packages/irs_soi/historic_table_2_state_broad_2023/source_package.yaml
+++ b/packages/irs_soi/historic_table_2_state_broad_2023/source_package.yaml
@@ -368,22 +368,22 @@ record_sets:
     expected_column_header: A00900
     value_scale: 1000
   - measure_id: net_capital_gains_returns
-    label: Returns with taxable net capital gains
+    label: Returns with net capital gain (less loss)
     ordinal: 17
     column: AK
     source_column_id: N01000
-    concept: irs_soi.returns_with_taxable_net_capital_gains
+    concept: irs_soi.returns_with_form_1040_capital_gain_or_loss
     unit: count
     aggregation: sum
     expected_cell_type: number
     expected_column_header_row: 1
     expected_column_header: N01000
   - measure_id: net_capital_gains_amount
-    label: Taxable net capital gains
+    label: Net capital gain (less loss)
     ordinal: 18
     column: AL
     source_column_id: A01000
-    concept: irs_soi.taxable_net_capital_gains
+    concept: irs_soi.form_1040_capital_gain_or_loss
     unit: usd
     aggregation: sum
     expected_cell_type: number

Re-running gen_ht2_2023.py on a tree with #304 produces the same 8 lines, because it copies each measure's label and concept from the TY2022 twin.

Snapshot. Once relabelled, the TY2023 HT2 US capital-gains rows stop sharing a semantic key with Table 1.4's TY2023 row. That should take 2 off this PR's Table 1.4 overlap count when you re-pin. I haven't verified that number; CI will confirm it.

@MaxGhenis

Copy link
Copy Markdown
Contributor Author

chronicle#304 is merged to main (ce43336). When this PR is rebased onto main (after #291 lands), apply the 8-line TY2023 relabel from the comment above. tests/test_chronicle_soi_capital_gain_concepts.py and this PR's own mirror test will fail until it's applied.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant