diff --git a/changelog.d/us-restamped-source-vintage-aging.fixed.md b/changelog.d/us-restamped-source-vintage-aging.fixed.md new file mode 100644 index 000000000..da28f71c5 --- /dev/null +++ b/changelog.d/us-restamped-source-vintage-aging.fixed.md @@ -0,0 +1 @@ +Read restamped Chronicle facts at their data year. The pinned US feed stamps five packages' pinned files as ty2023: 26,893 facts from the TY2020 W-2 table and the TY2022 congressional-district (PolicyEngine/chronicle#117), state and IRA tables. The compile detects a restamp either by its registered file digest or by a `raw_r2_key` year earlier than its period. It refuses the compile unless a reviewed `US_RESTAMPED_SOURCE_PACKAGES` entry matches. If one matches, aging starts from the data year. It also refuses where a restamp would outrank newer truthful data in latest-vintage selection, or would sit in an aging growth index. On the pinned feed the `national_state` W-2 Box 7 tips amount target moves from $28.28B, aged from 2023, to $34.29B, chained from TY2020. The surface keeps 5,694 targets, and its registry moves from `d315c75804ef` to `884bc45ef335`. On the full surface, 13,176 congressional-district-file dollar targets age from TY2022 under the current series policy: the measures on the AGI series move +3.05% and net capital gains -23.9%. diff --git a/docs/us-chronicle-feed-repin.md b/docs/us-chronicle-feed-repin.md index 14aee6d60..8d3a28ee0 100644 --- a/docs/us-chronicle-feed-repin.md +++ b/docs/us-chronicle-feed-repin.md @@ -181,6 +181,126 @@ and target id, and the register entry carries a reviewed `state_name` (`_geography_fallback_label` is UK-only, so the label cannot be derived from the FIPS code). +## Restamped packages + +Five Chronicle packages in this feed restamp their data: 26,893 observation +facts carry a later period than their publisher file describes. A Chronicle +package can pin one file with `artifact.artifact_year` and still render +`{year}` into its period, record ids and vintage from the build year. The +scope builds these five at 2023, so each emits its pinned file as `ty2023`. +PolicyEngine/chronicle#117 reported the congressional-district case in July. +microcosm#1030 found the other four, and PolicyEngine/chronicle#292 fixes the +stamp at source. + +| Package | File | Data year | Facts stamped ty2023 | +|---|---|---|---| +| `soi-congressional-district-2022` | `22incd.csv` | TY2022 | 26,880 | +| `soi-w2-statistics-2020` | `20in04w2all.xlsx` | TY2020 | 5 | +| `soi-state-2022` | `22in54us.xlsx` | TY2022 | 4 | +| `soi-ira-roth-contributions-2022` | `22in06ira.xlsx` | TY2022 | 2 | +| `soi-ira-traditional-contributions-2022` | `22in05ira.xlsx` | TY2022 | 2 | + +The data years come from the files themselves: the title cells ("Tax Year +2020", "Tax Year 2022"), cell-for-cell equality with the ty2020 W-2 twins, +and the TY2022 county file for the congressional-district data. + +**When it started.** The previous pin already carried 26,891 of these facts: +- every congressional-district, state and IRA row; +- the ty2023 W-2 taxpayer count, 401(k) and Roth 401(k) rows. + +The 2026-09-18 re-pin added the other two: the ty2023 tips amount and return +count. These are the "W-2 Social Security tips for 2023 (2)" cells listed +under [What moved and what did not](#what-moved-and-what-did-not). Before +them, the only tips amount was the ty2020 row, which would have aged on the +chained SOI wages bridge from 2020. The previous pin cannot compile on main, +so that is read from the code, not observed. The restamped row won +latest-vintage selection and aged on the direct CBO ratio from 2023, +reaching $28.28B at 2024 instead of the $34.29B it gets aged from its data +year. + +**What the compile does** (`us_runtime/source_vintage.py`). Two rules detect a +restamp: +- the fact's `source.source_sha256` is a registered file and its period is + after that file's data year; +- its `source.raw_r2_key` names an earlier year than its period. + +Each detected fact must match a reviewed `US_RESTAMPED_SOURCE_PACKAGES` entry, +or the compile refuses. For matching specs, the readings that aging and the +period contract use move to the data year: +- `source_period`; +- `uprating_to_period`, for specs rebased onto a restamped control. + +The stamp is kept in `source_vintage_stamped_period`. Values and target names +do not change. Two readings still see the stamp, because moving them would +change which fact wins: +- **Latest-vintage selection.** The compile refuses where a restamp would + outrank a truthful fact dated after its data year. +- **The choice of a rebase control.** + +A restamped fact inside an aging growth index is refused outright. + +The key-year rule is exact for single-year IRS files. A multi-year release +keyed by its publication year can carry later columns: BEA's 2024-keyed +`SAINC.zip` has a 2025 column. A truthful build of such a column needs a +`US_LATER_PERIOD_OBSERVATION_EXEMPTIONS` entry. + +**On this feed**, compiled as the release does: + +- **`national_state`** keeps 5,694 targets, and its registry moves from + `d315c75804ef` to `884bc45ef335`. + - One value moves: the W-2 Box 7 tips amount, from $28,280,884,269 to + $34,287,530,779 (+21.2%). + - 53 specs change metadata only: 51 Historic Table 2 net-capital-gains + return counts rebased onto the congressional-district US row, and the two + `state_2022` counts. Counts do not age. + - The 51 rebased counts now read `uprating_from_period` and + `uprating_to_period` 2022: a same-year rescale of TY2022 shares onto the + TY2022 CD file's total. +- **Full surface.** 13,176 congressional-district-file dollar targets age from + TY2022 instead of 2023. + - The 23 measures that age on the CBO AGI series (AGI, income tax, the EITC + amounts, taxable interest, dividends, pensions, SALT and the rest) move + +3.05%. + - Net capital gains move -23.9%, because the SOI Table 1.4 chain records + their TY2022 to TY2023 fall. + - Qualified dividends (+8.4%) and Schedule C and partnership income (-0.5%) + move mostly because their series changes: with no SOI bridge for them, the + aging model falls back to the AGI chain from 2022. + - These are changes relative to aging from the data year under the current + series policy, not realized growth. Taxable interest rose ×2.35 from + TY2022 to TY2023 in Table 1.4, far above the AGI link (#117). +- **Metadata only.** + - 13,614 count specs change `source_period`: 13,612 CD-file counts plus the + two `state_2022` counts above. + - 9 more rebased counts are CD-classified and full-surface only. +- **No target.** The 401(k), Roth 401(k), IRA, `state_2022` AGI and EITC + amount rows compile none. + +`test_pinned_feed_restamp_register_matches_the_feed` checks the register +against this feed in both directions. It runs where the feed is present and +skips in CI. A re-pin on a Chronicle commit that stamps truthfully leaves +entries that match nothing, and the test fails until they are deleted. At that +re-pin the scope moves these pairs to their data years (`ty2022`, and `ty2020` +for W-2). IRS has also published TY2023 Historic Table 2 and IRA tables, which +would give true `ty2023` facts. + +Two things need a decision before that re-pin: + +- **The Historic Table 2 capital-gains returns control.** The CD US row wins a + same-period tie on feed order. Stamped truthfully, it would lose to Table + 1.4 ty2023, which counts returns with a taxable net gain (a different + population), and the 51 `national_state` counts would fall 58.5%. Today the + 51 state counts sum to 2.39× the national Table 1.4 target of the same + compiled variable. +- **`SOI_CONGRESSIONAL_DISTRICT_RECORD_SET_ID`**, which names the `ty2023` + record set. + +One mis-stamp runs the other way and is out of reach of both rules. +`bea-regional-state-personal-income-components-2024` reads the 2024 column +under a `cy2023` stamp (416 facts). No target compiles from it, but +`bea_regional.state_wages_salaries` is a deferred parity family. Activating it +before chronicle#292 would calibrate 2024 wages as cy2023. + ## The consumer artifact is refused at this commit `chronicle build-consumer-artifact` validates every row against diff --git a/packages/microcosm-build/src/microcosm/build/us_runtime/fiscal_targets.py b/packages/microcosm-build/src/microcosm/build/us_runtime/fiscal_targets.py index 1cf1d0ec5..6bf86c309 100644 --- a/packages/microcosm-build/src/microcosm/build/us_runtime/fiscal_targets.py +++ b/packages/microcosm-build/src/microcosm/build/us_runtime/fiscal_targets.py @@ -28,6 +28,13 @@ from microcosm.build.us_runtime.congressional_district_vintage import ( translate_congressional_district_facts_to_current_vintage, ) +from microcosm.build.us_runtime.source_vintage import ( + RestampShadowsNewerVintageError, + SourceVintageCorrection, + apply_source_vintage_corrections, + check_restamps_stay_out_of_aging_indexes, + source_vintage_corrections, +) from microcosm.build.us_runtime.target_aging import ( age_us_dollar_targets, enforce_period_contract, @@ -1015,10 +1022,17 @@ def compile_us_fiscal_target_registry( materialized_facts, congressional_district_vintage_crosswalk, ) + # Restamped Chronicle facts (a pinned artifact labelled with a later + # build year, PolicyEngine/chronicle#117). Refuses a restamp that is not + # in the reviewed register, or one an aging index would read at its stamp. + restamps = source_vintage_corrections(materialized_facts) + if age_targets: + check_restamps_stay_out_of_aging_indexes(materialized_facts, restamps) references = ( *_dynamic_us_fiscal_target_references( materialized_facts, target_period=target_period, + restamps=restamps, ), *_references_for_target_period( US_JCT_TAX_EXPENDITURE_TARGET_REFERENCES, @@ -1051,6 +1065,10 @@ def compile_us_fiscal_target_registry( "target_period": target_period, }, ) + # From here on restamped facts are read at their data year: aging starts + # from it and the period contract checks it. Selection and the rebase + # controls above still saw the stamp. + registry = apply_source_vintage_corrections(registry, restamps) if age_targets: # Final nominal transform: age dollar amounts from their source period # to the build period on the fully within-surface-aligned registry @@ -2290,8 +2308,13 @@ def _dynamic_us_fiscal_target_references( facts: tuple[object, ...], *, target_period: int | str, + restamps: Mapping[str, SourceVintageCorrection] | None = None, ) -> tuple[LedgerTargetReference, ...]: - selected = _latest_dynamic_target_references(facts, target_period=target_period) + selected = _latest_dynamic_target_references( + facts, + target_period=target_period, + restamps=restamps, + ) _check_exclusion_vintage_scope(source_record_id for source_record_id, _ in selected) return tuple(reference for _, reference in selected) @@ -2300,11 +2323,17 @@ def _latest_dynamic_target_references( facts: Iterable[object], *, target_period: int | str, + restamps: Mapping[str, SourceVintageCorrection] | None = None, ) -> tuple[tuple[str, LedgerTargetReference], ...]: """Select one fact per model target shape: the latest eligible period. Returns ``(source_record_id, reference)`` pairs so the exclusion vintage guard and receipt can see which fact won each key. + + Selection compares stamped periods. With ``restamps`` (the compile's + :func:`source_vintage_corrections`), it refuses a key where a restamped + fact competes with a truthful fact dated after the restamp's data year: + the stamp would pick the older data (microcosm#1030 review). """ candidates: list[ @@ -2331,6 +2360,12 @@ def _latest_dynamic_target_references( reference, ) ) + if restamps: + _refuse_restamps_that_shadow_newer_vintages( + candidates, + restamps, + target_period_key=_period_key_from_value(target_period), + ) keys_with_positive_observations = { key for key, _, value, _, _ in candidates if value > 0 } @@ -2364,6 +2399,51 @@ def _latest_dynamic_target_references( ) +def _refuse_restamps_that_shadow_newer_vintages( + candidates: Iterable[ + tuple[tuple[str, ...], tuple[int, int, str], float, str, object] + ], + restamps: Mapping[str, SourceVintageCorrection], + *, + target_period_key: tuple[int, int, str], +) -> None: + """Refuse a restamped candidate that would outrank newer truthful data. + + A TY2020 cell stamped 2023 beats a truthful TY2021 fact of the same + target shape on the stamp alone, and the newer data would drop silently. + A truthful fact dated from the year after the restamp's data year up to + its stamp is shadowed; only candidates eligible at the target period + count. None exists on the pinned feed: the ty2020 W-2 twins sit at the + data year, not after it. + """ + + eligible = [ + candidate + for candidate in candidates + if _not_after_target_period(candidate[1], target_period_key) + ] + restamped_keys: dict[tuple[str, ...], list[SourceVintageCorrection]] = {} + for key, _, _, source_record_id, _ in eligible: + restamp = restamps.get(source_record_id) + if restamp is not None: + restamped_keys.setdefault(key, []).append(restamp) + if not restamped_keys: + return + conflicts: list[str] = [] + for key, period_key, _, source_record_id, _ in eligible: + if source_record_id in restamps or not period_key[0]: + continue + year = period_key[1] // 100 + for restamp in restamped_keys.get(key, ()): + if restamp.data_year < year <= restamp.stamped_year: + conflicts.append( + f"{restamp.source_record_id} ({restamp.label}) would outrank " + f"{source_record_id} ({year})" + ) + if conflicts: + raise RestampShadowsNewerVintageError(tuple(sorted(conflicts))) + + def _exclusion_vintage_bypasses( source_record_ids: Iterable[str], ) -> dict[str, tuple[str, ...]]: @@ -3410,9 +3490,11 @@ def _references_for_target_period( def _soi_target_role(fact: object, measure_id: str) -> str: # W-2 item facts (generic "amount" measure id, layout-routed via the # form_w2_item override) get a named role so target aging can pin them - # to the wages series: tips are a W-2 wage component, and the feed's - # TY2020 vintage needs the SOI wages actuals as its chain bridge into - # the CBO projection years (microcosm#451 item 3). + # to the wages series: tips are a W-2 wage component, and the data are + # TY2020 (the newest IRS W-2 table), so the SOI wages actuals are the + # chain bridge into the CBO projection years (microcosm#451 item 3). + # The feed's ty2023 rows restamp the same TY2020 cells; source_vintage + # reads them at 2020 before aging. if ( measure_id == "amount" and _str_at(fact, "layout", "groupby_dimension") diff --git a/packages/microcosm-build/src/microcosm/build/us_runtime/source_vintage.py b/packages/microcosm-build/src/microcosm/build/us_runtime/source_vintage.py new file mode 100644 index 000000000..4d07c82eb --- /dev/null +++ b/packages/microcosm-build/src/microcosm/build/us_runtime/source_vintage.py @@ -0,0 +1,574 @@ +"""Read restamped Chronicle facts at the year their data describes. + +A Chronicle source package can pin one publisher artifact with +``artifact.artifact_year`` while still rendering ``{year}`` into its period, +record ids and vintage labels from the build year. Built with a later +``--year``, it re-labels the pinned file's values as that year: the same +cells, the same bytes, a later period. The US feed pinned at Chronicle +``c5e5bf8`` carries five such packages, all built at 2023 by +``us/chronicle_feed_scope.json`` (the congressional-district case is +PolicyEngine/chronicle#117; microcosm#1030 found the other four). For example, +the ty2023 W-2 Box 7 tips amount is cell D13 of the TY2020 workbook +``20in04w2all.xlsx``, byte for byte the same as its ty2020 twin. + +Chronicle owns the fix: it records each fact's publisher reference period +(Chronicle ``AGENTS.md``), and PolicyEngine/chronicle#292 stops the restamp. +Until the feed is re-pinned on that fix, microcosm reads these facts at their +data year where aging and the period contract read a period: the +``source_period`` aging starts from +(:mod:`~microcosm.build.us_runtime.target_aging`), and the +``uprating_to_period`` a spec inherits from a restamped rebase control. Values +and target names do not change. Two readings still see the stamp, by design, +because moving them would change which fact wins: + +- Latest-vintage selection. It refuses the compile if a restamp would outrank + a truthful fact dated after the restamp's data year + (:class:`RestampShadowsNewerVintageError`). +- The choice of a rebase control. On the pinned feed, the congressional- + district US ``net_capital_gains_returns`` row wins a period-2023 tie over + Table 1.4 only because of its stamp. + +The aging bridge indexes must not contain a restamped fact at all +(:func:`check_restamps_stay_out_of_aging_indexes`). + +Detection has two rules, and each detected restamp must match a reviewed +:data:`US_RESTAMPED_SOURCE_PACKAGES` entry (same package, stamped period and +data year) or the compile refuses (:class:`UnreviewedRestampError`): + +- **Content.** A fact whose ``source.source_sha256`` is a registered file and + whose period is after that file's data year. This holds whatever shape the + storage key takes. +- **Structure.** An observation whose ``source.raw_r2_key`` + (``raw/////``) names an earlier year + than its period. Chronicle copies that key from the manifest entry it read, + and the entry is keyed by the pinned artifact year, but the key string is + authored and not validated at ``c5e5bf8`` (chronicle#227). For a + single-year IRS file this year is the data year, so the rule is exact + there. A multi-year release keyed by its publication year can hold later + columns: BEA's 2024-keyed ``SAINC.zip`` has a 2025 column. Once Chronicle + builds such a column truthfully, the rule flags it. Register it in + :data:`US_LATER_PERIOD_OBSERVATION_EXEMPTIONS` after reading the file. + +Out of reach of both rules: a package that reads a *later* column than its +stamp. ``bea-regional-state-personal-income-components-2024`` stamps CY2024 +BEA state income as cy2023, 416 facts. No target compiles from it today, but +``bea_regional.state_wages_salaries`` is a deferred parity family ("a real +us-data target"). Activating it before chronicle#292 would calibrate 2024 +wages as cy2023. + +The register is reviewed against the pinned feed +(``test_pinned_feed_restamp_register_matches_the_feed``, which runs where the +feed is present and skips in CI). After a re-pin on a Chronicle commit that +stamps truthfully, an entry matches nothing and that test fails until the +entry is deleted. An entry the content rule still matches is still needed. +""" + +from __future__ import annotations + +import json +from collections.abc import Iterable, Mapping +from dataclasses import dataclass, replace +from types import MappingProxyType + +from microcosm.build.us_runtime.target_aging import ( + _cbo_projection_series, + _period_year, + _soi_national_chain_series, +) +from microcosm.calibrate import TargetRegistry, TargetSpec + +__all__ = [ + "US_LATER_PERIOD_OBSERVATION_EXEMPTIONS", + "US_RESTAMPED_SOURCE_PACKAGES", + "LaterPeriodObservationExemption", + "RestampShadowsNewerVintageError", + "RestampedAgingIndexError", + "RestampedSourcePackage", + "SourceVintageCorrection", + "UnreviewedRestampError", + "apply_source_vintage_corrections", + "check_restamps_stay_out_of_aging_indexes", + "detect_restamped_facts", + "fact_artifact_year", + "source_vintage_corrections", +] + + +@dataclass(frozen=True) +class RestampedSourcePackage: + """A reviewed Chronicle package whose feed stamp postdates its data. + + Attributes: + package_id: The package segment of ``source.raw_r2_key``. + data_year: The year the pinned publisher file describes. It must equal + the artifact year in ``raw_r2_key`` where the key parses. + stamped_periods: The periods the pinned feed stamps on it. + source_file: The pinned publisher file, for the reader. + source_sha256: The pinned file's digest (``source.source_sha256``), + which the content rule matches. + reason: Why the stamp is wrong, with the evidence. + """ + + package_id: str + data_year: int + stamped_periods: frozenset[int] + source_file: str + source_sha256: str + reason: str + + +@dataclass(frozen=True) +class LaterPeriodObservationExemption: + """A reviewed package whose truthful observations postdate its key year. + + For example, a multi-year release keyed by its publication year that also + carries a later data column. Facts from the package at ``periods`` are not + restamps. + """ + + package_id: str + periods: frozenset[int] + reason: str + + +def _restamped( + package_id: str, + data_year: int, + source_file: str, + source_sha256: str, + reason: str, + *, + stamped_periods: Iterable[int] = (2023,), +) -> tuple[str, RestampedSourcePackage]: + return package_id, RestampedSourcePackage( + package_id=package_id, + data_year=data_year, + stamped_periods=frozenset(stamped_periods), + source_file=source_file, + source_sha256=source_sha256, + reason=reason, + ) + + +# Reviewed against consumer_facts_us_c5e5bf8.jsonl (Chronicle c5e5bf8, sha256 +# b8543739...): these are the 26,893 facts both rules detect. Each data year +# was confirmed from the file itself (title cell or content match), not from +# Chronicle labels. "Newest" claims were checked on irs.gov on 2026-09-25. +US_RESTAMPED_SOURCE_PACKAGES: Mapping[str, RestampedSourcePackage] = MappingProxyType( + dict( + ( + _restamped( + "soi-w2-statistics-2020", + 2020, + "20in04w2all.xlsx", + "1178d77618cc1d2f873506909eeec660f36e3599854f31337f9dcaec6cfc442f", + "IRS SOI Form W-2 statistics Table 4.B, titled 'Tax Year 2020'. " + "The ty2023 tips return_count, taxpayer_count and amount share " + "their source cells, bytes and values with the ty2020 twins; " + "the ty2023 401(k) and designated Roth 401(k) amounts are cells " + "D22 and D41 of the same table. As of 2026-09-25 TY2020 is the " + "newest W-2 table on irs.gov.", + ), + _restamped( + "soi-congressional-district-2022", + 2022, + "22incd.csv", + "137522878af78d624cfddc0e17cd36ae76a21b7bed166c7bb96ad3243a18a668", + "IRS SOI congressional district data for tax year 2022 (the " + "newest CD release as of 2026-09-25). The single-district " + "states match the TY2022 county file (for example WY N1 " + "280,740), and the US taxable interest is 92.7% of TY2022 " + "Table 1.4 against 39.4% of TY2023 (PolicyEngine/chronicle#117).", + ), + _restamped( + "soi-state-2022", + 2022, + "22in54us.xlsx", + "5e58050449c07c2e941f280c60d597784047f02ee1fa93dc1aed64fc1ea493f6", + "IRS SOI Historic Table 2 US totals, titled 'Tax Year 2022'.", + ), + _restamped( + "soi-ira-roth-contributions-2022", + 2022, + "22in06ira.xlsx", + "c1fb0894cb09d2486be4510b7e6ce0b8597787725b928ab8813f577d9fb13183", + "IRS SOI IRA Table 6 (Roth contributions), titled 'Tax Year 2022'.", + ), + _restamped( + "soi-ira-traditional-contributions-2022", + 2022, + "22in05ira.xlsx", + "31ce9dd2fe631b336170af339bd7e9de86901cd20dbfd4cb8fb1c9a1d55cd90f", + "IRS SOI IRA Table 5 (traditional contributions), titled " + "'Tax Year 2022'.", + ), + ) + ) +) + +# None on the pinned feed: no package there observes a year after its key. +US_LATER_PERIOD_OBSERVATION_EXEMPTIONS: Mapping[ + str, LaterPeriodObservationExemption +] = MappingProxyType({}) + + +@dataclass(frozen=True) +class SourceVintageCorrection: + """One restamped fact read at its data year.""" + + source_record_id: str + package_id: str + stamped_year: int + data_year: int + fact_key: str = "" + + @property + def label(self) -> str: + return ( + f"{self.package_id}: stamped {self.stamped_year}, " + f"data year {self.data_year}" + ) + + +class UnreviewedRestampError(ValueError): + """A fact is stamped later than its source file's year without review. + + Either the Chronicle feed restamped a package nobody reviewed, or a + reviewed package now carries a different stamp or data year. Read the + publisher file. Then add or amend the :data:`US_RESTAMPED_SOURCE_PACKAGES` + entry (a restamp), add a :data:`US_LATER_PERIOD_OBSERVATION_EXEMPTIONS` + entry (a truthful later observation), or fix the stamp in Chronicle and + re-pin. + """ + + def __init__(self, problems: tuple[str, ...]) -> None: + self.problems = problems + super().__init__( + f"{len(problems)} fact(s) carry a period later than their source " + "file's year without a matching reviewed " + f"US_RESTAMPED_SOURCE_PACKAGES entry: {_preview(problems)}" + ) + + +class RestampShadowsNewerVintageError(ValueError): + """A restamped fact would outrank a newer truthful fact in selection.""" + + def __init__(self, conflicts: tuple[str, ...]) -> None: + self.conflicts = conflicts + super().__init__( + f"{len(conflicts)} restamped fact(s) would win latest-vintage " + "selection over a truthful fact dated after the restamp's data " + "year, and drop the newer data. Re-pin on a Chronicle commit that " + f"stamps truthfully, or exclude the restamp: {_preview(conflicts)}" + ) + + +class RestampedAgingIndexError(ValueError): + """A restamped fact would feed an aging growth index at its stamp.""" + + def __init__(self, record_ids: tuple[str, ...]) -> None: + self.record_ids = record_ids + super().__init__( + f"{len(record_ids)} restamped fact(s) sit in a target-aging growth " + "index (the CBO projection series or the national SOI chain), " + "which reads fact periods as stamped. Every chained factor through " + f"that year would use the wrong level: {_preview(record_ids)}" + ) + + +def fact_artifact_year(fact: object) -> tuple[str, int] | None: + """``(package_id, year)`` from a fact's ``source.raw_r2_key``. + + Chronicle copies the key from the manifest entry it read, and a pinned + package's entry is keyed by its ``artifact_year``. The key is an authored + string, so this reads a label, not a validated fact. Returns ``None`` when + the key is absent or not ``raw////...``. + """ + + key = _str_at(fact, "source", "raw_r2_key") + parts = key.split("/") + if len(parts) < 5 or parts[0] != "raw" or not parts[2]: + return None + year = parts[3] + if not (year.isdigit() and len(year) == 4): + return None + return parts[2], int(year) + + +def detect_restamped_facts( + facts: Iterable[object], + *, + register: Mapping[str, RestampedSourcePackage] = US_RESTAMPED_SOURCE_PACKAGES, + exemptions: Mapping[ + str, LaterPeriodObservationExemption + ] = US_LATER_PERIOD_OBSERVATION_EXEMPTIONS, +) -> tuple[tuple[object, str, int, int], ...]: + """Every observation fact stamped later than its source file's year. + + Returns ``(fact, package_id, file_year, stamped_year)`` tuples. The file + year is the ``raw_r2_key`` year where the key parses, else the registered + data year. Two rules flag a fact (see the module docstring): its + ``source_sha256`` is a registered file and its period is after that + file's data year, or its key names an earlier year than its period. + + A publisher projection (``assertion == "source_projection"``) is exempt: + a 2025 JCT score of FY2027 legitimately describes a later year. A + projection that Chronicle types as an observation is flagged like a + restamp, so the compile refuses until Chronicle types it + ``source_projection``. + """ + + by_sha = {entry.source_sha256: entry for entry in register.values()} + detected: list[tuple[object, str, int, int]] = [] + for fact in facts: + if _str_at(fact, "assertion") not in ("", "observation"): + continue + stamped_year = _period_year(_at(fact, "period", "value")) + if stamped_year is None: + continue + artifact = fact_artifact_year(fact) + entry = by_sha.get(_str_at(fact, "source", "source_sha256")) + if entry is not None and stamped_year > entry.data_year: + file_year = artifact[1] if artifact is not None else entry.data_year + detected.append((fact, entry.package_id, file_year, stamped_year)) + continue + if artifact is None: + continue + package_id, artifact_year = artifact + if artifact_year >= stamped_year: + continue + exemption = exemptions.get(package_id) + if exemption is not None and stamped_year in exemption.periods: + continue + detected.append((fact, package_id, artifact_year, stamped_year)) + return tuple(detected) + + +def source_vintage_corrections( + facts: Iterable[object], + *, + register: Mapping[str, RestampedSourcePackage] = US_RESTAMPED_SOURCE_PACKAGES, + exemptions: Mapping[ + str, LaterPeriodObservationExemption + ] = US_LATER_PERIOD_OBSERVATION_EXEMPTIONS, +) -> Mapping[str, SourceVintageCorrection]: + """Map each restamped fact's ``source_record_id`` to its correction. + + Raises: + UnreviewedRestampError: A detected restamp has no register entry, or + its stamp or file year disagrees with the entry. + ValueError: One source record id is detected twice with different + stamps or data years. + """ + + corrections: dict[str, SourceVintageCorrection] = {} + problems: list[str] = [] + for fact, package_id, file_year, stamped_year in detect_restamped_facts( + facts, register=register, exemptions=exemptions + ): + source_record_id = _source_record_id(fact) + entry = register.get(package_id) + if entry is None: + problems.append( + f"{source_record_id} ({package_id}: artifact {file_year}, " + f"stamped {stamped_year}; no register entry)" + ) + continue + if stamped_year not in entry.stamped_periods: + problems.append( + f"{source_record_id} ({package_id}: stamped {stamped_year}, " + f"register reviews {sorted(entry.stamped_periods)})" + ) + continue + if file_year != entry.data_year: + problems.append( + f"{source_record_id} ({package_id}: artifact {file_year}, " + f"register data year {entry.data_year})" + ) + continue + correction = SourceVintageCorrection( + source_record_id=source_record_id, + package_id=package_id, + stamped_year=stamped_year, + data_year=entry.data_year, + fact_key=_fact_key(fact), + ) + existing = corrections.get(source_record_id) + if existing is not None and existing != correction: + raise ValueError( + f"Restamped source record {source_record_id!r} appears with " + f"two corrections: {existing.label} vs {correction.label}." + ) + corrections[source_record_id] = correction + if problems: + raise UnreviewedRestampError(tuple(sorted(problems))) + return MappingProxyType(corrections) + + +def check_restamps_stay_out_of_aging_indexes( + facts: Iterable[object], + corrections: Mapping[str, SourceVintageCorrection], +) -> None: + """Refuse a restamped fact in an aging growth index. + + ``_cbo_projection_series`` and ``_soi_national_chain_series`` index facts + by their stamped period, and every chained aging factor through that year + reads them. None is restamped on the pinned feed: the chain draws on + Table 1.1 and Table 1.4, and the projections are CBO's. + """ + + if not corrections: + return + facts = tuple(facts) + indexed = { + record_id + for index in (_cbo_projection_series(facts), _soi_national_chain_series(facts)) + for by_year in index.values() + for _, record_id in by_year.values() + } + restamped = tuple(sorted(indexed & set(corrections))) + if restamped: + raise RestampedAgingIndexError(restamped) + + +def apply_source_vintage_corrections( + registry: TargetRegistry, + corrections: Mapping[str, SourceVintageCorrection], +) -> TargetRegistry: + """Point the aging and period-contract readings at the data year. + + Two readings move: + + - A spec backed by a restamped fact: ``source_period`` becomes the data + year. That is where target aging starts and what the period contract + checks. ``source_vintage_stamped_period`` keeps the feed's stamp, and + ``ledger_fact_period`` is left as stamped. + - A spec rebased onto a restamped control: ``uprating_to_period`` and + ``uprating_index_source_period`` become the control's data year. Aging + starts a rebased spec from ``uprating_to_period``, so a dollar spec + rebased onto a TY2022 total must not be aged as if it were TY2023. + + Values are never touched here. The one number that moves afterwards is the + aging factor, computed from the corrected period. + + Raises: + ValueError: A restamped fact reaches a spec in a way the correction + cannot vouch for: at a period other than its stamp (a fiscal-year + comparable period), as a spec that was itself rebased, as one member + of a multi-fact spec, or inside a pooled uprating index. + """ + + if not corrections: + return registry + corrected_fact_keys = { + correction.fact_key + for correction in corrections.values() + if correction.fact_key + } + specs: list[TargetSpec] = [] + for spec in registry.specs: + metadata = dict(spec.metadata) + changed = False + members = _member_fact_keys(metadata) + if len(members) > 1 and not corrected_fact_keys.isdisjoint(members): + raise ValueError( + f"Target {spec.name!r} aggregates several facts, and some are " + "restamped; its source period is its representative's, so the " + "correction cannot vouch for it. Review it." + ) + backing = corrections.get(metadata.get("ledger_source_record_id", "")) + if backing is not None: + if _period_year(metadata.get("ledger_fact_period")) != backing.stamped_year: + raise ValueError( + f"Target {spec.name!r} is backed by restamped fact " + f"{backing.source_record_id!r} ({backing.label}) at period " + f"{metadata.get('ledger_fact_period')!r}, not its stamp; " + "review how the period was compared." + ) + if "uprating_factor" in metadata: + # No restamped fact is itself rebased today. If one ever is, + # the rebase ran on the stamped period and needs review + # before its aging start can be corrected. + raise ValueError( + f"Target {spec.name!r} is backed by restamped fact " + f"{backing.source_record_id!r} ({backing.label}) and was " + "rebased; review the rebase before correcting its period." + ) + metadata["source_period"] = str(backing.data_year) + metadata["source_vintage_stamped_period"] = str(backing.stamped_year) + metadata["source_vintage_correction"] = backing.label + changed = True + for control_id in _split_ids( + metadata.get("uprating_index_source_record_ids", "") + ): + if control_id in corrections: + # A pooled index (the EITC AGI-group uprating) mixes several + # controls under one period; none is restamped today. + raise ValueError( + f"Target {spec.name!r} is uprated by a multi-source " + f"index that includes restamped fact {control_id!r} " + f"({corrections[control_id].label}); review it." + ) + control = corrections.get(metadata.get("uprating_index_source_record_id", "")) + if control is not None and ( + _period_year(metadata.get("uprating_index_source_period")) + == control.stamped_year + ): + metadata["uprating_to_period"] = str(control.data_year) + metadata["uprating_index_source_period"] = str(control.data_year) + metadata["uprating_index_source_vintage_stamped_period"] = str( + control.stamped_year + ) + metadata["uprating_index_source_vintage_correction"] = control.label + changed = True + specs.append(replace(spec, metadata=metadata) if changed else spec) + return TargetRegistry(specs, country=registry.country) + + +def _member_fact_keys(metadata: Mapping[str, str]) -> tuple[str, ...]: + payload = metadata.get("ledger_member_fact_keys", "") + if not payload: + return () + return tuple(str(key) for key in json.loads(payload)) + + +def _split_ids(value: object) -> tuple[str, ...]: + return tuple(part for part in str(value or "").split(",") if part) + + +def _preview(items: tuple[str, ...]) -> str: + suffix = "" if len(items) <= 5 else f"; +{len(items) - 5} more" + return "; ".join(items[:5]) + suffix + + +def _fact_key(fact: object) -> str: + # Mirrors ledger_targets._fact_key, the key multi-fact specs record. + return ( + _str_at(fact, "aggregate_fact_key") + or _str_at(fact, "fact_key") + or _str_at(fact, "legacy_fact_key") + or _source_record_id(fact) + ) + + +def _source_record_id(fact: object) -> str: + return _str_at(fact, "source_record_id") or _str_at( + fact, "lineage", "source_record_id" + ) + + +def _at(obj: object, *path: str) -> object: + current: object = obj + for key in path: + if current is None: + return None + if isinstance(current, dict): + current = current.get(key) + else: + current = getattr(current, key, None) + return current + + +def _str_at(obj: object, *path: str) -> str: + value = _at(obj, *path) + return "" if value is None else str(value) diff --git a/packages/microcosm-build/src/microcosm/build/us_runtime/target_aging.py b/packages/microcosm-build/src/microcosm/build/us_runtime/target_aging.py index c422e9f5e..015778cee 100644 --- a/packages/microcosm-build/src/microcosm/build/us_runtime/target_aging.py +++ b/packages/microcosm-build/src/microcosm/build/us_runtime/target_aging.py @@ -133,7 +133,9 @@ "nipa_proprietors_income": "net_business_income", # W-2 Box 7 social security tips (microcosm#451 item 3): tips are a wage # component, and the fact's TY2020 vintage predates the CBO projection - # span, so the wages series' SOI actuals provide the chained bridge. + # span, so the wages series' SOI actuals provide the chained bridge. The + # feed's ty2023 tips row is the same TY2020 cell restamped; source_vintage + # sets its source_period back to 2020 so it takes this chain too. "w2_social_security_tips_total": "wages_and_salaries", } diff --git a/packages/microcosm-build/tests/test_us_fiscal_targets.py b/packages/microcosm-build/tests/test_us_fiscal_targets.py index 3559eb845..48f9fb819 100644 --- a/packages/microcosm-build/tests/test_us_fiscal_targets.py +++ b/packages/microcosm-build/tests/test_us_fiscal_targets.py @@ -1103,26 +1103,33 @@ def _load_repo_tool(name: str): @pytest.fixture(scope="module") -def pinned_feed_national_state_surface(): - """Compile the pinned feed as the release does: compile, Medicaid - substitutions, then ``--target-surface national_state``.""" +def pinned_feed_facts(): + """The pinned Chronicle feed's facts, digest-checked against the pin.""" feed_path = _load_repo_tool("build_us_target_parity_manifest").DEFAULT_FEED_PATH if not feed_path.exists(): pytest.skip(f"pinned feed not present at {feed_path}") from microcosm.build.ledger_artifact import load_ledger_consumer_artifact + from microcosm.build.us_runtime.chronicle_feed import load_us_chronicle_feed + + return load_ledger_consumer_artifact( + feed_path, + expected_facts_sha256=load_us_chronicle_feed().facts_sha256, + expected_manifest_sha256=None, + ).facts + + +@pytest.fixture(scope="module") +def pinned_feed_national_state_surface(pinned_feed_facts): + """Compile the pinned feed as the release does: compile, Medicaid + substitutions, then ``--target-surface national_state``.""" from microcosm.build.us_runtime import ( apply_us_medicaid_enrollment_substitutions, default_congressional_district_vintage_crosswalk_path, load_congressional_district_vintage_crosswalk, ) - from microcosm.build.us_runtime.chronicle_feed import load_us_chronicle_feed from microcosm.calibrate import TargetRegistry - facts = load_ledger_consumer_artifact( - feed_path, - expected_facts_sha256=load_us_chronicle_feed().facts_sha256, - expected_manifest_sha256=None, - ).facts + facts = pinned_feed_facts crosswalk = load_congressional_district_vintage_crosswalk( default_congressional_district_vintage_crosswalk_path() ) @@ -1179,15 +1186,16 @@ def test_pinned_feed_national_state_surface_restores_the_fences( ) -> None: """The release surface on the pinned feed: 32,842 compiled targets after Medicaid substitution, 5,694 national_state targets at registry - d315c75804ef, 32 CHIP rows none of them for an M-CHIP state, no + 884bc45ef335, 32 CHIP rows none of them for an M-CHIP state, no other-income row and no tips return count. Before microcosm#956 it was - 32,867 / 5,719 at d5f9d854fe11, and before decision d179 dropped the - ty2020 tips return count it was 32,843 / 5,695 at 386fac439e77 - (docs/us-chronicle-feed-repin.md).""" + 32,867 / 5,719 at d5f9d854fe11; before decision d179 dropped the ty2020 + tips return count it was 32,843 / 5,695 at 386fac439e77; and before the + restamped W-2 tips amount was read at TY2020 it was 32,842 / 5,694 at + d315c75804ef (docs/us-chronicle-feed-repin.md).""" registry, surface, _ = pinned_feed_national_state_surface assert len(registry.specs) == 32_842 assert len(surface.specs) == 5_694 - assert surface.version == "d315c75804ef" + assert surface.version == "884bc45ef335" chip = [ spec for spec in surface.specs @@ -1214,6 +1222,150 @@ def test_pinned_feed_national_state_surface_restores_the_fences( ) +_W2_TIPS_AMOUNT = ( + "irs_soi.ty{year}.form_w2_social_security_tips.box_7_social_security_tips.amount" +) + + +def test_pinned_feed_restamp_register_matches_the_feed(pinned_feed_facts) -> None: + """Every observation fact the pinned feed stamps later than its artifact's + year belongs to a reviewed register entry, with that entry's stamp and + data year, and every entry still matches the feed. A re-pin on a + Chronicle commit that fixes the stamp (chronicle#117) fails here until + the stale entries are deleted.""" + from collections import Counter + + from microcosm.build.us_runtime.source_vintage import ( + US_RESTAMPED_SOURCE_PACKAGES, + detect_restamped_facts, + source_vintage_corrections, + ) + + detected = detect_restamped_facts(pinned_feed_facts) + assert Counter(package_id for _, package_id, _, _ in detected) == { + "soi-congressional-district-2022": 26_880, + "soi-w2-statistics-2020": 5, + "soi-state-2022": 4, + "soi-ira-roth-contributions-2022": 2, + "soi-ira-traditional-contributions-2022": 2, + } + assert set(US_RESTAMPED_SOURCE_PACKAGES) == { + package_id for _, package_id, _, _ in detected + } + for fact, package_id, artifact_year, stamped_year in detected: + entry = US_RESTAMPED_SOURCE_PACKAGES[package_id] + assert artifact_year == entry.data_year + assert stamped_year in entry.stamped_periods + # Both detection rules agree on every row: the key names the file's + # year and the digest is the registered file. + assert fact["source"]["raw_r2_key"].split("/")[2:4] == [ + package_id, + str(entry.data_year), + ] + assert fact["source"]["source_sha256"] == entry.source_sha256 + assert len(source_vintage_corrections(pinned_feed_facts)) == 26_893 + + +def test_pinned_feed_w2_tips_amount_ages_from_tax_year_2020( + pinned_feed_facts, pinned_feed_national_state_surface +) -> None: + """The route-A tips target is the ty2023 row, which is the TY2020 cell + restamped. It ages from 2020 on the chained SOI wages bridge, so it + equals what its honest ty2020 twin would compile to. Aged from the stamp + (the direct 2023 CBO ratio) it was $28.28B, 17.5% lower.""" + from microcosm.build.us_runtime.target_aging import ( + _cbo_projection_series, + _soi_national_chain_series, + ) + + _, surface, _ = pinned_feed_national_state_surface + spec = _aged_spec_by_source_record_id(surface, _W2_TIPS_AMOUNT.format(year=2023)) + wages = _cbo_projection_series(pinned_feed_facts)["wages_and_salaries"] + soi_wages = _soi_national_chain_series(pinned_feed_facts)["wages_and_salaries"] + factor = (soi_wages[2023][0] / soi_wages[2020][0]) * ( + wages[2024][0] / wages[2023][0] + ) + assert spec.metadata["ledger_fact_period"] == "2023" + assert spec.metadata["source_period"] == "2020" + assert spec.metadata["source_vintage_stamped_period"] == "2023" + assert spec.metadata["aging_factor_source"].startswith( + "chained:irs_soi.ty2023.table_1_4.all.wages_salaries_amount+" + ) + assert float(spec.metadata["aging_factor"]) == pytest.approx(factor, rel=1e-12) + assert spec.value == pytest.approx(26_786_522_000 * factor, rel=1e-12) + assert spec.value == pytest.approx(34_287_530_779, abs=1) + + +def test_pinned_feed_corrects_the_translated_congressional_district_specs( + pinned_feed_national_state_surface, +) -> None: + """Most corrections land on congressional-district facts that the vintage + crosswalk re-keys (``current_cd.*``). Detection runs on the translated + facts, so every translated spec carries the correction, and the dollar + ones age on the chained bridge from TY2022.""" + from collections import Counter + + registry, _, _ = pinned_feed_national_state_surface + corrected = Counter( + ( + spec.metadata["source_vintage_correction"].split(":")[0], + ".current_cd." in spec.metadata["ledger_source_record_id"], + spec.metadata["aging_factor_source"].startswith("chained:"), + ) + for spec in registry.specs + if "source_vintage_correction" in spec.metadata + ) + assert corrected == { + ("soi-congressional-district-2022", True, True): 11_772, + ("soi-congressional-district-2022", True, False): 12_208, + ("soi-congressional-district-2022", False, True): 1_404, + ("soi-congressional-district-2022", False, False): 1_404, + ("soi-state-2022", False, False): 2, + ("soi-w2-statistics-2020", False, True): 1, + } + assert all( + spec.metadata["source_period"] == "2022" + for spec in registry.specs + if spec.metadata.get("source_vintage_correction", "").startswith( + "soi-congressional-district-2022" + ) + ) + assert ( + sum( + "uprating_index_source_vintage_correction" in spec.metadata + for spec in registry.specs + ) + == 60 + ) + + +def test_pinned_feed_capital_gains_returns_control_reads_its_data_year( + pinned_feed_national_state_surface, +) -> None: + """The Historic Table 2 net-capital-gains return counts rebase onto the + congressional-district file's US row, which is TY2022 data stamped + ty2023. The rebase lands them at the control's data year, 2022; the + counts never age, so no value moves. Which control those rows should use + is a separate concept question: the Table 1.4 ty2023 row counts a + different return population.""" + _, surface, _ = pinned_feed_national_state_surface + control_id = ( + "irs_soi.ty2023.congressional_district_2022.all_returns.us." + "net_capital_gains_returns" + ) + rebased = [ + spec + for spec in surface.specs + if spec.metadata.get("uprating_index_source_record_id") == control_id + ] + assert len(rebased) == 51 + for spec in rebased: + assert spec.metadata["uprating_to_period"] == "2022" + assert spec.metadata["uprating_index_source_vintage_stamped_period"] == "2023" + assert spec.metadata["uprating_factor"] == "0.97964474977721" + assert spec.metadata["aging_factor_source"] == "not_dollar_amount" + + def test_reviewed_zero_support_facts_are_not_active_targets() -> None: excluded_source_record_id = ( "hhs_acf_tanf.fy2024.cash_assistance.ar." @@ -5064,6 +5216,216 @@ def test_age_targets_chains_w2_tips_through_soi_wages_bridge() -> None: assert abs(spec.value - 26_786_522_000 * expected_factor) < 1.0 +def _w2_tips_chain_facts(*, tips_year: int, raw_r2_key: str | None = None): + """Tips amount plus the SOI wages bridge and CBO wages projections.""" + tips = _dynamic_ledger_fact( + source_record_id=_W2_TIPS_AMOUNT.format(year=tips_year), + source_name="irs_soi", + measure_id="amount", + value=26_786_522_000, + period_value=tips_year, + layout_record_set_id=f"irs_soi.ty{tips_year}.form_w2_social_security_tips", + groupby_dimension="irs_soi.form_w2_item", + groupby_value_id="box_7_social_security_tips", + ) + tips["assertion"] = "observation" + if raw_r2_key is not None: + tips["source"]["raw_r2_key"] = raw_r2_key + wages = [ + _dynamic_ledger_fact( + source_record_id=f"irs_soi.ty{year}.table_1_4.all.wages_salaries_amount", + source_name="irs_soi", + measure_id="wages_salaries_amount", + value=value, + period_value=year, + dimensions={"income_range": "all", "filing_status": "all"}, + layout_record_set_id=f"irs_soi.ty{year}.table_1_4", + groupby_dimension="us:statutes/26/62#adjusted_gross_income", + groupby_value_id="all", + ) + for year, value in ((2020, 8_416_495_535_000), (2023, 10_000_000_000_000)) + ] + return [ + *packaged_reference_facts(), + tips, + *wages, + _cbo_income_source_projection_fact( + 2023, "wages_and_salaries", value=10_200_000_000_000 + ), + _cbo_income_source_projection_fact( + 2024, "wages_and_salaries", value=10_700_000_000_000 + ), + ] + + +_W2_RAW_R2_KEY = ( + "raw/irs_soi/soi-w2-statistics-2020/2020/" + "1178d77618cc1d2f873506909eeec660f36e3599854f31337f9dcaec6cfc442f/" + "20in04w2all.xlsx" +) + + +def test_age_targets_reads_a_restamped_w2_tips_amount_at_its_data_year() -> None: + # chronicle#117: the feed's ty2023 tips amount is the TY2020 workbook + # cell stamped 2023. Its raw key names the 2020 artifact, so it must age + # exactly as its honest ty2020 twin does (differential), not from 2023. + restamped = compile_us_fiscal_target_registry( + _w2_tips_chain_facts(tips_year=2023, raw_r2_key=_W2_RAW_R2_KEY), + target_period=2024, + age_targets=True, + ) + honest = compile_us_fiscal_target_registry( + _w2_tips_chain_facts(tips_year=2020, raw_r2_key=_W2_RAW_R2_KEY), + target_period=2024, + age_targets=True, + ) + spec = _aged_spec_by_source_record_id(restamped, _W2_TIPS_AMOUNT.format(year=2023)) + twin = _aged_spec_by_source_record_id(honest, _W2_TIPS_AMOUNT.format(year=2020)) + assert spec.metadata["source_period"] == "2020" + assert spec.metadata["ledger_fact_period"] == "2023" + assert spec.metadata["source_vintage_correction"] == ( + "soi-w2-statistics-2020: stamped 2023, data year 2020" + ) + assert spec.metadata["aging_factor_source"].startswith("chained:") + assert spec.metadata["aging_factor"] == twin.metadata["aging_factor"] + assert spec.value == twin.value + assert "source_vintage_correction" not in twin.metadata + + # Without the raw key there is no evidence of the restamp, and the fact + # ages from its stamp on the direct CBO ratio. + unkeyed = compile_us_fiscal_target_registry( + _w2_tips_chain_facts(tips_year=2023), + target_period=2024, + age_targets=True, + ) + direct = _aged_spec_by_source_record_id(unkeyed, _W2_TIPS_AMOUNT.format(year=2023)) + assert direct.metadata["source_period"] == "2023" + assert direct.value == pytest.approx( + 26_786_522_000 * 10_700_000_000_000 / 10_200_000_000_000 + ) + + +def test_compile_refuses_a_restamp_that_would_outrank_newer_data() -> None: + # A truthful TY2021 tips fact would lose latest-vintage selection to the + # TY2020 cell stamped 2023; the compile refuses rather than drop it. + from microcosm.build.us_runtime.source_vintage import ( + RestampShadowsNewerVintageError, + ) + + newer = _dynamic_ledger_fact( + source_record_id=_W2_TIPS_AMOUNT.format(year=2021), + source_name="irs_soi", + measure_id="amount", + value=30_000_000_000, + period_value=2021, + layout_record_set_id="irs_soi.ty2021.form_w2_social_security_tips", + groupby_dimension="irs_soi.form_w2_item", + groupby_value_id="box_7_social_security_tips", + ) + facts = [ + *_w2_tips_chain_facts(tips_year=2023, raw_r2_key=_W2_RAW_R2_KEY), + newer, + ] + with pytest.raises(RestampShadowsNewerVintageError, match="ty2021"): + compile_us_fiscal_target_registry(facts, target_period=2024, age_targets=True) + # A truthful fact at the stamp year itself is shadowed too (the tie-break, + # not the data, would decide), and one after the target period is not a + # candidate at all. + at_stamp = dict(newer, value=31_000_000_000) + at_stamp["lineage"] = { + "source_record_id": _W2_TIPS_AMOUNT.format(year=2023) + "_truthful" + } + at_stamp["period"] = {"type": "calendar_year", "value": 2023} + for key in ("aggregate_fact_key", "semantic_fact_key", "legacy_fact_key"): + at_stamp[key] = str(newer[key]).replace("2021", "2023truthful") + with pytest.raises(RestampShadowsNewerVintageError, match="2023"): + compile_us_fiscal_target_registry( + [ + *_w2_tips_chain_facts(tips_year=2023, raw_r2_key=_W2_RAW_R2_KEY), + at_stamp, + ], + target_period=2024, + age_targets=True, + ) + # Selection eligibility applies first: at target 2022 neither the 2023 + # restamp nor the 2023 truthful fact is a candidate. + from microcosm.build.us_runtime.source_vintage import SourceVintageCorrection + + restamp = SourceVintageCorrection("restamped", "pkg", 2023, 2020) + candidates = [ + (("key",), (1, 202399, "2023"), 1.0, "restamped", None), + (("key",), (1, 202399, "2023"), 1.0, "truthful", None), + ] + fiscal_targets._refuse_restamps_that_shadow_newer_vintages( + candidates, + {"restamped": restamp}, + target_period_key=(1, 202299, "2022"), + ) + with pytest.raises(RestampShadowsNewerVintageError, match="truthful"): + fiscal_targets._refuse_restamps_that_shadow_newer_vintages( + candidates, + {"restamped": restamp}, + target_period_key=(1, 202499, "2024"), + ) + + # The honest twin at the data year is not newer data: no refusal. + twin = dict(newer, value=26_786_522_000) + twin["lineage"] = {"source_record_id": _W2_TIPS_AMOUNT.format(year=2020)} + twin["period"] = {"type": "calendar_year", "value": 2020} + twin["layout"] = dict( + newer["layout"], record_set_id="irs_soi.ty2020.form_w2_social_security_tips" + ) + for key in ("aggregate_fact_key", "semantic_fact_key", "legacy_fact_key"): + twin[key] = str(newer[key]).replace("2021", "2020twin") + compile_us_fiscal_target_registry( + [*_w2_tips_chain_facts(tips_year=2023, raw_r2_key=_W2_RAW_R2_KEY), twin], + target_period=2024, + age_targets=True, + ) + + +def test_compile_refuses_a_restamp_in_an_aging_index_only_when_aging( + monkeypatch, +) -> None: + # No registered restamp sits in the chain today; force one to prove the + # compile runs the check before aging, and skips it when not aging. + from microcosm.build.us_runtime.source_vintage import ( + RestampedAgingIndexError, + SourceVintageCorrection, + ) + + chain_id = "irs_soi.ty2023.table_1_4.all.wages_salaries_amount" + monkeypatch.setattr( + fiscal_targets, + "source_vintage_corrections", + lambda facts: { + chain_id: SourceVintageCorrection(chain_id, "soi-t14-2022", 2023, 2022) + }, + ) + facts = _w2_tips_chain_facts(tips_year=2020) + with pytest.raises(RestampedAgingIndexError, match=chain_id): + compile_us_fiscal_target_registry(facts, target_period=2024, age_targets=True) + compile_us_fiscal_target_registry( + facts, + target_period=2024, + age_targets=False, + allow_unaged_dollar_targets=True, + ) + + +def test_compile_refuses_an_unreviewed_restamp() -> None: + from microcosm.build.us_runtime.source_vintage import UnreviewedRestampError + + facts = _w2_tips_chain_facts( + tips_year=2023, + raw_r2_key=_W2_RAW_R2_KEY.replace( + "soi-w2-statistics-2020/2020", "soi-w2-statistics-2021/2021" + ), + ) + with pytest.raises(UnreviewedRestampError, match="soi-w2-statistics-2021"): + compile_us_fiscal_target_registry(facts, target_period=2024, age_targets=True) + + def test_ssa_ssi_age_band_counts_bind_as_person_age_indicator_targets() -> None: # microcosm#470: the SSA SSI Monthly age-band recipient counts bind as # national indicator counts of engine ssi receipt sliced by the fact's diff --git a/packages/microcosm-build/tests/test_us_source_vintage.py b/packages/microcosm-build/tests/test_us_source_vintage.py new file mode 100644 index 000000000..181676e24 --- /dev/null +++ b/packages/microcosm-build/tests/test_us_source_vintage.py @@ -0,0 +1,580 @@ +"""Restamped Chronicle facts are read at their data year (chronicle#117). + +The compile-level cases (a restamped W-2 tips amount ageing from TY2020, the +refusals of an unreviewed or shadowing restamp) and the pinned-feed register +check live in ``test_us_fiscal_targets.py`` beside the other W-2 aging tests. +""" + +from __future__ import annotations + +import json + +import pytest + +from microcosm.build.us_runtime.source_vintage import ( + US_LATER_PERIOD_OBSERVATION_EXEMPTIONS, + US_RESTAMPED_SOURCE_PACKAGES, + LaterPeriodObservationExemption, + RestampedAgingIndexError, + RestampedSourcePackage, + SourceVintageCorrection, + UnreviewedRestampError, + apply_source_vintage_corrections, + check_restamps_stay_out_of_aging_indexes, + detect_restamped_facts, + fact_artifact_year, + source_vintage_corrections, +) +from microcosm.calibrate import TargetRegistry, TargetSpec + +_W2 = "soi-w2-statistics-2020" +_W2_SHA = US_RESTAMPED_SOURCE_PACKAGES[_W2].source_sha256 +_OTHER_SHA = "0" * 64 +_TIPS_2023 = ( + "irs_soi.ty2023.form_w2_social_security_tips.box_7_social_security_tips.amount" +) + + +def _raw_key(package_id: str, artifact_year: int | str, file: str = "f.xlsx") -> str: + return f"raw/irs_soi/{package_id}/{artifact_year}/{_OTHER_SHA}/{file}" + + +def _fact( + source_record_id: str, + *, + period: int | str, + raw_r2_key: str | None, + assertion: str | None = "observation", + sha: str | None = None, +) -> dict[str, object]: + source: dict[str, object] = {} + if raw_r2_key is not None: + source["raw_r2_key"] = raw_r2_key + if sha is not None: + source["source_sha256"] = sha + fact: dict[str, object] = { + "aggregate_fact_key": f"ledger.aggregate_fact.v2:{source_record_id}@{period}", + "lineage": {"source_record_id": source_record_id}, + "period": {"type": "tax_year", "value": period}, + "source": source, + } + if assertion is not None: + fact["assertion"] = assertion + return fact + + +def _spec(name: str, value: float = 1.0, **metadata: str) -> TargetSpec: + return TargetSpec( + name=name, + entity="tax_unit", + value=value, + measure="tip_income", + period=2024, + source="test", + family="irs_soi", + metadata=metadata, + ) + + +def _entry(package_id: str, data_year: int, *stamps: int, sha: str = "s"): + return RestampedSourcePackage( + package_id, data_year, frozenset(stamps), "f.xlsx", sha, "reason" + ) + + +def test_fact_artifact_year_reads_the_raw_key_package_and_year() -> None: + assert fact_artifact_year( + _fact("x", period=2023, raw_r2_key=_raw_key(_W2, 2020)) + ) == (_W2, 2020) + for key in ( + None, + "", + "raw/irs_soi", + f"raw/irs_soi/{_W2}/source_capture/{_OTHER_SHA}/f.xlsx", + f"raw/irs_soi/{_W2}/20/{_OTHER_SHA}/f.xlsx", + f"staged/irs_soi/{_W2}/2020/{_OTHER_SHA}/f.xlsx", + f"raw/irs_soi//2020/{_OTHER_SHA}/f.xlsx", + ): + assert fact_artifact_year(_fact("x", period=2023, raw_r2_key=key)) is None + + +def test_detect_restamped_facts_flags_only_observations_stamped_after_their_data() -> ( + None +): + facts = [ + _fact("restamp", period=2023, raw_r2_key=_raw_key(_W2, 2020)), + _fact("legacy", period=2023, raw_r2_key=_raw_key(_W2, 2020), assertion=None), + _fact("same_year", period=2020, raw_r2_key=_raw_key(_W2, 2020)), + # A 2024 release carrying cy2023 rows is the ordinary direction. + _fact("later_release", period=2023, raw_r2_key=_raw_key("bea-nipa", 2024)), + # A publisher's projection of a later year is not a restamp. + _fact( + "projection", + period=2027, + raw_r2_key=_raw_key("jct-obbba", 2025), + assertion="source_projection", + ), + _fact("no_key", period=2023, raw_r2_key=None), + _fact("month_label", period="2024-12", raw_r2_key=_raw_key("cms", 2026)), + ] + detected = detect_restamped_facts(facts) + assert [row[0]["lineage"]["source_record_id"] for row in detected] == [ + "restamp", + "legacy", + ] + assert {row[1:] for row in detected} == {(_W2, 2020, 2023)} + + +def test_the_content_rule_detects_a_registered_file_under_any_key_shape() -> None: + # A Chronicle key-layout change (for example a country segment) must not + # hide a registered restamp: the pinned file's digest still names it. + moved_key = f"raw/us/irs_soi/{_W2}/2020/{_W2_SHA}/20in04w2all.xlsx" + detected = detect_restamped_facts( + [ + _fact(_TIPS_2023, period=2023, raw_r2_key=moved_key, sha=_W2_SHA), + _fact("no_key", period=2023, raw_r2_key=None, sha=_W2_SHA), + # At its data year the registered file is stamped truthfully. + _fact("honest", period=2020, raw_r2_key=None, sha=_W2_SHA), + ] + ) + assert [(row[0]["lineage"]["source_record_id"], *row[1:]) for row in detected] == [ + (_TIPS_2023, _W2, 2020, 2023), + ("no_key", _W2, 2020, 2023), + ] + assert len(source_vintage_corrections([row[0] for row in detected])) == 2 + + +def test_a_reviewed_exemption_admits_a_truthful_later_observation() -> None: + # BEA's 2024-keyed SAINC.zip carries a 2025 column; once Chronicle builds + # it truthfully, the structural rule alone would flag it. + fact = _fact( + "bea_regional.cy2025.x", period=2025, raw_r2_key=_raw_key("bea-sainc", 2024) + ) + assert detect_restamped_facts([fact]) + exemptions = { + "bea-sainc": LaterPeriodObservationExemption( + "bea-sainc", frozenset({2025}), "multi-year release" + ) + } + assert not detect_restamped_facts([fact], exemptions=exemptions) + assert not source_vintage_corrections([fact], exemptions=exemptions) + # The exemption covers its reviewed periods only. + later = _fact( + "bea_regional.cy2026.x", period=2026, raw_r2_key=_raw_key("bea-sainc", 2024) + ) + assert detect_restamped_facts([later], exemptions=exemptions) + assert US_LATER_PERIOD_OBSERVATION_EXEMPTIONS == {} + + +def test_register_entries_describe_one_earlier_year_each() -> None: + assert set(US_RESTAMPED_SOURCE_PACKAGES) == { + "soi-w2-statistics-2020", + "soi-congressional-district-2022", + "soi-state-2022", + "soi-ira-roth-contributions-2022", + "soi-ira-traditional-contributions-2022", + } + digests = [entry.source_sha256 for entry in US_RESTAMPED_SOURCE_PACKAGES.values()] + assert len(set(digests)) == len(digests) + for package_id, entry in US_RESTAMPED_SOURCE_PACKAGES.items(): + assert entry.package_id == package_id + assert entry.stamped_periods + assert all(stamp > entry.data_year for stamp in entry.stamped_periods) + # IRS SOI file names lead with the two-digit tax year (20in04w2all). + assert int(entry.source_file[:2]) == entry.data_year % 100 + assert package_id.endswith(str(entry.data_year)) + assert len(entry.source_sha256) == 64 + assert entry.reason + + +def test_corrections_map_each_registered_restamp_to_its_data_year() -> None: + corrections = source_vintage_corrections( + [ + _fact(_TIPS_2023, period=2023, raw_r2_key=_raw_key(_W2, 2020)), + _fact( + "irs_soi.ty2020.form_w2_social_security_tips.x.amount", + period=2020, + raw_r2_key=_raw_key(_W2, 2020), + ), + ] + ) + assert dict(corrections) == { + _TIPS_2023: SourceVintageCorrection( + source_record_id=_TIPS_2023, + package_id=_W2, + stamped_year=2023, + data_year=2020, + fact_key=f"ledger.aggregate_fact.v2:{_TIPS_2023}@2023", + ) + } + + +@pytest.mark.parametrize( + ("package_id", "artifact_year", "period", "reason"), + [ + ("soi-unreviewed-2021", 2021, 2023, "no register entry"), + (_W2, 2020, 2024, r"register reviews \[2023\]"), + (_W2, 2019, 2023, "register data year 2020"), + ], +) +def test_an_unreviewed_restamp_is_refused( + package_id: str, artifact_year: int, period: int, reason: str +) -> None: + with pytest.raises(UnreviewedRestampError, match=reason) as raised: + source_vintage_corrections( + [_fact("x", period=period, raw_r2_key=_raw_key(package_id, artifact_year))] + ) + assert len(raised.value.problems) == 1 + + +def test_a_registered_file_under_a_key_naming_another_year_is_refused() -> None: + fact = _fact(_TIPS_2023, period=2023, raw_r2_key=_raw_key(_W2, 2021), sha=_W2_SHA) + with pytest.raises(UnreviewedRestampError, match="register data year 2020"): + source_vintage_corrections([fact]) + + +def test_one_record_with_two_different_corrections_is_refused() -> None: + register = {"a": _entry("a", 2020, 2023, 2024)} + with pytest.raises(ValueError, match="two corrections"): + source_vintage_corrections( + [ + _fact("x", period=2023, raw_r2_key=_raw_key("a", 2020)), + _fact("x", period=2024, raw_r2_key=_raw_key("a", 2020)), + ], + register=register, + ) + + +def _chain_fact(year: int) -> dict[str, object]: + return { + "lineage": { + "source_record_id": f"irs_soi.ty{year}.table_1_4.all.wages_salaries_amount" + }, + "period": {"type": "tax_year", "value": year}, + "value": 1e13, + "geography": {"level": "country"}, + "layout": { + "record_set_id": f"irs_soi.ty{year}.table_1_4", + "groupby_value_id": "all", + }, + "observed_measure": { + "source_name": "irs_soi", + "source_measure_id": "wages_salaries_amount", + }, + } + + +def test_a_restamped_fact_in_an_aging_chain_index_is_refused() -> None: + # A pinned TY2022 Table 1.4 stamped 2023 would stand in for the TY2023 + # wages level in every chained factor that pivots on 2023. + chain = [_chain_fact(2020), _chain_fact(2023)] + restamped_id = "irs_soi.ty2023.table_1_4.all.wages_salaries_amount" + correction = SourceVintageCorrection(restamped_id, "soi-t14-2022", 2023, 2022) + with pytest.raises(RestampedAgingIndexError, match=restamped_id): + check_restamps_stay_out_of_aging_indexes(chain, {restamped_id: correction}) + check_restamps_stay_out_of_aging_indexes(chain, {_TIPS_2023: _CORRECTION}) + check_restamps_stay_out_of_aging_indexes(chain, {}) + + +def test_a_restamped_fact_in_the_cbo_projection_index_is_refused() -> None: + projection_id = ( + "cbo.revenue_projection.ty2023.income_by_source.wages_and_salaries." + "projected_amount" + ) + projection = { + "assertion": "source_projection", + "lineage": {"source_record_id": projection_id}, + "period": {"type": "tax_year", "value": 2023}, + "value": 1e13, + "layout": { + "record_set_id": "cbo.revenue_projection.ty2023.income_by_source", + "groupby_dimension": "cbo.income_source", + "groupby_value_id": "wages_and_salaries", + }, + "observed_measure": { + "source_name": "cbo", + "source_measure_id": "projected_amount", + }, + } + correction = SourceVintageCorrection(projection_id, "cbo-2026-02", 2023, 2022) + with pytest.raises(RestampedAgingIndexError, match="cbo.revenue_projection"): + check_restamps_stay_out_of_aging_indexes( + [projection], {projection_id: correction} + ) + + +_CORRECTION = SourceVintageCorrection( + _TIPS_2023, _W2, 2023, 2020, fact_key="ledger.aggregate_fact.v2:tips" +) + + +def test_apply_moves_only_the_backing_facts_source_period() -> None: + registry = TargetRegistry( + [ + _spec( + "restamped", + 28.0, + ledger_source_record_id=_TIPS_2023, + ledger_fact_period="2023", + source_period="2023", + ), + _spec( + "other", + 5.0, + ledger_source_record_id="irs_soi.ty2023.table_1_4.all.x", + ledger_fact_period="2023", + source_period="2023", + ), + ], + country="us", + ) + corrected = { + spec.name: spec + for spec in apply_source_vintage_corrections( + registry, {_TIPS_2023: _CORRECTION} + ).specs + } + restamped = corrected["restamped"] + assert restamped.value == 28.0 + assert restamped.metadata["source_period"] == "2020" + assert restamped.metadata["ledger_fact_period"] == "2023" + assert restamped.metadata["source_vintage_stamped_period"] == "2023" + assert restamped.metadata["source_vintage_correction"] == ( + "soi-w2-statistics-2020: stamped 2023, data year 2020" + ) + assert corrected["other"] == registry.specs[1] + + +def test_apply_refuses_a_backed_spec_compared_at_another_period() -> None: + # A fiscal-year fact is compared at its coverage start year, which can + # differ from the stamp the correction was made for; fail closed. + spec = _spec( + "fiscal", + ledger_source_record_id=_TIPS_2023, + ledger_fact_period="2022", + source_period="2023", + ) + with pytest.raises(ValueError, match="not its stamp"): + apply_source_vintage_corrections( + TargetRegistry([spec], country="us"), {_TIPS_2023: _CORRECTION} + ) + + +def test_apply_moves_a_rebased_specs_control_period_to_the_controls_data_year() -> None: + control_id = ( + "irs_soi.ty2023.congressional_district_2022.all_returns.us." + "net_capital_gains_returns" + ) + control = SourceVintageCorrection( + control_id, "soi-congressional-district-2022", 2023, 2022 + ) + spec = _spec( + "irs_soi.ty2022.historic_table_2.state_broad.ca.all.net_capital_gains_returns", + 3_774_052.0, + ledger_source_record_id=( + "irs_soi.ty2022.historic_table_2.state_broad.ca.all." + "net_capital_gains_returns" + ), + ledger_fact_period="2022", + source_period="2022", + uprating_factor="0.97964474977721", + uprating_from_period="2022", + uprating_to_period="2023", + uprating_index_source_period="2023", + uprating_index_source_record_id=control_id, + ) + (corrected,) = apply_source_vintage_corrections( + TargetRegistry([spec], country="us"), {control_id: control} + ).specs + assert corrected.value == spec.value + assert corrected.metadata["uprating_factor"] == "0.97964474977721" + assert corrected.metadata["uprating_to_period"] == "2022" + assert corrected.metadata["uprating_index_source_period"] == "2022" + assert corrected.metadata["uprating_index_source_vintage_stamped_period"] == "2023" + assert corrected.metadata["source_period"] == "2022" + + +def test_apply_refuses_a_restamped_fact_that_was_itself_rebased() -> None: + spec = _spec( + "restamped", + ledger_source_record_id=_TIPS_2023, + ledger_fact_period="2023", + source_period="2023", + uprating_factor="1.1", + uprating_to_period="2023", + ) + with pytest.raises(ValueError, match="was rebased"): + apply_source_vintage_corrections( + TargetRegistry([spec], country="us"), {_TIPS_2023: _CORRECTION} + ) + + +def test_apply_refuses_a_pooled_uprating_index_with_a_restamped_member() -> None: + spec = _spec( + "pooled", + ledger_source_record_id="irs_soi.ty2022.table_2_5.x", + ledger_fact_period="2022", + uprating_index_source_record_ids=f"irs_soi.ty2024.a,{_TIPS_2023}", + ) + with pytest.raises(ValueError, match="multi-source index"): + apply_source_vintage_corrections( + TargetRegistry([spec], country="us"), {_TIPS_2023: _CORRECTION} + ) + + +def test_apply_refuses_a_multi_fact_spec_with_a_restamped_member() -> None: + # The spec's source period is its representative's; a restamped member + # that is not the representative would go uncorrected. + spec = _spec( + "summed", + ledger_source_record_id="irs_soi.ty2023.representative", + ledger_fact_period="2023", + ledger_member_fact_keys=json.dumps( + ["ledger.aggregate_fact.v2:representative", _CORRECTION.fact_key] + ), + ) + with pytest.raises(ValueError, match="aggregates several facts"): + apply_source_vintage_corrections( + TargetRegistry([spec], country="us"), {_TIPS_2023: _CORRECTION} + ) + + +def test_correction_invariants_hold_for_generated_registries() -> None: + """For any registry and any corrections: values, names, periods and order + never move; the correction is idempotent; a spec changes if and only if + it is backed by, or rebased onto, a corrected fact at its stamped period; + and every moved period lands on that fact's data year. Hypothesis is a + workspace dependency; the wheels job installs no test extras, so it + skips there.""" + pytest.importorskip("hypothesis") + from hypothesis import given, settings + from hypothesis import strategies as st + + record_ids = [f"irs_soi.ty2023.r{i}" for i in range(5)] + years = st.integers(2015, 2026) + + @st.composite + def cases(draw): + corrections = {} + for record_id in draw(st.sets(st.sampled_from(record_ids))): + stamped = draw(years) + data = draw(st.integers(2010, stamped - 1)) + corrections[record_id] = SourceVintageCorrection( + record_id, "pkg", stamped, data + ) + specs = [] + for index in range(draw(st.integers(0, 8))): + record_id = draw(st.sampled_from(record_ids)) + # A backed spec is always compared at its fact's stamp; the + # mismatch case is refused (tested above). + fact_period = ( + corrections[record_id].stamped_year + if record_id in corrections + else draw(years) + ) + metadata = { + "ledger_source_record_id": record_id, + "ledger_fact_period": str(fact_period), + "source_period": str(draw(years)), + } + if draw(st.booleans()): + metadata["uprating_index_source_record_id"] = draw( + st.sampled_from(record_ids) + ) + metadata["uprating_index_source_period"] = str(draw(years)) + metadata["uprating_to_period"] = metadata[ + "uprating_index_source_period" + ] + specs.append( + _spec( + f"spec{index}", + draw(st.floats(0, 1e12, allow_nan=False)), + **metadata, + ) + ) + return specs, corrections + + def backed(spec, corrections): + return spec.metadata["ledger_source_record_id"] in corrections + + def rebased(spec, corrections): + correction = corrections.get( + spec.metadata.get("uprating_index_source_record_id", "") + ) + return correction is not None and str(correction.stamped_year) == ( + spec.metadata.get("uprating_index_source_period") + ) + + @settings(max_examples=300, deadline=None) + @given(cases()) + def check(case): + specs, corrections = case + registry = TargetRegistry(specs, country="us") + once = apply_source_vintage_corrections(registry, corrections) + twice = apply_source_vintage_corrections(once, corrections) + assert once.specs == twice.specs + assert [spec.name for spec in once.specs] == [spec.name for spec in specs] + for before, after in zip(specs, once.specs, strict=True): + assert after.value == before.value + assert after.period == before.period + assert ( + after.metadata["ledger_fact_period"] + == (before.metadata["ledger_fact_period"]) + ) + is_backed = backed(before, corrections) + is_rebased = rebased(before, corrections) + assert (after != before) == (is_backed or is_rebased) + if is_backed: + correction = corrections[before.metadata["ledger_source_record_id"]] + assert after.metadata["source_period"] == str(correction.data_year) + else: + assert ( + after.metadata["source_period"] + == (before.metadata["source_period"]) + ) + if is_rebased: + correction = corrections[ + before.metadata["uprating_index_source_record_id"] + ] + assert after.metadata["uprating_to_period"] == str(correction.data_year) + + check() + + +def test_detection_matches_its_two_rules_for_generated_facts() -> None: + """A fact is detected exactly when it is an observation (or carries no + assertion) and either its digest is a registered file stamped after that + file's data year, or its raw key parses to a year before its period.""" + pytest.importorskip("hypothesis") + from hypothesis import given, settings + from hypothesis import strategies as st + + @settings(max_examples=400, deadline=None) + @given( + artifact_year=st.integers(2000, 2030), + period=st.one_of( + st.integers(2000, 2030), + st.integers(2000, 2030).map(lambda year: f"{year}-12"), + st.integers(2000, 2030).map(lambda year: f"ty{year}"), + ), + assertion=st.sampled_from([None, "observation", "source_projection"]), + keyed=st.booleans(), + registered_file=st.booleans(), + ) + def check(artifact_year, period, assertion, keyed, registered_file): + fact = _fact( + "x", + period=period, + raw_r2_key=_raw_key("some-package", artifact_year) if keyed else None, + assertion=assertion, + sha=_W2_SHA if registered_file else None, + ) + period_year = int(str(period).removeprefix("ty")[:4]) + observation = assertion != "source_projection" + by_content = registered_file and period_year > 2020 + by_key = keyed and artifact_year < period_year + assert bool(detect_restamped_facts([fact])) == ( + observation and (by_content or by_key) + ) + + check() diff --git a/packages/microcosm-build/tests/test_us_spine_blindness.py b/packages/microcosm-build/tests/test_us_spine_blindness.py index a18f026f5..5c0b4fe78 100644 --- a/packages/microcosm-build/tests/test_us_spine_blindness.py +++ b/packages/microcosm-build/tests/test_us_spine_blindness.py @@ -305,6 +305,9 @@ "snap_take_up.py", "source_coverage.py", "source_runtime.py", + # Restamped-fact period correction over compiled target specs + # (chronicle#117); reads Ledger facts only, no population treatment. + "source_vintage.py", "sources.py", "spine_agreement.py", "spine_assembly.py",