Skip to content

Audit period strips by reviewed template, not an exact spelling list - #309

Open
MaxGhenis wants to merge 1 commit into
codex/thesis-ledger-factsfrom
fix/catalog-period-invariant
Open

MaxGhenis wants to merge 1 commit into
codex/thesis-ledger-factsfrom
fix/catalog-period-invariant

Conversation

@MaxGhenis

Copy link
Copy Markdown
Contributor

Problem

test_committed_catalog_is_current_and_valid pinned the exact list of stripped period spellings in the committed catalog. Every Thesis resolver append that records the next weekly claims print adds a spelling (week_2026-08-29, week_2026-09-05, …), so every append failed the required "Arch checks" status. The resolver merges through the API with an admin token (it waits only for "Append gate"), so the red check was ignored:

What the list protected

The builder can't tell a real period label from a statute, cohort, or edition label that happens to spell the row's own period. stripped_segments exists so a curator can catch that (build_series_catalog.py module docstring; classify_segments). The exact list forced a human look at every new spelling. It also failed when a spelling disappeared.

It had a gap: only spellings were pinned, not concepts. A new concept starting to strip an already-listed spelling (for example 2026_07) passed unseen.

The invariant now checked

build_series_catalog.stripped_segment_review_problems(stripped_segments, stripped_kinds, reviewed) audits every (spelling, concept) strip. The reviewed map is tests/fixtures/series_catalog/reviewed_stripped_segments.json: the catalog at 55bbf3d that #244 reviewed, recorded with each strip's classify_segments kinds.

A strip that is not in the reviewed map passes only when both of these hold:

  1. its concept already has a reviewed strip with the same period_spelling_template (digits become 9, month words become <month>, separators and literal words are kept), so it is the next week or month written the same way;
  2. every current observation behind it spells its own declared period directly (kind derived, never overlap).

Everything else is still reported:

Event Old exact list Now
next weekly/monthly print, same concept and spelling style fail (every append) pass
new concept strips for the first time (new series, or a label that spells its row's period) fail fail
existing concept strips a new template (e.g. week_ending_… where it used week_…) fail fail
overlap strip (identifier finer than the declared period, the June-13 defect) fail only if the spelling was new fail
reviewed strip starts being reached as an overlap pass fail
new concept strips an already-listed spelling pass fail
reviewed strip disappears fail fail
non-period spelling / catalog–plan disagreement n/a fail

build_catalog now exposes each strip's kinds as plan["stripped_kinds"]. The catalog bytes do not change: --check is green, the byte-regeneration assertion passes, and no catalog keys are added, so Thesis's frozen CATALOG_TOP_LEVEL_KEYS/CATALOG_ROW_KEYS are unaffected.

week_2026-06-13 vs week_2026_06_13

Not a naming bug. Both come from the one 2026-06-13 initial-claims row (ledger line 44):

  • its source_record_id is us.dol.initial_claims.sa.week_2026-06-13 (ISO hyphens);
  • its measure.concept is us.dol.initial_claims.sa.week_2026_06_13 (underscores).

The builder strips period tokens from both fields and records each distinct spelling. Both strip as overlap because the row declares month 2026-06, the cadence defect already in docs/catalog-curation-backlog.md ("Initial claims cadence metadata"). This is now documented in the fixture notes and the backlog, and test_reviewed_stripped_segments_fixture_is_wellformed asserts it.

Evidence on real data

I overlaid this PR's builder, test and fixture on each refused proposal head and ran tests/test_build_series_catalog.py:

Head Before After
#283 78308f8 1 failed (pin) 171 passed
#297 7917dbe (job 108675953661) 1 failed (pin) 171 passed
#298 7bc1595 1 failed (pin) 171 passed
this branch's base 598f394 (merged append) 1 failed (pin) 171 passed
#289 177c461 2 failed (pin, anchor) 2 failed: '2026_08' on 'abs.labour.unemployment_rate': first '9999_99' strip for this concept + the existing anchor test

#289 is a real curation event. The docket placeholder abs.labour.unemployment_rate took its first print beside an already-observed abs.labour.unemployment_rate.australia, which looks like a concept-spelling duplicate. So a failure there is the check doing its job.

Deliberately corrupted copies of #297's head all fail test_committed_catalog_is_current_and_valid:

  • Statute label on a new concept. I appended treasury.debt_limit.suspension_act.2026_09.first_print (month 2026-09) and regenerated, so the catalog is current. Result: '2026_09' on 'treasury.debt_limit.suspension_act': first '9999_99' strip….
  • Weekly print declared as a month. I appended us.dol.initial_claims.sa.week_2026-09-12 with period month 2026-09 and regenerated. Result: 'week_2026-09-12' on 'us.dol.initial_claims.sa': unreviewed overlap strip….
  • Hand-edited catalog. Deleting after_mpc_june_2026 from stripped_segments fails the byte-regeneration assertion.

Invariants (tested)

Property tests use Hypothesis, derandomized so this required check stays deterministic:

  • T1: a spelling's template depends only on how it is written, never on the period it names.
  • T2: different spelling families get different templates. Full and abbreviated month names share one template on purpose, because may is both.
  • T3: every generated family spells its own period directly.
  • A1, routine continuation passes: end to end through the real builder, appending further periods of observed series written as they already are never produces a problem. A non-vacuity assertion checks that the new strips really are unreviewed.
  • A2, soundness: end to end, a new concept, a new template, or an overlap strip is always reported by name. The overlap case covers both a fresh spelling and a reviewed spelling that gains a kind.
  • A3, persistence: a reviewed pair that is not in the catalog is reported as gone.
  • A4, monotone in review: reviewing more live strips never adds problems, and reviewing every live strip leaves none.
  • A5: the output is sorted and does not depend on input order.

Other tests:

  • Differential: test_stripped_kinds_match_an_independent_replay re-derives every strip's kinds from the ledger with classify_segments. It resolves each observation to its catalog row through that row's concept and aliases, and requires plan["stripped_kinds"] and the catalog's stripped_segments to agree exactly.
  • Real ledger, next period: for all ≥100 current observations with direct period spellings, append the next period written the same way. They land on the existing identities with no mints, supersedes or drops, and add ≥50 new strips, none needing review.
  • Mutation check: I applied 11 mutants to the audit and template code (dropping each check, identity or coarser templates, reporting every kind as derived, unsorted output). Every one was killed by at least one test.

Scope and coordination

  • Resolve catalog supersede links against every row's effective assertion id #306 (open, green) adds Hypothesis with the byte-identical pyproject.toml and uv.lock change, so those hunks merge cleanly in either order. Resolve catalog supersede links against every row's effective assertion id #306 also re-pins the exact list (week_2026-08-29/09-05/09-12). Whichever PR lands second resolves that one hunk by keeping this PR's audit, since the new spellings pass as routine continuations.
  • Unchanged pins: the row-count pin (== 228), suspect_segments == ["LNU02374597"] and test_identity_uuid_map_matches_reviewed_anchor fire only on identity events (mints, enrichments, supersedes). The repo documents those as needing curation. No routine append this month tripped them; only Record 8 first-print observation(s) via resolve_pending.py #289, a real curation question, did.
  • CI steps run locally: build_series_catalog.py --check, pytest tests/test_build_series_catalog.py (171 passed, about 8s), --verify-registry-append-only against the base registry, and the full "Lint Arch surface" ruff list. I did not run the Arch-surface pytest list, the DB build or the wheel steps locally (the host was overloaded). No module outside this test imports the builder, and CI runs them.

axiom: n/a: catalog test and tooling change, no policy logic.

🤖 Generated with Claude Code

test_committed_catalog_is_current_and_valid pinned the exact list of
stripped period spellings in the committed catalog. Every resolver append
that recorded the next weekly claims print added a spelling
(week_2026-08-29, week_2026-09-05, ...) and failed the required "Arch
checks" status, so proposals merged past it with the resolver's admin
token (#240 on 2026-09-03, 598f394 on 2026-09-29) and the check stopped
meaning anything.

The list stood for one question: has a curator decided that this strip is
a period label and not a statute, cohort, or edition label that happens
to spell its row's period? The test now asks that per (spelling, concept)
pair:

- build_series_catalog.stripped_segment_review_problems audits the
  catalog's stripped_segments against a reviewed map
  (tests/fixtures/series_catalog/reviewed_stripped_segments.json, the
  catalog at 55bbf3d that #244 reviewed, with each strip's kinds).
- A strip outside the map passes only if its concept already has a
  reviewed strip of the same period_spelling_template (the same way of
  writing a period) and every observation behind it spells its own
  declared period directly (kind derived).
- New concepts, new templates, unreviewed overlap strips, reviewed strips
  that gain the overlap kind, reviewed strips that disappear, non-period
  spellings, and catalog/plan disagreements all still fail. A concept
  that starts stripping an already-listed spelling used to pass unseen;
  it now fails.

build_catalog exposes each strip's classify_segments kinds as
plan["stripped_kinds"]; catalog bytes are unchanged (--check green).

week_2026-06-13 and week_2026_06_13 are not a naming bug: both come from
the one 2026-06-13 initial-claims row (ledger line 44), whose
source_record_id uses ISO hyphens and whose measure.concept uses
underscores. Both strip as overlap because that row declares a month (the
cadence defect in docs/catalog-curation-backlog.md). Documented in the
fixture notes and the backlog, and asserted by the fixture test.

Hypothesis joins the dev extras with the same pyproject and uv.lock change
as #306.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant