Skip to content

Receipts: record receipt_origin's contemporaneous/correction/back-port split beside 'snippets verified' (#2933) - #4441

Open
realmarcin wants to merge 5 commits into
mainfrom
fix/2933-receipt-origin-block
Open

realmarcin wants to merge 5 commits into
mainfrom
fix/2933-receipt-origin-block

Conversation

@realmarcin

@realmarcin realmarcin commented Oct 5, 2026 •

Copy link
Copy Markdown
Collaborator

Refs #2933 (second PR; the parts under "Left open" stay open).

Receipt origin in the provenance record and beside "snippets verified" (#2933, part 2)

#3034 added d4d receipts origin. It reads a native or direct run's transcript and classifies each coverage-receipt snippet as contemporaneous, phase1_correction or phase3_backport. It is report-only, so no record, gate or summary carried the split. As a result, "snippets N/N verified" still mixed evidence written while reading with evidence added after the draft: 53 of 329 snippets on the three direct-arm CHORUS attempts. This PR records the split under receipts.origin and prints it beside "snippets verified".

The record contract takes it unchanged

The pinned contract (src/data_sheets_schema/schema/d4d_generation_record.yaml, compiled by record_schema.py) types receipts as AnyBlock, which compiles to additionalProperties: true. GenerationRecord itself compiles to additionalProperties: false. So receipts.origin validates with no contract change, while a new top-level receipt_origin block would be rejected. test_the_record_contract_takes_the_block_inside_receipts_only pins both. No schema file changes.

What changes

receipt_origin_record.py (new) defines what a record keeps of receipt_origin.origin's block:

  • every key except local paths: transcripts by basename, sha256 and line count, and the receipt by sha256 (the receipts block beside it names its path);
  • the instrument (receipt_origin v6 …) and its NON_CHECKS, beside the counts;
  • no snippet text and no tool payloads, like the instrument's own output.

Which block a writer records (for_record):

split() returns counts only for a checked origin measured on the receipt the block checked (same sha256). A split of other bytes is never shown beside the block's counts.

d4d receipts check gains --transcript (repeatable, first invocation first), --receipt-at-run and --full-at-run. With --write, the measured block goes into the record.

  • The split prints on the line after "snippets M/M verified", but only where a transcript was read, now or earlier. If the receipt has since changed, the line says unknown and gives the reason; it is never left out silently.
  • With no transcript and no measurement in the record, the output is byte-identical to main. That is the playbook's mid-run call. Checked on one fixed fixture under both trees: no record yet; a record without --write, with --write twice, and with --strict; the receipts block minus origin; and the claims sidecar.
  • One difference, which is a fix. When apply keeps a checked block over an unchecked recomputation (Review-instrument fixes: identity receipt join (#899), review-aware selection (#660), spelling v3 (#836/#859), dispositions (#903) #907), main still printed "✓ receipts block written". It now says the block was not written. Where --transcript was given, it also says the origin it read was not written either.

backfill_checks.compute, and so d4d provenance record and backfill-checks --blocks receipts, writes receipts.origin beside the receipts check: the record's measurement kept, or unknown. Neither command holds a transcript.

  • provenance record rewrites the record from scratch before its inline checks run. It therefore reads the origin first and passes it through by keyword: _inline_checks(path, *, origin_prior=…) → compute(origin_prior=…).
  • summarise appends (C contemporaneous, P post-draft), or (origin unknown) where a measurement stopped applying. It appends nothing for a record never measured, so the recorder's line during a run is unchanged.

Canary.

  • d4d api batch's per-run line gains two reported-only rows next to snippets unverified: snippets contemporaneous and snippets post-draft. They read — wherever no transcript was read, which is every API-path run.
  • d4d api verdict, the offline gate for a record of any arm, prints the origin line after its rows.
  • Neither is a verdict row or a floor, and verdict() is identical with and without an origin (tested). Post-draft snippets stay unaccepted as semantic support until independent review (Clarify Phase 1 receipt correction versus terminal native evidence-check stops #2067).

What does not change

  • No pinned file is touched: the schema directory, receipt_origin.py, receipts.py, api_runner.py, playbooks, agents, prompts, and the native and direct controllers.
  • --strict, receipt_floors and verdict() read nothing new.
  • Records written by the pinned API runner (api_runner._receipts_block) carry no origin key. The summaries read that as not measured.

No corpus backfill

No committed record under data/d4d_concatenated/** is rewritten. None of the 286 provenance records there carries an origin block, and none comes from the native or direct arms. Measuring a run is an owner decision, made per run with its preserved transcript(s):

poetry run d4d receipts check --method <method> --label <label> --project <PROJECT> --write \
    --transcript <transcript.jsonl> [--transcript <resumed.jsonl> ...] \
    [--receipt-at-run <receipt path as the transcript spelled it>] \
    [--full-at-run <full-record path as the transcript spelled it>]

backfill-checks --blocks receipts --overwrite --execute measures nothing, because it reads no transcript. In every receipts block it rewrites, it writes origin: unknown or keeps a measurement the record already carries. It also recomputes the whole receipts block, so running it over the corpus is a separate decision.

Tests

  • tests/test_receipt_origin_record.py: 21 tests, synthetic transcripts only; none walks the corpus. They cover:
    • a contemporaneous snippet, a Phase 1 correction and a Phase 3 back-port;
    • an all-contemporaneous run;
    • the issue's re-address-only case (0 added, 0 removed, 3 re-addressed);
    • no transcript, an API-path record, and a missing or truncated transcript;
    • a kept measurement, and the prior chain after the receipt changes;
    • split() and line();
    • the canary rows, with verdict() equal with and without the origin;
    • d4d api verdict;
    • compute, apply and summarise;
    • receipts check --write --transcript end to end, after the receipt moved, and on a withheld write;
    • a d4d provenance record re-record that keeps a measured origin;
    • the contract check.
  • Mutation: 28 distinct mutants on git archive copies, each counted killed only by a failure in this file. 27 ran at 5afd1fc. At 2376d2c the recorder hand-off mutants were re-run and one new one was added for the changed call site. All 28 were killed.
  • The 110 test files that import a changed module or the root CLI group (-m "not corpus" -n 2, at 5afd1fc): 3,714 passed, 29 skipped, 3 failed.
    • The three failures were this branch's: test_agentic_render_to_record_preserves_selected_inputs[command|header|none] stubbed _inline_checks with lambda path: None, and the branch passed it a second argument.
    • Fixed in 2376d2c. origin_prior is now keyword-only, and that stub takes *a, **k, like the helper's stubs in test_profiles.py.
  • After the fix, at 2376d2c:
    • every file that calls or stubs _inline_checks, plus the failing one: 178 passed;
    • the 25 further files that drive provenance record: 1,378 passed, 1 skipped.
  • Corpus lane (-m corpus -n 2, at 2376d2c, run after the PR opened): 116 passed, exit 0.

Left open (#2933 stays open)

🤖 Generated with Claude Code

realmarcin and others added 5 commits October 5, 2026 03:35
…t split beside 'snippets verified' (#2933)

The pinned record contract's `receipts` slot is an AnyBlock (additionalProperties:
true), so `receipts.origin` validates without a contract change; a new top-level
block would not (tested).

- receipt_origin_record (new): the block a record keeps of receipt_origin.origin --
  every key but local paths (transcripts by basename and sha256, the receipt by
  sha256), with the instrument and NON_CHECKS beside the counts. A writer with no
  transcript keeps a measurement the record carries while the receipt still has
  the sha256 it was measured on (#907: a recomputation that cannot read the
  evidence never erases a measurement); where the receipt changed, the block is
  unknown and the measurement is kept under `prior`; with none, unknown with the
  reason (an API-path record says why it has no transcript). Never contemporaneous.
- d4d receipts check: --transcript (with --receipt-at-run/--full-at-run) measures
  the origin and --write records it; the split prints under "snippets M/M
  verified" only where a transcript was read, so the playbook's mid-run output is
  unchanged. A withheld write now says so instead of ticking "written".
- backfill_checks.compute (so `provenance record` and `backfill-checks --blocks
  receipts`): the receipts block carries `origin`; the recorder hands a re-record
  the origin it is about to rewrite. summarise shows the split only when measured.
- canary: two reported-only rows (snippets contemporaneous, snippets post-draft)
  in the batch summary; `d4d api verdict` prints the origin line after its rows.
  Never gated: post-draft snippets stay unaccepted for semantic support (#2067).

No committed record is rewritten; a corpus backfill is an owner decision.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… the instrument (#2933)

The module docstring said a block read from no transcript carries neither
the instrument nor NON_CHECKS; unknown() names the instrument, and the
tests pin that. Docstring only.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
line() described any checked origin that split() refused as measured on
other bytes; a checked block whose counts are not integers (a hand-edited
record) now says so instead. One helper decides 'other bytes' for both.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…rinted, not silent (#2933)

`receipts check` and the recorder's summary line printed the origin only
for a block that is itself a measurement. Once the receipt changes after
its origin was measured, the block is `unknown` and keeps the measurement
under `prior`, and both went silent, as for a record never measured.
`ever_measured` covers both cases, so the line now says why the split no
longer applies.

The withheld-write note said an origin "read here" was not written
whenever the block carried a measurement, including one only kept from
the record with no --transcript. It now says so only when --transcript
was given. The three new `check` parameters lose their Python defaults,
like the command's other parameters.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… by keyword (#2933)

`record` called `_inline_checks(path, origin_prior)` with two positional
arguments. tests/test_generation_manifest_identity.py stubs the helper
with `lambda path: None`, so all three parametrizations of
test_agentic_render_to_record_preserves_selected_inputs failed with a
TypeError on this branch; the earlier commits ran only their own test
file. `origin_prior` is now keyword-only and passed by name, and that
stub takes `*a, **k`, as the helper's other stubs in test_profiles.py
already do.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant