Receipts: record receipt_origin's contemporaneous/correction/back-port split beside 'snippets verified' (#2933) - #4441
Open
realmarcin wants to merge 5 commits into
Open
realmarcin wants to merge 5 commits into
realmarcin wants to merge 5 commits into
Conversation
…t split beside 'snippets verified' (#2933) The pinned record contract's `receipts` slot is an AnyBlock (additionalProperties: true), so `receipts.origin` validates without a contract change; a new top-level block would not (tested). - receipt_origin_record (new): the block a record keeps of receipt_origin.origin -- every key but local paths (transcripts by basename and sha256, the receipt by sha256), with the instrument and NON_CHECKS beside the counts. A writer with no transcript keeps a measurement the record carries while the receipt still has the sha256 it was measured on (#907: a recomputation that cannot read the evidence never erases a measurement); where the receipt changed, the block is unknown and the measurement is kept under `prior`; with none, unknown with the reason (an API-path record says why it has no transcript). Never contemporaneous. - d4d receipts check: --transcript (with --receipt-at-run/--full-at-run) measures the origin and --write records it; the split prints under "snippets M/M verified" only where a transcript was read, so the playbook's mid-run output is unchanged. A withheld write now says so instead of ticking "written". - backfill_checks.compute (so `provenance record` and `backfill-checks --blocks receipts`): the receipts block carries `origin`; the recorder hands a re-record the origin it is about to rewrite. summarise shows the split only when measured. - canary: two reported-only rows (snippets contemporaneous, snippets post-draft) in the batch summary; `d4d api verdict` prints the origin line after its rows. Never gated: post-draft snippets stay unaccepted for semantic support (#2067). No committed record is rewritten; a corpus backfill is an owner decision. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… the instrument (#2933) The module docstring said a block read from no transcript carries neither the instrument nor NON_CHECKS; unknown() names the instrument, and the tests pin that. Docstring only. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
line() described any checked origin that split() refused as measured on other bytes; a checked block whose counts are not integers (a hand-edited record) now says so instead. One helper decides 'other bytes' for both. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…rinted, not silent (#2933) `receipts check` and the recorder's summary line printed the origin only for a block that is itself a measurement. Once the receipt changes after its origin was measured, the block is `unknown` and keeps the measurement under `prior`, and both went silent, as for a record never measured. `ever_measured` covers both cases, so the line now says why the split no longer applies. The withheld-write note said an origin "read here" was not written whenever the block carried a measurement, including one only kept from the record with no --transcript. It now says so only when --transcript was given. The three new `check` parameters lose their Python defaults, like the command's other parameters. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… by keyword (#2933) `record` called `_inline_checks(path, origin_prior)` with two positional arguments. tests/test_generation_manifest_identity.py stubs the helper with `lambda path: None`, so all three parametrizations of test_agentic_render_to_record_preserves_selected_inputs failed with a TypeError on this branch; the earlier commits ran only their own test file. `origin_prior` is now keyword-only and passed by name, and that stub takes `*a, **k`, as the helper's other stubs in test_profiles.py already do. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
This was referenced Oct 5, 2026
Open
Open
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Refs #2933 (second PR; the parts under "Left open" stay open).
Receipt origin in the provenance record and beside "snippets verified" (#2933, part 2)
#3034 added
d4d receipts origin. It reads a native or direct run's transcript and classifies each coverage-receipt snippet ascontemporaneous,phase1_correctionorphase3_backport. It is report-only, so no record, gate or summary carried the split. As a result, "snippets N/N verified" still mixed evidence written while reading with evidence added after the draft: 53 of 329 snippets on the three direct-arm CHORUS attempts. This PR records the split underreceipts.originand prints it beside "snippets verified".The record contract takes it unchanged
The pinned contract (
src/data_sheets_schema/schema/d4d_generation_record.yaml, compiled byrecord_schema.py) typesreceiptsasAnyBlock, which compiles toadditionalProperties: true.GenerationRecorditself compiles toadditionalProperties: false. Soreceipts.originvalidates with no contract change, while a new top-levelreceipt_originblock would be rejected.test_the_record_contract_takes_the_block_inside_receipts_onlypins both. No schema file changes.What changes
receipt_origin_record.py(new) defines what a record keeps ofreceipt_origin.origin's block:receipt_origin v6 …) and itsNON_CHECKS, beside the counts;Which block a writer records (
for_record):unknown, with a reason naming both sha256s, and the measurement is kept underprior. A later writer keeps thatprior, and restoring the measured bytes restores the measurement.unknownwith the reason, and no classification; nevercontemporaneous. An API-path record (one carryingapi_usage) says it has no tool-call transcript and names API v7: slots.without_receipt mixes phase-1 unreceipted values with reconcile rewrites; verified snippets can point at removed values #807 and v8 defect (CM4AI canary): a receipt addressed to a path the record does not carry ('subject' for a value placed under keywords) #952 as that path's own accounts.split()returns counts only for acheckedorigin measured on the receipt the block checked (same sha256). A split of other bytes is never shown beside the block's counts.d4d receipts checkgains--transcript(repeatable, first invocation first),--receipt-at-runand--full-at-run. With--write, the measured block goes into the record.unknownand gives the reason; it is never left out silently.--write, with--writetwice, and with--strict; the receipts block minusorigin; and the claims sidecar.applykeeps a checked block over an unchecked recomputation (Review-instrument fixes: identity receipt join (#899), review-aware selection (#660), spelling v3 (#836/#859), dispositions (#903) #907), main still printed "✓ receipts block written". It now says the block was not written. Where--transcriptwas given, it also says the origin it read was not written either.backfill_checks.compute, and sod4d provenance recordandbackfill-checks --blocks receipts, writesreceipts.originbeside the receipts check: the record's measurement kept, orunknown. Neither command holds a transcript.provenance recordrewrites the record from scratch before its inline checks run. It therefore reads the origin first and passes it through by keyword:_inline_checks(path, *, origin_prior=…)→compute(origin_prior=…).summariseappends(C contemporaneous, P post-draft), or(origin unknown)where a measurement stopped applying. It appends nothing for a record never measured, so the recorder's line during a run is unchanged.Canary.
d4d api batch's per-run line gains two reported-only rows next tosnippets unverified:snippets contemporaneousandsnippets post-draft. They read—wherever no transcript was read, which is every API-path run.d4d api verdict, the offline gate for a record of any arm, prints the origin line after its rows.verdict()is identical with and without an origin (tested). Post-draft snippets stay unaccepted as semantic support until independent review (Clarify Phase 1 receipt correction versus terminal native evidence-check stops #2067).What does not change
receipt_origin.py,receipts.py,api_runner.py, playbooks, agents, prompts, and the native and direct controllers.--strict,receipt_floorsandverdict()read nothing new.api_runner._receipts_block) carry nooriginkey. The summaries read that as not measured.No corpus backfill
No committed record under
data/d4d_concatenated/**is rewritten. None of the 286 provenance records there carries an origin block, and none comes from the native or direct arms. Measuring a run is an owner decision, made per run with its preserved transcript(s):backfill-checks --blocks receipts --overwrite --executemeasures nothing, because it reads no transcript. In every receipts block it rewrites, it writesorigin: unknownor keeps a measurement the record already carries. It also recomputes the whole receipts block, so running it over the corpus is a separate decision.Tests
tests/test_receipt_origin_record.py: 21 tests, synthetic transcripts only; none walks the corpus. They cover:priorchain after the receipt changes;split()andline();verdict()equal with and without the origin;d4d api verdict;compute,applyandsummarise;receipts check --write --transcriptend to end, after the receipt moved, and on a withheld write;d4d provenance recordre-record that keeps a measured origin;git archivecopies, each counted killed only by a failure in this file. 27 ran at 5afd1fc. At 2376d2c the recorder hand-off mutants were re-run and one new one was added for the changed call site. All 28 were killed.-m "not corpus" -n 2, at 5afd1fc): 3,714 passed, 29 skipped, 3 failed.test_agentic_render_to_record_preserves_selected_inputs[command|header|none]stubbed_inline_checkswithlambda path: None, and the branch passed it a second argument.origin_prioris now keyword-only, and that stub takes*a, **k, like the helper's stubs intest_profiles.py._inline_checks, plus the failing one: 178 passed;provenance record: 1,378 passed, 1 skipped.-m corpus -n 2, at 2376d2c, run after the PR opened): 116 passed, exit 0.Left open (#2933 stays open)
notes/matched_cborg_2026-09-13/run_api_canary.py) changes only with a new registration.native_supervisor_gates.receipt_gate, used bynative_execution_gatesandnative_attempt_supervisor) check sealed, captured bytes, whilereceipt_origin.originreads files by path.d4d receipts readdress, the issue's optional part 3 (d4d receipts readdress: native counterpart of the API path's full_readdress (#2933 part 3, optional) #4446).d4d receipts originat receipt_origin v6. receipt origin v1: classify each receipt snippet as contemporaneous, Phase 1 correction or Phase 3 back-port from the transcript (report-only) (#2933 PR1) #3034 reproduced the issue's table at v1 (40ea81a). The transcripts are untracked in other worktrees, which this PR did not read. Which runs to measure is Measure receipts.origin for the preserved direct-arm and native attempts (owner decision) #4444.receipts.originor the new options (Document receipts.origin andreceipts check --transcriptin CLAUDE.md's Coverage Receipts section #4447); the loop does not edit CLAUDE.md.🤖 Generated with Claude Code