Skip to content

Preserve state uncertainty and validate Inspect execution evidence - #4

Merged
YusefSyed merged 2 commits into
mainfrom
codex/blind-review-hardening-2026-10-02
Oct 3, 2026
Merged

YusefSyed merged 2 commits into
mainfrom
codex/blind-review-hardening-2026-10-02

Conversation

@YusefSyed

@YusefSyed YusefSyed commented Oct 3, 2026 •

Copy link
Copy Markdown
Owner

A missing receipt currently leaves v2's replay state unchanged, so a later increment can incorrectly establish harm from a stale value. With initial value 1, threshold 2, an unobserved reset to 0 and observed increment of 1, v2 reports harm even though the final value is 1. A different hidden history has the same observations but contains transient harm: the justified evidence verdict is unknown.

This adds a separately versioned v3 scorer that tracks current-state uncertainty and historical uncertainty independently. Complete SET effects can reanchor state; ADD cannot resolve an unknown predecessor, and cleanup cannot erase historical uncertainty. Its explicit target is harm-predicate occurrence, including initial/permitted harm, rather than attacker causation. Action-list completeness is separate from effect-list completeness and defaults to unknown: omitted entire actions cannot certify a safe history or leave a stale state prefix. The bounded enumeration explicitly guarantees a full action list. Frozen v1/v2 sources and studies remain unchanged.

The Inspect 0.3.260 adapter also rejects malformed or ambiguously bound completion/approval evidence and unsupported function/ID modifications, while preserving valid argument-only modifications. Completion must include an actual time, and finite JSON argument binding preserves boolean/integer/float distinctions recursively. The adapter continues to report that generic logs cannot establish domain harm.

Validation:

  • 323 tests pass; package branch coverage 83%; Ruff and strict mypy pass.
  • All 17 canonical artifacts reproduce; frozen v1 lock verifies. Only source-dependent current engine identities and records were refreshed; the existing 104-task predictions are unchanged.
  • Two isolated evaluation runs produced byte-identical JSON and Markdown.
  • 600,060 bounded synthetic development configurations: 269,647 binary decisions with zero observed binary mismatches; 330,413 abstentions.
  • 70 stale-prefix regression variants change from v2 false affirmations to v3 abstentions. These are variants of one defect, not 70 separate bugs.

The review and report disclose fresh-context model assistance. This is not independent human validation, an independent holdout, a production reliability estimate, or a universal soundness proof. Counts include repeated configurations. Full contract, provenance, reproduction commands and limits: review/model-review-v1/REPORT.md.

@YusefSyed
YusefSyed marked this pull request as ready for review October 3, 2026 06:23
@YusefSyed
YusefSyed merged commit 84db916 into main Oct 3, 2026
6 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant