Repository navigation
Preserve state uncertainty and validate Inspect execution evidence - #4
Merged
Merged
Conversation
YusefSyed
marked this pull request as ready for review
October 3, 2026 06:23
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
A missing receipt currently leaves v2's replay state unchanged, so a later increment can incorrectly establish harm from a stale value. With initial value 1, threshold 2, an unobserved reset to 0 and observed increment of 1, v2 reports harm even though the final value is 1. A different hidden history has the same observations but contains transient harm: the justified evidence verdict is unknown.
This adds a separately versioned v3 scorer that tracks current-state uncertainty and historical uncertainty independently. Complete SET effects can reanchor state; ADD cannot resolve an unknown predecessor, and cleanup cannot erase historical uncertainty. Its explicit target is harm-predicate occurrence, including initial/permitted harm, rather than attacker causation. Action-list completeness is separate from effect-list completeness and defaults to unknown: omitted entire actions cannot certify a safe history or leave a stale state prefix. The bounded enumeration explicitly guarantees a full action list. Frozen v1/v2 sources and studies remain unchanged.
The Inspect 0.3.260 adapter also rejects malformed or ambiguously bound completion/approval evidence and unsupported function/ID modifications, while preserving valid argument-only modifications. Completion must include an actual time, and finite JSON argument binding preserves boolean/integer/float distinctions recursively. The adapter continues to report that generic logs cannot establish domain harm.
Validation:
The review and report disclose fresh-context model assistance. This is not independent human validation, an independent holdout, a production reliability estimate, or a universal soundness proof. Counts include repeated configurations. Full contract, provenance, reproduction commands and limits:
review/model-review-v1/REPORT.md.