Skip to content

feat(vgr): add fail-closed decision policy - #717

Open
gburachas wants to merge 1 commit into
vgr/review/01-capabilitiesfrom
vgr/review/02-policy
Open

gburachas wants to merge 1 commit into
vgr/review/01-capabilitiesfrom
vgr/review/02-policy

Conversation

@gburachas

@gburachas gburachas commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

What

Adds the decision policy for verification-gated routing, building on the capabilities introduced in #716. It combines verifier scores, yes/no verdicts, and tool-result history to decide whether to accept the local attempt or escalate to the capable tier.

It also adds verifier request builders, verdict parsing, and a probability readout from OpenAI Chat logprobs.

Why

Each verification branch needs explicit acceptance rules. A factual answer can be checked directly, while an agentic result also needs evidence that tool failures were resolved. Keeping these rules separate from execution makes the policy reviewable without following model calls or streaming behavior.

Notes for reviewers

Review this against #716. Start with decide in crates/libsy/src/algorithms/vgr/decide.rs, then read readout.rs and rungs.rs.

  • Acceptance rules: Readout thresholds are 0.7 for answer/default verification, 0.3 for chat, and 0.2 for agentic work. Alternative verifier paths differ by branch. Coding and unknown inputs always escalate; this layer has no sandboxed code checker.
  • Agentic recovery: Runs with tool errors and at most 100 results can qualify through local verification if the final result is clean. Longer runs additionally require the configured clean-result tail and an affirmative capable-tier judge.
  • Verdict handling: Missing or unrecognized verdicts do not count as approval. The parser requires a final yes/no line and rejects replies with an explicit non-completion stop reason. Truncated local attempts remain available as evidence.
  • Existing APIs: Reuses the shared protocol request/response types, feat(vgr): add trusted capability derivation #716’s redaction and clipping helpers, and the existing response-preservation mechanism to read logprobs. The translation change forwards logprobs alongside the already-supported top_logprobs.
  • Readout semantics: Scores normalize the returned yes/no token probabilities. If only affirmative alternatives appear, the score is 1.0; if neither appears, it is unavailable. These are verifier scores, not calibrated correctness probabilities.
  • Tests: Cover recovery boundaries, missing and indeterminate evidence, verdict parsing, verifier request isolation, and logprob scoring.

This PR defines policy and verifier inputs. #718 connects them to runtime execution.

@gburachas
gburachas requested a review from a team as a code owner September 16, 2026 02:06
@github-actions

github-actions Bot commented Sep 16, 2026 •

Copy link
Copy Markdown
PR Preview Action v1.8.1

🚀 View preview at
https://NVIDIA-NeMo.github.io/Switchyard/pr-preview/pr-717/

Built to branch gh-pages at 2026-09-30 19:20 UTC.
Preview will be ready when the GitHub Pages deployment is complete.

@afourniernv afourniernv reopened this Sep 27, 2026
Signed-off-by: adhaile <adhaile@nvidia.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants