Skip to content

feat(vgr): expose verified routing surface - #719

Open
gburachas wants to merge 1 commit into
vgr/review/03-runtimefrom
vgr/review/04-public-surface
Open

gburachas wants to merge 1 commit into
vgr/review/03-runtimefrom
vgr/review/04-public-surface

Conversation

@gburachas

@gburachas gburachas commented Sep 16, 2026 •

Copy link
Copy Markdown
Contributor

What

Makes VGR available through deployment TOML as type = "vgr", wiring local, cloud, and optional verifier targets into the existing runner and server.

It also adds live context-window discovery for VGR routes backed by llama.cpp, two provider compatibility fixes, and a configuration guide.

Why

The runtime from #718 needs a deployment configuration and an end-to-end server path. Clients also need the local backend’s actual context capacity, and context-overflow errors need to reach the existing fallback handling.

Notes for reviewers

Review this against #718. Start with VgrRouteConfig and build_vgr in crates/switchyard-runner/src/algorithm.rs, then follow route construction in config.rs and the server integration test.

  • Existing patterns: Uses AlgorithmSpec, target-name resolution, model categories, and the existing route/client construction path. Local and cloud targets populate the efficient/capable categories; verifier targets are registered as judge dependencies.
  • Configuration: Exposes serving mode, verification deadline, task typing, circuit-breaker settings, and recovery confirmation. The default mode remains off. Construction checks target references, requires distinct local/cloud model IDs, and delegates runtime validation—including active-mode approval—to Vgr::new.
  • Context discovery: /v1/models probes the VGR local backend’s /props endpoint and reads default_generation_settings.n_ctx. Each probe has a two-second timeout; unavailable or invalid properties leave the configured capabilities unchanged. The request reuses backend authentication and extra headers, with redirects disabled.
  • Provider fixes: Recognizes llama.cpp’s “exceeds the available context size” error as context overflow. For Anthropic-format requests to inference-api.nvidia.com, it removes the unsupported tool strict field. These changes affect the shared client, so they merit review separately from the VGR wiring.
  • Tests: Add TOML construction and rejection cases, authenticated property discovery, provider compatibility checks, and an HTTP integration test covering local acceptance, cloud escalation, the existing selected-model response header, and advertised context length.

The configuration example and serving-mode descriptions are in docs/routing_algorithms/vgr_routing.md. #881 builds on this surface with agentic handoff controls.

@gburachas
gburachas requested a review from a team as a code owner September 16, 2026 02:09
@github-actions

github-actions Bot commented Sep 16, 2026 •

Copy link
Copy Markdown
PR Preview Action v1.8.1

🚀 View preview at
https://NVIDIA-NeMo.github.io/Switchyard/pr-preview/pr-719/

Built to branch gh-pages at 2026-09-30 19:23 UTC.
Preview will be ready when the GitHub Pages deployment is complete.

@afourniernv afourniernv reopened this Sep 27, 2026
Add the vgr route type to runner config, publish the local tier's live context
window through /v1/models, strip Anthropic strict tool fields on NVIDIA
Inference API, and document the route.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants