Skip to content

feat(server): add OpenAI Chat buffered projections - #2413

Merged
Pouyanpi merged 34 commits into
developfrom
pouyanpi/openai-chat-buffered-projections-3
Oct 9, 2026
Merged

Pouyanpi merged 34 commits into
developfrom
pouyanpi/openai-chat-buffered-projections-3

Conversation

@Pouyanpi

@Pouyanpi Pouyanpi commented Sep 26, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Adds authoritative OpenAI Chat request and buffered-response models with typed policy declarations, runtime bindings, and provider provenance.

Why

The Python models already own runtime acceptance. Their field policy and object-level coverage belong alongside those types, so the integration can derive metadata rather than maintain a separate YAML policy.

What Changed

  • Adds provider-neutral payload declarations, exact guarded targets, strict JSON parsing, and projection coverage checks.
  • Expresses guarded/constrained/disabled/opaque field policy in model annotations, with explicit defaults and ObjectPolicy for object-level metadata.
  • Derives reviewed-field inventories and guarded text locations from the model declarations; request/response bindings retain their runtime responsibilities.
  • Includes the pinned OpenAI source, boundary guide, and offline provenance tests formerly reviewed in feat(server): define OpenAI Chat contract #2412.
  • Provides payload-schema export support; the endpoint-bound operation export and artifact are in feat(server): integrate OpenAI Chat buffered proxy #2414.
  • Tests policy declarations, exact targeting, preservation, unknown-field handling, unsupported shapes, and import boundaries.

Review Notes

Review this PR against #2411. The provenance and boundary files from #2412 are now part of this diff; no separate handwritten operation contract remains.

Nullable response annotations remain nullable, and recognizing stream does not enable streaming. Format validation does not establish provider conformance or arbitrary Python/JSON Schema equivalence.

Review questions carried from #2412 remain open for explicit policy decisions; folding the PR does not resolve them:

Nullable stream content is handled by the stream declarations in #2435.

Validation

Offline experimental tests and pre-commit passed on this branch. Provider source metadata tests do not fetch the upstream document or verify its digest against downloaded bytes.

AI Assistance

  • No AI tools were used.
  • AI tools were used; a human reviewed and can explain every change.

Codex assisted with implementation, tests, documentation, and stack maintenance.

Checklist

  • I've read the CONTRIBUTING guidelines.
  • This is maintainer-led work tracked internally.
  • The PR title follows the project commit convention.
  • Applicable documentation and tests are included.
  • Generated changelog files were not edited manually.
  • Automated and human review comments are addressed or answered.
  • The responsible reviewer or team is mentioned.

Stack Position

Part 2 of 3.

Stack Context

The Python projection models and runtime bindings are authoritative. This buffered stack declares their policy alongside the types and exports YAML for inspection and documentation. Request handling does not load the exported contract.

The open chain is #2411 → #2413 → #2414 → #2434 → #2435 → #2436 → #2437. #2412 is closed; its provider provenance and boundary documentation are included in #2413.

Streaming continues in #2434–#2437. The shared exporter currently emits request and buffered-response policy; streaming contract export remains separate work.

Review each PR against its listed base branch.

Order PR Branch Base
1 #2411 pouyanpi/guard-contract-foundation-1 develop
2 #2413 pouyanpi/openai-chat-buffered-projections-3 pouyanpi/guard-contract-foundation-1
3 #2414 pouyanpi/openai-chat-buffered-integration-4 pouyanpi/openai-chat-buffered-projections-3

Summary by CodeRabbit

  • New Features

    • Added experimental support for validating and encoding JSON object payloads, with clear errors for invalid or unsupported input.
    • Added guarded request and response handling for buffered OpenAI Chat Completions, including single-text content validation and replacement controls.
    • Added documentation describing supported payloads, provider boundaries, and limitations.
  • Tests

    • Added offline checks for payload validation, provider contracts, and request and response behavior.

@github-actions github-actions Bot added status: needs triage New issues that have not yet been reviewed or categorized. size: XL labels Sep 26, 2026
@Pouyanpi Pouyanpi self-assigned this Sep 26, 2026
@Pouyanpi Pouyanpi added this to the v0.25.0 milestone Sep 26, 2026
@codecov

codecov Bot commented Sep 26, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@Pouyanpi
Pouyanpi marked this pull request as ready for review September 26, 2026 20:41
@Pouyanpi Pouyanpi added status: triaged Triaged by a maintainer; eligible for automated review (CodeRabbit/Greptile). and removed status: needs triage New issues that have not yet been reviewed or categorized. labels Sep 26, 2026
@greptile-apps

greptile-apps Bot commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

[Medium impact] This PR appears safe to merge based on the reviewed changes.

Summary

This PR adds buffered OpenAI Chat request and response projections, derives policy metadata and guarded text locations from typed models, and adds provider provenance and offline tests.

  • The changes since the previous review are documentation-only.
  • No new actionable issue was identified.

Reviews (14) · Last reviewed commit: "style: tighten docstrings" · Reviewed by Greptile

Comment thread nemoguardrails/server/experimental/_json_payload.py
Comment thread nemoguardrails/server/experimental/provider/payload.py Outdated
@Pouyanpi
Pouyanpi force-pushed the pouyanpi/openai-chat-buffered-projections-3 branch from e871266 to 65ea4a9 Compare September 26, 2026 20:53
Comment thread tests/server/experimental/test_openai_chat_projections.py Outdated
@Pouyanpi
Pouyanpi force-pushed the pouyanpi/openai-chat-buffered-projections-3 branch 2 times, most recently from 4547b0a to cde5161 Compare October 5, 2026 14:58
@Pouyanpi
Pouyanpi force-pushed the pouyanpi/openai-chat-buffered-projections-3 branch from cde5161 to 1973def Compare October 6, 2026 11:05
@Pouyanpi
Pouyanpi force-pushed the pouyanpi/openai-chat-buffered-projections-3 branch from 1973def to eae248d Compare October 7, 2026 09:00
@Pouyanpi
Pouyanpi changed the base branch from pouyanpi/openai-chat-contract-2 to pouyanpi/guard-contract-foundation-1 October 8, 2026 14:14
@Pouyanpi
Pouyanpi force-pushed the pouyanpi/openai-chat-buffered-projections-3 branch from 7056fe0 to 620d536 Compare October 8, 2026 17:52
@Pouyanpi
Pouyanpi force-pushed the pouyanpi/openai-chat-buffered-projections-3 branch from 620d536 to 850c36e Compare October 9, 2026 06:58
@coderabbitai

coderabbitai Bot commented Oct 9, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Repository: NVIDIA-NeMo/Guardrails/.coderabbit.yaml
  • Review profile: CHILL
  • Plan: Enterprise
  • Run ID: a5969ea5-c400-4029-8e05-d1934e0c2109
📥 Commits

Reviewing files that changed from the base of the PR and between 7de04a5 and 850c36e.

📒 Files selected for processing (19)
  • nemoguardrails/server/experimental/_json_payload.py
  • nemoguardrails/server/experimental/contracts/README.md
  • nemoguardrails/server/experimental/contracts/openai/README.md
  • nemoguardrails/server/experimental/contracts/openai/source.yaml
  • nemoguardrails/server/experimental/provider/payload.py
  • nemoguardrails/server/experimental/provider/projection_policy.py
  • nemoguardrails/server/experimental/provider/types.py
  • nemoguardrails/server/experimental/providers/__init__.py
  • nemoguardrails/server/experimental/providers/openai/__init__.py
  • nemoguardrails/server/experimental/providers/openai/chat_completions/README.md
  • nemoguardrails/server/experimental/providers/openai/chat_completions/__init__.py
  • nemoguardrails/server/experimental/providers/openai/chat_completions/request_binding.py
  • nemoguardrails/server/experimental/providers/openai/chat_completions/request_projection.py
  • nemoguardrails/server/experimental/providers/openai/chat_completions/response_binding.py
  • nemoguardrails/server/experimental/providers/openai/chat_completions/response_projection.py
  • tests/server/experimental/test_openai_chat_contract.py
  • tests/server/experimental/test_openai_chat_projections.py
  • tests/server/experimental/test_projection_policy.py
  • tests/server/experimental/test_provider_payload.py

Included review availability: This review used your included allowance. Your plan provides up to 12 included reviews per hour; 10 remain after this review.


📝 Walkthrough

Walkthrough

Adds strict JSON utilities and reusable payload projection policies. Defines OpenAI Chat Completions request and buffered-response projections, derives their contracts and guarded text locations, and adds documentation and tests for the boundary.

Changes

OpenAI buffered single-text boundary

Layer / File(s) Summary
JSON payload and guarded text handling
nemoguardrails/server/experimental/provider/types.py, nemoguardrails/server/experimental/_json_payload.py, nemoguardrails/server/experimental/provider/payload.py, tests/server/experimental/test_provider_payload.py
Adds strict JSON object parsing and encoding, payload contract validation, guarded text location and replacement utilities, and tests for scalar and array-based text locations.
Projection policy and schema derivation
nemoguardrails/server/experimental/provider/projection_policy.py, tests/server/experimental/test_projection_policy.py
Adds policy declarations and validation, nested-model coverage and contract derivation, schema export, and guarded text-location derivation. Tests cover policy constraints and schema behavior.
OpenAI request projection and binding
nemoguardrails/server/experimental/contracts/README.md, nemoguardrails/server/experimental/contracts/openai/*, nemoguardrails/server/experimental/providers/*, nemoguardrails/server/experimental/providers/openai/chat_completions/*, tests/server/experimental/test_openai_chat_contract.py, tests/server/experimental/test_openai_chat_projections.py
Adds the pinned provider source metadata, request projection and binding, boundary documentation, and tests for request validation, JSON parsing, metadata, and binding contracts.
OpenAI response projection and binding
nemoguardrails/server/experimental/providers/openai/chat_completions/response_*, tests/server/experimental/test_openai_chat_projections.py
Adds the guarded assistant response projection and binding. Tests cover supported response shapes and annotation-based replacement rules.

Priority: ➖ Normal

Estimated code review effort: 4 (Complex) | ~45 minutes

Change: Feature

Suggested reviewers: tgasser-nv

Merge Risk: ⚪ Minimal · up to 850c3

The staged projections have no established current integration failure. Confirm provider compatibility when connecting them to an endpoint.

🚥 Pre-merge checks | ✅ 5 | ❌ 1

❌ Failed checks (1 warning)

Check name Status Explanation Resolution
Docstring Coverage Warning Docstring coverage is 76.34% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 93 functions across 15 files. (4 skipped:… Write docstrings for the functions missing them to satisfy the coverage threshold.
✅ Passed checks (5 passed)
Check name Status Explanation
Linked Issues check Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check Passed Check skipped because no linked issues were found for this pull request.
Test Results For Major Changes Passed The PR contains major feature and refactoring changes, with 2,050 added lines across payload parsing, projection policies, OpenAI request/response models, bindings, and four new test modules. The desc…
Description Check Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check Passed The title clearly identifies the main change: adding buffered OpenAI Chat projections in the server.
Full details: Docstring Coverage

Explanation

Docstring coverage is 76.34% which is insufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 93 functions across 15 files. (4 skipped: 4 unsupported.)

  • Fix all pre-merge checks with AI
✨ Finishing Touches 💡 1
📝 Generate docstrings 💡
  • Commit to this branch
  • Create a new PR
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR
  • Autofix · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Comment @coderabbitai help to get the list of available commands.

Some OpenAI-compatible servers match JSON member names
case-insensitively; Go's encoding/json, used by Ollama, is one. A
payload such as {"content": "hello", "Content": "attack"} lets the proxy
inspect one value while the provider uses the other. A lone "Tools" can
also slip past the disabled "tools" field.

Reject case-folded duplicate members while parsing. Reject extra members
whose names match a reviewed field or opaque inventory entry only
case-insensitively.

Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
The buffered choice and assistant message allowed unreviewed members.
Members such as message.reasoning or provider_specific_fields.reasoning,
sent by OpenAI-compatible servers, could therefore carry text to the
client that output rails never inspected.

Forbid extra members on both objects so the response accepts exactly the
reviewed OpenAI fields. Fields specific to compatible servers belong to
their own reviewed extensions rather than to the OpenAI policy. The
export now marks both objects closed.

Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
The pinned OpenAI schema types response message tool_calls as an array
with no minimum length, so an empty list is a valid OpenAI value. It
carries no tool content. The disabled, null-only declaration was
stricter than the safety policy requires and rejected such responses
with a 502 when output inspection was on.

Constrain response tool_calls to null or an empty list; any tool call is
still rejected. The base OpenAI policy must accept this itself, because
compatible-server extensions may add fields but never loosen an OpenAI
field. Older vLLM releases send this value, which is how the gap was
noticed.

Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
Request logprobs and top_logprobs were forwarded as opaque settings,
while buffered response logprobs are null-only. A request with
logprobs: true therefore paid for an upstream call that then failed with
a 502 when output inspection was on.

Log probabilities carry token text that rails do not inspect. Constrain
request logprobs to false or null and top_logprobs to null, so the
request fails before dispatch.

Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
guarded() and constrained() passed any Pydantic constraint to Field.
Numeric constraints such as ge and le export as minimum and maximum,
which the guard contract field vocabulary does not define, so a future
declaration would export a document that fails contract validation.

Accept only min_length, max_length, and pattern, whose exported keywords
exist in the contract schema, and reject anything else when the field is
declared.

Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
The request and response bindings each repeated the OpenAI document URL,
revision, version, and digest from contracts/openai/source.yaml, and no
code read them.

Move the pin to providers/openai/source.py and add a test that it equals
source.yaml. The bindings no longer repeat it.

Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
The request root and its user message accepted unreviewed members.
OpenAI-compatible servers render several such members into the prompt
without input rails seeing them: llama.cpp copies chat_template_kwargs
over its template inputs, vLLM renders documents and
chat_template_kwargs.tools, and vLLM replaces the prompt with
kv_transfer_params.prompt_token_ids.

None of these fields is in the pinned OpenAI schema, so they are
compatible-server extensions rather than base policy. Every
CreateChatCompletionRequest property is already declared, so closing the
request rejects no valid OpenAI request. Declare the reviewed opaque
settings as fields and forbid extra members on the request and the user
message.

Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
The response root accepted unreviewed members. Compatible servers place
generated or prompt text there, for example vLLM prompt_text and
prompt_logprobs or llama.cpp __verbose.content, and output rails never
inspect it. That is not a bypass while a blocked response is replaced
whole, but it becomes one once text replacement is supported.

None of these members is in the pinned OpenAI schema. Every
CreateChatCompletionResponse property is already declared, so declare
the reviewed opaque metadata as fields and forbid extra members on the
response root.

Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
Response annotations accepted any items. OpenAI's url_citation items
carry provider text such as titles, and arbitrary items can carry any
text. Output rails inspect only message content, so this text reached
the client unchecked.

annotations is an OpenAI field, but a non-empty value can carry content
the rails do not inspect, so the base policy narrows it. Web search is
disabled on the request, and an OpenAI response without citations has
null or empty annotations. Constrain annotations to null or an empty
list. The declared replacement blocker is kept for when citation content
is supported.

Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
messages[0].name was opaque and accepted any value. OpenAI's schema
describes it as information that helps the model tell participants
apart, and compatible servers such as vLLM render it into the prompt.
Input rails inspect only the message content, so text in name reached
the model unchecked. A dict value also passed despite the schema typing
name as a string.

name is an OpenAI field, but a value can carry model-visible text that
the rails do not inspect, so the base policy narrows it. Disable it as a
null-only field.

Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
reasoning_effort was opaque and accepted any value, although OpenAI's
schema defines it as an enum. Compatible servers such as llama.cpp pass
the value into the chat template, so a free-form string could reach the
model without input rail inspection.

reasoning_effort is an OpenAI field, and a value outside the schema is
rejected. Constrain it to OpenAI's enumerated efforts or null.

Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
The parser rejected member names that differ only by case at every
nesting level. That also rejected legitimate opaque provider data, such
as metadata with both "Env" and "env" keys.

Every reviewed Chat object is now closed, so a case variant of a
reviewed field is rejected there as an unknown member, and the
case-variant check still covers open policy models. Keep rejecting exact
duplicates, which parsers resolve differently, and leave case variants
to the projection models.

Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
The OpenAI contract summary called response tool_calls the one exception
to null-only unsupported fields. Request logprobs and response
annotations are now constrained the same way, to values that OpenAI's
schema allows and that carry no content. List all three.

Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
text_location copied replacement_blocked_by into the runtime location
without checking it. A misspelled blocker, such as "annotatons", never
matched a provider member, so replacement stayed allowed even when the
blocking content was present.

Require the blocker to be a declared field or reviewed opaque field of
the subject's object when the location is derived.

Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
No test checked the exported Chat request and response policy against
guard-contract.schema.json, so a declaration that exports a keyword the
contract does not define, such as minimum from Field(ge=0), would go
unnoticed.

Wrap both exports in a minimal operation contract and validate it with
the contract schema. Also cover an unexportable constraint to show the
check catches it.

Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
Reasons, object sources, and opaque inventories were accepted in any
form and exported unchanged. A free-text reason, a source that already
had a JSON Pointer prefix, or a "*" opaque entry produced a document
that guard-contract.schema.json rejects, and nothing failed until that
document was validated.

Apply the contract's formats when the policy is declared. Reasons must
be structured, such as core_capability.tool_content. Sources must be
bare component names. Opaque inventories must name their fields.

Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
Earlier commits closed the OpenAI objects with a Pydantic extra="forbid"
config on each class. That bypassed the existing unknown-field
vocabulary: it removed the configurable markers that the contract
compiler and later provider contracts read, turned reviewed opaque
inventories into declared fields, and left the runtime FORBID path
unused. Meanwhile the runtime default was ALLOW, so configurable
objects were open, and roots were never checked.

Enforce the policy in PolicyModel for every object, including roots.
ObjectPolicy.unknown_fields is now "forbid" by default or
"configurable". A member outside the declared fields and opaque
inventory is rejected unless the object is configurable and validation
explicitly allows unknown content fields. That is reserved for trusted
configuration, and validation now defaults to FORBID. Case variants of
reviewed names are always rejected. The export derives
additionalProperties from the same policy.

The OpenAI objects return to declaring only policy. The roots use the
forbid default, while the user message, choice, and assistant message
keep their configurable markers. All of them are closed today.

Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
OpenAI's schema allows assistant content to be an empty string or null,
for example when generation stops at a length limit or a content
filter. The response projection required non-empty content, so such
valid responses failed with a 502 when output inspection was on.

Allow empty or null response content. A guarded subject that is nullable
or has no minimum length now yields a location that allows empty text,
and GuardedMessageTarget.has_text reports whether there is text to
check. This is safe only because the closed message keeps every other
reviewed field content-free: null content with refusal, tool call,
annotation, or reasoning text is still rejected, because that text is
untrusted model output the rails would not see. Request content still
requires non-empty text.

The buffered integration can then relay a response without text instead
of calling output rails; that one-line change belongs to the integration.

Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
Closed objects exported additionalProperties: false while listing their
reviewed opaque names only in the object-level opaque_fields metadata.
Read as a schema, the export therefore rejected members the runtime
accepts, such as model and temperature. A generator that lowers
additionalProperties: false to a forbidding model would reject them too.
The contract compiler already reads opaque names from opaque-classified
properties, not from that list.

Keep opaque values as runtime extras, but export each reviewed opaque
name as a property classified opaque, so additionalProperties describes
only unreviewed members. The replacement-blocker check now relies on
those properties.

Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
Checking the export against the contract format does not show that it
describes the members the runtime accepts. Add a differential test. It
derives acceptance from the exported Chat schemas alone and compares the
result with the handwritten runtime under FORBID and trusted ALLOW,
covering reviewed opaque names, unknown members on closed and
configurable objects, and case variants, including non-ASCII folds.
Removing the case rule from the derivation makes the test fail.

This does not establish compiler equivalence, which needs tests against
generated models. Document the case-variant rule as contract semantics
that runtimes and generators must implement: the contract vocabulary has
no propertyNames keyword to state it.

Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
The OpenAI boundary summary said the integration relays responses
without text. This layer only reports has_text; the buffered integration
does not use it yet. Say what the projection provides and leave the
relay to the integration.

Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
Replace the regex approximation with an export-derived casefold check. Cover multi-character folds, Kelvin and long-s variants, exact reviewed names, and trailing newlines under FORBID and trusted ALLOW. Clarify exact-name semantics and the legacy opaque inventory in the contract documentation.

Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
Patch coverage left 48 changed lines untested, mostly declaration and
mismatch branches in the provider-neutral payload and policy layers.

Cover coverage and contract metadata validation, contract and direction
mismatches, error descriptions that must not echo provider values, and
target reads and replacements. Also cover array-selection locations,
unresolved paths, text alternatives, payload bindings, alias and
nested-model checks, model name collisions, recursive exports,
non-string subjects, and the parser's non-finite, non-bytes, and
encoding paths.

Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
Keep generic policy invariants on local models, move Chat behavior and export comparisons into the OpenAI projection tests, and move complete Chat contract checks into the contract tests. Share the export-derived member-policy validator through a provider-neutral fixture while preserving all test names and parametrizations.

Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
@Pouyanpi
Pouyanpi merged commit 83100a6 into develop Oct 9, 2026
17 checks passed
@Pouyanpi
Pouyanpi deleted the pouyanpi/openai-chat-buffered-projections-3 branch October 9, 2026 09:58
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size: XL status: triaged Triaged by a maintainer; eligible for automated review (CodeRabbit/Greptile).

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants