Skip to content

feat(server): integrate OpenAI Chat buffered proxy - #2414

Merged
Pouyanpi merged 27 commits into
developfrom
pouyanpi/openai-chat-buffered-integration-4
Oct 9, 2026
Merged

Pouyanpi merged 27 commits into
developfrom
pouyanpi/openai-chat-buffered-integration-4

Conversation

@Pouyanpi

@Pouyanpi Pouyanpi commented Sep 26, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Integrates the authoritative OpenAI Chat projections into the buffered provider-native proxy path and exports their guard contract.

Why

The integration must preserve original provider data while connecting exact guarded targets, checker execution, native errors, route ownership, and OpenAPI documentation. Its exported contract describes the models actually bound to the endpoint.

What Changed

  • Adds the guarded POST /v1/chat/completions endpoint and reserves its API-version path family.
  • Carries original bytes, validated payload, and exact guarded target through one request state.
  • Uses one OpenAI error mapping for runtime rendering and OpenAPI response documentation.
  • Preserves non-success provider errors and rejects incompatible successful responses.
  • Adds provider revision bindings and mocked HTTP integration coverage.
  • Adds a shared buffered exporter that reads the models bound to an existing endpoint, with format validation and an exact drift check for contracts/openai/_generated/chat-completions.buffered.guard.yaml.

Review Notes

Review this PR against #2413. The exported YAML is for inspection and documentation; request handling does not load it. The export contains request and buffered-response policy. The shared CLI consumes an existing endpoint declaration; Python models and bindings remain handwritten.

Streaming requests are recognized but rejected before dispatch. Replacement eligibility does not enable executing replacement outcomes, which remain unsupported here.

create_openai_chat_router(...) returns an APIRouter, not a standalone configured server. The embedding application supplies ContentChecker and HttpDispatch; concrete outbound transport, deployment composition, and trusted Rails configuration resolution remain separate work.

Validation

Offline experimental tests passed, including export and mocked HTTP integration tests. Pre-commit and the exported-contract drift check passed. No live provider calls were used.

AI Assistance

  • No AI tools were used.
  • AI tools were used; a human reviewed and can explain every change.

Codex assisted with implementation, tests, documentation, and stack maintenance.

Checklist

  • I've read the CONTRIBUTING guidelines.
  • This is maintainer-led work tracked internally.
  • The PR title follows the project commit convention.
  • Applicable documentation and tests are included.
  • Generated changelog files were not edited manually.
  • Automated and human review comments are addressed or answered.
  • The responsible reviewer or team is mentioned.

Stack Position

Part 3 of 3.

Stack Context

The Python projection models and runtime bindings are authoritative. This buffered stack declares their policy alongside the types and exports YAML for inspection and documentation. Request handling does not load the exported contract.

The open chain is #2411 → #2413 → #2414 → #2434 → #2435 → #2436 → #2437. #2412 is closed; its provider provenance and boundary documentation are included in #2413.

Streaming continues in #2434–#2437. The shared exporter currently emits request and buffered-response policy; streaming contract export remains separate work.

Review each PR against its listed base branch.

Order PR Branch Base
1 #2411 pouyanpi/guard-contract-foundation-1 develop
2 #2413 pouyanpi/openai-chat-buffered-projections-3 pouyanpi/guard-contract-foundation-1
3 #2414 pouyanpi/openai-chat-buffered-integration-4 pouyanpi/openai-chat-buffered-projections-3

Summary by CodeRabbit

  • New Features
    • Added guarded proxy support for OpenAI Chat Completions, including request and response inspection, provider-compatible error responses, and API revision validation.
    • Added buffered contract export to YAML, with options to write output or check existing artifacts for differences.
    • Added OpenAPI documentation for guarded routes, revision parameters, and provider error responses.
  • Documentation
    • Documented contract export options, generated artifact details, and current runtime limitations.

@github-actions github-actions Bot added size: XL status: needs triage New issues that have not yet been reviewed or categorized. labels Sep 26, 2026
@Pouyanpi Pouyanpi self-assigned this Sep 26, 2026
@Pouyanpi Pouyanpi added this to the v0.25.0 milestone Sep 26, 2026
@codecov

codecov Bot commented Sep 26, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 99.05482% with 5 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
...ls/server/experimental/provider/contract_export.py 96.70% 3 Missing ⚠️
...uardrails/server/experimental/provider/endpoint.py 97.82% 1 Missing ⚠️
.../server/experimental/provider/projection_policy.py 98.96% 1 Missing ⚠️

📢 Thoughts on this report? Let us know!

@Pouyanpi
Pouyanpi marked this pull request as ready for review September 26, 2026 20:41
@Pouyanpi Pouyanpi added status: triaged Triaged by a maintainer; eligible for automated review (CodeRabbit/Greptile). and removed status: needs triage New issues that have not yet been reviewed or categorized. labels Sep 26, 2026
@greptile-apps

greptile-apps Bot commented Sep 26, 2026 •

Copy link
Copy Markdown
Contributor

RetriggerConfidence Score: 5/5

[High impact] The PR appears safe to merge based on the reviewed changes.

Summary

The PR connects the OpenAI Chat projections to a buffered guarded proxy and exports their endpoint-bound contract.

  • Adds guarded route ownership, request preparation, response inspection, and OpenAI-compatible errors.
  • Adds contract export, OpenAPI declarations, and mocked integration coverage.

Diagram

%%{init: {'theme': 'neutral'}}%%
flowchart LR
  A[Chat request] --> B[Validate revision and JSON projection]
  B --> C[Input check]
  C --> D[Provider dispatch]
  D --> E[Validate buffered response projection]
  E --> F[Output check when text is present]
  F --> G[Relay provider response]
Loading

Reviews (18) · Last reviewed commit: "test(server): cover buffered integration..." · Reviewed by Greptile

Comment thread nemoguardrails/server/experimental/provider/transport.py Outdated
Comment thread nemoguardrails/server/experimental/providers/openai/errors.py
@Pouyanpi
Pouyanpi force-pushed the pouyanpi/openai-chat-buffered-integration-4 branch from 7162ba0 to 43c9841 Compare September 26, 2026 20:53
@Pouyanpi
Pouyanpi force-pushed the pouyanpi/openai-chat-buffered-integration-4 branch 2 times, most recently from e91c3e8 to c4126a8 Compare September 26, 2026 21:19
@Pouyanpi
Pouyanpi force-pushed the pouyanpi/openai-chat-buffered-integration-4 branch from c4126a8 to 4c14da9 Compare October 5, 2026 12:58
@Pouyanpi
Pouyanpi force-pushed the pouyanpi/openai-chat-buffered-integration-4 branch from 4c14da9 to dad6b4a Compare October 5, 2026 14:58
@Pouyanpi
Pouyanpi force-pushed the pouyanpi/openai-chat-buffered-integration-4 branch from dad6b4a to 0c5c264 Compare October 6, 2026 11:05
Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
@Pouyanpi
Pouyanpi force-pushed the pouyanpi/openai-chat-buffered-integration-4 branch from 43ef5a4 to 852674e Compare October 9, 2026 09:58
After the closed Chat projections merged, the buffered integration
fixtures relied on an unreviewed top-level "opaque" member, which the
request and response roots now reject. The unsupported-response example
used tool_calls: [], which is now a valid empty value. The checked-in
contract export no longer matched the policy models.

Use reviewed opaque fields (metadata and usage) to show byte-preserving
relay, use an actual tool call as the unsupported response, and
regenerate the export with the contract export CLI. The export and the
runtime now both accept only null or empty annotations.

Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
The response projection accepts empty or null assistant content, as
OpenAI's schema allows, for example when generation stops at a length
limit or a content filter. The buffered proxy still read the target's
message, which raises for null content and runs output rails on an
empty string.

After shape validation, check has_text and return
ContentInspectionNotApplicable when there is no text, so the kernel
relays the original response without output checks. Validation still
runs first: a missing message or missing content key is rejected, and
null content with refusal, reasoning, tool, or citation content is
rejected.

Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
Clients such as the OpenAI SDK send Accept-Encoding: gzip, and the proxy
forwarded it. A provider that compressed the response then failed output
inspection with a 502, because encoded successful responses cannot be
inspected.

Replace Accept-Encoding with identity on guarded provider requests. A
provider that still encodes the response is rejected rather than relayed
without inspection.

Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
The OpenAI boundary summary said relaying responses without text was
left to the integration, which now does it. The buffered integration
docs did not describe the relay or the identity-encoding request, and
still showed a Poetry command for the contract export CLI.

State that empty or null content is relayed without output checks only
after the closed projection passes. Document the identity-encoding
request and the rejection of encoded responses. Use uv for the export
command, and note that exported annotations allow only an empty list.

Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
Declare Unicode case-alias rejection and closed-by-default unknown-field policy in the shared contract vocabulary. Regenerate the buffered Chat export and verify exported acceptance against the handwritten runtime under both trusted field policies.

Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
The guarded Chat glue replaced Accept-Encoding with identity on every
guarded request. The plain body is needed only when output checks run.
Without them the response is relayed unchanged, so the client's
compression preference can be kept.

Move the header change into the buffered kernel, which knows the
resolved inspection policy. It applies to any guarded operation, not
just OpenAI Chat. The guarded Chat glue forwards the request as before.
A still-encoded response under output inspection remains rejected.

Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
TestRaiseLLMCallException fails on some CI orderings. The exception
helper prefers model and provider names from llm_call_info_var. An
earlier test on the same worker leaves that variable set to a fake
model, so these tests read model=fake instead of their own test-model.

Reset llm_call_info_var around each test in the class. This is test
isolation only and does not change runtime behavior.

Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
Patch coverage reported about 51 changed lines without coverage. 24 of
them were the contract export CLI, which the tests exercised only in
subprocesses that coverage does not measure. The rest were endpoint,
revision binding, adapter, and error rendering branches.

Run the CLI in process for stdout, output, check, mismatch, and export
failure, while keeping the subprocess tests for the module entry point.
Cover endpoint route and binding declarations, revision binding values,
the guarded operation's revision check and its OpenAPI parameter, paired
request adapters, unique error status codes, HTTP failure rendering
without exception text, and policy models used as mapping values.

The lines left uncovered cannot be reached through the declarations:
the profile-mismatch checks (one capability profile exists), a
PolicyModel check that export performs first, the object-schema guard
(PolicyModel cannot be a root model), and the __main__ entry line.

Signed-off-by: Pouyanpi <13303554+Pouyanpi@users.noreply.github.com>
@Pouyanpi
Pouyanpi merged commit 4512d08 into develop Oct 9, 2026
19 checks passed
@Pouyanpi
Pouyanpi deleted the pouyanpi/openai-chat-buffered-integration-4 branch October 9, 2026 13:16
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size: XL status: triaged Triaged by a maintainer; eligible for automated review (CodeRabbit/Greptile).

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants