Skip to content

docs(integrations): explain Claude caching and thinking behind OpenAI-compatible gateways - #874

Draft
elyasmnvidian wants to merge 3 commits into
mainfrom
emehtabuddin/switch-1632-claude-gateway-docs
Draft

elyasmnvidian wants to merge 3 commits into
mainfrom
emehtabuddin/switch-1632-claude-gateway-docs

Conversation

@elyasmnvidian

@elyasmnvidian elyasmnvidian commented Sep 30, 2026 •

Copy link
Copy Markdown
Contributor

What

This PR changes only docs:

  • docs/integrations/pi.md gets a new section, "Claude targets behind an OpenAI-compatible gateway". It opens with a two-row table that picks the Claude LLM client by who holds the gateway key. Three subsections follow:
    • Prompt caching recommends openai_chat or anthropic_messages for Claude targets and shows how to check a gateway with --routing-log-file: send the same prompt of at least 5,000 tokens twice, then read cached_tokens and cache_creation_tokens. Its Estimate the cost part prices those token counts with your provider's rates. Its jq command reads a prices.json file and prints a warning, not $0, for a model that has no price.
    • Thinking explains the HTTP 400 and gives two fixes: an anthropic_messages LLM client, or omit_body_fields on an openai_chat or openai_responses target.
    • Forwarded keys shows how pi sends the gateway key ("apiKey": "$GATEWAY_API_KEY") and which setups work when a route forwards it.
  • docs/integrations/oh_my_pi.md gets the same section. It links to the pi guide and covers what differs for Oh My Pi: the thinking mode on anthropic-messages, and apiKey: GATEWAY_API_KEY in place of auth: none. The "Check the routing" section now says that omp also sends no session header on the Responses API.
  • docs/reference/toml_schema.md gets an omit_body_fields row in the [targets.<name>] table. The config parser already accepts the key, but the reference did not list it.

Why

On a gateway that serves Claude on /v1/chat/completions, /v1/responses, and /v1/messages with one API key, such as a LiteLLM proxy, the two OpenAI-format Claude setups each fail, even though Switchyard's config check accepts both:

  • format = "openai_responses": the gateway never caches Claude prompts. On the gateway we tested, a 22,607-token prompt sent twice to /v1/responses read 0 tokens from the cache both times, so every request counted the whole prompt as uncached input. The same prompt sent through an openai_chat target read 22,605 tokens from the cache on the second request.

  • format = "openai_chat" with thinking on: every request returns HTTP 400. pi sends reasoning_effort on Chat Completions; Oh My Pi sends reasoning_effort on Chat Completions and reasoning.effort on Responses. Switchyard passes the effort to an openai_chat target as reasoning_effort, and the gateway returns HTTP 400 for Claude Opus 5.5 and Sonnet 5. The error shows that the gateway turned the effort into Anthropic's older thinking: {type: "enabled"}, which these models refuse. An openai_responses target fails the same way with reasoning.effort:

    "thinking.type.enabled" is not supported for this model. Use "thinking.type.adaptive" and "output_config.effort" to control thinking behavior.
    

An anthropic_messages LLM client avoids both problems: Switchyard turns the effort into thinking: {type: "adaptive"} plus output_config.effort, and the gateway caches the prompt. This fix stops working when the route forwards the caller's key (forward_auth = true) to the anthropic_messages client:

  • A route cannot forward the caller's key to both an anthropic_messages client and an OpenAI-format client. The server does not start and prints route <name> cannot forward both Anthropic and OpenAI caller credentials.
  • A route that forwards the key to an anthropic_messages client accepts requests only on /v1/messages, and the pi guide tells pi users not to use that API.

So a route with a GPT judge on openai_responses has two working setups:

  • Forward the caller's key to every LLM client, and keep Claude on openai_chat with omit_body_fields = ["reasoning_effort"]. The caller's thinking level then has no effect on Claude, and Claude thinks at its default effort.
  • Forward the key only to the judge, and put Claude on an anthropic_messages client with api_key_env. The caller's thinking level works, but every caller's Claude requests use the server-owned gateway key.

Two client settings were also missing from the guides:

  • When Oh My Pi calls Switchyard on anthropic-messages with thinking on, the gateway returns the same 400. Switchyard sends omp's thinking object to an anthropic_messages target unchanged, and omp does not recognize a route ID as a Claude model, so it sends thinking: {type: "enabled", budget_tokens: 8192}. Adding thinking: {mode: anthropic-adaptive, efforts: [low, medium, high]} to the model entry makes omp send adaptive thinking.
  • When a route forwards the key, the gateway returns 401. The Configure sections set pi's apiKey to a placeholder and Oh My Pi's provider to auth: none, so pi sends the placeholder and omp sends no key. The guides now say how each client sends the real key.

Relative cost of each Claude client

The routing log records token counts, not dollars. The guide estimates a request's cost by multiplying each token count by the matching price: uncached input (prompt_tokens - cached_tokens - cache_creation_tokens) at the input price, cached_tokens at the cache-read price, cache_creation_tokens at the cache-write price, and completion_tokens at the output price. Anthropic's published pricing sets cache reads at 0.1 times the input price, and cache writes at 1.25 times it for a 5-minute cache or 2 times for a 1-hour cache.

The table applies those ratios to the 5,765-token prompt in Run 3. Each value is in units of what the prompt costs as uncached input, with output tokens left out. It is an estimate, not the gateway's bill: a gateway may charge its own prices, and the routing log does not record which cache lifetime was used.

Claude LLM client First request Each later request that reuses the prompt
openai_responses (/v1/responses) about 1 about 1
openai_chat (/v1/chat/completions) about 2 (1-hour cache) about 0.1
anthropic_messages (/v1/messages) about 1.25 (5-minute cache) about 0.1

The cache lifetimes come from the usage objects the gateway returned in Run 5. On /v1/chat/completions, the gateway wrote almost the whole prompt to the 1-hour cache. On the /v1/messages requests from Switchyard, it wrote the whole prompt to the 5-minute cache. With a 1-hour cache, two requests cost about 2.1 units instead of 2, and every later request saves about 0.9 units. With a 5-minute cache, two requests cost about 1.35 units instead of 2.

Notes for reviewers

Start with the pi guide section.

Related PR. #873 lets a route that forwards the caller's key use both OpenAI-format and Anthropic-format LLM clients when all of them point at the same host. Whichever PR merges second should add that exception to the Forwarded keys sections of both guides.

The built site contains every new anchor that the guides link to (pi/#thinking, pi/#forwarded-keys, pi/#estimate-the-cost, oh_my_pi/#forwarded-keys, toml_schema/#targetsname).

Evidence

All runs used switchyard-server built from this branch or its base against a LiteLLM gateway; the branch changes only docs. The gateway URL and model IDs below are placeholders: gateway.example.com replaces the gateway host, and claude-sonnet-5 and claude-opus-5-5 replace the gateway's own model IDs.

Run 1: Responses caller, 22,607-token prompt, caller's key forwarded

Every request went to /v1/responses, the way Oh My Pi calls Switchyard, and used the same prompt of 22,607 tokens.

schema_version = 1

[llm_clients.gateway_chat]
format = "openai_chat"
base_url = "https://gateway.example.com/v1"
forward_auth = true

[llm_clients.gateway_responses]
format = "openai_responses"
base_url = "https://gateway.example.com/v1"
forward_auth = true

[llm_clients.gateway_chat_omit]
format = "openai_chat"
base_url = "https://gateway.example.com/v1"
forward_auth = true

[targets.sonnet_chat]
id = "claude-sonnet-5"
llm_client = "gateway_chat"

[targets.sonnet_responses]
id = "claude-sonnet-5"
llm_client = "gateway_responses"

[targets.sonnet_chat_omit]
id = "claude-sonnet-5"
llm_client = "gateway_chat_omit"
omit_body_fields = ["reasoning_effort"]

[routes.claude_chat]
id = "claude-chat"
type = "passthrough"
target = "sonnet_chat"

[routes.claude_responses]
id = "claude-responses"
type = "passthrough"
target = "sonnet_responses"

[routes.claude_chat_omit]
id = "claude-chat-omit"
type = "passthrough"
target = "sonnet_chat_omit"

The records written by --routing-log-file, filtered with jq -c '{route_id, prompt_tokens, cached_tokens, cache_creation_tokens}':

{"route_id":"claude-chat","prompt_tokens":22607,"cached_tokens":0,"cache_creation_tokens":22605}
{"route_id":"claude-chat","prompt_tokens":22607,"cached_tokens":22605,"cache_creation_tokens":0}
{"route_id":"claude-responses","prompt_tokens":22607,"cached_tokens":0,"cache_creation_tokens":0}
{"route_id":"claude-responses","prompt_tokens":22607,"cached_tokens":0,"cache_creation_tokens":0}
{"route_id":"claude-chat-omit","prompt_tokens":22607,"cached_tokens":22605,"cache_creation_tokens":0}

The chat route wrote the prompt to the cache and read it on the second request. The responses route never did.

The same request with "reasoning": {"effort": "medium"}:

  • claude-chat returned 400 with the "thinking.type.enabled" is not supported for this model error shown above.
  • claude-chat-omit returned 200. Its answer was correct, and its routing record is the last line above.

An anthropic_messages client with a server-owned key (api_key_env) also returned 200 for the same Responses request with "reasoning": {"effort": "medium"}.

A random route that forwards the caller's key to both an openai_responses client and an anthropic_messages client fails --dry-run:

route mixed cannot forward both Anthropic and OpenAI caller credentials

Run 2: Chat Completions caller, 11,424-token prompt, server-owned key

Run 2 used the same three routes, but each client read a server-owned key from api_key_env instead of forwarding the caller's key. Every request went to /v1/chat/completions, pi's default API. The run sent 8 requests. The one 400 response wrote no routing record, so the log has 7 records:

{"route_id":"claude-responses","prompt_tokens":11424,"cached_tokens":0,"cache_creation_tokens":0}
{"route_id":"claude-responses","prompt_tokens":11424,"cached_tokens":0,"cache_creation_tokens":0}
{"route_id":"claude-chat","prompt_tokens":11424,"cached_tokens":0,"cache_creation_tokens":11422}
{"route_id":"claude-chat","prompt_tokens":11424,"cached_tokens":11422,"cache_creation_tokens":0}
{"route_id":"claude-chat-omit","prompt_tokens":11424,"cached_tokens":11422,"cache_creation_tokens":0}
{"route_id":"claude-chat","prompt_tokens":5134,"cached_tokens":0,"cache_creation_tokens":5132}
{"route_id":"claude-chat","prompt_tokens":5134,"cached_tokens":5132,"cache_creation_tokens":0}
  • The responses route again never read from the cache, and the chat route wrote 11,422 tokens and then read 11,422.
  • With reasoning_effort set to medium, claude-chat returned 400 with the same "thinking.type.enabled" error. claude-chat-omit returned 200 and read 11,422 cached tokens.
  • A 5,134-token prompt was also cached: 5,132 tokens written, then 5,132 read.

The same run checked the forwarded-key rules against the real binary. The server refused each bad config or request itself, so nothing reached the gateway:

  • In one config, the table [routes.mixed_held] (route id mixed-held) forwards the key to both families. --dry-run fails with route mixed_held cannot forward both Anthropic and OpenAI caller credentials, so the name in the error is the table key, not the route id.
  • A route that forwards the key to an anthropic_messages client returned 400 route claude-forwarded forwards an Anthropic login; call it through /v1/messages on both /v1/chat/completions and /v1/responses.
  • A second, separate config reused the route id mixed-held for a route with a forwarded openai_responses client and an anthropic_messages client that uses a server-owned key. It passes --dry-run, and it returned 400 route mixed-held forwards an OpenAI login; call it through /v1/chat/completions or /v1/responses on /v1/messages.

Run 3: anthropic_messages caching and the Opus 5.5 400

A later run used a server-owned key and passthrough routes to claude-sonnet-5 on each client format. Each route got its own 5,765-token prompt, sent twice from a Chat Completions caller and then twice from a Responses caller:

{"route_id":"claude-chat","prompt_tokens":5765,"cached_tokens":0,"cache_creation_tokens":5763}
{"route_id":"claude-chat","prompt_tokens":5765,"cached_tokens":5763,"cache_creation_tokens":0}
{"route_id":"claude-chat","prompt_tokens":5765,"cached_tokens":0,"cache_creation_tokens":5763}
{"route_id":"claude-chat","prompt_tokens":5765,"cached_tokens":5763,"cache_creation_tokens":0}
{"route_id":"claude-responses","prompt_tokens":5765,"cached_tokens":0,"cache_creation_tokens":0}
{"route_id":"claude-responses","prompt_tokens":5765,"cached_tokens":0,"cache_creation_tokens":0}
{"route_id":"claude-responses","prompt_tokens":5765,"cached_tokens":0,"cache_creation_tokens":0}
{"route_id":"claude-responses","prompt_tokens":5765,"cached_tokens":0,"cache_creation_tokens":0}
{"route_id":"claude-messages","prompt_tokens":5765,"cached_tokens":0,"cache_creation_tokens":5763}
{"route_id":"claude-messages","prompt_tokens":5765,"cached_tokens":5763,"cache_creation_tokens":0}
{"route_id":"claude-messages","prompt_tokens":5765,"cached_tokens":0,"cache_creation_tokens":5763}
{"route_id":"claude-messages","prompt_tokens":5765,"cached_tokens":5763,"cache_creation_tokens":0}

On the /v1/messages requests, Switchyard added one cache_control marker to the request body; the other two endpoints had none.

The same run sent a Responses request with "reasoning": {"effort": "medium"} to a claude-responses route. Switchyard sent "reasoning": {"effort": "medium"} to /v1/responses, and the gateway returned the same 400. With omit_body_fields = ["reasoning"] on that target, the request returned 200.

Oh My Pi 18.2.11 with api: anthropic-messages, a model entry without a thinking setting, and --thinking medium, against an anthropic_messages target for claude-opus-5-5. omp sent this thinking object, and Switchyard sent it to the gateway unchanged:

{"type":"enabled","budget_tokens":8192,"display":"summarized"}

The gateway returned 400:

"thinking.type.enabled" is not supported for this model. Use "thinking.type.adaptive" and "output_config.effort" to control thinking behavior.

The same request to claude-sonnet-5 returned the same 400.

Run 4: the documented client settings

pi 0.84.3 and Oh My Pi 18.2.11 ran with the configs that the edited guides show, against claude-opus-5-5 and claude-sonnet-5, one short prompt each:

  • omp -p --model switchyard/switchyard --thinking medium, with api: anthropic-messages and thinking: {mode: anthropic-adaptive, efforts: [low, medium, high]} on the model entry, to an anthropic_messages target for claude-opus-5-5 with a server-owned key. Result: 200. Switchyard sent "thinking": {"type": "adaptive"} and "output_config": {"effort": "medium"} to /v1/messages. The routing record carried omp's session id and "cache_creation_tokens": 31171.
  • pi -p --provider switchyard --model switchyard --thinking medium, with "apiKey": "$GATEWAY_API_KEY", to a route that forwards the key to an openai_chat target for claude-sonnet-5 with omit_body_fields = ["reasoning_effort"]. Result: 200. The request to the gateway carried an authorization header and no reasoning_effort.
  • The same pi run with "apiKey": "GATEWAY_API_KEY" (no $) returned 401 invalid_credential from the gateway, because pi sent the variable name as the key.
  • A /v1/messages request to that forwarding route returned 400 route switchyard forwards an OpenAI login; call it through /v1/chat/completions or /v1/responses, and nothing reached the gateway.

Run 5: cache lifetime and the cost recipe

Two passthrough routes to claude-sonnet-5 used a server-owned key: claude-chat (openai_chat, with omit_body_fields = ["reasoning_effort"]) and claude-messages (anthropic_messages). Each got its own prompt of about 8,100 tokens, sent twice to Switchyard's /v1/chat/completions. The usage objects that the gateway returned on the first request of each pair, trimmed to the cache fields:

{"path":"/v1/chat/completions","prompt_tokens_details":{"cache_creation_tokens":8048,"cache_creation_token_details":{"ephemeral_5m_input_tokens":12,"ephemeral_1h_input_tokens":8036}}}
{"path":"/v1/messages","cache_creation_input_tokens":8031,"cache_creation":{"ephemeral_5m_input_tokens":8031,"ephemeral_1h_input_tokens":0}}

The second request of each pair read the whole written prefix from the cache (8,048 and 8,031 tokens).

The output of the guide's jq command, run with jq 1.7.1 on routing records from Runs 3, 4, and 5. prices.json held the guide's illustrative rates under the Sonnet 5 model ID only, so the command printed a warning for the Opus 5.5 record. The last two records come from Run 3 and have no completion_tokens field:

claude-chat claude-sonnet-5 estimated $0.10109
claude-chat claude-sonnet-5 estimated $0.008538
claude-messages claude-sonnet-5 estimated $0.100878
claude-messages claude-sonnet-5 estimated $0.008521
warning: no price for model claude-opus-5-5; add it to prices.json
switchyard claude-sonnet-5 estimated $0.028383
claude-messages claude-sonnet-5 estimated $0.072058
claude-messages claude-sonnet-5 estimated $0.005783

Signed-off-by: Elyas Mehtabuddin <emehtabuddin@nvidia.com>
Signed-off-by: Elyas Mehtabuddin <emehtabuddin@nvidia.com>
@github-actions

github-actions Bot commented Sep 30, 2026 •

Copy link
Copy Markdown
PR Preview Action v1.8.1

🚀 View preview at
https://NVIDIA-NeMo.github.io/Switchyard/pr-preview/pr-874/

Built to branch gh-pages at 2026-10-01 20:08 UTC.
Preview will be ready when the GitHub Pages deployment is complete.

…and key setup steps

Signed-off-by: Elyas Mehtabuddin <emehtabuddin@nvidia.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant