docs(integrations): explain Claude caching and thinking behind OpenAI-compatible gateways - #874
Draft
elyasmnvidian wants to merge 3 commits into
Draft
elyasmnvidian wants to merge 3 commits into
elyasmnvidian wants to merge 3 commits into
Conversation
Signed-off-by: Elyas Mehtabuddin <emehtabuddin@nvidia.com>
Signed-off-by: Elyas Mehtabuddin <emehtabuddin@nvidia.com>
|
…and key setup steps Signed-off-by: Elyas Mehtabuddin <emehtabuddin@nvidia.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
This PR changes only docs:
docs/integrations/pi.mdgets a new section, "Claude targets behind an OpenAI-compatible gateway". It opens with a two-row table that picks the Claude LLM client by who holds the gateway key. Three subsections follow:openai_chatoranthropic_messagesfor Claude targets and shows how to check a gateway with--routing-log-file: send the same prompt of at least 5,000 tokens twice, then readcached_tokensandcache_creation_tokens. Its Estimate the cost part prices those token counts with your provider's rates. Itsjqcommand reads aprices.jsonfile and prints a warning, not $0, for a model that has no price.anthropic_messagesLLM client, oromit_body_fieldson anopenai_chatoropenai_responsestarget."apiKey": "$GATEWAY_API_KEY") and which setups work when a route forwards it.docs/integrations/oh_my_pi.mdgets the same section. It links to the pi guide and covers what differs for Oh My Pi: thethinkingmode onanthropic-messages, andapiKey: GATEWAY_API_KEYin place ofauth: none. The "Check the routing" section now says thatompalso sends no session header on the Responses API.docs/reference/toml_schema.mdgets anomit_body_fieldsrow in the[targets.<name>]table. The config parser already accepts the key, but the reference did not list it.Why
On a gateway that serves Claude on
/v1/chat/completions,/v1/responses, and/v1/messageswith one API key, such as a LiteLLM proxy, the two OpenAI-format Claude setups each fail, even though Switchyard's config check accepts both:format = "openai_responses": the gateway never caches Claude prompts. On the gateway we tested, a 22,607-token prompt sent twice to/v1/responsesread 0 tokens from the cache both times, so every request counted the whole prompt as uncached input. The same prompt sent through anopenai_chattarget read 22,605 tokens from the cache on the second request.format = "openai_chat"with thinking on: every request returns HTTP 400. pi sendsreasoning_efforton Chat Completions; Oh My Pi sendsreasoning_efforton Chat Completions andreasoning.efforton Responses. Switchyard passes the effort to anopenai_chattarget asreasoning_effort, and the gateway returns HTTP 400 for Claude Opus 5.5 and Sonnet 5. The error shows that the gateway turned the effort into Anthropic's olderthinking: {type: "enabled"}, which these models refuse. Anopenai_responsestarget fails the same way withreasoning.effort:An
anthropic_messagesLLM client avoids both problems: Switchyard turns the effort intothinking: {type: "adaptive"}plusoutput_config.effort, and the gateway caches the prompt. This fix stops working when the route forwards the caller's key (forward_auth = true) to theanthropic_messagesclient:anthropic_messagesclient and an OpenAI-format client. The server does not start and printsroute <name> cannot forward both Anthropic and OpenAI caller credentials.anthropic_messagesclient accepts requests only on/v1/messages, and the pi guide tells pi users not to use that API.So a route with a GPT judge on
openai_responseshas two working setups:openai_chatwithomit_body_fields = ["reasoning_effort"]. The caller's thinking level then has no effect on Claude, and Claude thinks at its default effort.anthropic_messagesclient withapi_key_env. The caller's thinking level works, but every caller's Claude requests use the server-owned gateway key.Two client settings were also missing from the guides:
anthropic-messageswith thinking on, the gateway returns the same 400. Switchyard sendsomp'sthinkingobject to ananthropic_messagestarget unchanged, andompdoes not recognize a route ID as a Claude model, so it sendsthinking: {type: "enabled", budget_tokens: 8192}. Addingthinking: {mode: anthropic-adaptive, efforts: [low, medium, high]}to the model entry makesompsend adaptive thinking.apiKeyto a placeholder and Oh My Pi's provider toauth: none, so pi sends the placeholder andompsends no key. The guides now say how each client sends the real key.Relative cost of each Claude client
The routing log records token counts, not dollars. The guide estimates a request's cost by multiplying each token count by the matching price: uncached input (
prompt_tokens - cached_tokens - cache_creation_tokens) at the input price,cached_tokensat the cache-read price,cache_creation_tokensat the cache-write price, andcompletion_tokensat the output price. Anthropic's published pricing sets cache reads at 0.1 times the input price, and cache writes at 1.25 times it for a 5-minute cache or 2 times for a 1-hour cache.The table applies those ratios to the 5,765-token prompt in Run 3. Each value is in units of what the prompt costs as uncached input, with output tokens left out. It is an estimate, not the gateway's bill: a gateway may charge its own prices, and the routing log does not record which cache lifetime was used.
openai_responses(/v1/responses)openai_chat(/v1/chat/completions)anthropic_messages(/v1/messages)The cache lifetimes come from the usage objects the gateway returned in Run 5. On
/v1/chat/completions, the gateway wrote almost the whole prompt to the 1-hour cache. On the/v1/messagesrequests from Switchyard, it wrote the whole prompt to the 5-minute cache. With a 1-hour cache, two requests cost about 2.1 units instead of 2, and every later request saves about 0.9 units. With a 5-minute cache, two requests cost about 1.35 units instead of 2.Notes for reviewers
Start with the pi guide section.
Related PR. #873 lets a route that forwards the caller's key use both OpenAI-format and Anthropic-format LLM clients when all of them point at the same host. Whichever PR merges second should add that exception to the Forwarded keys sections of both guides.
The built site contains every new anchor that the guides link to (
pi/#thinking,pi/#forwarded-keys,pi/#estimate-the-cost,oh_my_pi/#forwarded-keys,toml_schema/#targetsname).Evidence
All runs used
switchyard-serverbuilt from this branch or its base against a LiteLLM gateway; the branch changes only docs. The gateway URL and model IDs below are placeholders:gateway.example.comreplaces the gateway host, andclaude-sonnet-5andclaude-opus-5-5replace the gateway's own model IDs.Run 1: Responses caller, 22,607-token prompt, caller's key forwarded
Every request went to
/v1/responses, the way Oh My Pi calls Switchyard, and used the same prompt of 22,607 tokens.The records written by
--routing-log-file, filtered withjq -c '{route_id, prompt_tokens, cached_tokens, cache_creation_tokens}':{"route_id":"claude-chat","prompt_tokens":22607,"cached_tokens":0,"cache_creation_tokens":22605} {"route_id":"claude-chat","prompt_tokens":22607,"cached_tokens":22605,"cache_creation_tokens":0} {"route_id":"claude-responses","prompt_tokens":22607,"cached_tokens":0,"cache_creation_tokens":0} {"route_id":"claude-responses","prompt_tokens":22607,"cached_tokens":0,"cache_creation_tokens":0} {"route_id":"claude-chat-omit","prompt_tokens":22607,"cached_tokens":22605,"cache_creation_tokens":0}The chat route wrote the prompt to the cache and read it on the second request. The responses route never did.
The same request with
"reasoning": {"effort": "medium"}:claude-chatreturned 400 with the"thinking.type.enabled" is not supported for this modelerror shown above.claude-chat-omitreturned 200. Its answer was correct, and its routing record is the last line above.An
anthropic_messagesclient with a server-owned key (api_key_env) also returned 200 for the same Responses request with"reasoning": {"effort": "medium"}.A
randomroute that forwards the caller's key to both anopenai_responsesclient and ananthropic_messagesclient fails--dry-run:Run 2: Chat Completions caller, 11,424-token prompt, server-owned key
Run 2 used the same three routes, but each client read a server-owned key from
api_key_envinstead of forwarding the caller's key. Every request went to/v1/chat/completions, pi's default API. The run sent 8 requests. The one 400 response wrote no routing record, so the log has 7 records:{"route_id":"claude-responses","prompt_tokens":11424,"cached_tokens":0,"cache_creation_tokens":0} {"route_id":"claude-responses","prompt_tokens":11424,"cached_tokens":0,"cache_creation_tokens":0} {"route_id":"claude-chat","prompt_tokens":11424,"cached_tokens":0,"cache_creation_tokens":11422} {"route_id":"claude-chat","prompt_tokens":11424,"cached_tokens":11422,"cache_creation_tokens":0} {"route_id":"claude-chat-omit","prompt_tokens":11424,"cached_tokens":11422,"cache_creation_tokens":0} {"route_id":"claude-chat","prompt_tokens":5134,"cached_tokens":0,"cache_creation_tokens":5132} {"route_id":"claude-chat","prompt_tokens":5134,"cached_tokens":5132,"cache_creation_tokens":0}reasoning_effortset tomedium,claude-chatreturned 400 with the same"thinking.type.enabled"error.claude-chat-omitreturned 200 and read 11,422 cached tokens.The same run checked the forwarded-key rules against the real binary. The server refused each bad config or request itself, so nothing reached the gateway:
[routes.mixed_held](routeidmixed-held) forwards the key to both families.--dry-runfails withroute mixed_held cannot forward both Anthropic and OpenAI caller credentials, so the name in the error is the table key, not the routeid.anthropic_messagesclient returned 400route claude-forwarded forwards an Anthropic login; call it through /v1/messageson both/v1/chat/completionsand/v1/responses.idmixed-heldfor a route with a forwardedopenai_responsesclient and ananthropic_messagesclient that uses a server-owned key. It passes--dry-run, and it returned 400route mixed-held forwards an OpenAI login; call it through /v1/chat/completions or /v1/responseson/v1/messages.Run 3:
anthropic_messagescaching and the Opus 5.5 400A later run used a server-owned key and passthrough routes to
claude-sonnet-5on each client format. Each route got its own 5,765-token prompt, sent twice from a Chat Completions caller and then twice from a Responses caller:{"route_id":"claude-chat","prompt_tokens":5765,"cached_tokens":0,"cache_creation_tokens":5763} {"route_id":"claude-chat","prompt_tokens":5765,"cached_tokens":5763,"cache_creation_tokens":0} {"route_id":"claude-chat","prompt_tokens":5765,"cached_tokens":0,"cache_creation_tokens":5763} {"route_id":"claude-chat","prompt_tokens":5765,"cached_tokens":5763,"cache_creation_tokens":0} {"route_id":"claude-responses","prompt_tokens":5765,"cached_tokens":0,"cache_creation_tokens":0} {"route_id":"claude-responses","prompt_tokens":5765,"cached_tokens":0,"cache_creation_tokens":0} {"route_id":"claude-responses","prompt_tokens":5765,"cached_tokens":0,"cache_creation_tokens":0} {"route_id":"claude-responses","prompt_tokens":5765,"cached_tokens":0,"cache_creation_tokens":0} {"route_id":"claude-messages","prompt_tokens":5765,"cached_tokens":0,"cache_creation_tokens":5763} {"route_id":"claude-messages","prompt_tokens":5765,"cached_tokens":5763,"cache_creation_tokens":0} {"route_id":"claude-messages","prompt_tokens":5765,"cached_tokens":0,"cache_creation_tokens":5763} {"route_id":"claude-messages","prompt_tokens":5765,"cached_tokens":5763,"cache_creation_tokens":0}On the
/v1/messagesrequests, Switchyard added onecache_controlmarker to the request body; the other two endpoints had none.The same run sent a Responses request with
"reasoning": {"effort": "medium"}to aclaude-responsesroute. Switchyard sent"reasoning": {"effort": "medium"}to/v1/responses, and the gateway returned the same 400. Withomit_body_fields = ["reasoning"]on that target, the request returned 200.Oh My Pi 18.2.11 with
api: anthropic-messages, a model entry without athinkingsetting, and--thinking medium, against ananthropic_messagestarget forclaude-opus-5-5.ompsent thisthinkingobject, and Switchyard sent it to the gateway unchanged:{"type":"enabled","budget_tokens":8192,"display":"summarized"}The gateway returned 400:
The same request to
claude-sonnet-5returned the same 400.Run 4: the documented client settings
pi 0.84.3 and Oh My Pi 18.2.11 ran with the configs that the edited guides show, against
claude-opus-5-5andclaude-sonnet-5, one short prompt each:omp -p --model switchyard/switchyard --thinking medium, withapi: anthropic-messagesandthinking: {mode: anthropic-adaptive, efforts: [low, medium, high]}on the model entry, to ananthropic_messagestarget forclaude-opus-5-5with a server-owned key. Result: 200. Switchyard sent"thinking": {"type": "adaptive"}and"output_config": {"effort": "medium"}to/v1/messages. The routing record carriedomp's session id and"cache_creation_tokens": 31171.pi -p --provider switchyard --model switchyard --thinking medium, with"apiKey": "$GATEWAY_API_KEY", to a route that forwards the key to anopenai_chattarget forclaude-sonnet-5withomit_body_fields = ["reasoning_effort"]. Result: 200. The request to the gateway carried anauthorizationheader and noreasoning_effort."apiKey": "GATEWAY_API_KEY"(no$) returned 401invalid_credentialfrom the gateway, because pi sent the variable name as the key./v1/messagesrequest to that forwarding route returned 400route switchyard forwards an OpenAI login; call it through /v1/chat/completions or /v1/responses, and nothing reached the gateway.Run 5: cache lifetime and the cost recipe
Two passthrough routes to
claude-sonnet-5used a server-owned key:claude-chat(openai_chat, withomit_body_fields = ["reasoning_effort"]) andclaude-messages(anthropic_messages). Each got its own prompt of about 8,100 tokens, sent twice to Switchyard's/v1/chat/completions. The usage objects that the gateway returned on the first request of each pair, trimmed to the cache fields:{"path":"/v1/chat/completions","prompt_tokens_details":{"cache_creation_tokens":8048,"cache_creation_token_details":{"ephemeral_5m_input_tokens":12,"ephemeral_1h_input_tokens":8036}}} {"path":"/v1/messages","cache_creation_input_tokens":8031,"cache_creation":{"ephemeral_5m_input_tokens":8031,"ephemeral_1h_input_tokens":0}}The second request of each pair read the whole written prefix from the cache (8,048 and 8,031 tokens).
The output of the guide's
jqcommand, run with jq 1.7.1 on routing records from Runs 3, 4, and 5.prices.jsonheld the guide's illustrative rates under the Sonnet 5 model ID only, so the command printed a warning for the Opus 5.5 record. The last two records come from Run 3 and have nocompletion_tokensfield: