From 84c50f44c90b04a58658b88772e3785bbd43dbb4 Mon Sep 17 00:00:00 2001 From: Elyas Mehtabuddin Date: Tue, 29 Sep 2026 23:13:25 -0700 Subject: [PATCH 1/3] docs(integrations): add Claude gateway caching and thinking notes Signed-off-by: Elyas Mehtabuddin --- docs/integrations/oh_my_pi.md | 22 +++++++++++++ docs/integrations/pi.md | 59 +++++++++++++++++++++++++++++++++++ 2 files changed, 81 insertions(+) diff --git a/docs/integrations/oh_my_pi.md b/docs/integrations/oh_my_pi.md index 0e0d9f89d..1ba5674e5 100644 --- a/docs/integrations/oh_my_pi.md +++ b/docs/integrations/oh_my_pi.md @@ -94,3 +94,25 @@ Port 4000 is also the default port for `omp`'s `litellm` provider and for as a LiteLLM proxy. In that case, run Switchyard on another port or set `LITELLM_BASE_URL`. Set `cost` on the model entry if you want `omp` to show a non-zero cost. + +### Claude targets behind an OpenAI-compatible gateway + +The [pi guide's notes on Claude targets behind a gateway](pi.md#claude-targets-behind-an-openai-compatible-gateway) +apply to `omp` too. The target's LLM client `format` decides prompt caching and +thinking, not the `api` that `omp` uses. Prefer `format = "openai_chat"` or +`"anthropic_messages"` for Claude targets, because a gateway may not cache Claude +prompts on `/v1/responses`. + +With thinking on, `omp` sends a reasoning effort on `openai-completions` and +`openai-responses`. Switchyard passes it to an `openai_chat` target as +`reasoning_effort`, which some gateways reject for Claude Opus 5.5 and Sonnet 5 with +HTTP 400. Switchyard turns the effort into adaptive thinking only when it translates a +Chat Completions or Responses request for an `anthropic_messages` target. It forwards an +`anthropic-messages` request to that target with `omp`'s own `thinking` settings. + +A route that forwards the caller's key to an `anthropic_messages` client accepts +requests only on `/v1/messages`, and it cannot also forward the key to an +`openai_chat` or `openai_responses` client. If such a route also needs an OpenAI-format +target, such as a GPT judge on `openai_responses`, keep the Claude targets on +`openai_chat` with `omit_body_fields = ["reasoning_effort"]`. The models then think at +their default effort, and `--thinking` has no effect on them. diff --git a/docs/integrations/pi.md b/docs/integrations/pi.md index 1560d14a2..3af94fa95 100644 --- a/docs/integrations/pi.md +++ b/docs/integrations/pi.md @@ -96,3 +96,62 @@ clients keep the local id `switchyard` instead. Set `cost` on the model entry if you want pi to show a non-zero cost. [`benchmark/run-baseline.sh`](../../benchmark/README.md) runs Terminal-Bench tasks with pi through Switchyard when you pass `--agent pi`. + +### Claude targets behind an OpenAI-compatible gateway + +Some gateways, such as a LiteLLM proxy, serve Claude models on `/v1/chat/completions`, +`/v1/responses`, and `/v1/messages` with one API key. The target's LLM client `format` +decides which endpoint Switchyard calls. That choice can change prompt caching and +thinking for Claude. + +**Prompt caching.** On the LiteLLM gateway tested for this page, a repeated Claude prompt +was read from the cache on `/v1/chat/completions` and `/v1/messages`, but never on +`/v1/responses`. Every request to `/v1/responses` paid for the whole prompt again. Use +`format = "openai_chat"` or `"anthropic_messages"` for Claude targets. To check your +gateway, start the server with `--routing-log-file PATH` and send the same long prompt +twice. Claude does not cache short prompts, so use one of at least 5,000 tokens. Then +read the records: + +```bash +jq -c '{route_id, prompt_tokens, cached_tokens, cache_creation_tokens}' PATH +``` + +If the gateway caches the prompt, the first record shows it in `cache_creation_tokens` +and the second shows `cached_tokens` close to `prompt_tokens`. Writing the cache makes +the first request cost more, and each later request that reuses the prompt costs much +less. If the second record shows `"cached_tokens": 0`, the gateway billed the whole +prompt again. + +**Thinking.** Claude Opus 5.5 and Sonnet 5 accept only adaptive thinking. Some gateways +turn the Chat Completions `reasoning_effort` field into Anthropic's older +`thinking: {type: "enabled"}`, and the model then returns HTTP 400: + +```text +"thinking.type.enabled" is not supported for this model. Use "thinking.type.adaptive" and "output_config.effort" to control thinking behavior. +``` + +With `reasoning: true`, pi sends `reasoning_effort`, and Switchyard passes it to an +`openai_chat` target. You have two options: + +- Use an `anthropic_messages` target. Switchyard turns the requested effort into + `thinking: {type: "adaptive"}` and `output_config.effort`. +- Keep the `openai_chat` target and drop the field: + + ```toml + [targets.claude] + id = "claude-opus-5-5" + llm_client = "gateway_chat" # format = "openai_chat" + omit_body_fields = ["reasoning_effort"] + ``` + + The request then succeeds, and the model thinks at its default effort. pi's + `--thinking` level has no effect on this target. + +**Forwarded keys.** Every LLM client in one route that sets `forward_auth = true` must +use the same API family: `openai_chat` and `openai_responses`, or `anthropic_messages`. +Otherwise the server does not start and prints +`route cannot forward both Anthropic and OpenAI caller credentials`. A route that +forwards the caller's key to an `anthropic_messages` client also accepts requests only +on `/v1/messages`, and pi should not use that API. So when Switchyard forwards pi's key, +use `openai_chat` Claude targets with `omit_body_fields`. If the server holds the key +through `api_key_env`, an `anthropic_messages` target works with every request API. From dd145220aefb6361d68d208e790f915125dac440 Mon Sep 17 00:00:00 2001 From: Elyas Mehtabuddin Date: Tue, 29 Sep 2026 23:36:20 -0700 Subject: [PATCH 2/3] docs(integrations): clarify when each Claude gateway setup applies Signed-off-by: Elyas Mehtabuddin --- docs/integrations/oh_my_pi.md | 34 ++++++++++++++++++++++------------ docs/integrations/pi.md | 20 ++++++++++++++------ 2 files changed, 36 insertions(+), 18 deletions(-) diff --git a/docs/integrations/oh_my_pi.md b/docs/integrations/oh_my_pi.md index 1ba5674e5..e9a180c00 100644 --- a/docs/integrations/oh_my_pi.md +++ b/docs/integrations/oh_my_pi.md @@ -98,21 +98,31 @@ cost. ### Claude targets behind an OpenAI-compatible gateway The [pi guide's notes on Claude targets behind a gateway](pi.md#claude-targets-behind-an-openai-compatible-gateway) -apply to `omp` too. The target's LLM client `format` decides prompt caching and -thinking, not the `api` that `omp` uses. Prefer `format = "openai_chat"` or -`"anthropic_messages"` for Claude targets, because a gateway may not cache Claude -prompts on `/v1/responses`. +apply to `omp` too. The target's LLM client `format`, not the `api` that `omp` uses, +decides which gateway endpoint Switchyard calls and so whether the gateway caches the +prompt. Prefer `format = "openai_chat"` or `"anthropic_messages"` for Claude targets, +because a gateway may not cache Claude prompts on `/v1/responses`. Thinking depends on +both the `format` and the `api`, as the next paragraph explains. With thinking on, `omp` sends a reasoning effort on `openai-completions` and `openai-responses`. Switchyard passes it to an `openai_chat` target as -`reasoning_effort`, which some gateways reject for Claude Opus 5.5 and Sonnet 5 with -HTTP 400. Switchyard turns the effort into adaptive thinking only when it translates a -Chat Completions or Responses request for an `anthropic_messages` target. It forwards an -`anthropic-messages` request to that target with `omp`'s own `thinking` settings. +`reasoning_effort`. Some gateways turn that field into a thinking setting that Claude +Opus 5.5 and Sonnet 5 reject with HTTP 400. Switchyard turns the effort into adaptive +thinking only when it translates a Chat Completions or Responses request for an +`anthropic_messages` target. When `omp` uses `anthropic-messages`, Switchyard forwards +the request to that target with `omp`'s own `thinking` settings unchanged. A route that forwards the caller's key to an `anthropic_messages` client accepts requests only on `/v1/messages`, and it cannot also forward the key to an -`openai_chat` or `openai_responses` client. If such a route also needs an OpenAI-format -target, such as a GPT judge on `openai_responses`, keep the Claude targets on -`openai_chat` with `omit_body_fields = ["reasoning_effort"]`. The models then think at -their default effort, and `--thinking` has no effect on them. +`openai_chat` or `openai_responses` client. So a route with an OpenAI-format target, +such as a GPT judge on `openai_responses`, cannot forward the caller's key to +both the judge and Claude targets on `anthropic_messages`. Choose one of two setups: + +- Forward the key to every client, and keep the Claude targets on `openai_chat` with + `omit_body_fields = ["reasoning_effort"]`. The Claude models then think at their + default effort, and `--thinking` has no effect on them. +- Forward the key only to the GPT judge, and put the Claude targets on an + `anthropic_messages` client that reads a key held by the server from `api_key_env`. + Switchyard then turns the effort into adaptive thinking. The route accepts only + `/v1/chat/completions` and `/v1/responses`, so use `openai-completions` or + `openai-responses`. diff --git a/docs/integrations/pi.md b/docs/integrations/pi.md index 3af94fa95..b9aed795d 100644 --- a/docs/integrations/pi.md +++ b/docs/integrations/pi.md @@ -109,8 +109,8 @@ was read from the cache on `/v1/chat/completions` and `/v1/messages`, but never `/v1/responses`. Every request to `/v1/responses` paid for the whole prompt again. Use `format = "openai_chat"` or `"anthropic_messages"` for Claude targets. To check your gateway, start the server with `--routing-log-file PATH` and send the same long prompt -twice. Claude does not cache short prompts, so use one of at least 5,000 tokens. Then -read the records: +twice. Claude does not cache short prompts. A prompt of at least 5,000 tokens is a safe +size; it is not Claude's exact minimum. Then read the records: ```bash jq -c '{route_id, prompt_tokens, cached_tokens, cache_creation_tokens}' PATH @@ -150,8 +150,16 @@ With `reasoning: true`, pi sends `reasoning_effort`, and Switchyard passes it to **Forwarded keys.** Every LLM client in one route that sets `forward_auth = true` must use the same API family: `openai_chat` and `openai_responses`, or `anthropic_messages`. Otherwise the server does not start and prints -`route cannot forward both Anthropic and OpenAI caller credentials`. A route that +`route cannot forward both Anthropic and OpenAI caller credentials`, where +`` is the route's `[routes.]` table key, not its `id`. A route that forwards the caller's key to an `anthropic_messages` client also accepts requests only -on `/v1/messages`, and pi should not use that API. So when Switchyard forwards pi's key, -use `openai_chat` Claude targets with `omit_body_fields`. If the server holds the key -through `api_key_env`, an `anthropic_messages` target works with every request API. +on `/v1/messages`, and pi should not use that API. So when Switchyard forwards pi's key +to the Claude targets, use `openai_chat` Claude targets with `omit_body_fields`. + +An `anthropic_messages` client that reads the key from `api_key_env` works with every +request API, as long as no other LLM client in the route sets `forward_auth = true`. If +one does, the forwarded client limits the route's endpoints. For example, a route with a +forwarded `openai_responses` client accepts only `/v1/chat/completions` and +`/v1/responses`, and returns HTTP 400 on `/v1/messages`. pi uses those APIs, so such a +route can forward pi's key to a GPT judge on `openai_responses` while its Claude targets +use `anthropic_messages` with a key that the server holds. From 18d388a33647de0289b9649de40a3cfe4380ea23 Mon Sep 17 00:00:00 2001 From: Elyas Mehtabuddin Date: Thu, 1 Oct 2026 13:07:10 -0700 Subject: [PATCH 3/3] docs(integrations): estimate Claude cache costs and fix the thinking and key setup steps Signed-off-by: Elyas Mehtabuddin --- docs/integrations/oh_my_pi.md | 96 ++++++++++++------ docs/integrations/pi.md | 181 +++++++++++++++++++++++++--------- docs/reference/toml_schema.md | 1 + 3 files changed, 202 insertions(+), 76 deletions(-) diff --git a/docs/integrations/oh_my_pi.md b/docs/integrations/oh_my_pi.md index e9a180c00..69d9b6a7a 100644 --- a/docs/integrations/oh_my_pi.md +++ b/docs/integrations/oh_my_pi.md @@ -29,7 +29,8 @@ providers: - `auth: none` marks the provider as keyless. Switchyard ignores client keys unless an LLM client sets `forward_auth = true`. Without `auth: none`, `omp` refuses to send a - request. + request. To send your gateway key through a route that forwards it, see + [Forwarded keys](#forwarded-keys). - `models[].id` must equal a route `id` from your TOML file. `contextWindow` and `maxTokens` set `omp`'s compaction limit and output cap. `reasoning: true` turns on the `--thinking` flag. @@ -68,8 +69,8 @@ the default model. ## Check the routing The checks in [Use Switchyard with pi](pi.md#check-the-routing) work the same way for -`omp`. On the Chat Completions API, `omp` sends no session header for a custom provider, -so the routing log records `"session_id": null`. Routes with +`omp`. On the Chat Completions and Responses APIs, `omp` sends no session header for a +custom provider, so the routing log records `"session_id": null`. Routes with `classify_trigger = "user_turn"` or `"new_session"`, advisor budgets, and the stage router's `capable_hold_turns` then treat each request as its own session. If you need per-session routing, use `anthropic-messages`. On that API `omp` sends the @@ -95,34 +96,65 @@ as a LiteLLM proxy. In that case, run Switchyard on another port or set `LITELLM_BASE_URL`. Set `cost` on the model entry if you want `omp` to show a non-zero cost. -### Claude targets behind an OpenAI-compatible gateway - -The [pi guide's notes on Claude targets behind a gateway](pi.md#claude-targets-behind-an-openai-compatible-gateway) -apply to `omp` too. The target's LLM client `format`, not the `api` that `omp` uses, -decides which gateway endpoint Switchyard calls and so whether the gateway caches the -prompt. Prefer `format = "openai_chat"` or `"anthropic_messages"` for Claude targets, -because a gateway may not cache Claude prompts on `/v1/responses`. Thinking depends on -both the `format` and the `api`, as the next paragraph explains. - -With thinking on, `omp` sends a reasoning effort on `openai-completions` and -`openai-responses`. Switchyard passes it to an `openai_chat` target as -`reasoning_effort`. Some gateways turn that field into a thinking setting that Claude -Opus 5.5 and Sonnet 5 reject with HTTP 400. Switchyard turns the effort into adaptive -thinking only when it translates a Chat Completions or Responses request for an -`anthropic_messages` target. When `omp` uses `anthropic-messages`, Switchyard forwards -the request to that target with `omp`'s own `thinking` settings unchanged. - -A route that forwards the caller's key to an `anthropic_messages` client accepts -requests only on `/v1/messages`, and it cannot also forward the key to an -`openai_chat` or `openai_responses` client. So a route with an OpenAI-format target, -such as a GPT judge on `openai_responses`, cannot forward the caller's key to -both the judge and Claude targets on `anthropic_messages`. Choose one of two setups: - -- Forward the key to every client, and keep the Claude targets on `openai_chat` with +## Claude targets behind an OpenAI-compatible gateway + +The pi guide's section +[Claude targets behind an OpenAI-compatible gateway](pi.md#claude-targets-behind-an-openai-compatible-gateway) +applies to `omp` too. Use its table to choose the Claude LLM client by who holds the +gateway key. Switchyard calls the gateway endpoint that matches the target's LLM client +`format`, whatever `api` `omp` uses, so the `format` decides whether the gateway caches +the prompt. Use `format = "openai_chat"` or `"anthropic_messages"` for Claude targets, +because a gateway may not cache Claude prompts on `/v1/responses`. To estimate what the +requests cost from the routing log, see [Estimate the cost](pi.md#estimate-the-cost) in +the pi guide. This section covers what differs for `omp`, checked with Oh My Pi 18.2.11. + +### Thinking + +Thinking depends on both the LLM client `format` and the `api`: + +- On `openai-completions`, `omp` sends `reasoning_effort`, and on `openai-responses` it + sends `reasoning.effort`. Switchyard passes the effort to an `openai_chat` target as + `reasoning_effort` and to an `openai_responses` target as `reasoning.effort`. Some + gateways turn either field into a thinking setting that Claude Opus 5.5 and Sonnet 5 + refuse with HTTP 400. Switchyard turns the effort into adaptive thinking only for an + `anthropic_messages` target. The pi guide's [Thinking](pi.md#thinking) section shows + how to remove the field with `omit_body_fields` instead. +- On `anthropic-messages`, Switchyard sends `omp`'s own `thinking` settings to an + `anthropic_messages` target unchanged. `omp` does not recognize a route id such as + `switchyard` as a Claude model, so with thinking on it sends + `thinking: {type: "enabled"}`, and Claude Opus 5.5 and Sonnet 5 return HTTP 400. Tell + `omp` to use adaptive thinking on the model entry: + + ```yaml + - id: switchyard + reasoning: true + thinking: + mode: anthropic-adaptive + efforts: [low, medium, high] + ``` + + `omp` requires `efforts` next to `mode`. It then sends `thinking: {type: "adaptive"}` + and `output_config.effort`. + +### Forwarded keys + +To forward `omp`'s key, remove `auth: none` and set `apiKey: GATEWAY_API_KEY`, the name +of the environment variable that holds your gateway key. Unlike pi, `omp` reads the name +without a leading `$`. + +A route that forwards the key to an `anthropic_messages` client accepts requests only on +`/v1/messages`, and it cannot also forward the key to an `openai_chat` or +`openai_responses` client. So a route with a GPT judge on `openai_responses` cannot +forward the caller's key to both the judge and Claude targets on `anthropic_messages`. +Choose one of two setups: + +- Forward the key to every LLM client, and keep the Claude targets on `openai_chat` with `omit_body_fields = ["reasoning_effort"]`. The Claude models then think at their default effort, and `--thinking` has no effect on them. -- Forward the key only to the GPT judge, and put the Claude targets on an - `anthropic_messages` client that reads a key held by the server from `api_key_env`. - Switchyard then turns the effort into adaptive thinking. The route accepts only - `/v1/chat/completions` and `/v1/responses`, so use `openai-completions` or - `openai-responses`. +- Forward the key only to the GPT judge, and give the Claude targets an + `anthropic_messages` client with `api_key_env`, so they use a server-owned key. + Switchyard then turns the effort into adaptive thinking. + +Both setups forward the key to an OpenAI-format LLM client, so the route accepts only +`/v1/chat/completions` and `/v1/responses` and returns HTTP 400 on `/v1/messages`. Use +`openai-completions` or `openai-responses`. diff --git a/docs/integrations/pi.md b/docs/integrations/pi.md index b9aed795d..69a990b92 100644 --- a/docs/integrations/pi.md +++ b/docs/integrations/pi.md @@ -41,7 +41,9 @@ with route id `switchyard`. - `models[].id` must equal a route `id` from your TOML file. Add one entry per route. - `apiKey` is a placeholder. Switchyard ignores client keys unless an LLM client sets - `forward_auth = true`. pi still needs some value here before it lists the model. + `forward_auth = true`. pi still needs some value here before it lists the model. To + send your gateway key through a route that forwards it, see + [Forwarded keys](#forwarded-keys). - `contextWindow` and `maxTokens` set pi's compaction limit and output cap. pi does not read these values from the server. Use the smallest context window among the route's targets. @@ -97,45 +99,119 @@ Set `cost` on the model entry if you want pi to show a non-zero cost. [`benchmark/run-baseline.sh`](../../benchmark/README.md) runs Terminal-Bench tasks with pi through Switchyard when you pass `--agent pi`. -### Claude targets behind an OpenAI-compatible gateway +## Claude targets behind an OpenAI-compatible gateway Some gateways, such as a LiteLLM proxy, serve Claude models on `/v1/chat/completions`, -`/v1/responses`, and `/v1/messages` with one API key. The target's LLM client `format` -decides which endpoint Switchyard calls. That choice can change prompt caching and -thinking for Claude. - -**Prompt caching.** On the LiteLLM gateway tested for this page, a repeated Claude prompt -was read from the cache on `/v1/chat/completions` and `/v1/messages`, but never on -`/v1/responses`. Every request to `/v1/responses` paid for the whole prompt again. Use -`format = "openai_chat"` or `"anthropic_messages"` for Claude targets. To check your -gateway, start the server with `--routing-log-file PATH` and send the same long prompt -twice. Claude does not cache short prompts. A prompt of at least 5,000 tokens is a safe -size; it is not Claude's exact minimum. Then read the records: +`/v1/responses`, and `/v1/messages` with one API key. Switchyard calls the endpoint that +matches the target's LLM client `format`, whatever `api` pi uses. On the gateway tested +for this page, the `format` decided whether Claude prompts were cached and whether pi's +`--thinking` level worked. + +Choose the Claude LLM client by who holds the gateway key: + +| Who holds the gateway key | Claude LLM client | Result | +|---|---|---| +| The server, through `api_key_env` | `format = "anthropic_messages"` | Prompt caching and pi's `--thinking` level both work. Every caller's Claude requests use the server-owned key. | +| pi sends it as `apiKey`, and the Claude LLM client forwards it with `forward_auth = true` | `format = "openai_chat"`, with `omit_body_fields = ["reasoning_effort"]` on the target | Prompt caching works. pi's `--thinking` level has no effect, and Claude thinks at its default effort. | + +Do not use `format = "openai_responses"` for Claude targets on such a gateway. The +gateway tested for this page never cached Claude prompts on `/v1/responses`, and it +returned HTTP 400 when thinking was on (see [Thinking](#thinking)). In both setups, a GPT +judge on `openai_responses` can still use pi's forwarded key (see +[Forwarded keys](#forwarded-keys)). + +### Prompt caching + +On the LiteLLM gateway tested for this page, a repeated Claude prompt was read from the +cache on `/v1/chat/completions` and `/v1/messages`, but never on `/v1/responses`. Every +`/v1/responses` request counted the whole prompt as uncached input. + +To check your gateway, start the server with `--routing-log-file PATH` and send the same +prompt twice. Claude does not cache short prompts, so use a prompt of at least 5,000 +tokens. In the tests for this page, prompts of about 5,000 tokens were cached on Claude +Opus 5.5 and Sonnet 5; the tests did not find Claude's exact minimum. Then read the +records: ```bash jq -c '{route_id, prompt_tokens, cached_tokens, cache_creation_tokens}' PATH ``` If the gateway caches the prompt, the first record shows it in `cache_creation_tokens` -and the second shows `cached_tokens` close to `prompt_tokens`. Writing the cache makes -the first request cost more, and each later request that reuses the prompt costs much -less. If the second record shows `"cached_tokens": 0`, the gateway billed the whole -prompt again. +and the second shows `cached_tokens` close to `prompt_tokens`. If the second record shows +`"cached_tokens": 0`, the gateway read nothing from the cache, and the whole prompt +counts as uncached input again. + +#### Estimate the cost + +The routing log records token counts, not prices. To estimate what a request cost, +multiply each token count in its record by the matching price, and add the results: + +| Tokens in the record | Price | +|---|---| +| Uncached input: `prompt_tokens - cached_tokens - cache_creation_tokens` | Input price | +| `cached_tokens` | Cache-read price | +| `cache_creation_tokens` | Cache-write price | +| `completion_tokens` | Output price | + +Anthropic's published pricing sets the cache-read price at 0.1 times the input price. It +sets the cache-write price at 1.25 times the input price for a 5-minute cache, or 2 times +for a 1-hour cache. The routing log does not record which cache lifetime the gateway +used, and a gateway may charge its own prices, so the result is an estimate, not the +gateway's bill. + +Put your prices in a `prices.json` file, in USD per million tokens. Key each entry by the +`model` value from the routing log. The rates below are illustrative: they only follow +Anthropic's published ratios, with cache reads at 0.1 times and 5-minute cache writes at +1.25 times the input price. Replace them with your provider's current prices. + +```json +{ + "claude-opus-5-5": {"input": 10.00, "cache_read": 1.00, "cache_write": 12.50, "output": 50.00} +} +``` + +The command below prints one estimated cost per record. For a record whose model has no +entry in `prices.json`, it prints a warning instead of a cost: + +```bash +jq -r --slurpfile prices prices.json ' + . as $r + | ($prices[0][$r.model // ""]) as $p + | if $p == null then + "warning: no price for model \($r.model); add it to prices.json" + else + ((($r.prompt_tokens // 0) - ($r.cached_tokens // 0) - ($r.cache_creation_tokens // 0)) * $p.input + + ($r.cached_tokens // 0) * $p.cache_read + + ($r.cache_creation_tokens // 0) * $p.cache_write + + ($r.completion_tokens // 0) * $p.output) / 1000000 + | "\($r.route_id) \($r.model) estimated $\(. * 1000000 | round / 1000000)" + end' PATH +``` -**Thinking.** Claude Opus 5.5 and Sonnet 5 accept only adaptive thinking. Some gateways -turn the Chat Completions `reasoning_effort` field into Anthropic's older -`thinking: {type: "enabled"}`, and the model then returns HTTP 400: +### Thinking + +Claude Opus 5.5 and Sonnet 5 accept only adaptive thinking. With `reasoning: true`, pi +sends `reasoning_effort`, even when you do not pass `--thinking`, and Switchyard +passes that field unchanged to an `openai_chat` target. Some gateways, including the one +tested for this page, turn `reasoning_effort` into Anthropic's older +`thinking: {type: "enabled"}` and return HTTP 400: ```text "thinking.type.enabled" is not supported for this model. Use "thinking.type.adaptive" and "output_config.effort" to control thinking behavior. ``` -With `reasoning: true`, pi sends `reasoning_effort`, and Switchyard passes it to an -`openai_chat` target. You have two options: +An `openai_responses` target fails the same way. Switchyard sends the effort to it as +`reasoning.effort`, and the gateway returns the same error. + +You have two options: -- Use an `anthropic_messages` target. Switchyard turns the requested effort into - `thinking: {type: "adaptive"}` and `output_config.effort`. -- Keep the `openai_chat` target and drop the field: +- Use an `anthropic_messages` LLM client. Switchyard turns pi's effort into + `thinking: {type: "adaptive"}` and `output_config.effort`, so pi's `--thinking` level + still applies. +- Keep the OpenAI-format LLM client and remove the effort field from requests to the + target with [`omit_body_fields`](../reference/toml_schema.md#targetsname). Use the + field name of the target's format: `reasoning_effort` on `openai_chat`, or `reasoning` + on `openai_responses`. ```toml [targets.claude] @@ -144,22 +220,39 @@ With `reasoning: true`, pi sends `reasoning_effort`, and Switchyard passes it to omit_body_fields = ["reasoning_effort"] ``` - The request then succeeds, and the model thinks at its default effort. pi's - `--thinking` level has no effect on this target. - -**Forwarded keys.** Every LLM client in one route that sets `forward_auth = true` must -use the same API family: `openai_chat` and `openai_responses`, or `anthropic_messages`. -Otherwise the server does not start and prints -`route cannot forward both Anthropic and OpenAI caller credentials`, where -`` is the route's `[routes.]` table key, not its `id`. A route that -forwards the caller's key to an `anthropic_messages` client also accepts requests only -on `/v1/messages`, and pi should not use that API. So when Switchyard forwards pi's key -to the Claude targets, use `openai_chat` Claude targets with `omit_body_fields`. - -An `anthropic_messages` client that reads the key from `api_key_env` works with every -request API, as long as no other LLM client in the route sets `forward_auth = true`. If -one does, the forwarded client limits the route's endpoints. For example, a route with a -forwarded `openai_responses` client accepts only `/v1/chat/completions` and -`/v1/responses`, and returns HTTP 400 on `/v1/messages`. pi uses those APIs, so such a -route can forward pi's key to a GPT judge on `openai_responses` while its Claude targets -use `anthropic_messages` with a key that the server holds. + The request then succeeds, and Claude thinks at its default effort. pi's `--thinking` + level has no effect on this target. + +### Forwarded keys + +An LLM client with `forward_auth = true` sends the caller's key to the gateway. An LLM +client with `api_key_env` sends a server-owned key, which the server reads from an +environment variable (see +[`[llm_clients.]`](../reference/toml_schema.md#llm_clientsname)). To forward pi's +key, replace the `apiKey` placeholder with the name of an environment variable that holds +your gateway key, with a leading `$`: `"apiKey": "$GATEWAY_API_KEY"`. Without the `$`, +pi sends the name itself as the key. + +Two rules limit forwarding in one route: + +- Every LLM client in the route that sets `forward_auth = true` must use the same API + family: `openai_chat` and `openai_responses`, or `anthropic_messages`. Otherwise the + server does not start and prints + `route cannot forward both Anthropic and OpenAI caller credentials`. `` is + the route's `[routes.]` table key, not its `id`. +- A route that forwards the key to an `anthropic_messages` client accepts requests only on + `/v1/messages`, and pi should not use that API (see [Which request API](#which-request-api)). + +So if the route forwards pi's key to its Claude targets, put them on `openai_chat` with +`omit_body_fields`. + +To keep pi's `--thinking` level, give the Claude targets an `anthropic_messages` client +with `api_key_env`. Which endpoints the route accepts then depends on the route's other +LLM clients: + +- If no LLM client in the route sets `forward_auth = true`, the route accepts every + request API. +- If an OpenAI-format LLM client forwards the key, for example a GPT judge on + `openai_responses`, the route accepts only `/v1/chat/completions` and `/v1/responses` + and returns HTTP 400 on `/v1/messages`. pi uses those two APIs, so this setup works + with pi. diff --git a/docs/reference/toml_schema.md b/docs/reference/toml_schema.md index 8bf382a08..3e73dc872 100644 --- a/docs/reference/toml_schema.md +++ b/docs/reference/toml_schema.md @@ -105,6 +105,7 @@ endpoint, before it calls an upstream. | `llm_client` | Yes | — | Key under `[llm_clients]`. | | `system_prompt` | No | unset | System prompt prepended when this target serves a completion. | | `extra_body` | No | `{}` | Values merged into the upstream request when the request does not already set that key. | +| `omit_body_fields` | No | `[]` | Top-level fields removed from every request body that Switchyard sends to this target. Switchyard removes them after it translates the request to the LLM client's `format`, so use that format's field names, for example `reasoning_effort` on `openai_chat` or `reasoning` on `openai_responses`. Switchyard applies `extra_body` and `reasoning_effort` after the removal, so either can set a removed field again. | | `reasoning_effort` | No | unset | Reasoning effort forced on every request to this target, replacing the value the caller sent (`reasoning.effort` on `openai_responses`, `reasoning_effort` on `openai_chat`). Rejected on `anthropic_messages` clients. Use it to run one target at a different effort than the client asked for, for example a strong tier at `max` behind a client that sends `high`. Targets with different effort settings need distinct model IDs when used within one route. Separate routes may use the same model ID with separate `llm_clients` entries (same endpoint, different name). | Within one route, callable targets with the same model ID must use the same `llm_client`.