diff --git a/docs/integrations/oh_my_pi.md b/docs/integrations/oh_my_pi.md index 0e0d9f89d..69d9b6a7a 100644 --- a/docs/integrations/oh_my_pi.md +++ b/docs/integrations/oh_my_pi.md @@ -29,7 +29,8 @@ providers: - `auth: none` marks the provider as keyless. Switchyard ignores client keys unless an LLM client sets `forward_auth = true`. Without `auth: none`, `omp` refuses to send a - request. + request. To send your gateway key through a route that forwards it, see + [Forwarded keys](#forwarded-keys). - `models[].id` must equal a route `id` from your TOML file. `contextWindow` and `maxTokens` set `omp`'s compaction limit and output cap. `reasoning: true` turns on the `--thinking` flag. @@ -68,8 +69,8 @@ the default model. ## Check the routing The checks in [Use Switchyard with pi](pi.md#check-the-routing) work the same way for -`omp`. On the Chat Completions API, `omp` sends no session header for a custom provider, -so the routing log records `"session_id": null`. Routes with +`omp`. On the Chat Completions and Responses APIs, `omp` sends no session header for a +custom provider, so the routing log records `"session_id": null`. Routes with `classify_trigger = "user_turn"` or `"new_session"`, advisor budgets, and the stage router's `capable_hold_turns` then treat each request as its own session. If you need per-session routing, use `anthropic-messages`. On that API `omp` sends the @@ -94,3 +95,66 @@ Port 4000 is also the default port for `omp`'s `litellm` provider and for as a LiteLLM proxy. In that case, run Switchyard on another port or set `LITELLM_BASE_URL`. Set `cost` on the model entry if you want `omp` to show a non-zero cost. + +## Claude targets behind an OpenAI-compatible gateway + +The pi guide's section +[Claude targets behind an OpenAI-compatible gateway](pi.md#claude-targets-behind-an-openai-compatible-gateway) +applies to `omp` too. Use its table to choose the Claude LLM client by who holds the +gateway key. Switchyard calls the gateway endpoint that matches the target's LLM client +`format`, whatever `api` `omp` uses, so the `format` decides whether the gateway caches +the prompt. Use `format = "openai_chat"` or `"anthropic_messages"` for Claude targets, +because a gateway may not cache Claude prompts on `/v1/responses`. To estimate what the +requests cost from the routing log, see [Estimate the cost](pi.md#estimate-the-cost) in +the pi guide. This section covers what differs for `omp`, checked with Oh My Pi 18.2.11. + +### Thinking + +Thinking depends on both the LLM client `format` and the `api`: + +- On `openai-completions`, `omp` sends `reasoning_effort`, and on `openai-responses` it + sends `reasoning.effort`. Switchyard passes the effort to an `openai_chat` target as + `reasoning_effort` and to an `openai_responses` target as `reasoning.effort`. Some + gateways turn either field into a thinking setting that Claude Opus 5.5 and Sonnet 5 + refuse with HTTP 400. Switchyard turns the effort into adaptive thinking only for an + `anthropic_messages` target. The pi guide's [Thinking](pi.md#thinking) section shows + how to remove the field with `omit_body_fields` instead. +- On `anthropic-messages`, Switchyard sends `omp`'s own `thinking` settings to an + `anthropic_messages` target unchanged. `omp` does not recognize a route id such as + `switchyard` as a Claude model, so with thinking on it sends + `thinking: {type: "enabled"}`, and Claude Opus 5.5 and Sonnet 5 return HTTP 400. Tell + `omp` to use adaptive thinking on the model entry: + + ```yaml + - id: switchyard + reasoning: true + thinking: + mode: anthropic-adaptive + efforts: [low, medium, high] + ``` + + `omp` requires `efforts` next to `mode`. It then sends `thinking: {type: "adaptive"}` + and `output_config.effort`. + +### Forwarded keys + +To forward `omp`'s key, remove `auth: none` and set `apiKey: GATEWAY_API_KEY`, the name +of the environment variable that holds your gateway key. Unlike pi, `omp` reads the name +without a leading `$`. + +A route that forwards the key to an `anthropic_messages` client accepts requests only on +`/v1/messages`, and it cannot also forward the key to an `openai_chat` or +`openai_responses` client. So a route with a GPT judge on `openai_responses` cannot +forward the caller's key to both the judge and Claude targets on `anthropic_messages`. +Choose one of two setups: + +- Forward the key to every LLM client, and keep the Claude targets on `openai_chat` with + `omit_body_fields = ["reasoning_effort"]`. The Claude models then think at their + default effort, and `--thinking` has no effect on them. +- Forward the key only to the GPT judge, and give the Claude targets an + `anthropic_messages` client with `api_key_env`, so they use a server-owned key. + Switchyard then turns the effort into adaptive thinking. + +Both setups forward the key to an OpenAI-format LLM client, so the route accepts only +`/v1/chat/completions` and `/v1/responses` and returns HTTP 400 on `/v1/messages`. Use +`openai-completions` or `openai-responses`. diff --git a/docs/integrations/pi.md b/docs/integrations/pi.md index 1560d14a2..69a990b92 100644 --- a/docs/integrations/pi.md +++ b/docs/integrations/pi.md @@ -41,7 +41,9 @@ with route id `switchyard`. - `models[].id` must equal a route `id` from your TOML file. Add one entry per route. - `apiKey` is a placeholder. Switchyard ignores client keys unless an LLM client sets - `forward_auth = true`. pi still needs some value here before it lists the model. + `forward_auth = true`. pi still needs some value here before it lists the model. To + send your gateway key through a route that forwards it, see + [Forwarded keys](#forwarded-keys). - `contextWindow` and `maxTokens` set pi's compaction limit and output cap. pi does not read these values from the server. Use the smallest context window among the route's targets. @@ -96,3 +98,161 @@ clients keep the local id `switchyard` instead. Set `cost` on the model entry if you want pi to show a non-zero cost. [`benchmark/run-baseline.sh`](../../benchmark/README.md) runs Terminal-Bench tasks with pi through Switchyard when you pass `--agent pi`. + +## Claude targets behind an OpenAI-compatible gateway + +Some gateways, such as a LiteLLM proxy, serve Claude models on `/v1/chat/completions`, +`/v1/responses`, and `/v1/messages` with one API key. Switchyard calls the endpoint that +matches the target's LLM client `format`, whatever `api` pi uses. On the gateway tested +for this page, the `format` decided whether Claude prompts were cached and whether pi's +`--thinking` level worked. + +Choose the Claude LLM client by who holds the gateway key: + +| Who holds the gateway key | Claude LLM client | Result | +|---|---|---| +| The server, through `api_key_env` | `format = "anthropic_messages"` | Prompt caching and pi's `--thinking` level both work. Every caller's Claude requests use the server-owned key. | +| pi sends it as `apiKey`, and the Claude LLM client forwards it with `forward_auth = true` | `format = "openai_chat"`, with `omit_body_fields = ["reasoning_effort"]` on the target | Prompt caching works. pi's `--thinking` level has no effect, and Claude thinks at its default effort. | + +Do not use `format = "openai_responses"` for Claude targets on such a gateway. The +gateway tested for this page never cached Claude prompts on `/v1/responses`, and it +returned HTTP 400 when thinking was on (see [Thinking](#thinking)). In both setups, a GPT +judge on `openai_responses` can still use pi's forwarded key (see +[Forwarded keys](#forwarded-keys)). + +### Prompt caching + +On the LiteLLM gateway tested for this page, a repeated Claude prompt was read from the +cache on `/v1/chat/completions` and `/v1/messages`, but never on `/v1/responses`. Every +`/v1/responses` request counted the whole prompt as uncached input. + +To check your gateway, start the server with `--routing-log-file PATH` and send the same +prompt twice. Claude does not cache short prompts, so use a prompt of at least 5,000 +tokens. In the tests for this page, prompts of about 5,000 tokens were cached on Claude +Opus 5.5 and Sonnet 5; the tests did not find Claude's exact minimum. Then read the +records: + +```bash +jq -c '{route_id, prompt_tokens, cached_tokens, cache_creation_tokens}' PATH +``` + +If the gateway caches the prompt, the first record shows it in `cache_creation_tokens` +and the second shows `cached_tokens` close to `prompt_tokens`. If the second record shows +`"cached_tokens": 0`, the gateway read nothing from the cache, and the whole prompt +counts as uncached input again. + +#### Estimate the cost + +The routing log records token counts, not prices. To estimate what a request cost, +multiply each token count in its record by the matching price, and add the results: + +| Tokens in the record | Price | +|---|---| +| Uncached input: `prompt_tokens - cached_tokens - cache_creation_tokens` | Input price | +| `cached_tokens` | Cache-read price | +| `cache_creation_tokens` | Cache-write price | +| `completion_tokens` | Output price | + +Anthropic's published pricing sets the cache-read price at 0.1 times the input price. It +sets the cache-write price at 1.25 times the input price for a 5-minute cache, or 2 times +for a 1-hour cache. The routing log does not record which cache lifetime the gateway +used, and a gateway may charge its own prices, so the result is an estimate, not the +gateway's bill. + +Put your prices in a `prices.json` file, in USD per million tokens. Key each entry by the +`model` value from the routing log. The rates below are illustrative: they only follow +Anthropic's published ratios, with cache reads at 0.1 times and 5-minute cache writes at +1.25 times the input price. Replace them with your provider's current prices. + +```json +{ + "claude-opus-5-5": {"input": 10.00, "cache_read": 1.00, "cache_write": 12.50, "output": 50.00} +} +``` + +The command below prints one estimated cost per record. For a record whose model has no +entry in `prices.json`, it prints a warning instead of a cost: + +```bash +jq -r --slurpfile prices prices.json ' + . as $r + | ($prices[0][$r.model // ""]) as $p + | if $p == null then + "warning: no price for model \($r.model); add it to prices.json" + else + ((($r.prompt_tokens // 0) - ($r.cached_tokens // 0) - ($r.cache_creation_tokens // 0)) * $p.input + + ($r.cached_tokens // 0) * $p.cache_read + + ($r.cache_creation_tokens // 0) * $p.cache_write + + ($r.completion_tokens // 0) * $p.output) / 1000000 + | "\($r.route_id) \($r.model) estimated $\(. * 1000000 | round / 1000000)" + end' PATH +``` + +### Thinking + +Claude Opus 5.5 and Sonnet 5 accept only adaptive thinking. With `reasoning: true`, pi +sends `reasoning_effort`, even when you do not pass `--thinking`, and Switchyard +passes that field unchanged to an `openai_chat` target. Some gateways, including the one +tested for this page, turn `reasoning_effort` into Anthropic's older +`thinking: {type: "enabled"}` and return HTTP 400: + +```text +"thinking.type.enabled" is not supported for this model. Use "thinking.type.adaptive" and "output_config.effort" to control thinking behavior. +``` + +An `openai_responses` target fails the same way. Switchyard sends the effort to it as +`reasoning.effort`, and the gateway returns the same error. + +You have two options: + +- Use an `anthropic_messages` LLM client. Switchyard turns pi's effort into + `thinking: {type: "adaptive"}` and `output_config.effort`, so pi's `--thinking` level + still applies. +- Keep the OpenAI-format LLM client and remove the effort field from requests to the + target with [`omit_body_fields`](../reference/toml_schema.md#targetsname). Use the + field name of the target's format: `reasoning_effort` on `openai_chat`, or `reasoning` + on `openai_responses`. + + ```toml + [targets.claude] + id = "claude-opus-5-5" + llm_client = "gateway_chat" # format = "openai_chat" + omit_body_fields = ["reasoning_effort"] + ``` + + The request then succeeds, and Claude thinks at its default effort. pi's `--thinking` + level has no effect on this target. + +### Forwarded keys + +An LLM client with `forward_auth = true` sends the caller's key to the gateway. An LLM +client with `api_key_env` sends a server-owned key, which the server reads from an +environment variable (see +[`[llm_clients.]`](../reference/toml_schema.md#llm_clientsname)). To forward pi's +key, replace the `apiKey` placeholder with the name of an environment variable that holds +your gateway key, with a leading `$`: `"apiKey": "$GATEWAY_API_KEY"`. Without the `$`, +pi sends the name itself as the key. + +Two rules limit forwarding in one route: + +- Every LLM client in the route that sets `forward_auth = true` must use the same API + family: `openai_chat` and `openai_responses`, or `anthropic_messages`. Otherwise the + server does not start and prints + `route cannot forward both Anthropic and OpenAI caller credentials`. `` is + the route's `[routes.]` table key, not its `id`. +- A route that forwards the key to an `anthropic_messages` client accepts requests only on + `/v1/messages`, and pi should not use that API (see [Which request API](#which-request-api)). + +So if the route forwards pi's key to its Claude targets, put them on `openai_chat` with +`omit_body_fields`. + +To keep pi's `--thinking` level, give the Claude targets an `anthropic_messages` client +with `api_key_env`. Which endpoints the route accepts then depends on the route's other +LLM clients: + +- If no LLM client in the route sets `forward_auth = true`, the route accepts every + request API. +- If an OpenAI-format LLM client forwards the key, for example a GPT judge on + `openai_responses`, the route accepts only `/v1/chat/completions` and `/v1/responses` + and returns HTTP 400 on `/v1/messages`. pi uses those two APIs, so this setup works + with pi. diff --git a/docs/reference/toml_schema.md b/docs/reference/toml_schema.md index 8bf382a08..3e73dc872 100644 --- a/docs/reference/toml_schema.md +++ b/docs/reference/toml_schema.md @@ -105,6 +105,7 @@ endpoint, before it calls an upstream. | `llm_client` | Yes | — | Key under `[llm_clients]`. | | `system_prompt` | No | unset | System prompt prepended when this target serves a completion. | | `extra_body` | No | `{}` | Values merged into the upstream request when the request does not already set that key. | +| `omit_body_fields` | No | `[]` | Top-level fields removed from every request body that Switchyard sends to this target. Switchyard removes them after it translates the request to the LLM client's `format`, so use that format's field names, for example `reasoning_effort` on `openai_chat` or `reasoning` on `openai_responses`. Switchyard applies `extra_body` and `reasoning_effort` after the removal, so either can set a removed field again. | | `reasoning_effort` | No | unset | Reasoning effort forced on every request to this target, replacing the value the caller sent (`reasoning.effort` on `openai_responses`, `reasoning_effort` on `openai_chat`). Rejected on `anthropic_messages` clients. Use it to run one target at a different effort than the client asked for, for example a strong tier at `max` behind a client that sends `high`. Targets with different effort settings need distinct model IDs when used within one route. Separate routes may use the same model ID with separate `llm_clients` entries (same endpoint, different name). | Within one route, callable targets with the same model ID must use the same `llm_client`.