Skip to content
Draft
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
70 changes: 67 additions & 3 deletions docs/integrations/oh_my_pi.md
Original file line number Diff line number Diff line change
Expand Up @@ -29,7 +29,8 @@ providers:

- `auth: none` marks the provider as keyless. Switchyard ignores client keys unless an
LLM client sets `forward_auth = true`. Without `auth: none`, `omp` refuses to send a
request.
request. To send your gateway key through a route that forwards it, see
[Forwarded keys](#forwarded-keys).
- `models[].id` must equal a route `id` from your TOML file. `contextWindow` and
`maxTokens` set `omp`'s compaction limit and output cap. `reasoning: true` turns on the
`--thinking` flag.
Expand Down Expand Up @@ -68,8 +69,8 @@ the default model.
## Check the routing

The checks in [Use Switchyard with pi](pi.md#check-the-routing) work the same way for
`omp`. On the Chat Completions API, `omp` sends no session header for a custom provider,
so the routing log records `"session_id": null`. Routes with
`omp`. On the Chat Completions and Responses APIs, `omp` sends no session header for a
custom provider, so the routing log records `"session_id": null`. Routes with
`classify_trigger = "user_turn"` or `"new_session"`, advisor budgets, and the stage
router's `capable_hold_turns` then treat each request as its own session. If you need
per-session routing, use `anthropic-messages`. On that API `omp` sends the
Expand All @@ -94,3 +95,66 @@ Port 4000 is also the default port for `omp`'s `litellm` provider and for
as a LiteLLM proxy. In that case, run Switchyard on another port or set
`LITELLM_BASE_URL`. Set `cost` on the model entry if you want `omp` to show a non-zero
cost.

## Claude targets behind an OpenAI-compatible gateway

The pi guide's section
[Claude targets behind an OpenAI-compatible gateway](pi.md#claude-targets-behind-an-openai-compatible-gateway)
applies to `omp` too. Use its table to choose the Claude LLM client by who holds the
gateway key. Switchyard calls the gateway endpoint that matches the target's LLM client
`format`, whatever `api` `omp` uses, so the `format` decides whether the gateway caches
the prompt. Use `format = "openai_chat"` or `"anthropic_messages"` for Claude targets,
because a gateway may not cache Claude prompts on `/v1/responses`. To estimate what the
requests cost from the routing log, see [Estimate the cost](pi.md#estimate-the-cost) in
the pi guide. This section covers what differs for `omp`, checked with Oh My Pi 18.2.11.

### Thinking

Thinking depends on both the LLM client `format` and the `api`:

- On `openai-completions`, `omp` sends `reasoning_effort`, and on `openai-responses` it
sends `reasoning.effort`. Switchyard passes the effort to an `openai_chat` target as
`reasoning_effort` and to an `openai_responses` target as `reasoning.effort`. Some
gateways turn either field into a thinking setting that Claude Opus 5.5 and Sonnet 5
refuse with HTTP 400. Switchyard turns the effort into adaptive thinking only for an
`anthropic_messages` target. The pi guide's [Thinking](pi.md#thinking) section shows
how to remove the field with `omit_body_fields` instead.
- On `anthropic-messages`, Switchyard sends `omp`'s own `thinking` settings to an
`anthropic_messages` target unchanged. `omp` does not recognize a route id such as
`switchyard` as a Claude model, so with thinking on it sends
`thinking: {type: "enabled"}`, and Claude Opus 5.5 and Sonnet 5 return HTTP 400. Tell
`omp` to use adaptive thinking on the model entry:

```yaml
- id: switchyard
reasoning: true
thinking:
mode: anthropic-adaptive
efforts: [low, medium, high]
```

`omp` requires `efforts` next to `mode`. It then sends `thinking: {type: "adaptive"}`
and `output_config.effort`.

### Forwarded keys

To forward `omp`'s key, remove `auth: none` and set `apiKey: GATEWAY_API_KEY`, the name
of the environment variable that holds your gateway key. Unlike pi, `omp` reads the name
without a leading `$`.

A route that forwards the key to an `anthropic_messages` client accepts requests only on
`/v1/messages`, and it cannot also forward the key to an `openai_chat` or
`openai_responses` client. So a route with a GPT judge on `openai_responses` cannot
forward the caller's key to both the judge and Claude targets on `anthropic_messages`.
Choose one of two setups:

- Forward the key to every LLM client, and keep the Claude targets on `openai_chat` with
`omit_body_fields = ["reasoning_effort"]`. The Claude models then think at their
default effort, and `--thinking` has no effect on them.
- Forward the key only to the GPT judge, and give the Claude targets an
`anthropic_messages` client with `api_key_env`, so they use a server-owned key.
Switchyard then turns the effort into adaptive thinking.

Both setups forward the key to an OpenAI-format LLM client, so the route accepts only
`/v1/chat/completions` and `/v1/responses` and returns HTTP 400 on `/v1/messages`. Use
`openai-completions` or `openai-responses`.
162 changes: 161 additions & 1 deletion docs/integrations/pi.md
Original file line number Diff line number Diff line change
Expand Up @@ -41,7 +41,9 @@ with route id `switchyard`.

- `models[].id` must equal a route `id` from your TOML file. Add one entry per route.
- `apiKey` is a placeholder. Switchyard ignores client keys unless an LLM client sets
`forward_auth = true`. pi still needs some value here before it lists the model.
`forward_auth = true`. pi still needs some value here before it lists the model. To
send your gateway key through a route that forwards it, see
[Forwarded keys](#forwarded-keys).
- `contextWindow` and `maxTokens` set pi's compaction limit and output cap. pi does not
read these values from the server. Use the smallest context window among the route's
targets.
Expand Down Expand Up @@ -96,3 +98,161 @@ clients keep the local id `switchyard` instead.
Set `cost` on the model entry if you want pi to show a non-zero cost.
[`benchmark/run-baseline.sh`](../../benchmark/README.md) runs Terminal-Bench tasks with
pi through Switchyard when you pass `--agent pi`.

## Claude targets behind an OpenAI-compatible gateway

Some gateways, such as a LiteLLM proxy, serve Claude models on `/v1/chat/completions`,
`/v1/responses`, and `/v1/messages` with one API key. Switchyard calls the endpoint that
matches the target's LLM client `format`, whatever `api` pi uses. On the gateway tested
for this page, the `format` decided whether Claude prompts were cached and whether pi's
`--thinking` level worked.

Choose the Claude LLM client by who holds the gateway key:

| Who holds the gateway key | Claude LLM client | Result |
|---|---|---|
| The server, through `api_key_env` | `format = "anthropic_messages"` | Prompt caching and pi's `--thinking` level both work. Every caller's Claude requests use the server-owned key. |
| pi sends it as `apiKey`, and the Claude LLM client forwards it with `forward_auth = true` | `format = "openai_chat"`, with `omit_body_fields = ["reasoning_effort"]` on the target | Prompt caching works. pi's `--thinking` level has no effect, and Claude thinks at its default effort. |

Do not use `format = "openai_responses"` for Claude targets on such a gateway. The
gateway tested for this page never cached Claude prompts on `/v1/responses`, and it
returned HTTP 400 when thinking was on (see [Thinking](#thinking)). In both setups, a GPT
judge on `openai_responses` can still use pi's forwarded key (see
[Forwarded keys](#forwarded-keys)).

### Prompt caching

On the LiteLLM gateway tested for this page, a repeated Claude prompt was read from the
cache on `/v1/chat/completions` and `/v1/messages`, but never on `/v1/responses`. Every
`/v1/responses` request counted the whole prompt as uncached input.

To check your gateway, start the server with `--routing-log-file PATH` and send the same
prompt twice. Claude does not cache short prompts, so use a prompt of at least 5,000
tokens. In the tests for this page, prompts of about 5,000 tokens were cached on Claude
Opus 5.5 and Sonnet 5; the tests did not find Claude's exact minimum. Then read the
records:

```bash
jq -c '{route_id, prompt_tokens, cached_tokens, cache_creation_tokens}' PATH
```

If the gateway caches the prompt, the first record shows it in `cache_creation_tokens`
and the second shows `cached_tokens` close to `prompt_tokens`. If the second record shows
`"cached_tokens": 0`, the gateway read nothing from the cache, and the whole prompt
counts as uncached input again.

#### Estimate the cost

The routing log records token counts, not prices. To estimate what a request cost,
multiply each token count in its record by the matching price, and add the results:

| Tokens in the record | Price |
|---|---|
| Uncached input: `prompt_tokens - cached_tokens - cache_creation_tokens` | Input price |
| `cached_tokens` | Cache-read price |
| `cache_creation_tokens` | Cache-write price |
| `completion_tokens` | Output price |

Anthropic's published pricing sets the cache-read price at 0.1 times the input price. It
sets the cache-write price at 1.25 times the input price for a 5-minute cache, or 2 times
for a 1-hour cache. The routing log does not record which cache lifetime the gateway
used, and a gateway may charge its own prices, so the result is an estimate, not the
gateway's bill.

Put your prices in a `prices.json` file, in USD per million tokens. Key each entry by the
`model` value from the routing log. The rates below are illustrative: they only follow
Anthropic's published ratios, with cache reads at 0.1 times and 5-minute cache writes at
1.25 times the input price. Replace them with your provider's current prices.

```json
{
"claude-opus-5-5": {"input": 10.00, "cache_read": 1.00, "cache_write": 12.50, "output": 50.00}
}
```

The command below prints one estimated cost per record. For a record whose model has no
entry in `prices.json`, it prints a warning instead of a cost:

```bash
jq -r --slurpfile prices prices.json '
. as $r
| ($prices[0][$r.model // ""]) as $p
| if $p == null then
"warning: no price for model \($r.model); add it to prices.json"
else
((($r.prompt_tokens // 0) - ($r.cached_tokens // 0) - ($r.cache_creation_tokens // 0)) * $p.input
+ ($r.cached_tokens // 0) * $p.cache_read
+ ($r.cache_creation_tokens // 0) * $p.cache_write
+ ($r.completion_tokens // 0) * $p.output) / 1000000
| "\($r.route_id) \($r.model) estimated $\(. * 1000000 | round / 1000000)"
end' PATH
```

### Thinking

Claude Opus 5.5 and Sonnet 5 accept only adaptive thinking. With `reasoning: true`, pi
sends `reasoning_effort`, even when you do not pass `--thinking`, and Switchyard
passes that field unchanged to an `openai_chat` target. Some gateways, including the one
tested for this page, turn `reasoning_effort` into Anthropic's older
`thinking: {type: "enabled"}` and return HTTP 400:

```text
"thinking.type.enabled" is not supported for this model. Use "thinking.type.adaptive" and "output_config.effort" to control thinking behavior.
```

An `openai_responses` target fails the same way. Switchyard sends the effort to it as
`reasoning.effort`, and the gateway returns the same error.

You have two options:

- Use an `anthropic_messages` LLM client. Switchyard turns pi's effort into
`thinking: {type: "adaptive"}` and `output_config.effort`, so pi's `--thinking` level
still applies.
- Keep the OpenAI-format LLM client and remove the effort field from requests to the
target with [`omit_body_fields`](../reference/toml_schema.md#targetsname). Use the
field name of the target's format: `reasoning_effort` on `openai_chat`, or `reasoning`
on `openai_responses`.

```toml
[targets.claude]
id = "claude-opus-5-5"
llm_client = "gateway_chat" # format = "openai_chat"
omit_body_fields = ["reasoning_effort"]
```

The request then succeeds, and Claude thinks at its default effort. pi's `--thinking`
level has no effect on this target.

### Forwarded keys

An LLM client with `forward_auth = true` sends the caller's key to the gateway. An LLM
client with `api_key_env` sends a server-owned key, which the server reads from an
environment variable (see
[`[llm_clients.<name>]`](../reference/toml_schema.md#llm_clientsname)). To forward pi's
key, replace the `apiKey` placeholder with the name of an environment variable that holds
your gateway key, with a leading `$`: `"apiKey": "$GATEWAY_API_KEY"`. Without the `$`,
pi sends the name itself as the key.

Two rules limit forwarding in one route:

- Every LLM client in the route that sets `forward_auth = true` must use the same API
family: `openai_chat` and `openai_responses`, or `anthropic_messages`. Otherwise the
server does not start and prints
`route <name> cannot forward both Anthropic and OpenAI caller credentials`. `<name>` is
the route's `[routes.<name>]` table key, not its `id`.
- A route that forwards the key to an `anthropic_messages` client accepts requests only on
`/v1/messages`, and pi should not use that API (see [Which request API](#which-request-api)).

So if the route forwards pi's key to its Claude targets, put them on `openai_chat` with
`omit_body_fields`.

To keep pi's `--thinking` level, give the Claude targets an `anthropic_messages` client
with `api_key_env`. Which endpoints the route accepts then depends on the route's other
LLM clients:

- If no LLM client in the route sets `forward_auth = true`, the route accepts every
request API.
- If an OpenAI-format LLM client forwards the key, for example a GPT judge on
`openai_responses`, the route accepts only `/v1/chat/completions` and `/v1/responses`
and returns HTTP 400 on `/v1/messages`. pi uses those two APIs, so this setup works
with pi.
1 change: 1 addition & 0 deletions docs/reference/toml_schema.md
Original file line number Diff line number Diff line change
Expand Up @@ -105,6 +105,7 @@ endpoint, before it calls an upstream.
| `llm_client` | Yes | — | Key under `[llm_clients]`. |
| `system_prompt` | No | unset | System prompt prepended when this target serves a completion. |
| `extra_body` | No | `{}` | Values merged into the upstream request when the request does not already set that key. |
| `omit_body_fields` | No | `[]` | Top-level fields removed from every request body that Switchyard sends to this target. Switchyard removes them after it translates the request to the LLM client's `format`, so use that format's field names, for example `reasoning_effort` on `openai_chat` or `reasoning` on `openai_responses`. Switchyard applies `extra_body` and `reasoning_effort` after the removal, so either can set a removed field again. |
| `reasoning_effort` | No | unset | Reasoning effort forced on every request to this target, replacing the value the caller sent (`reasoning.effort` on `openai_responses`, `reasoning_effort` on `openai_chat`). Rejected on `anthropic_messages` clients. Use it to run one target at a different effort than the client asked for, for example a strong tier at `max` behind a client that sends `high`. Targets with different effort settings need distinct model IDs when used within one route. Separate routes may use the same model ID with separate `llm_clients` entries (same endpoint, different name). |

Within one route, callable targets with the same model ID must use the same `llm_client`.
Expand Down
Loading