Repository navigation
No prompt caching on /provider/v1/chat/completions (GOAT): prompt_cache_key never sent, cache reads ~0% #132
Description
Activity
Thanks for the thorough trace — this is actionable, and you're right that the #83 conclusion doesn't cover the Provider API path.
Confirming the client-side half from our side: on the
openai-completionsbranch the key indeed never leaves the extension, for the three reasons you listed. So regardless of what the gateway does with it, we're currently sending nothing. That's on us.Before writing code, one question that decides between your options (a) and (b): does the
/chat/completionsgateway forwardprompt_cache_keywhen it IS present? The #83 evidence covers/responses(forwarded) and/alpha/generate(dropped);/chat/completionsis untested. Could you run a minimal keyed-vs-keyless pair directly againstPOST /provider/v1/chat/completions— same model, same prompt, immediate repeat, explicitprompt_cache_keyon the keyed leg — and post the cache-read tokens on the second call of each leg? If the gateway forwards it there too, option (b) (inject viaonPayload) becomes a small contained fix; if it drops it, (a) via the Responses wire (see also #124) is the way.Assigning myself for tracking — happy to build on your numbers.
Independent confirmation with billed numbers, recomputed against
src/pricing.ts($/M):meta/muse-spark-1.3-contributor (input 0.10 / output 0.20 / cacheRead 0.002):
- 89,352 in + 955 out → 0.00894 + 0.00019 = $0.00913 billed $0.00912 → 0% cache, to the cent
- 89,126 in + 97 out → $0.00893 billed $0.00892 → 0% cache
- 85,972 in + 84 out → $0.00861 billed $0.0086 → 0% cache
- one outlier: 89,252 in billed $0.00211, between fresh ($0.0093) and full-cache ($0.0006) → partial hit, i.e. the coin-flip behavior from Missing MODEL_COSTS for muse-spark-1.3 + no prompt_cache_key sent (cache collapse on Contrib) #83
deepseek/deepseek-v4.1-flash (input 0.15 / output 0.60 / cacheRead 0.003):
- 16,508 in + 4,074 out → cached input: 0.00005 + 0.00244 = $0.00249 billed $0.00253 → ~100% input cache hits
- the other three calls land on the cache-read price the same way
So the symptom matches this issue exactly: the Meta/Contrib pool bills full fresh-input price essentially every turn (~50x vs cache-read on long sessions), while the DeepSeek pool caches normally. Same pi setup, so this is pool-side stickiness (no key sent, no stable routing), not host misreading.
- addedbugSomething isn't workingSomething isn't workingproviderProvider-related issuesProvider-related issues
on Oct 3, 2026 Field confirmation, running the Responses wire locally (#133) on meta/muse-spark-1.3-contributor — same session, back-to-back turns (input $0.10/M, output $0.20/M, cacheRead $0.002/M):
Before (Chat Completions path):
- 108,185 in + 2,135 out billed $0.0112 = 1081850.10 + 21350.20 = $0.01125 → full fresh-input price, 0% cache
After (Responses path):
- 113,744 in + 127 out billed $0.000839 → input at ~$0.007/M effective, i.e. ~95% of the prefix at cache-read
- 113,505 in + 2,530 out billed $0.000844 → input at ~$0.003/M effective, i.e. ~99% of the prefix at cache-read
Same workload size, ~13x cheaper. The wire change fixes the real-world case, not just the probe.
Hi, first of all thanks for the extension — the GOAT transport work in 0.6.0 and the pricing fixes in 0.7.0 are much appreciated.
I'm opening this as a follow-up to #83. That issue nailed the pricing half (fixed in 0.7.0, confirmed) but the caching half was closed as "gateway ignores the key, nothing we can do client-side". I think there's new information that reopens it — specifically for the Provider API transport (
/provider/v1/chat/completions), which is what GOAT accounts actually use. The #83 probe tested/alpha/generate, but that's the Go fallback, not the GOAT path.Environment
pi-commandcode-provider0.7.3 (npm)provider, nevergenerate)meta/muse-spark-1.3-contributor(also affects the other Meta models and anything else on the openai-completions branch)Symptom
Running long agentic sessions (lots of tool calls, many turns), the reported cache usage on commandcode models sits permanently under 1%. Same workload on another OpenAI-compatible provider in the same pi setup reports cache hits near 100%. So this isn't pi misreading the usage — the cache reads genuinely aren't happening, and every turn re-bills the full prefix at fresh-input price ($0.10/M instead of $0.002/M on contrib — roughly 50x). On long sessions that adds up fast.
Root cause (traced through the code)
For GOAT,
transport.tspinsstreamProvider— the request goes toPOST https://api.commandcode.ai/provider/v1/chat/completionsthrough pi-ai's nativeopenai-completionsstreaming. And in that path, pi-ai'sbuildParams(pi-ai/dist/api/openai-completions.js) gates the cache key like this:Three independent reasons this is always
undefinedfor commandcode models:model.baseUrlishttps://api.commandcode.ai/provider/v1— never containsapi.openai.com, so the first clause is dead.cacheRetentiondefaults to"short"(only"long"withPI_CACHE_RETENTION=long), so the second clause is dead by default.PI_CACHE_RETENTION=long, the extension registers openai-branch models withsupportsStore: falseand nosupportsLongCacheRetention(index.ts,createProviderConfig), soprompt_cache_retentionstaysundefinedanyway.On top of that, no other cache/stickiness signal is sent on this path:
getCompatCacheControlrequirescacheControlFormat === "anthropic"(never true on the openai branch), and session-affinity headers are only emitted for openrouter (sendSessionAffinityHeaders). So the request carries literally nothing the gateway could use for sticky routing to a warm Meta backend — and the Contrib pool without sticky routing is exactly what @Star-233 and @beyondhumanwork documented in #83 (coin-flip hits, fresh keys hitting other pairs' warm cache, i.e. no key participation in cache identity).Why the #83 conclusion doesn't cover this path
The #83 A/B tested
/alpha/generateand correctly found the gateway drops unknown fields there. But GOAT never touches/alpha/generate—transport.tsonly falls back there on403 upgrade_required. The Provider API is a different gateway surface, and per the last comment on #83 from @urawazakun,prompt_cache_keyis forwarded on/provider/v1/responses(just not on/chat/completions): keyed requests onmeta/muse-spark-1.3-contributorhit ~99% from the 2nd call vs 0/5 keyless. Small sample, but it points exactly where the fix should go.Also worth noting: pi-ai's
openai-responsesbranch sendsprompt_cache_keyunconditionally (cacheRetention === "none" ? undefined : clamp(sessionId)— noapi.openai.comgate), which is consistent with that observation.Suggested directions (whatever you think fits best)
openai-responsesapi mapping for the Meta models (and any other models CommandCode serves on/provider/v1/responses), so the request goes through the endpoint that actually forwards the key.apiForModelIdcurrently only knowsopenai-completionsvsanthropic-messages.prompt_cache_key(stable per session, e.g. fromoptions.sessionId) viaonPayloadon the existingchat/completionspath, and verify whether CommandCode's gateway forwards it there too.supportsLongCacheRetentionfor models where CommandCode honorsprompt_cache_retention, soPI_CACHE_RETENTION=longhas an effect.Happy to test any branch against the live GOAT endpoint and report back cache-read numbers, same as @beyondhumanwork did for #83. I can run A/B pairs (keyed vs control, immediate-repeat and isolated fixtures) and post the
cacheReadTokensdistributions.Thanks again for maintaining this — and for the thorough #83 discussion that made this trace possible.