Skip to content

feat: price OpenAI fast/priority and ultrafast by actual service tier - #4532

Merged
dgageot merged 1 commit into
mainfrom
feat/openai-actual-service-tier-cost
Oct 7, 2026
Merged

dgageot merged 1 commit into
mainfrom
feat/openai-actual-service-tier-cost

Conversation

@dgageot

@dgageot dgageot commented Oct 7, 2026

Copy link
Copy Markdown
Member

Cost estimates for service_tier: fast, priority, and the new ultrafast requests have been using catalogue pricing regardless of what OpenAI actually billed. Since a request can be downgraded to standard pricing even after Fast or Ultrafast was asked for, the previous approach could overstate costs, and there was no coverage at all for the new Ultrafast mode on GPT-6 Astra.

Cost accounting now reads the service tier the provider actually served from the response, instead of trusting the one requested, and applies the right multiplier at settlement time. This is wired through Chat Completions and Responses, over both SSE and WebSocket, including incomplete responses. The allowlist of models eligible for premium pricing is explicit: Fast and priority apply a 2x multiplier to gpt-5.6 and its Sol/Terra/Luna aliases plus gpt-6-astra, gpt-6-sol, gpt-6-luna, and gpt-6.1-sol; Ultrafast applies 6x, and only to gpt-6-astra. Multipliers cover fresh input, cache read, cache write, output, and the long-context band. A response that comes back as default keeps standard pricing even if a premium tier was requested, and the zero-cost case is handled explicitly rather than falling through.

Custom endpoints (including OPENAI_BASE_URL), rule-based routers, and OpenAI-compatible providers are left on catalogue pricing, and gateway model IDs that are already priced as -fast are not multiplied a second time. The model catalogue is cloned before these adjustments so the shared, cached catalogue stays immutable. As part of this, the WebSocket transport now resolves its base URL the same way the HTTP transport does, fixed once at client construction, so both paths agree on which endpoint is in play.

This only changes how fast/ultrafast tiers are costed; it does not touch request routing, semantic RAG, or any other stream accounting behavior. Docs and examples/openai-service-tier.yaml were updated to describe Ultrafast mode and the new pricing behavior.

Cost accounting now uses the tier the provider actually served instead
of the one requested, covering chat, responses, and websocket
transports. Custom endpoints, routers, and cost overrides are left
alone. WebSocket and HTTP paths now resolve the base URL the same way,
fixed at client construction for both transports.

Assisted-By: docker-agent
@dgageot
dgageot requested a review from a team as a code owner October 7, 2026 09:04
@dgageot
dgageot added this pull request to the merge queue Oct 7, 2026
Merged via the queue into main with commit f111a9a Oct 7, 2026
22 checks passed
@dgageot
dgageot deleted the feat/openai-actual-service-tier-cost branch October 7, 2026 09:27
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants