Skip to content

Fix pre-request cost estimation for base64 media inputs - #135

Closed
Lupc9102 wants to merge 2 commits into
hackclub:mainfrom
Lupc9102:main
Closed

Lupc9102 wants to merge 2 commits into
hackclub:mainfrom
Lupc9102:main

Conversation

@Lupc9102

@Lupc9102 Lupc9102 commented Aug 6, 2026

Copy link
Copy Markdown
Contributor

Estimate media (video/image/audio) cost from base64 payload size using provider tokenization rates and the model's actual OpenRouter per-token pricing instead of counting raw base64 as input tokens, which inflated estimates and rejected legitimate video requests with a 429 quota error. (Fix written by Qwen 3.8 Max, Im not confident in js so please double check everything 🥹 )

Estimate media (video/image/audio) cost from base64 payload size using
provider tokenization rates and the model's actual OpenRouter per-token
pricing instead of counting raw base64 as input tokens, which inflated
estimates and rejected legitimate video requests with a 429 quota error.
@kyto-agent

Copy link
Copy Markdown
Contributor

Verified the root cause from Slack support: the old JSON.stringify-everything estimate counted a base64 image as text (3 chars/token), so e.g. a 4MB base64 image on google/gemini-3-pro-image (/M prompt) inflated the reservation by ~.8 plus completion and tripped the $3/day spending-limit 429 even at $0 spent. The 1024-token/image flat rate here fixes that. Users were seeing this exact 'big image edits 429, smaller ones work' pattern. One nit: the model page still has no image-editing example — I opened #142 for that (docs-only).

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants