Conversation
…out-46f0dd refactor(ui): move every page header onto the shared PageHeader
fix(proxy): stop cache eviction errors from failing /key/update
…rify fix(aiohttp): honor global ssl_verify on the aiohttp_openai handler path
The Default Budget Duration field in Team Member Settings only offered daily, weekly and monthly, so a team member budget could never be set to never reset. It now uses the shared BudgetDurationDropdown, and /team/update writes an explicitly null duration through to the member budget row along with its reset time. Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
… resets The member duration dropdown reused its placeholder as "Never resets", so a team with no member budget yet showed "Never resets" while sending nothing and inheriting the team's own reset period. Use the dropdown's never-resets sentinel for an explicit null and label the untouched state as inheriting.
…back_flush fix(logging_worker): rescue dequeued logging tasks lost at event loop close
…g them to Gemini voices
…ompt version POST /prompts silently stored an empty template when litellm_params.prompt_id was combined with prompt_data keyed by template name, because the loader wrapped the already-keyed dict under prompt_id a second time. The loader now wraps only a flat template (a dict carrying a content key), and create, update, and patch reject the ambiguous keyed+prompt_id combination with a 400 that names both valid shapes. The API also returned version null on every create and lost version, environment, and created_by on registry reload; both now carry through. Versioned ids like my-prompt.v1, which the create API itself returns, now resolve to their base template on the SDK prompt hooks, and a flat DB prompt with no litellm_params.prompt_id registers under its base API id instead of garbage.
… backfill Anthropic re-export entries
…update_budget The three member-budget tests patched litellm internals and asserted only on the mock, which tripped the TQ002 and TQ008 test-quality ratchet. Fake the prisma budget table on the shared client and assert on the row that reaches the database plus the returned team payload.
fix(caching): require the namespace delimiter when checking already-namespaced redis keys
…s registration endpoint
…he fixed create_prompt example
…and honor ignore_prompt_manager_model On /v1/responses the prompt template ran inside litellm.aresponses, after the router had already resolved a deployment and injected its api_key/api_base, so a prompt whose metadata.model pointed at another provider sent the old deployment's credentials cross-provider (401). The proxy now runs the prompt template for aresponses in the pre-call hook, before routing, so the router picks the deployment that matches the swapped model. As a backstop, the SDK refuses a cross-provider swap when explicit credentials are already present instead of forwarding them. ignore_prompt_manager_model and ignore_prompt_manager_optional_params saved on a prompt were only read by the generic manager, so dotprompt prompts ignored them on every endpoint. PromptManagementBase now merges the prompt spec's flags with the per-request ones for every manager, and the generic manager no longer drops caller flags when no spec is present.
…ude 4.x re-export entries
The committed snapshot behind /openapi.json for unloaded lazy features had drifted on 30 of 31 fragments and never had one for a2a_registration or gemini_agents, so those routes showed as placeholder GET stubs or old docstrings until traffic loaded them. Regenerate the snapshot and schema.d.ts, make the check-ui-api-types job and make check regenerate the snapshot and fail on drift, and make the generator refuse to write a snapshot when any feature fails to import so a broken import cannot silently drop fragments.
Gemini 2.5 Flash Preview TTS, Gemini 2.5 Pro Preview TTS, and the three gemini-2.5-flash-native-audio entries carried rates copied from the text models, so audio output was billed 2x to 6x under Google's published prices. Set the published per-token rates on all ten keys, add output_cost_per_audio_token to the native-audio entries, and drop the long-context tier rates Google does not publish for Pro TTS.
…pe-discipline gate
…_nest fix(prompts): reject keyed prompt_data with prompt_id and populate prompt version
fix(cost_calculator): resolve real cost key when model_name alias contains '/'
…in_tokens fix(cost-map): correct prompt_cache_min_tokens for Claude Fable 5 and backfill Anthropic re-export entries
…_retry_tests test(azure-ai): pin the 422 retry that drops the field the provider rejected
) The dashboard resolved the complexity-router tier set three different ways: a private TIER_KEYS in build_complexity_router_config.ts, TIER_ORDER in complexity_router_tiers.ts, and TIER_KEYS in ComplexityRouterConfig.tsx. The edit modal went further and re-implemented the whole create payload builder, kept in sync only by a comment reading "Mirrors buildComplexityRouterConfig". tier_rows.ts now owns the tier set. Every consumer reads activeTierRows(value) and a row carries its own id, so the plan-mode floor and per-model params point at a row rather than at a position, and the leaves that already wanted entries (buildAutoRouterTestTargets, getRequiredModels, model_info_view) take them. buildUpdatedComplexityRouterConfig becomes preserve-unmanaged-keys around the shared builder instead of a second copy of it. Also drops the literal ", ]" that renders as visible text in two DialogFooter blocks on the auto-router routing-test and connection-test dialogs, left over from a JSX array-to-fragment conversion. No behaviour change: all 566 tests over the touched modules pass with fixture shape changes only, no assertion edited.
reload_search_tools_from_db is a read-modify-write of the shared llm_router global: it reads the whole table, merges the config tools in, and replaces router.search_tools wholesale. Two of those interleaving lets the older snapshot's assignment land last and put back a tool the newer one deleted, so a revoked tool keeps serving on the provider key it carried until the next reload. Take MODEL_RECONCILE_LOCK, which add_deployment already uses to serialize the same shape of work on the same global. It has to go on this entry point rather than in _init_search_tools_in_db, because _init_non_llm_objects_in_db calls that while already holding the lock and asyncio.Lock is not reentrant. A separate search-tools-only lock would not close the race: the periodic reconcile reaches _init_search_tools_in_db under MODEL_RECONCILE_LOCK, so only that same lock orders an endpoint refresh against a cron tick. Ordering across workers is unchanged and still reconciles on the next tick.
…synthesis fix(transcription): synthesize srt/vtt output for adapters without native subtitle formats
…_tier fix(streaming): preserve provider service-tier metadata so Vertex flex streams bill at flex rates
…tokens fix(realtime): bill Gemini Live native-audio output tokens at the audio rate
…_passthrough fix(anthropic): carry tool_reference tool results through the guardrail translation round trip
fix(anthropic-adapter): pass provider-native and OpenAI-format tools through on /v1/messages
test(e2e): serve the vision image from our own fixture
feat(together_ai): add zai-org/GLM-5.3-Flash to the model registry
…_hardening fix(ui): stop server-searched comboboxes from clobbering picks and queries
…t levels (#38481) Kimi K3 accepts exactly low, high and max, defaults to max, and always thinks. The map could not say that: medium and high have no supports_*_reasoning_effort flag because every other reasoning model takes them, so the ten kimi-k3 entries carried supports_reasoning alone and resolved to unknown. The dashboard then fell back to a capability-blind level list that deliberately omits max, which is why a kimi-k3 tier cannot be set to max thinking today. Add reasoning_effort_levels, an array key in the shape the map already uses for supported_endpoints and supported_modalities. Where present it is read first and wins whole; every other entry keeps answering through the per-level flags, unchanged. It is deliberately a different name from the computed ModelGroupInfo.supported_reasoning_efforts, which stays derived from a group's deployments and is never seeded from one deployment's model_info. The levels are per entry rather than per model, because the deployments differ: Moonshot, Together, Fireworks and Azure Foundry all forward the level unchanged and get the model's own low/high/max, while Perplexity documents a six-value enum it maps down internally and gets that. The /v1/messages degradation chain consults the same declaration, so the level the map advertises is the level that path forwards.
…llback_model_access
…e target (#38533) /v1/messages forwarded `thinking` verbatim for a Claude-family model and then returned, carrying `output_config.effort` only when the model string started with a Bedrock prefix. Every other bridged provider got a bare adaptive thinking block, so the caller's effort did nothing: max and minimal produced byte-identical upstream bodies. Send those targets the tier as `reasoning_effort`, which is the param they take. Bedrock keeps taking `output_config`, since the two are not interchangeable there: an application inference profile ARN resolves to no chat config, so `reasoning_effort` is dropped and the tier vanishes, and a provider that rebuilds `output_config` from it overwrites a caller-set `thinking.display` on the way. The tier stays a plain string, the summary already travelling inside the forwarded `thinking` block. Adaptive with no tier, and budgeted thinking, both stay exactly as they were.
* feat(alerting): add native Microsoft Teams alerting destination Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(alerting): preserve active destinations on MS Teams save and confirm health test delivery Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix(ui): read persisted alerting destinations at MS Teams save time Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: yassin <yassin@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…els (#38587) Two lazily loaded models changed without their generated artifacts being regenerated, so check-ui-api-types has been red on every branch off staging. The snapshot that /openapi.json serves for unloaded features was missing ChatCompletionToolReferenceObject, and the dashboard types were missing aws_external_id. The snapshot step runs first and short-circuits, so only the first one was visible until it was fixed. Both files are regenerated with `python -m litellm.proxy._lazy_openapi_snapshot` and `npm run gen:api`, no hand edits.
…xity_router_config (#38570) A complexity-router setting placed beside complexity_router_config, or inside a tier entry's litellm_params, is read by nobody: the router loads its settings only from litellm_params.complexity_router_config. It does not stay inert. The alias-marker forwarding and the per-tier param spread carry every unrecognized key onto the outbound request, and all_litellm_params only knows the outer names, so the key reaches the provider as an unknown body field and every call through that model group fails with an error naming an internal config key. Guard the whole set, derived from ComplexityRouterConfig.model_fields so a field added later is covered, and scoped to complexity-router deployments because the names only mean this there (embedding_model is a legitimate flat param on an s3_vectors vector store). Scope is read from the same merged field view the naming check is judged on, so a router named only by its default model is in scope and a field added to the required-field table is covered without another edit. The write endpoints reject with a 400 naming the keys and where they belong, config.yaml refuses to start for the same reason max_agentic_loops does, and a tier entry is judged by the config model itself. An already-stored deployment keeps loading, so an upgrade cannot take a running gateway down over a row that was written before the gate existed.
* feat(ui): session-level cache observability in request logs Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * fix: guard cache_hit filter against non-string defaults in direct calls Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> * refactor(ui): drop redundant cache_hit field comment Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com> --------- Co-authored-by: yassin <yassin@berri.ai> Co-authored-by: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
Co-Authored-By: Devin AI <158243242+devin-ai-integration[bot]@users.noreply.github.com>
…itellm_fallback_access_group_check
…irtual_keys_link fix(ui): link Virtual Keys hint through the migrated /ui route
…reasoning_effort (#38592) The /v1/messages bridge decided a Claude target could take `reasoning_effort` from the model name, which says nothing about the params the provider in front of it accepts. Snowflake serves Claude over the Anthropic dialect and declares `thinking` alone, so `get_optional_params` raised `UnsupportedParamsError` before the request reached the wire: every adaptive request carrying an effort tier turned a 200 into a 400 for all seven of its Claude entries. The tier is now offered only where the target declares the param, reading the same `get_supported_openai_params` the sibling `_supports_prompt_cache_key` reads twelve lines up. A target declaring neither carrier keeps its bare `thinking` block, which is what this bridge sent before it carried a tier at all. Without a resolved provider the tier stays behind rather than being offered blind. Resolving one from the model's prefix instead would run an OAuth device flow for github_copilot and chatgpt, blocking for minutes, and one of the two callers in that position is a logging callback. The copilot case is pinned by a test.
…blocks do not fail (#38483) * fix(presidio): chunk oversized text before /analyze so large content blocks do not fail The Presidio PII guardrail sent each content block to the analyzer as a single /analyze call with no size check. Analyzer deployments commonly cap the request body (the reporting deployment rejects bodies over 1,000,000 bytes with HTTP 413), so large blocks failed closed, and analyzer latency grew linearly with payload size. analyze_text now splits texts larger than presidio_analyze_chunk_size_bytes (default 500,000 UTF-8 bytes, configurable per guardrail) into overlapping chunks, analyzes them concurrently, remaps each detection's start/end onto the original text, and deduplicates detections from the overlap regions. Anonymization, blocked-entity checks, score filtering, numbered-token unmasking, telemetry, and the dashboard entity positions all consume the remapped global offsets unchanged. Resolves LIT-4785 Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(presidio): review-round hardening for chunked analyze - measure the chunk budget on the JSON-serialized text (non-ASCII escapes expand beyond raw UTF-8, so a raw-byte budget could still exceed the analyzer body limit) - share the chunk fan-out semaphore per event loop and instance instead of per call, so many oversized blocks cannot multiply concurrent analyzer calls - apply configured score thresholds and deny list per chunk BEFORE overlap resolution, so a below-threshold span cannot displace a detection the thresholds keep Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…as errors (#38476) Route records below WARNING to stdout (WARNING and above stay on stderr), emit ANSI color codes only when both streams are a TTY (honoring NO_COLOR), and parse JSON_LOGS strictly so JSON_LOGS=false no longer enables JSON logs.
…ving it (#38595) * feat(ui): dry-run an auto-router config against the backend before saving it Both auto-router forms built a payload and posted it, so anything the write gate refused came back as a raw 400 with the backend's message buried in it. They now POST the exact payload to /auto_router/validate_complexity_router_config first and surface its verdict inline. One dryRunRejection owns the gate, and it reads valid alone. The verdict's two fields arrive independently, so gating on the error message would let a rejection that carried none through to the write. A transport failure fails open as valid, leaving the write gate authoritative rather than blocking a save on a flaky network. Applies to every auto-router, built-in tiers included. * fix(ui): hold the auto-router create closed for the full dry-run and create sequence A second submit while the dry-run round-trip was pending started another create against the non-idempotent /model/new. The submit handler now refuses re-entry and the button disables for the whole sequence, matching the edit modal's loading guard. Also drops the explanatory comments this PR had added. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
…8568) * fix(guardrails): add fail-open mode to CrowdStrike AIDR guardrail Add a fail_on_error param (default True, preserving existing behaviour) to the CrowdStrike AIDR guardrail, mirroring model_armor and generic_guardrail_api. When fail_on_error=False the guard fails open only on server errors (5xx) and connectivity failures, so the request proceeds unmodified. Caller-controlled 4xx responses and result.blocked policy blocks always fail closed. The applied-guardrails header is recorded even on the fail-open path. * fix(guardrails): fail open AIDR 4xx * refactor(guardrails): isolate AIDR fail-open * style(guardrails): format AIDR fail-open * ci: satisfy unit workflow timeout invariant * refactor(guardrails): accept AIDR mappings * test(guardrails): inject AIDR HTTP client * fix(guardrails): harden AIDR fail-open against delivered verdicts and record fail-open status Reads the blocked verdict from the raw body before guard_output validation so schema drift or a changed verdict type cannot fail open past a delivered block. A transformed response that cannot be parsed fails closed so delivered redactions are never dropped. Fail-open runs record guardrail_status guardrail_failed_to_respond with timings instead of success. Restores the fail-open behavior tests dropped mid-PR and reverts the payload Mapping widening * test(guardrails): cover fail_on_error wiring and fail-closed default for CrowdStrike AIDR * chore(guardrails): annotate the transformed-drift detail payload for the LIT002 budget --------- Co-authored-by: abrekhov <abrekhov@users.noreply.github.com> Co-authored-by: mateo-berri <277851410+mateo-berri@users.noreply.github.com>
…fallback stamp keys (#38586)
…check feat(proxy): opt-in enforce_fallback_model_access authorizes router fallbacks against the calling key
…tead of failing requests (#38582) * fix(langfuse): warn and drop invalid LANGFUSE_TRACING_ENVIRONMENT instead of failing requests * fix(langfuse): treat a dynamic environment equal to the raw deployment value as redundant
…nking-extra-body fix(tencent): route thinking through extra_body in chat completions
…#38590) The endpoint built messages=[{"role": "user", "content": prompt}], so a dry run could not carry prior turns, the caller's system prompt, or the tool definitions a request advertises. A real agentic turn reduced to its last sentence classified as trivial, which is why a config sweep reported savings for every configuration. Accept messages, system and tools, and forward them to the same pre-routing hook untranslated, with the raw-body snapshot built by the serving path's own owner, refresh_proxy_server_request_body_snapshot. Loose types are deliberate: the hook reads whatever dialect the surface produced, so validating against one surface's schema would reject the others. prompt stays as the single-ask shorthand, normalized into one user turn inside the request model so the handler carries no mode branch.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
TLDR
Problem this solves:
How it solves it:
User Flow
Relevant issues
Linear ticket
Pre-Submission checklist
Please complete all items before asking a LiteLLM maintainer to review your PR
uv run pytest tests/test_litellm/<your_test_file>.py -v. Leave the suites (make test-unit-*,make test-unit) to CI: it finishes in ~15 minutes where a laptop takes an hour or more@greptileaito re-request a review after pushing changes)Delays in PR merge?
If you're seeing a delay in your PR being merged, ping the LiteLLM Team on Slack (#pr-review).
Screenshots / Proof of Fix
Type
🆕 New Feature
🐛 Bug Fix
🧹 Refactoring
📖 Documentation
🚄 Infrastructure
✅ Test
Caveats (if any)
QA runbook
Final Attestation