docs(headroom): clarify that single-turn requests are never compressed - #1007
Open
devin-ai-integration[bot] wants to merge 1 commit into
Open
docs(headroom): clarify that single-turn requests are never compressed#1007devin-ai-integration[bot] wants to merge 1 commit into
devin-ai-integration[bot] wants to merge 1 commit into
Conversation
Contributor
Author
🤖 Devin AI EngineerI'll be helping with this pull request! Here's what you should know: ✅ I will automatically:
Note: I can only respond to comments from users who have write access to this repository. ⚙️ Control Options:
|
|
The latest updates on your projects. Learn more about Vercel for GitHub.
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
The Headroom quick start told people to test with a system message plus one user message, which can never compress anything. LiteLLM holds back the system rows, the last user row, and the last assistant row (
get_protected_indicesinlitellm/compression/compress.py, applied by the Headroom guardrail's_protected_indices), so on that payloadcompressibleis empty and the guardrail returns before ever calling/v1/compress. Anyone following the page verbatim sees Headroom do nothing and reasonably concludes the integration is broken.Reproduced on a live proxy (
guardrails: ["headroom-compression"], also attached to a virtual key): the doc's own curl returns 200 with zero requests reaching the Headroom service, while the same call with one earlier user/assistant turn sends 3 rows to/v1/compress.Changes: the quick-start examples (chat completions and
/v1/messages) now carry an earlier turn, a paragraph after them explains the protected rows and the single-turn no-op, and the "Whyrequests_compressedcan be 0" section now starts with that check before the Headroom-container defaults.Also corrected two claims about
x-litellm-applied-guardrails. That header is written from the guardrails the request opted into (_process_guardrail_metadatainlitellm/proxy/utils.py), so it is present even when Headroom was skipped; it confirms opt-in, not that any message was rewritten.guardrail_informationon the spend log row is the signal for the latter.Link to Devin session: https://app.devin.ai/sessions/835118e39ed94d4cbeafc573bfd2d078