Add challenge 121: Fused Logit Penalties (Medium) - #334
Open
claude[bot] wants to merge 1 commit into
Open
claude[bot] wants to merge 1 commit into
claude[bot] wants to merge 1 commit into
Conversation
Implements the logit penalty stage that LLM serving stacks (vLLM, SGLang, TGI) run before sampling: per-request repetition, frequency and presence penalties applied to a B x V logit batch given each request's prompt and generated token histories. The interesting part is work distribution — V is far larger than the token histories, so solvers must build per-request occurrence counts with a scatter instead of rescanning the histories per vocabulary entry. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
claude
Bot
requested review from
ishaan-arya,
kunal-mansukhani and
shxjames
as code owners
October 2, 2026 08:23
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Adds challenge 121: Fused Logit Penalties (Medium) — the logit-processor stage that vLLM, SGLang and HF TGI run on every decode step, right before sampling.
Given a batch of
Bin-flight requests, each with aV-wide logit row, its prompt/generated token histories (-1-padded), and its own penalty coefficients, the solver produces the penalized logits:z /= rifz > 0, elsez *= rz -= frequency_penalty * countz -= presence_penaltyifcount > 0Why this is a worthwhile GPU problem
V(up to 131,072) dwarfs the token histories, so the obvious "scan the history for each vocabulary entry" approach is hopelessly slow. A good solution builds the per-request occurrence table with a scatter (atomics or privatized counters), then streams the logits in a single bandwidth-bound pass. Padding slots and the asymmetry between the three penalties (two different "seen" predicates) add real bookkeeping. It is not an element-wise map.No overlap with existing sampling challenges (29 top-k, 60 top-p, 104 min-p) — those select/renormalize, this one is a gather/scatter histogram problem — and no overlap with any open PR topic or number.
Contents
challenge.py— reference impl (standard PyTorchscatter_add_/where, CUDA + XLA safe), 10 functional tests (edge sizes 1–4, powers of 2, 100/255 non-powers of 2, zeros, negatives, all-padding rows, penalty no-ops, realistic 32k-vocab decode), performance test atB=256, V=128,256, P=1,024, G=512(~264 MB of tensors)challenge.html— description, SVG walk-through of the scatter → count table → apply pipeline, worked example, constraintssolve, medium convention)Validation
python scripts/validate_challenges.py→ 0 errorspre-commit run --all-files→ all hooks pass--action run✓ and--action submit✓ ("All tests passed"). The solution file is not committed.🤖 Generated with Claude Code