An inspectable Qwen2/Qwen2.5 inference engine for studying how memory management, scheduling, and GPU kernels affect serving latency. The decoder, paged KV cache, request scheduler, and Triton kernels are implemented in this repository.
Quick start · Architecture · T4 benchmark evidence · Tests · Serving API
- Memory and scheduling: block-allocated KV caching, continuous batching, chunked prefill, and explicit preemption/recomputation accounting.
- GPU implementation: a PyTorch reference path plus Triton paged attention and fused INT8 linear layers, with focused parity tests.
- Measurement: seeded open-loop workloads, per-request latency records, component profiles, and retained results that include failures and regressions.
The model path loads Qwen checkpoints directly; Transformers supplies the tokenizer. This is an experimental systems implementation, with an OpenAI-compatible completions API and observability stack for exercising the engine.
The repository retains Tesla T4 experiments using Qwen2.5-1.5B, PyTorch 2.11.0, and CUDA 12.8. The batching rows below report medians over five repetitions of the same right-skewed workload at maximum batch size 8:
| Scheduler | Output throughput | p95 time to first token |
|---|---|---|
| Static batching | 61.3 tokens/s | 4,741 ms |
| Continuous batching | 94.5 tokens/s | 2,863 ms |
These are comparisons within this engine and workload. The raw checkpoint 3 results record the other batch sizes, prefill ablations, source revision, and request outcomes.
The INT8 experiment records a memory/latency tradeoff: fused INT8 reduced stored model bytes from 4.03 GB to 2.72 GB, but was slower than FP16 in the measured end-to-end workloads. The legacy reference INT8 path ran out of memory on T4; its failed row is retained. The accuracy artifact covers a 2,048-token WikiText validation slice and is a narrow comparison, not a general quality evaluation.
See benchmark methodology for matching rules and retained artifacts for JSON, CSV, and checksums. New local outputs are ignored by Git by default; selected checkpoint evidence is tracked.
- Benchmark harness: deterministic open-loop arrivals; clipped log-normal prompt and output lengths; per-request TTFT, TPOT, p50/p95/p99 end-to-end latency; output throughput; NVML utilization and memory high-water sampling; JSON and CSV output.
- Sequential baseline: one request at a time and a full-prefix forward pass for every generated token. A lock prevents concurrent callers from accidentally turning it into a batched baseline.
- Paged KV cache: fixed-size physical blocks, per-sequence block tables, an O(1) free-list allocator, memory accounting, and explicit recomputation counters after preemption. PyTorch and Triton attention backends traverse physical blocks with online softmax.
- Continuous batching and chunked prefill: new requests join between model iterations; configurable prefill chunks share the step budget with decode work. Static batching is available as the control. FCFS, SJF, and priority policies share the scheduler and model path.
- Weight-only quantization: symmetric per-output-channel INT8 and packed INT4 reference layers, plus a fused Triton INT8 backend. Evaluation tools report WikiText perplexity, HellaSwag accuracy, matched decode throughput, and model size. This is not GPTQ or AWQ.
- Speculative decoding: greedy draft proposals, one target verification pass per proposal block, bonus tokens, and acceptance/target-pass reporting.
- Serving surface: OpenAI-compatible
POST /v1/completions, SSE streaming, Prometheus metrics, a provisioned Grafana dashboard, Docker Compose, and a constant-arrival-rate k6 workload mirroring the benchmark distributions.
flowchart LR
C["Client / benchmark load generator"] --> Q["Incoming request queue"]
Q --> S["Iteration scheduler: FCFS / SJF / priority"]
S --> M["Qwen2-style PyTorch decoder"]
M --> A["Block-wise paged attention"]
A <--> K["Block tables + KV block allocator"]
M --> T["Token callbacks / SSE"]
T --> C
K --> P["Preemption + recomputation accounting"]
The uncached reference path uses PyTorch scaled-dot-product attention. The serving path is separate: each layer projects Q/K/V, writes K/V to a physical cache slot, then performs numerically stable online softmax while traversing that sequence's block table, using the selected PyTorch or Triton backend. See architecture.md.
Python 3.10–3.13 is supported.
uv sync --extra all
uv run pytest
# Fast CPU smoke test with the bundled tiny random model.
uv run llmserve benchmark \
--config continuous \
--requests 8 \
--arrival-rate 20 \
--output artifacts/smoke.jsonThe random model tests engine behavior; it does not produce meaningful language or accuracy.
These retained figures describe the earlier PyTorch reference path: 4 requests, the bundled random 4-layer model, PyTorch 2.13.0 on an arm64 CPU, one run, no warm-up. They prove that the harness produces a number for every configuration and, usefully, show a failure: readable Python block traversal is slower than PyTorch's optimized dense kernel.
| Smoke config | Throughput (tok/s) | p95 TTFT | p95 TPOT | Cache memory reduction vs contiguous |
|---|---|---|---|---|
| Naive sequential | 337.8 | 386.9 ms | 3.18 ms | n/a |
| Paged, batch 1 | 102.2 | 1,914.8 ms | 7.07 ms | 75.0% |
| Static, batch 4 | 110.6 | 1,449.0 ms | 15.81 ms | 85.9% |
| Continuous, batch 4 | 112.7 | 1,052.2 ms | 19.57 ms | 85.5% |
The matched tiny-model quantization smoke test reduced stored model bytes by 59.7% with INT8 and 70.4% with INT4. INT4 worsened perplexity by 4.89 and both modes reduced CPU throughput because the reference layers dequantize before GEMM. Those regressions are documented instead of renamed wins.
The direct checkpoint loader maps Qwen2/Qwen2.5 safetensors into this repository's decoder modules.
It does not call transformers for model inference. The tokenizer is the only Transformers runtime
component.
uv sync --extra all
for config in naive paged static continuous; do
uv run llmserve benchmark \
--model Qwen/Qwen2.5-1.5B \
--device cuda \
--config "$config" \
--policy fcfs \
--profile configs/benchmark-sharegpt.yaml \
--output "artifacts/${config}.json"
done
# Scheduler fairness/tail-latency comparison.
for policy in fcfs sjf; do
uv run llmserve benchmark \
--model Qwen/Qwen2.5-1.5B \
--device cuda \
--config continuous \
--policy "$policy" \
--output "artifacts/continuous-${policy}.json"
doneAll runs reuse the seed and load profile from configs/benchmark-sharegpt.yaml. Do not compare two
configurations produced with different profiles, model revisions, warm-up rules, or hardware. Full
methodology and failure-regime experiments are in benchmarking.md.
uv run llmserve quantize-eval \
--model Qwen/Qwen2.5-1.5B \
--device cuda \
--eval-tokens 8192 \
--hellaswag-samples 100 \
--output artifacts/quantization.json
uv run python scripts/plot_quantization.py \
artifacts/quantization.json \
--output artifacts/accuracy-throughput.pngThe command above evaluates the reference INT8/INT4 layers, which dequantize weights before
PyTorch linear operations. The separate fused INT8 backend reads quantized weight tiles directly.
To inspect or reproduce the GPU comparison, use benchmark_checkpoint4.py
and evaluate_checkpoint4.py; both expose their options with --help.
For example, select the accelerator backends in a serving benchmark:
uv run llmserve benchmark \
--model Qwen/Qwen2.5-1.5B --device cuda --config continuous \
--paged-attention-backend triton --linear-backend triton-int8 \
--prefill-chunk-size 16 --output artifacts/local-triton-int8.json# Tiny local smoke service.
uv run llmserve serve
# Real checkpoint on one GPU.
uv run llmserve serve --model Qwen/Qwen2.5-1.5B --device cuda
curl http://localhost:8000/v1/completions \
-H 'Content-Type: application/json' \
-d '{"model":"Qwen/Qwen2.5-1.5B","prompt":"Paged attention is","max_tokens":32,"stream":false}'For the full observability stack:
docker compose up --build
k6 run loadtests/completions.jsThe API is at :8000, Prometheus at :9090, and Grafana at :3000.
- The PyTorch path emphasizes inspectability; the optional Triton paths require supported CUDA hardware. Kernel and end-to-end performance depend on model, shape, workload, and device.
- Chunked prefill is implemented. Prefix sharing is not implemented.
- Direct checkpoint loading currently targets the Qwen2 architecture and safetensors format.
- Speculative verification is greedy. Sampling-correct rejection/resampling is future work.
- Retained benchmarks measure this implementation on named hardware. They do not establish parity with production serving systems or predict performance on other devices.
src/llmserve/
benchmark/ open-loop workload, GPU monitor, metrics, reports
model.py Qwen2-style decoder and block-wise paged attention
kv_cache.py allocator, block tables, eviction, memory accounting
scheduler.py static/continuous admission and scheduling policies
engine.py naive reference and asynchronous serving engine
quantization.py reference INT8/INT4 and fused INT8 dispatch
triton_kernels/ paged attention and fused INT8 linear kernels
evaluation.py WikiText perplexity and HellaSwag accuracy
speculative.py draft/target verification
service/ FastAPI, SSE, and Prometheus instrumentation
MIT