A daily log of LLM / ML-systems efficiency papers — inference serving, KV cache, speculative decoding, quantization, and whatever else makes models cheaper to run. Every paper is read in full text, not from the abstract.
9 papers read · 0 notes · 2 active days · 2-day streak
8 papers read
- 📄 Efficient Long-Context Language Model Training by Core Attention Disaggregation (DistCA)
- 📄 TeleRAG: Efficient Retrieval-Augmented Generation Inference with Lookahead Retrieval
- 📄 A Contract-Grade Verifier for LLM-Generated GPU Kernels, and a Native Blackwell Backward for the Gated-Linear-Recurrence Family
- 📄 Simple Is Better: Multiplication May Be All You Need for LLM Request Scheduling (LMetric)
- 📄 Certifying Compressed Language Models: An Audit and a Statistical Toolkit
- 📄 A Full-Stack Characterization of High-Bandwidth Flash for KV-Centric LLM Serving
- 📄 Spec Sheets Are Not Kernels: An ISA- and Source-Level Audit of INT8 Availability on NVIDIA Blackwell Ultra
- 📄 TokenWeave: Efficient Compute-Communication Overlap for Distributed LLM Inference
1 paper read


