━━━━━━━━━━ ✦ ━━━━━━━━━━
I am a Master's student at the Key Laboratory of Data Intelligence and Advanced Computing, Soochow University, supervised by Prof. Juntao Li.
My work focuses on efficient inference for large language models — from KV-cache compression to sparse attention mechanisms.
- 🧠 Efficient LLM Inference — pushing the limits of fast, economical LLM serving
- 💾 KV-Cache Compression — shrinking the memory footprint of long-context inference
- ⚡ Sparse Attention — focusing computation on what matters most
- 🔬 Model Optimization — keeping models lean without sacrificing quality
━━━━━━━━━━ ✦ ━━━━━━━━━━
"Advancing AI through innovative research and optimization."
© 2026 Quantong Qiu
