You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Deploy Qwen3.6-35B-A3B (Q4_K_XL) + MTP speculative decoding on a single NVIDIA L4 24GB — GCP g2-standard-8 — via the official llama.cpp Docker image. Decode-optimized to ~91–99 tok/s (min ~91 chat, max ~99 math), lossless (full GPU residency + ECC-off).
Qwen3.8-27B (Unsloth UD-Q4_K_XL GGUF) served at 37.4 tok/s decode over a 226,048-token context on a single NVIDIA L4 24 GB: llama.cpp with DFlash 2 block-diffusion speculative decoding, a one-line CUDA kernel-routing patch worth 16%, one-command GCP provisioning, quality verification, and a measured account of where the ceiling is.
Benchmark study: NVIDIA L4 + Llama 3.1 8B FP8 on Red Hat OpenShift AI. Concurrent load benchmarks, quality evaluation, and deployment advisor grounded in real measured data.