Skip to content
#

nvidia-l4

Here are 5 public repositories matching this topic...

Language: All
Filter by language

Deploy Qwen3.6-35B-A3B (Q4_K_XL) + MTP speculative decoding on a single NVIDIA L4 24GB — GCP g2-standard-8 — via the official llama.cpp Docker image. Decode-optimized to ~91–99 tok/s (min ~91 chat, max ~99 math), lossless (full GPU residency + ECC-off).

  • Updated Aug 24, 2026
  • Shell

Qwen3.8-27B (Unsloth UD-Q4_K_XL GGUF) served at 37.4 tok/s decode over a 226,048-token context on a single NVIDIA L4 24 GB: llama.cpp with DFlash 2 block-diffusion speculative decoding, a one-line CUDA kernel-routing patch worth 16%, one-command GCP provisioning, quality verification, and a measured account of where the ceiling is.

  • Updated Aug 20, 2026
  • Shell

Improve this page

Add a description, image, and links to the nvidia-l4 topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the nvidia-l4 topic, visit your repo's landing page and select "manage topics."

Learn more