GLM-5.2-NVFP4-REAP-469B serving on SM120 (4× RTX PRO 6000 Blackwell) — one-command vLLM launch recipe, 250K context, DeepSeek Sparse Attention + MTP speculative decode
-
Updated
Jun 19, 2026 - Shell
GLM-5.2-NVFP4-REAP-469B serving on SM120 (4× RTX PRO 6000 Blackwell) — one-command vLLM launch recipe, 250K context, DeepSeek Sparse Attention + MTP speculative decode
MLX-compatible REAP for pruning MoE models on Apple Silicon
A reusable, model-agnostic methodology for evaluating pruned/compressed Mixture-of-Experts (MoE) model releases: rubric, red-flag catalog, runnable behavioral tests, scorecard + producer checklists, and worked examples.
Deploy the GLM-5.2-469B model on four RTX PRO 6000 Blackwell GPUs using a turnkey vLLM Docker configuration to enable high-speed sparse attention and inference.
How I fit 35B-parameter MoE models into 16 GB of consumer VRAM: REAP expert pruning + NVFP4 quantization on an RTX 5070 Ti (SM120), served on vLLM, with measured evals.
REAP pruning + NVFP4 quantization of Ornith-1.5-35B-A3B for 16GB VRAM
To associate your repository with the reap topic, visit your repo's landing page and select "manage topics."