Skip to content

[Blog] Optimizing DeepSeek V4.1 Flash on 8× H20: EP8 MoE Dispatch and vLLM Serving Performance - #350

Open
luoyuctl wants to merge 1 commit into
vllm-project:mainfrom
luoyuctl:draft/deepseek-v41-h20
Open

luoyuctl wants to merge 1 commit into
vllm-project:mainfrom
luoyuctl:draft/deepseek-v41-h20

Conversation

@luoyuctl

Copy link
Copy Markdown

Summary

  • Add an 8×H20 deployment case study tracing the MXFP4/Humming W4A8 path through FP8 conversion, Engram offload, and routed-row-aware MoE dispatch.
  • Report separate isolated operator results (up to 3.24×) and three-run whole-recipe serving comparisons against Humming: Decode c256 4,725 → 6,459 output tok/s; fixed-1K Prefill c64 9,875 → 15,048 input tok/s. These gains have different scopes and must not be multiplied.
  • Include figures, per-run data, paired 500-item GSM8K sanity-check results, methods, and environment/configuration details.

Related work and duplication check

No open PR in this blog repository was found for this DeepSeek V4.1 Flash on H20 case study. This is a blog post, not a duplicate of FlashInfer PR #5560 or vLLM issue #58799.

Validation

  • Production Jekyll build: passed.
  • node --check assets/repro/deepseek-v41-h20/quality_eval.mjs: passed.
  • Referenced article assets and JSON/JSONL parsing: passed.
  • git diff --check: passed.
  • Reported serving results are means of three runs per concurrency tier; paired GSM8K numeric accuracy was 476/494 (Humming) versus 475/494 (tuned).

AI assistance was used to draft and organize the article and reproduction materials; the human submitter reviewed the changes and reported measurements.

Co-authored-by: OpenAI Codex <codex@openai.com>
Signed-off-by: zack <51604064+luoyuctl@users.noreply.github.com>
@luoyuctl

Copy link
Copy Markdown
Author

@ywang96 @mgoin @esmeetu Could you review this 8×H20 DeepSeek V4.1 Flash serving case study when you have time? It includes a three-run Humming W4A8 vs FP8 recipe comparison, an isolated EP8 MoE operator ablation, and the reproduction data. I would especially appreciate feedback on the benchmark methodology and how the operator and whole-recipe results are attributed. Thank you!

@luoyuctl luoyuctl changed the title [Blog] DeepSeek V4.1 Flash serving on 8 H20 GPUs [Blog] Optimizing DeepSeek V4.1 Flash on 8× H20: EP8 MoE Dispatch and vLLM Serving Performance Sep 26, 2026
@luoyuctl

Copy link
Copy Markdown
Author

@NickLucche Would you have time to review this post? It is a serving case study for DeepSeek-V4.1-Flash on 8×H20 (three-run recipe comparison plus an isolated EP8 MoE dispatch ablation). The post is self-contained; the main thing I'd like a second opinion on is whether the whole-recipe and operator-only results are attributed clearly enough. Happy to shorten it if that helps.

This branch was successfully deployed

1 active deployment
Preview — 294642da Deployed Sep 26, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant