Conversation
Co-authored-by: OpenAI Codex <codex@openai.com> Signed-off-by: zack <51604064+luoyuctl@users.noreply.github.com>
|
@ywang96 @mgoin @esmeetu Could you review this 8×H20 DeepSeek V4.1 Flash serving case study when you have time? It includes a three-run Humming W4A8 vs FP8 recipe comparison, an isolated EP8 MoE operator ablation, and the reproduction data. I would especially appreciate feedback on the benchmark methodology and how the operator and whole-recipe results are attributed. Thank you! |
|
@NickLucche Would you have time to review this post? It is a serving case study for DeepSeek-V4.1-Flash on 8×H20 (three-run recipe comparison plus an isolated EP8 MoE dispatch ablation). The post is self-contained; the main thing I'd like a second opinion on is whether the whole-recipe and operator-only results are attributed clearly enough. Happy to shorten it if that helps. |
Summary
Related work and duplication check
No open PR in this blog repository was found for this DeepSeek V4.1 Flash on H20 case study. This is a blog post, not a duplicate of FlashInfer PR #5560 or vLLM issue #58799.
Validation
node --check assets/repro/deepseek-v41-h20/quality_eval.mjs: passed.git diff --check: passed.AI assistance was used to draft and organize the article and reproduction materials; the human submitter reviewed the changes and reported measurements.