Skip to content

[Klaud Cold] Update dsr1-fp4-b200-dynamo-sglang-mtp SGLang image to v0.5.19-cu130 with Dynamo wheel 1.5.0.dev20260910 / 将 dsr1-fp4-b200-dynamo-sglang-mtp 的 SGLang 镜像更新至 v0.5.19-cu130 并搭配 Dynamo wheel 1.5.0.dev20260910 - #3013

Closed
Klaud-Cold wants to merge 2 commits into
mainfrom
klaud/auto-39cf82a2bb6e3c04-3cf2069b9d51c51a

Conversation

@Klaud-Cold

Copy link
Copy Markdown
Collaborator

Refresh the dsr1-fp4-b200-dynamo-sglang-mtp disaggregated MTP family from lmsysorg/sglang:v0.5.12.post1 to the current SGLang release lmsysorg/sglang:v0.5.19-cu130, switching its coupled Dynamo install from the source hash 5b4bc1dd (1.3.0, pinned to SGLang 0.5.12) to the published ai-dynamo wheel 1.5.0.dev20260910 (pinned to SGLang 0.5.19) in the nine recipe YAMLs and the master router.version. Model, precision, topologies, MTP settings, workloads and concurrency points are unchanged.

Baseline

  • Published date: 2026-07-08 (Single-turn 8k/1k, DeepSeek-R1-0528 FP4, B200, dynamo-sglang, MTP, disaggregated, NIXL)
  • Old image: lmsysorg/sglang:v0.5.12.post1 (Docker Hub manifest digest sha256:ceaf8b16e02d165143633ac228bbb994a05fe77d7e0526cf035ae4bbf4eacc36, identical to v0.5.12.post1-cu130), Dynamo source hash 5b4bc1dd70965017a737c71b19db5a0aeaa88727
  • Producer run: 28916523710, head e0e00cdeb11e580bb6cd0e9fbe92b6a3ef7f3671, changelog PR #2113
  • API queries: /api/v1/workflow-info?date=2026-07-08, /api/v1/benchmarks?model=DeepSeek-R1-0528&date=2026-07-08&exact=true&sequence=8k%2F1k (rows filtered to framework=dynamo-sglang, precision=fp4, spec_method=mtp, disagg=true, isl=8192, osl=1024), /api/v1/evaluations?model=DeepSeek-R1-0528&date=2026-07-08
  • Topology / concurrency (13 points; tput = total tok/s per GPU, out = output tok/s per GPU, median TTFT in s, median TPOT in ms):
Topology (P / D) Conc tput/GPU out/GPU TTFT TPOT
1P tp4 / 5D tp8 (low-lat) 4 191.2 23.4 0.745 3.32
1P tp4 / 5D tp8 (low-lat) 8 366.0 45.3 0.858 3.67
1P tp4 / 5D tp8 (low-lat) 16 648.2 79.0 0.920 4.00
1P tp4 / 5D tp8 (low-lat) 32 950.5 116.9 2.087 4.62
1P tp4 / 3D tp8 (low-lat) 32 1382.7 180.4 1.458 5.77
1P tp4 / 3D tp8 (low-lat) 64 1527.9 197.7 6.689 6.11
1P tp4 / 1D tp8 (low-lat) 32 2281.4 382.6 1.172 9.36
1P dep4 / 1D dep8 (MTP2) 512 3944.2 657.3 79.222 10.98
2P dep4 / 1D dep8 (MTP2) 768 5838.0 1296.8 54.615 13.51
3P dep4 / 1D dep8 (MTP2) 1024 6963.5 1934.5 44.133 17.63
4P dep4 / 1D dep8 (MTP2) 512 7343.4 2447.6 5.181 20.00
4P dep4 / 1D dep8 (c2048 tune) 2048 8534.8 2847.2 60.718 21.72
5P dep4 / 1D dep8 (MTP2) 2048 8242.6 3208.1 48.389 26.27
  • Published evals (gsm8k em_strict, n=1319, same producer run): 1P/5D c32 0.9568, 1P/3D c64 0.9553, 1P/1D c32 0.9553, 1P/1D dep c512 0.9575, 2P/1D c768 0.9530, 3P/1D c1024 0.9522, 4P/1D c512 0.9538, 4P/1D c2048 0.9545, 5P/1D c2048 0.9530
  • Source links: SGLang v0.5.12.post1v0.5.19; Dynamo 5b4bc1dd → wheel ai-dynamo 1.5.0.dev20260910; srt-slurm launcher pin a98738de (unchanged)

dsr1-fp4-b200-dynamo-sglang-mtp 分离式 MTP 系列的镜像从 lmsysorg/sglang:v0.5.12.post1 更新至当前 SGLang 发布版 lmsysorg/sglang:v0.5.19-cu130,并在九个 recipe YAML 与主配置的 router.version 中,将配套的 Dynamo 安装源从源码哈希 5b4bc1dd(1.3.0,绑定 SGLang 0.5.12)切换为已发布的 ai-dynamo wheel 1.5.0.dev20260910(绑定 SGLang 0.5.19)。模型、精度、拓扑、MTP 设置、负载与并发点均保持不变。

基线

  • 发布日期:2026-07-08(单轮 8k/1k,DeepSeek-R1-0528 FP4,B200,dynamo-sglang,MTP,分离式,NIXL)
  • 旧镜像:lmsysorg/sglang:v0.5.12.post1(Docker Hub manifest 摘要 sha256:ceaf8b16e02d165143633ac228bbb994a05fe77d7e0526cf035ae4bbf4eacc36,与 v0.5.12.post1-cu130 相同),Dynamo 源码哈希 5b4bc1dd70965017a737c71b19db5a0aeaa88727
  • 生产运行:28916523710,head e0e00cdeb11e580bb6cd0e9fbe92b6a3ef7f3671,changelog PR #2113
  • API 查询:/api/v1/workflow-info?date=2026-07-08/api/v1/benchmarks?model=DeepSeek-R1-0528&date=2026-07-08&exact=true&sequence=8k%2F1k(按 framework=dynamo-sglang、precision=fp4、spec_method=mtp、disagg=true、isl=8192、osl=1024 过滤),/api/v1/evaluations?model=DeepSeek-R1-0528&date=2026-07-08
  • 拓扑 / 并发:共 13 个点,数值见上表(tput = 每 GPU 总 tok/s,out = 每 GPU 输出 tok/s,TTFT 中位数单位秒,TPOT 中位数单位毫秒)
  • 已发布评测(gsm8k em_strict,n=1319,同一生产运行):1P/5D c32 0.9568,1P/3D c64 0.9553,1P/1D c32 0.9553,1P/1D dep c512 0.9575,2P/1D c768 0.9530,3P/1D c1024 0.9522,4P/1D c512 0.9538,4P/1D c2048 0.9545,5P/1D c2048 0.9530
  • 源码链接:SGLang v0.5.12.post1v0.5.19;Dynamo 5b4bc1dd → wheel ai-dynamo 1.5.0.dev20260910;srt-slurm 启动器固定提交 a98738de(未变)

🤖 Generated with Claude Code

…u130

Update the master image from lmsysorg/sglang:v0.5.12.post1 to
lmsysorg/sglang:v0.5.19-cu130 and match model.container in the nine
referenced srt-slurm recipe YAMLs. Switch the coupled Dynamo install from
source hash 5b4bc1dd (1.3.0, pinned to SGLang 0.5.12 and importing the
removed sglang.srt.server_args_config_parser module) to the published
ai-dynamo wheel 1.5.0.dev20260910 (pinned to SGLang 0.5.19), mirrored in
the master router.version. Topology, MTP settings, workloads and
concurrency points are unchanged.

将主配置镜像从 lmsysorg/sglang:v0.5.12.post1 更新为
lmsysorg/sglang:v0.5.19-cu130,并同步九个被引用的 srt-slurm recipe YAML 的
model.container。配套的 Dynamo 安装源从源码哈希 5b4bc1dd(1.3.0,绑定
SGLang 0.5.12,且导入了已被移除的 sglang.srt.server_args_config_parser
模块)切换为已发布的 ai-dynamo wheel 1.5.0.dev20260910(绑定 SGLang
0.5.19),并在主配置 router.version 中同步。拓扑、MTP 设置、负载与并发点
保持不变。

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase As a PR reviewer and CODEOWNER, I have reviewed this and have.

For PR verification, add the full-sweep-fail-fast label (strongly recommended) to this PR — the benchmark sweep only runs on labeled PRs. Use full-sweep-enabled only if you need matrix jobs to keep running past a failure.

PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs


感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 As a PR reviewer and CODEOWNER, I have reviewed this and have

如需进行 PR 验证,请为此 PR 添加 full-sweep-fail-fast 标签(强烈推荐)— 基准测试 sweep 仅在带有标签的 PR 上运行。仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled

PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档

@Klaud-Cold

Klaud-Cold commented Sep 11, 2026

Copy link
Copy Markdown
Collaborator Author

Initial attempt — failed (engine crash), superseded by Repair 1/5

  • Image: lmsysorg/sglang:v0.5.19-cu130 (Docker Hub manifest digest sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9, amd64 layer sha256:37bbbd34…, same manifest as the plain v0.5.19 tag; SGLang tag v0.5.19, Dockerfile base CUDA 13.0.3, sglang-kernel 0.4.6.post1, FlashInfer 0.6.18, NIXL cu13). The digest is recorded here rather than in the image string because the b200-nscale compat launcher passes the string straight to enroot import docker://…, which does not parse tag@digest.
  • Head: 40084c6186e90ab7fe197cb98516cb9a2eb5deaf on PR [Klaud Cold] Update dsr1-fp4-b200-dynamo-sglang-mtp SGLang image to v0.5.19-cu130 with Dynamo wheel 1.5.0.dev20260910 / 将 dsr1-fp4-b200-dynamo-sglang-mtp 的 SGLang 镜像更新至 v0.5.19-cu130 并搭配 Dynamo wheel 1.5.0.dev20260910 #3013 (draft, no sweep labels)
  • Run: 34593285808e2e-tests.yml on main, test-config --config-files configs/nvidia-master.yaml --config-keys dsr1-fp4-b200-dynamo-sglang-mtp --trim-conc, fail-fast, Klaud background priority; 9 deployment shapes at their lowest concurrency (2–6 nodes each) plus the generator's default eval selection. Smoke evidence only, not a full curve.
  • Changes: master imagev0.5.19-cu130; nine recipe YAMLs model.container matched; Dynamo install source hash: 5b4bc1ddwheel: "1.5.0.dev20260910" in the same recipes, mirrored in master router.version. No other fields touched.
  • Why the Dynamo change is required: Dynamo 5b4bc1dd (1.3.0, pyproject pins sglang==0.5.12.post1) imports sglang.srt.server_args_config_parser at worker start-up; SGLang v0.5.19 moved that module to sglang/srt/utils/server_args_config_parser.py and removed ServerArgs.use_mla_backend, so the shipped pair would fail at import. The published ai-dynamo 1.5.0.dev20260910 wheel (metadata pins sglang==0.5.19) carries dynamo/sglang/_compat.py handling both module paths, and every sglang.srt.managers.io_struct name and server_args.* attribute it touches exists in v0.5.19. srt-slurm a98738de (launcher pin, unchanged) prefetches the exact ai-dynamo / ai-dynamo-runtime wheels on the head node and installs them in-container with --no-deps --no-index, so the image's SGLang is untouched. Same image + wheel pair validated in [Klaud Cold] Update dsv4-fp4-gb300-dynamo-sglang-agentic-agg SGLang image to v0.5.19-cu130 / 将 dsv4-fp4-gb300-dynamo-sglang-agentic-agg 的 SGLang 镜像更新到 v0.5.19-cu130 #2969 (GB300).
  • SGLang flag audit (v0.5.12.post1v0.5.19): all recipe flags still parse. --cuda-graph-max-bs is a deprecated alias for --cuda-graph-max-bs-decode; --prefill-round-robin-balance and --disable-cuda-graph are deprecated warning-only flags; --enable-flashinfer-allreduce-fusion is retained. trtllm_mla, modelopt_fp4, flashinfer_trtllm (MoE and FP4 GEMM), fp8_e4m3, nixl, round_robin and EAGLE remain valid choices. Defaults for the used flags are unchanged. The SGLANG_ENABLE_SPEC_V2 env var was removed upstream (V2 is always on; the recipes' setting now only logs a warning). No runtime patching in the selected launch path.
  • Result: the Dynamo wheel installed and both workers started (server log confirms ai-dynamo 1.5.0.dev20260910 installed from staged wheels and only the expected deprecation warnings). The first job to reach a forward pass (1P tp4 / 1D tp8 c32, eval-only) crashed in the prefill worker, exit 137: AttributeError: 'MergedColumnParallelLinear' object has no attribute 'weight_swiglu_interleaved' at sglang/srt/models/deepseek_v2.py:342 (shared-experts DeepseekV2MLP.forward). Fail-fast cancelled the eval matrix; Klaud Cold cancelled the remaining benchmark jobs after confirming the cause, since all eight recipes with the same setting would fail identically.
  • Diagnosis: v0.5.19 deepseek_v2.py enables the NVFP4 GEMM+SwiGLU fusion for shared experts whenever the platform is SM100, the linear method is ModelOpt FP4 w4a4 and N % 128 == 0, independent of the FP4 GEMM backend. In modelopt_quant.py the flashinfer_trtllm branch of process_weights_after_loading returns before the _interleave_for_swiglu_fusion block that creates weight_swiglu_interleaved; the CUTLASS/CuTe-DSL path reaches it. main has the same structure as of 2026-09-11, so no newer release image avoids it. Eight of the nine recipes pin fp4-gemm-backend: flashinfer_trtllm; the c2048 recipe relies on the default and was not affected.
  • Next: Repair 1/5 — switch those eight recipes to the upstream default fp4-gemm-backend: "auto" (CuTe DSL on SM100). No engine patching.

初次尝试 —— 失败(引擎崩溃),由修复 1/5 接替

  • 镜像:lmsysorg/sglang:v0.5.19-cu130(Docker Hub manifest 摘要 sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9,amd64 层 sha256:37bbbd34…,与纯 v0.5.19 标签同一 manifest;SGLang 标签 v0.5.19,Dockerfile 基础 CUDA 13.0.3,sglang-kernel 0.4.6.post1,FlashInfer 0.6.18,NIXL cu13)。摘要记录在此而不写入镜像字符串,因为 b200-nscale 兼容启动器将字符串直接传给 enroot import docker://…,而 enroot 不解析 tag@digest
  • Head:PR [Klaud Cold] Update dsr1-fp4-b200-dynamo-sglang-mtp SGLang image to v0.5.19-cu130 with Dynamo wheel 1.5.0.dev20260910 / 将 dsr1-fp4-b200-dynamo-sglang-mtp 的 SGLang 镜像更新至 v0.5.19-cu130 并搭配 Dynamo wheel 1.5.0.dev20260910 #3013 上的 40084c6186e90ab7fe197cb98516cb9a2eb5deaf(草稿,无 sweep 标签)
  • 运行:34593285808 —— 在 main 上的 e2e-tests.ymltest-config --config-files configs/nvidia-master.yaml --config-keys dsr1-fp4-b200-dynamo-sglang-mtp --trim-conc,fail-fast,Klaud 后台优先级;9 个部署形态各取最低并发(每个 2–6 节点)加生成器默认评测选择。仅为冒烟证据,不是完整曲线。
  • 变更:主配置 imagev0.5.19-cu130;九个 recipe YAML 的 model.container 同步;同一批 recipe 的 Dynamo 安装源 hash: 5b4bc1ddwheel: "1.5.0.dev20260910",并同步主配置 router.version。未改动其他字段。
  • 为何必须更换 Dynamo:Dynamo 5b4bc1dd(1.3.0,pyproject 绑定 sglang==0.5.12.post1)在 worker 启动时导入 sglang.srt.server_args_config_parser;SGLang v0.5.19 已将该模块移至 sglang/srt/utils/server_args_config_parser.py 并移除 ServerArgs.use_mla_backend,原组合会在导入阶段失败。已发布的 ai-dynamo 1.5.0.dev20260910 wheel(元数据绑定 sglang==0.5.19)自带 dynamo/sglang/_compat.py 兼容两个模块路径,其用到的所有 sglang.srt.managers.io_struct 名称与 server_args.* 属性在 v0.5.19 中均存在。srt-slurm a98738de(启动器固定提交,未变)在头节点预取精确的 ai-dynamo / ai-dynamo-runtime wheel,并以 --no-deps --no-index 在容器内安装,不触碰镜像自带的 SGLang。同一镜像 + wheel 组合已在 [Klaud Cold] Update dsv4-fp4-gb300-dynamo-sglang-agentic-agg SGLang image to v0.5.19-cu130 / 将 dsv4-fp4-gb300-dynamo-sglang-agentic-agg 的 SGLang 镜像更新到 v0.5.19-cu130 #2969(GB300)中验证通过。
  • SGLang 参数审计(v0.5.12.post1v0.5.19):recipe 中所有参数仍可解析。--cuda-graph-max-bs--cuda-graph-max-bs-decode 的弃用别名;--prefill-round-robin-balance--disable-cuda-graph 为仅告警的弃用参数;--enable-flashinfer-allreduce-fusion 保留。trtllm_mlamodelopt_fp4flashinfer_trtllm(MoE 与 FP4 GEMM)、fp8_e4m3nixlround_robinEAGLE 仍为有效选项。所用参数默认值未变。上游已移除环境变量 SGLANG_ENABLE_SPEC_V2(V2 始终开启;recipe 中的设置现仅打印告警)。所选启动路径无任何运行时补丁。
  • 结果:Dynamo wheel 安装成功,两个 worker 均已启动(服务器日志确认从暂存 wheel 安装了 ai-dynamo 1.5.0.dev20260910,仅有预期的弃用告警)。首个进入前向计算的作业(1P tp4 / 1D tp8 c32,eval-only)在 prefill worker 崩溃,退出码 137:AttributeError: 'MergedColumnParallelLinear' object has no attribute 'weight_swiglu_interleaved',位置 sglang/srt/models/deepseek_v2.py:342(共享专家 DeepseekV2MLP.forward)。fail-fast 取消了评测矩阵;Klaud Cold 在确认原因后取消了其余基准作业,因为使用相同设置的八个 recipe 会以同样方式失败。
  • 诊断:v0.5.19 的 deepseek_v2.py 只要平台为 SM100、线性层为 ModelOpt FP4 w4a4 且 N % 128 == 0,就为共享专家启用 NVFP4 GEMM+SwiGLU 融合,与 FP4 GEMM 后端无关。而 modelopt_quant.pyprocess_weights_after_loadingflashinfer_trtllm 分支在创建 weight_swiglu_interleaved_interleave_for_swiglu_fusion 代码块之前就返回了;CUTLASS/CuTe-DSL 路径则会执行到该块。截至 2026-09-11,main 结构相同,没有更新的发布镜像可以规避。九个 recipe 中有八个固定 fp4-gemm-backend: flashinfer_trtllm;c2048 recipe 使用默认值,不受影响。
  • 下一步:修复 1/5 —— 将这八个 recipe 切换为上游默认值 fp4-gemm-backend: "auto"(SM100 上为 CuTe DSL)。不做任何引擎补丁。

…p on SGLang v0.5.19

SGLang v0.5.19 enables the NVFP4 GEMM+SwiGLU fusion for DeepSeek shared
experts on SM100 (deepseek_v2.py sets _interleave_for_swiglu_fusion), but
ModelOptFp4LinearMethod.process_weights_after_loading returns from the
flashinfer_trtllm branch before building weight_swiglu_interleaved, so the
prefill worker crashes with AttributeError at its first forward. Switch the
eight recipes that pinned --fp4-gemm-backend flashinfer_trtllm to the
upstream default "auto" (flashinfer_cutedsl on SM100), which reaches the
interleave path. MoE runner backend, attention backend, quantization,
topology, MTP and concurrency settings are unchanged.

SGLang v0.5.19 会在 SM100 上为 DeepSeek 共享专家启用 NVFP4 GEMM+SwiGLU 融合
(deepseek_v2.py 设置 _interleave_for_swiglu_fusion),但
ModelOptFp4LinearMethod.process_weights_after_loading 在 flashinfer_trtllm
分支中提前返回,未构建 weight_swiglu_interleaved,导致 prefill worker 首次前向
即抛出 AttributeError。将八个固定 --fp4-gemm-backend flashinfer_trtllm 的
recipe 切换为上游默认值 "auto"(SM100 上解析为 flashinfer_cutedsl),该路径会
执行交织处理。MoE runner 后端、attention 后端、量化、拓扑、MTP 与并发设置均不变。

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@Klaud-Cold

Klaud-Cold commented Sep 11, 2026

Copy link
Copy Markdown
Collaborator Author

Repair 1/5 — FP4 GEMM backend flashinfer_trtllmauto: engine fixed, run stopped by the main result-collection blocker

  • Head: 186502243dff7051fb1e2d2baa069f8db7e1d1f0 on PR [Klaud Cold] Update dsr1-fp4-b200-dynamo-sglang-mtp SGLang image to v0.5.19-cu130 with Dynamo wheel 1.5.0.dev20260910 / 将 dsr1-fp4-b200-dynamo-sglang-mtp 的 SGLang 镜像更新至 v0.5.19-cu130 并搭配 Dynamo wheel 1.5.0.dev20260910 #3013 (draft, no sweep labels); image unchanged lmsysorg/sglang:v0.5.19-cu130, Dynamo wheel unchanged 1.5.0.dev20260910
  • Run: 34594623538 — same test-config … --trim-conc smoke as the initial attempt, fail-fast, Klaud background priority
  • Change: in the eight recipes that pinned fp4-gemm-backend: "flashinfer_trtllm" (prefill and decode), set fp4-gemm-backend: "auto", the v0.5.19 default that fp4_utils.py resolves to flashinfer_cutedsl on SM100. The 8k1k_mtp_4p1d_c2048.yaml recipe already used the default and is untouched. moe-runner-backend flashinfer_trtllm, attention-backend trtllm_mla, quantization modelopt_fp4, topologies, MTP and concurrencies are unchanged. The validated single-node dsr1-fp4-b200-sglang scripts on the same image ([Klaud Cold] Update dsr1-fp4-b200-sglang SGLang image to v0.5.19-cu130 / 将 dsr1-fp4-b200-sglang 的 SGLang 镜像更新至 v0.5.19-cu130 #2989) also leave the FP4 GEMM backend at its default.
  • Why this is in scope: the crash in the initial attempt comes from the shipped engine's flashinfer_trtllm dense FP4 GEMM path skipping the SwiGLU-interleave weight preparation that v0.5.19's DeepSeek shared-experts fusion requires. Selecting a supported upstream backend value in the recipe YAML lets the image run as shipped, with no engine or launcher patching.
  • Engine result: no AttributeError in any worker. The 4P dep4 / 1D dep8 c2048 benchmark (untouched recipe, 3 nodes) completed SA-Bench with 20480/20480 requests, and the 1P tp4 / 1D tp8 c32 eval-only job (modified recipe, the shape that crashed before) became healthy and passed gsm8k: exact_match,strict-match 0.9568, flexible-extract 0.9583 (published baseline for this point: 0.9553 / 0.9560; validator threshold 0.91). The other seven shapes did not reach a forward pass before cancellation.
  • Smoke delta for the one completed benchmark point (from the server-log artifact's SA-Bench JSON, not a pipeline result; single smoke point, not a curve; tput = total tok/s per GPU over 24 GPUs, out = output tok/s per decode GPU over 8 GPUs):
Point Metric Baseline 2026-07-08 Repair 1/5 Delta
4P dep4 / 1D dep8 c2048 tput/GPU 8534.8 8783.7 +2.9%
4P dep4 / 1D dep8 c2048 out/GPU 2847.2 2930.3 +2.9%
4P dep4 / 1D dep8 c2048 median TTFT (s) 60.718 55.498 −8.6%
4P dep4 / 1D dep8 c2048 median TPOT (ms) 21.72 25.55 +17.6%
4P dep4 / 1D dep8 c2048 median E2E (s) 79.82 77.98 −2.3%

Other 12 points: N/A (not run; benchmark matrix cancelled by fail-fast after the collection failure below).

  • Blocker: that c2048 job then failed in copy_fixed_sequence_results (runners/slurm_utils.sh:94, shared launcher code on main since fix(klaud): enforce completion and improve sweeps / 完善收尾与 sweep 可靠性 #2933) because utils/result_filename.py is resolved through a relative BASH_SOURCE path after the launcher cd'd into srt-slurm/; see the milestone. Fail-fast cancelled the eight other benchmark jobs. Klaud Cold cancelled the remaining eval jobs (three running, five queued) at 12:29 UTC after the 1P/1D c32 eval evidence landed, since no further job could produce a collectable benchmark result.
  • Outcome: engine repair confirmed; candidate closed as readiness-blocked (shared launcher defect outside this family's scope). On retry after the main fix, apply the same recipe change (Dynamo wheel 1.5.0.dev20260910, fp4-gemm-backend: "auto").

修复 1/5 —— FP4 GEMM 后端 flashinfer_trtllmauto:引擎已修复,运行因 main 的结果收集阻塞而停止

  • Head:PR [Klaud Cold] Update dsr1-fp4-b200-dynamo-sglang-mtp SGLang image to v0.5.19-cu130 with Dynamo wheel 1.5.0.dev20260910 / 将 dsr1-fp4-b200-dynamo-sglang-mtp 的 SGLang 镜像更新至 v0.5.19-cu130 并搭配 Dynamo wheel 1.5.0.dev20260910 #3013 上的 186502243dff7051fb1e2d2baa069f8db7e1d1f0(草稿,无 sweep 标签);镜像不变 lmsysorg/sglang:v0.5.19-cu130,Dynamo wheel 不变 1.5.0.dev20260910
  • 运行:34594623538 —— 与初次尝试相同的 test-config … --trim-conc 冒烟,fail-fast,Klaud 后台优先级
  • 变更:在固定 fp4-gemm-backend: "flashinfer_trtllm"(prefill 与 decode)的八个 recipe 中,改为 fp4-gemm-backend: "auto",即 v0.5.19 的默认值,fp4_utils.py 在 SM100 上解析为 flashinfer_cutedsl8k1k_mtp_4p1d_c2048.yaml 已使用默认值,未改动。moe-runner-backend flashinfer_trtllmattention-backend trtllm_mlaquantization modelopt_fp4、拓扑、MTP 与并发均不变。同一镜像上已验证的单节点 dsr1-fp4-b200-sglang 脚本([Klaud Cold] Update dsr1-fp4-b200-sglang SGLang image to v0.5.19-cu130 / 将 dsr1-fp4-b200-sglang 的 SGLang 镜像更新至 v0.5.19-cu130 #2989)同样保留 FP4 GEMM 后端默认值。
  • 为何在范围内:初次尝试的崩溃源于镜像自带引擎的 flashinfer_trtllm 稠密 FP4 GEMM 路径跳过了 v0.5.19 DeepSeek 共享专家融合所需的 SwiGLU 交织权重准备。在 recipe YAML 中选择上游支持的后端取值,可让镜像按原样运行,无需任何引擎或启动器补丁。
  • 引擎结果:所有 worker 均无 AttributeError。4P dep4 / 1D dep8 c2048 基准(未改动的 recipe,3 节点)完成 SA-Bench,20480/20480 请求;1P tp4 / 1D tp8 c32 eval-only 作业(被修改的 recipe,即此前崩溃的形态)服务健康并通过 gsm8k:exact_match,strict-match 0.9568,flexible-extract 0.9583(该点已发布基线:0.9553 / 0.9560;校验阈值 0.91)。其余七个形态在取消前未进入前向计算。
  • 唯一完成的基准点的冒烟差异(来自服务器日志工件中的 SA-Bench JSON,非流水线结果;单个冒烟点,非完整曲线;tput = 24 GPU 上每 GPU 总 tok/s,out = 8 个 decode GPU 上每 GPU 输出 tok/s),数值见上表。其余 12 个点:N/A(未运行;基准矩阵在下述收集失败后被 fail-fast 取消)。
  • 阻塞:该 c2048 作业随后在 copy_fixed_sequence_resultsrunners/slurm_utils.sh:94,自 fix(klaud): enforce completion and improve sweeps / 完善收尾与 sweep 可靠性 #2933 起位于 main 的共享启动器代码)失败,原因是启动器 cd 进入 srt-slurm/ 后仍通过相对 BASH_SOURCE 路径解析 utils/result_filename.py;详见里程碑。fail-fast 取消了其余八个基准作业。在 1P/1D c32 评测证据到达后,Klaud Cold 于 12:29 UTC 取消了剩余评测作业(三个运行中、五个排队中),因为后续作业均无法产出可收集的基准结果。
  • 结果:引擎修复已确认;候选以 readiness-blocked 关闭(共享启动器缺陷超出本系列范围)。main 修复后重试时,请应用同样的 recipe 变更(Dynamo wheel 1.5.0.dev20260910fp4-gemm-backend: "auto")。

@Klaud-Cold

Copy link
Copy Markdown
Collaborator Author

Milestone — Repair 1/5 engine fix works; benchmark result collection is blocked by shared launcher code on main

  • In run 34594623538 (Repair 1/5) the 4P dep4 / 1D dep8 c2048 point started all five workers on v0.5.19-cu130 + Dynamo 1.5.0.dev20260910, and SA-Bench finished (Benchmark completed successfully, 20480/20480 requests, run 13976). Its result file results_concurrency_2048_gpus_24_ctx_16_gen_8.json exists in the uploaded server-log artifact.
  • The job then failed in result collection, not in the engine: runners/slurm_utils.sh:94 (copy_fixed_sequence_results, added on main 2026-09-09 by fix(klaud): enforce completion and improve sweeps / 完善收尾与 sweep 可靠性 #2933) calls python3 "$(dirname "${BASH_SOURCE[0]}")/../utils/result_filename.py". The launcher is started as bash ./runners/launch_….sh, so BASH_SOURCE is relative, and the compat launcher has already cd'd into srt-slurm/; the log shows python3: can't open file '…/InferenceX/srt-slurm/./runners/../utils/result_filename.py'. The helper's output is empty, the result is copied to the workspace root with no name, and the job ends with Run failed: No benchmark result files found. The same helper is used by the b300-dsxe, gb300-nv, h100-dgxc and h200-dgxc launchers. Fail-fast then cancelled the eight other benchmark jobs.
  • This is outside the candidate's edit scope (shared launcher code), and no open PR touches runners/slurm_utils.sh or utils/result_filename.py as of 2026-09-11 12:25 UTC. Until main resolves that helper path once at source time (absolute path), no fixed-sequence multinode benchmark on this path can produce a result, whatever image is used.
  • Plan: let the already-running 1P tp4 / 1D tp8 c32 eval-only job finish as evidence that the auto FP4 GEMM backend also fixes the eight modified recipes, then cancel the remaining eval jobs and close this attempt as readiness-blocked so the candidate can be retried once the launcher fix lands. The recipe change needed on retry is the one in Repair 1/5.

里程碑 —— 修复 1/5 的引擎修复有效;基准结果收集被 main 上的共享启动器代码阻塞

  • 运行 34594623538修复 1/5)中,4P dep4 / 1D dep8 c2048 点在 v0.5.19-cu130 + Dynamo 1.5.0.dev20260910 上启动了全部五个 worker,SA-Bench 顺利完成(Benchmark completed successfully,20480/20480 请求,作业 13976)。结果文件 results_concurrency_2048_gpus_24_ctx_16_gen_8.json 存在于上传的服务器日志工件中。
  • 作业随后在结果收集阶段失败,而非引擎:runners/slurm_utils.sh:94copy_fixed_sequence_results,2026-09-09 由 fix(klaud): enforce completion and improve sweeps / 完善收尾与 sweep 可靠性 #2933 加入 main)调用 python3 "$(dirname "${BASH_SOURCE[0]}")/../utils/result_filename.py"。启动器以 bash ./runners/launch_….sh 启动,BASH_SOURCE 为相对路径,而兼容启动器此时已 cd 进入 srt-slurm/;日志显示 python3: can't open file '…/InferenceX/srt-slurm/./runners/../utils/result_filename.py'。辅助脚本输出为空,结果被复制到工作区根目录且没有文件名,作业以 Run failed: No benchmark result files found 结束。b300-dsxe、gb300-nv、h100-dgxc、h200-dgxc 启动器也使用同一辅助函数。fail-fast 随后取消了其余八个基准作业。
  • 这超出本候选的编辑范围(共享启动器代码),截至 2026-09-11 12:25 UTC 没有任何开放 PR 涉及 runners/slurm_utils.shutils/result_filename.py。在 main 将该辅助脚本路径改为在 source 时一次性解析为绝对路径之前,此路径上的任何固定序列多节点基准都无法产出结果,与镜像无关。
  • 计划:让已在运行的 1P tp4 / 1D tp8 c32 eval-only 作业完成,以证明 auto FP4 GEMM 后端同样修复了八个被修改的 recipe;随后取消其余评测作业,并以 readiness-blocked 结束本次尝试,以便启动器修复合入后重试本候选。重试时所需的 recipe 变更即修复 1/5 中的变更。

@Klaud-Cold

Copy link
Copy Markdown
Collaborator Author

Klaud Cold: readiness-blocked. Finishing cleanup; owned child runs will be stopped and checked before closure.


Klaud Cold:readiness-blocked。正在完成清理;将先停止并确认自有子运行的状态,再关闭 PR。

@Klaud-Cold Klaud-Cold closed this Sep 11, 2026
@Klaud-Cold
Klaud-Cold deleted the klaud/auto-39cf82a2bb6e3c04-3cf2069b9d51c51a branch September 11, 2026 12:33
@Klaud-Cold

Copy link
Copy Markdown
Collaborator Author

Klaud Cold: readiness-blocked. All owned runs are terminal. Repairs: 1. Runs: 34593285808, 34594623538.

PR closed; branch deleted for retry.


Klaud Cold:readiness-blocked。所有自有运行均已结束。修复次数:1。运行:34593285808, 34594623538

PR 已关闭;分支已删除,后续可以重试。

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Development

Successfully merging this pull request may close these issues.

1 participant