Skip to content

[Klaud Cold] Update dsr1-fp4-b200-sglang SGLang image to v0.5.19-cu130 / 将 dsr1-fp4-b200-sglang 的 SGLang 镜像更新至 v0.5.19-cu130 - #2989

Open
Klaud-Cold wants to merge 2 commits into
mainfrom
klaud/auto-d92e15065c55779d-5f146332463a3df7
Open

[Klaud Cold] Update dsr1-fp4-b200-sglang SGLang image to v0.5.19-cu130 / 将 dsr1-fp4-b200-sglang 的 SGLang 镜像更新至 v0.5.19-cu130#2989
Klaud-Cold wants to merge 2 commits into
mainfrom
klaud/auto-d92e15065c55779d-5f146332463a3df7

Conversation

@Klaud-Cold

@Klaud-Cold Klaud-Cold commented Sep 11, 2026

Copy link
Copy Markdown
Collaborator

Update the dsr1-fp4-b200-sglang master image from lmsysorg/sglang:v0.5.16-cu130 to the current SGLang release lmsysorg/sglang:v0.5.19-cu130 (Docker Hub digest sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9, build commit sgl-project/sglang@0bcd822 = tag v0.5.19). Model, TP4/EP1 and TP4/EP4 DP-attention topologies, the 8k1k workload, concurrency ranges and benchmarks/single_node/fixed_seq_len/dsr1_fp4_b200.sh are unchanged.

Baseline

  • Published date: 2026-08-06 (workflow-info?date=2026-08-06; benchmarks?model=DeepSeek-R1-0528&date=2026-08-06&exact=true&sequence=8k/1k, filtered to hardware b200, framework sglang, precision fp4, spec none, non-disagg, ISL/OSL 8192/1024; evaluations?model=DeepSeek-R1-0528&date=2026-08-06&exact=true)
  • Old image: lmsysorg/sglang:v0.5.16-cu130 (Docker Hub digest sha256:7b6a35df9839fd593a94a1eaee82d7777f472225d9f3ad1f8a2e0cb2bd1785d0, build commit sgl-project/sglang@fdebc93 = tag v0.5.16, CUDA 13.0.1, FlashInfer 0.6.14, sgl-kernel 0.4.5)
  • Workload / topology: single-node B200 (cluster:b200-nscale), nvidia/DeepSeek-R1-0528-FP4-V2, SGLang NVFP4, no speculative decoding, fixed-seq-len 8k1k (ISL 8192 / OSL 1024), random dataset; TP4/EP1 at concurrency 1–32 and TP4/EP4 DP-attention at concurrency 64–256 (9 points)
  • Producer: all 9 points come from run 30952773184 attempt 1 (head d4363bd7fd5bda1391d2d0a46d834cecc363c0d6, changelog PR #2492); benchmark result IDs 438719 … (curve snapshot curve_workflow_run_id 2261 is a logical snapshot, not the producer)
  • Published evals: N/A — the evaluations feed has no gsm8k row for this identity dated 2026-08-06; the newest published gsm8k for this identity is 2026-03-28 on an older image and is not comparable
Point tput/GPU (tok/s) output tput/GPU (tok/s) median TPOT (ms) median TTFT (ms)
TP4/EP1 c1 389.7 43.5 5.52 207
TP4/EP1 c2 672.2 75.4 6.37 245
TP4/EP1 c4 1134.2 126.2 7.48 245
TP4/EP1 c8 1778.0 199.9 9.49 269
TP4/EP1 c16 2558.7 283.5 13.27 494
TP4/EP1 c32 3428.7 383.4 19.75 769
TP4/EP4 DPA c64 4603.2 510.7 29.73 857
TP4/EP4 DPA c128 6249.9 692.3 44.39 831
TP4/EP4 DPA c256 7955.2 884.2 66.38 4092

dsr1-fp4-b200-sglang 的主配置镜像从 lmsysorg/sglang:v0.5.16-cu130 更新至当前 SGLang 发布版 lmsysorg/sglang:v0.5.19-cu130(Docker Hub 摘要 sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9,构建提交 sgl-project/sglang@0bcd822 = 标签 v0.5.19)。模型、TP4/EP1 与 TP4/EP4 DP-attention 拓扑、8k1k 工作负载、并发范围以及 benchmarks/single_node/fixed_seq_len/dsr1_fp4_b200.sh 均保持不变。

基线

  • 发布日期: 2026-08-06(workflow-info?date=2026-08-06benchmarks?model=DeepSeek-R1-0528&date=2026-08-06&exact=true&sequence=8k/1k,按硬件 b200、框架 sglang、精度 fp4、无投机解码、非分离式、ISL/OSL 8192/1024 过滤;evaluations?model=DeepSeek-R1-0528&date=2026-08-06&exact=true
  • 旧镜像: lmsysorg/sglang:v0.5.16-cu130(Docker Hub 摘要 sha256:7b6a35df9839fd593a94a1eaee82d7777f472225d9f3ad1f8a2e0cb2bd1785d0,构建提交 sgl-project/sglang@fdebc93 = 标签 v0.5.16,CUDA 13.0.1,FlashInfer 0.6.14,sgl-kernel 0.4.5)
  • 工作负载 / 拓扑: 单节点 B200(cluster:b200-nscale),nvidia/DeepSeek-R1-0528-FP4-V2,SGLang NVFP4,无投机解码,固定序列长度 8k1k(ISL 8192 / OSL 1024),随机数据集;TP4/EP1 并发 1–32,TP4/EP4 DP-attention 并发 64–256(共 9 个点)
  • 生产运行: 全部 9 个点均来自 run 30952773184 attempt 1(head d4363bd7fd5bda1391d2d0a46d834cecc363c0d6,changelog PR #2492);曲线快照 curve_workflow_run_id 2261 仅为逻辑快照,不是生产运行
  • 已发布评测: N/A —— 评测接口中没有该配置在 2026-08-06 的 gsm8k 记录;该配置最新的已发布 gsm8k 为 2026-03-28、基于更旧镜像,不可比较
点位 每 GPU 吞吐 (tok/s) 每 GPU 输出吞吐 (tok/s) TPOT 中位数 (ms) TTFT 中位数 (ms)
TP4/EP1 c1 389.7 43.5 5.52 207
TP4/EP1 c2 672.2 75.4 6.37 245
TP4/EP1 c4 1134.2 126.2 7.48 245
TP4/EP1 c8 1778.0 199.9 9.49 269
TP4/EP1 c16 2558.7 283.5 13.27 494
TP4/EP1 c32 3428.7 383.4 19.75 769
TP4/EP4 DPA c64 4603.2 510.7 29.73 857
TP4/EP4 DPA c128 6249.9 692.3 44.39 831
TP4/EP4 DPA c256 7955.2 884.2 66.38 4092

🤖 Generated with Claude Code


Note

Low Risk
Config and changelog-only change; no benchmark script or topology edits, only a pinned container image version update.

Overview
Bumps the dsr1-fp4-b200-sglang perf config container image from lmsysorg/sglang:v0.5.16-cu130 to lmsysorg/sglang:v0.5.19-cu130 in configs/nvidia-master.yaml. The new image brings CUDA 13.0.3, FlashInfer 0.6.18, and sgl-kernel 0.4.6.post1 (vs 13.0.1 / 0.6.14 / 0.4.5 on the old tag).

A matching perf-changelog.yaml entry records the digest, build commit, and states that model, B200 topology (TP4/EP1 and TP4/EP4 DP-attention), 8k1k fixed-seq-len grid, and benchmarks/single_node/fixed_seq_len/dsr1_fp4_b200.sh are unchanged—so this is a runtime refresh for comparable re-benchmarking only.

Reviewed by Cursor Bugbot for commit 25f8f98. Bugbot is set up for automated code reviews on this repo. Configure here.

@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase As a PR reviewer and CODEOWNER, I have reviewed this and have.

For PR verification, add the full-sweep-fail-fast label (strongly recommended) to this PR — the benchmark sweep only runs on labeled PRs. Use full-sweep-enabled only if you need matrix jobs to keep running past a failure.

PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs


感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 As a PR reviewer and CODEOWNER, I have reviewed this and have

如需进行 PR 验证,请为此 PR 添加 full-sweep-fail-fast 标签(强烈推荐)— 基准测试 sweep 仅在带有标签的 PR 上运行。仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled

PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档

@Klaud-Cold

Klaud-Cold commented Sep 11, 2026

Copy link
Copy Markdown
Collaborator Author

Initial attempt

  • Image / SHA: lmsysorg/sglang:v0.5.19-cu130 (Docker Hub digest sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9, build commit sgl-project/sglang@0bcd822 = tag v0.5.19, image created 2026-09-04) at PR head 2e86dbaa83a43e468d874965c3e16b787f157331
  • Change: master image only (configs/nvidia-master.yaml, key dsr1-fp4-b200-sglang); benchmarks/single_node/fixed_seq_len/dsr1_fp4_b200.sh unchanged, no runtime patching; the dsr1-fp4-b200-sglang-mtp sibling is not touched
  • Targeted run: e2e run 34548446011 (e2e-tests.yml from main, ref=2e86dbaa, test-config --config-files configs/nvidia-master.yaml --config-keys dsr1-fp4-b200-sglang --trim-conc, fail-fast, Klaud background priority; dispatched 2026-09-11T00:54Z) — completed, success (all jobs terminal: both benchmark jobs and both eval-only jobs succeeded; collect-results, collect-evals and calc-success-rate green; finished 2026-09-11T03:04Z). Points: TP4/EP1 concurrency 1 and TP4/EP4 DP-attention concurrency 64, both with the gsm8k smoke eval. Startup/compatibility check only, not the full curve.
  • Upstream source comparison (v0.5.16 fdebc93v0.5.19 0bcd822, image labels confirm both build commits):
    • python/sglang/srt/server_args.py: every flag the script passes still exists as a ServerArgs field in both versions (enable_dp_attention, enable_dp_attention_local_control_broadcast, enable_dp_lm_head, schedule_conservativeness, enable_prefill_delayer, scheduler_recv_interval, enable_symm_mem, moe_runner_backend, attention_backend, kv_cache_dtype, mem_fraction_static, max_running_requests, ep_size, quantization, stream_interval, chunked_prefill_size, enable_flashinfer_allreduce_fusion, context_length). --cuda-graph-max-bs and --disable-piecewise-cuda-graph are deprecated aliases for --cuda-graph-max-bs-decode and --cuda-graph-backend-prefill=disabled in both versions, with identical mapping, so no rename affects the recipe. SGLANG_RADIX_FORCE_MISS is still defined in environ.py and consumed in managers/schedule_policy.py in both versions.
    • Bundled stack (python/pyproject.toml and image layers): torch 2.11.0 → 2.13.0, FlashInfer 0.6.14 → 0.6.18, sgl-kernel 0.4.5 → 0.4.6.post1, DeepGEMM 0.1.4.post1 → 0.1.7, CUDA 13.0.1 → 13.0.3, cuDNN 9.13 → 9.14; torchao removed.
    • Behavior changes on this recipe's path, from the v0.5.17, v0.5.18 and v0.5.19 notes: MoE deferred finalize is on by default for NVFP4 + flashinfer_trtllm DeepSeek-V3-family models (#33618); FlashInfer MNNVL pure allreduce is auto-enabled for DeepSeek-V3-family models (#30700); piecewise prefill graph is avoided for trtllm_mla (#32785); breakable prefill CUDA graph became the DP-attention default (#31682) but the recipe pins the prefill graph backend to disabled; the unified radix tree is now the default cache (#35081) while the recipe forces cache misses; DP-attention decode→extend prefix off-by-one fix (#37505). These are default/implementation changes, not option removals, so the launch command is kept as is.
    • Cluster evidence: qwen3.5-fp4-b200-sglang already runs lmsysorg/sglang:v0.5.19-cu130 on cluster:b200-nscale with a green full sweep (#2954).
  • Benchmark results (smoke points, results_bmk/agg_bmk.json, image lmsysorg/sglang:v0.5.19-cu130, vs. the frozen 2026-08-06 baseline; per-GPU throughput and median latencies):
Point tput/GPU new (tok/s) tput/GPU base Δ tput median TPOT new (ms) base Δ TPOT median TTFT new (ms) base Δ TTFT
TP4/EP1 c1 406.9 389.7 +4.4% 5.27 5.52 −4.5% 209 207 +0.6%
TP4/EP4 DPA c64 4942.5 4603.2 +7.4% 27.83 29.73 −6.4% 757 857 −11.7%
  • Eval results: TP4/EP4 DPA c64 gsm8k exact_match strict 0.9560 (flexible 0.9560, n=1319, stderr 0.0056), meta_env.json matches TP4/EP4/DPA conc 64; published baseline eval N/A (no gsm8k row for this identity dated 2026-08-06). TP4/EP1 c1 gsm8k exact_match strict 0.9515 (flexible 0.9538, n=1319, stderr 0.0059), meta_env.json matches TP4/EP1 conc 1. Both evals ran against the smoke server's own image; eval_results_all/agg_eval_all.json lists both rows with infrastructure_success: true.
  • Outcome: targeted smoke passed (startup, both deployment shapes, both smoke evals). This is targeted success only, not final validation.
  • Next step: append the perf-changelog entry, recheck capacity and start the final full sweep (full-sweep-enabled) on the new PR head.

初始尝试

  • 镜像 / SHA: lmsysorg/sglang:v0.5.19-cu130(Docker Hub 摘要 sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9,构建提交 sgl-project/sglang@0bcd822 = 标签 v0.5.19,镜像创建于 2026-09-04),PR head 2e86dbaa83a43e468d874965c3e16b787f157331
  • 改动: 仅主配置镜像(configs/nvidia-master.yaml,键 dsr1-fp4-b200-sglang);benchmarks/single_node/fixed_seq_len/dsr1_fp4_b200.sh 未改动,无运行时补丁;未触碰 dsr1-fp4-b200-sglang-mtp 同族配方
  • 定向运行: e2e run 34548446011(从 main 触发 e2e-tests.ymlref=2e86dbaatest-config --config-files configs/nvidia-master.yaml --config-keys dsr1-fp4-b200-sglang --trim-conc,fail-fast,Klaud 后台优先级;2026-09-11T00:54Z 触发)—— 已完成,成功(所有作业均已结束:两个基准作业与两个仅评测作业均成功;collect-results、collect-evals 与 calc-success-rate 均为绿色;2026-09-11T03:04Z 结束)。点位:TP4/EP1 并发 1 与 TP4/EP4 DP-attention 并发 64,均带 gsm8k 冒烟评测。仅为启动/兼容性检查,不是完整曲线。
  • 上游源码对比v0.5.16 fdebc93v0.5.19 0bcd822,两个镜像的标签均确认了构建提交):
    • python/sglang/srt/server_args.py:脚本传入的所有参数在两个版本中都仍是 ServerArgs 字段。--cuda-graph-max-bs--disable-piecewise-cuda-graph 在两个版本中都是 --cuda-graph-max-bs-decode--cuda-graph-backend-prefill=disabled 的弃用别名,映射一致,没有重命名影响本配方。SGLANG_RADIX_FORCE_MISS 在两个版本的 environ.py 中仍有定义,并在 managers/schedule_policy.py 中被使用。
    • 捆绑组件(python/pyproject.toml 与镜像层):torch 2.11.0 → 2.13.0,FlashInfer 0.6.14 → 0.6.18,sgl-kernel 0.4.5 → 0.4.6.post1,DeepGEMM 0.1.4.post1 → 0.1.7,CUDA 13.0.1 → 13.0.3,cuDNN 9.13 → 9.14;移除了 torchao。
    • 影响本配方路径的行为变化(来自 v0.5.17v0.5.18v0.5.19 发布说明):NVFP4 + flashinfer_trtllm 的 DeepSeek-V3 系列默认开启 MoE deferred finalize(#33618);DeepSeek-V3 系列自动启用 FlashInfer MNNVL 纯 allreduce(#30700);trtllm_mla 不再使用分段 prefill 图(#32785);DP-attention 默认改为可中断 prefill CUDA 图(#31682),但本配方将 prefill 图后端固定为 disabled;统一 radix 树成为默认缓存(#35081),而本配方强制缓存未命中;修复 DP-attention decode→extend 前缀差一错误(#37505)。这些都是默认值/实现层面的变化,并非删除选项,因此启动命令保持不变。
    • 集群证据:qwen3.5-fp4-b200-sglang 已在 cluster:b200-nscale 上使用 lmsysorg/sglang:v0.5.19-cu130 并通过完整扫描(#2954)。
  • 基准结果(冒烟点位,results_bmk/agg_bmk.json,镜像 lmsysorg/sglang:v0.5.19-cu130,对比冻结的 2026-08-06 基线;每 GPU 吞吐与中位延迟):
点位 新每 GPU 吞吐 (tok/s) 基线 Δ 吞吐 新 TPOT 中位数 (ms) 基线 Δ TPOT 新 TTFT 中位数 (ms) 基线 Δ TTFT
TP4/EP1 c1 406.9 389.7 +4.4% 5.27 5.52 −4.5% 209 207 +0.6%
TP4/EP4 DPA c64 4942.5 4603.2 +7.4% 27.83 29.73 −6.4% 757 857 −11.7%
  • 评测结果: TP4/EP4 DPA c64 gsm8k exact_match strict 0.9560(flexible 0.9560,n=1319,标准误 0.0056),meta_env.json 与 TP4/EP4/DPA 并发 64 一致;已发布基线评测 N/A(该配置没有 2026-08-06 的 gsm8k 记录)。TP4/EP1 c1 gsm8k exact_match strict 0.9515(flexible 0.9538,n=1319,标准误 0.0059),meta_env.json 与 TP4/EP1 并发 1 一致。两次评测均针对冒烟服务自身的镜像运行;eval_results_all/agg_eval_all.json 包含两行且 infrastructure_success: true
  • 结果: 定向冒烟通过(启动、两种部署形态、两次冒烟评测)。这只是定向成功,不是最终验证。
  • 下一步: 追加 perf-changelog 条目,重新检查容量,并在新的 PR head 上启动最终完整扫描(full-sweep-enabled)。

Klaud-Cold and others added 2 commits September 11, 2026 03:05
Move the B200 DeepSeek-R1-0528 NVFP4 single-node SGLang recipe from
lmsysorg/sglang:v0.5.16-cu130 (build commit
sgl-project/sglang@fdebc93, CUDA 13.0.1,
FlashInfer 0.6.14, sgl-kernel 0.4.5) to the v0.5.19 release image
lmsysorg/sglang:v0.5.19-cu130 (Docker Hub digest
sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9,
build commit sgl-project/sglang@0bcd822
= tag v0.5.19, CUDA 13.0.3, FlashInfer 0.6.18, sgl-kernel 0.4.6.post1).
Model, TP4/EP1 and TP4/EP4 DP-attention topologies, the 8k1k workload,
concurrency ranges and benchmarks/single_node/fixed_seq_len/dsr1_fp4_b200.sh
are unchanged; the dsr1-fp4-b200-sglang-mtp sibling is not touched.

将 B200 DeepSeek-R1-0528 NVFP4 单节点 SGLang 配方的镜像从
lmsysorg/sglang:v0.5.16-cu130(构建提交
sgl-project/sglang@fdebc93,CUDA 13.0.1,
FlashInfer 0.6.14,sgl-kernel 0.4.5)切换到 v0.5.19 发布版
lmsysorg/sglang:v0.5.19-cu130(Docker Hub 摘要
sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9,
构建提交 sgl-project/sglang@0bcd822
= 标签 v0.5.19,CUDA 13.0.3,FlashInfer 0.6.18,sgl-kernel 0.4.6.post1)。
模型、TP4/EP1 与 TP4/EP4 DP-attention 拓扑、8k1k 工作负载、并发范围以及
benchmarks/single_node/fixed_seq_len/dsr1_fp4_b200.sh 均保持不变;
未改动 dsr1-fp4-b200-sglang-mtp 同族配方。

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Append the perf-changelog entry for moving the B200 DeepSeek-R1-0528 NVFP4
single-node SGLang recipe to lmsysorg/sglang:v0.5.19-cu130 (PR #2989).
The entry selects the whole family without scenario, eval-selection or
append-only modifiers; all prior bytes are preserved.

为将 B200 DeepSeek-R1-0528 NVFP4 单节点 SGLang 配方切换到
lmsysorg/sglang:v0.5.19-cu130(PR #2989)追加 perf-changelog 条目。
该条目选择整个配方族,不带 scenario、评测选择或 append-only 修饰符;
所有既有字节均保持不变。

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@Klaud-Cold
Klaud-Cold force-pushed the klaud/auto-d92e15065c55779d-5f146332463a3df7 branch from 2e86dba to 25f8f98 Compare September 11, 2026 03:05
@Klaud-Cold

Klaud-Cold commented Sep 11, 2026

Copy link
Copy Markdown
Collaborator Author

Final full sweep

  • Image / SHA: lmsysorg/sglang:v0.5.19-cu130 (digest sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9) at PR head 25f8f98fa258680daf485f6d6b60084eb961d24b (rebased onto main a5f23f46 so the perf-changelog diff is purely additive; utils/validate_perf_changelog.py --base-ref origin/main passes and all prior changelog bytes are preserved)
  • Changes since the Initial attempt: appended the dsr1-fp4-b200-sglang perf-changelog entry (whole family, no scenario/eval-selection/append-only modifiers); the image line is unchanged
  • Run: run-sweep 34557177019 (full-sweep-enabled, PR kept draft; started 2026-09-11T03:06Z, completed 04:32Z) — completed, success: all 9 benchmark jobs, all 3 eval jobs, check-changelog, collect-results, collect-evals, compare-results, calc-success-rate and upload-changelog-metadata succeeded; no failures or retries. klaud-sweep-manifest records head 25f8f98f, run 34557177019 attempt 1, full-sweep: true. An earlier push-triggered run 34557151441 was cancelled at setup by the workflow concurrency group before any GPU work when this labeled run started. Matrix: 9 points on cluster:b200-nscale — TP4/EP1 concurrency 1, 2, 4, 8, 16, 32 and TP4/EP4 DP-attention concurrency 64, 128, 256, with default gsm8k evals at TP4/EP1 c32 and TP4/EP4 c128/c256
  • Benchmark results (results_bmk/agg_bmk.json, 9 of 9 points, image lmsysorg/sglang:v0.5.19-cu130, vs. the frozen 2026-08-06 baseline from run 30952773184; same topology, concurrency and 8k1k random dataset per point):
Point tput/GPU new (tok/s) base Δ median TPOT new (ms) base Δ median TTFT new (ms) base Δ median E2EL new (s) base Δ
TP4/EP1 c1 407.4 389.7 +4.6% 5.27 5.52 −4.5% 208 207 +0.1% 4.99 5.22 −4.3%
TP4/EP1 c2 704.4 672.2 +4.8% 6.06 6.37 −4.8% 250 245 +2.3% 5.91 6.26 −5.5%
TP4/EP1 c4 1179.5 1134.2 +4.0% 7.21 7.48 −3.6% 247 245 +0.8% 6.79 7.08 −4.0%
TP4/EP1 c8 1828.3 1778.0 +2.8% 9.26 9.49 −2.5% 267 269 −0.6% 8.88 9.09 −2.2%
TP4/EP1 c16 2584.5 2558.7 +1.0% 13.14 13.27 −0.9% 491 494 −0.5% 12.46 12.57 −0.9%
TP4/EP1 c32 3487.5 3428.7 +1.7% 19.35 19.75 −2.0% 764 769 −0.5% 18.52 19.06 −2.8%
TP4/EP4 DPA c64 4918.8 4603.2 +6.9% 27.91 29.73 −6.1% 780 857 −9.0% 26.67 28.46 −6.3%
TP4/EP4 DPA c128 6637.6 6249.9 +6.2% 41.90 44.39 −5.6% 762 831 −8.3% 39.25 41.73 −5.9%
TP4/EP4 DPA c256 8365.6 7955.2 +5.2% 62.93 66.38 −5.2% 3445 4092 −15.8% 62.28 65.66 −5.1%

Throughput improves on every point (+1.0% to +6.9%); TPOT and E2EL improve everywhere. The only regressions are median TTFT at TP4/EP1 c1 (+0.1%), c2 (+2.3%) and c4 (+0.8%), all within a few milliseconds.

  • Eval results (eval_results_all/agg_eval_all.json, 3 of 3 default gsm8k evals, infrastructure_success: true, n=1319 each): TP4/EP1 c32 em_strict 0.9545 (flexible 0.9545); TP4/EP4 DPA c128 em_strict 0.9545 (flexible 0.9553); TP4/EP4 DPA c256 em_strict 0.9538 (flexible 0.9553). Published baseline eval N/A (no gsm8k row for this identity dated 2026-08-06); the smoke evals on this image scored 0.9515 (c1) and 0.9560 (c64).
  • Outcome: final validation passed on the exact PR head. This is final validation of the family sweep, not global PR approval; CODEOWNER review and merge remain human steps.
  • Next step: run finish to verify matrix/result coverage and mark the PR ready for review.

最终完整扫描

  • 镜像 / SHA: lmsysorg/sglang:v0.5.19-cu130(摘要 sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9),PR head 25f8f98fa258680daf485f6d6b60084eb961d24b(已重新基于 main a5f23f46,使 perf-changelog 差异纯为追加;utils/validate_perf_changelog.py --base-ref origin/main 通过,所有既有 changelog 字节保持不变)
  • 相对初始尝试的改动: 追加了 dsr1-fp4-b200-sglang 的 perf-changelog 条目(整个配方族,不带 scenario/评测选择/append-only 修饰符);镜像行未变
  • 运行: run-sweep 34557177019full-sweep-enabled,PR 保持草稿;2026-09-11T03:06Z 开始,04:32Z 完成)—— 已完成,成功:9 个基准作业、3 个评测作业、check-changelog、collect-results、collect-evals、compare-results、calc-success-rate 与 upload-changelog-metadata 全部成功;无失败或重试。klaud-sweep-manifest 记录 head 25f8f98f、run 34557177019 attempt 1、full-sweep: true。此前由推送触发的 run 34557151441 在本次带标签运行开始时被工作流并发组于 setup 阶段取消,未产生任何 GPU 工作。矩阵:cluster:b200-nscale 上 9 个点 —— TP4/EP1 并发 1、2、4、8、16、32 与 TP4/EP4 DP-attention 并发 64、128、256,默认 gsm8k 评测位于 TP4/EP1 c32 及 TP4/EP4 c128/c256
  • 基准结果results_bmk/agg_bmk.json,9/9 个点,镜像 lmsysorg/sglang:v0.5.19-cu130,对比来自 run 30952773184 的冻结 2026-08-06 基线;各点拓扑、并发与 8k1k 随机数据集一致):
点位 新每 GPU 吞吐 (tok/s) 基线 Δ 新 TPOT 中位数 (ms) 基线 Δ 新 TTFT 中位数 (ms) 基线 Δ 新 E2EL 中位数 (s) 基线 Δ
TP4/EP1 c1 407.4 389.7 +4.6% 5.27 5.52 −4.5% 208 207 +0.1% 4.99 5.22 −4.3%
TP4/EP1 c2 704.4 672.2 +4.8% 6.06 6.37 −4.8% 250 245 +2.3% 5.91 6.26 −5.5%
TP4/EP1 c4 1179.5 1134.2 +4.0% 7.21 7.48 −3.6% 247 245 +0.8% 6.79 7.08 −4.0%
TP4/EP1 c8 1828.3 1778.0 +2.8% 9.26 9.49 −2.5% 267 269 −0.6% 8.88 9.09 −2.2%
TP4/EP1 c16 2584.5 2558.7 +1.0% 13.14 13.27 −0.9% 491 494 −0.5% 12.46 12.57 −0.9%
TP4/EP1 c32 3487.5 3428.7 +1.7% 19.35 19.75 −2.0% 764 769 −0.5% 18.52 19.06 −2.8%
TP4/EP4 DPA c64 4918.8 4603.2 +6.9% 27.91 29.73 −6.1% 780 857 −9.0% 26.67 28.46 −6.3%
TP4/EP4 DPA c128 6637.6 6249.9 +6.2% 41.90 44.39 −5.6% 762 831 −8.3% 39.25 41.73 −5.9%
TP4/EP4 DPA c256 8365.6 7955.2 +5.2% 62.93 66.38 −5.2% 3445 4092 −15.8% 62.28 65.66 −5.1%

所有点的吞吐均有提升(+1.0% 至 +6.9%);TPOT 与 E2EL 全部改善。仅有的回退是 TP4/EP1 c1(+0.1%)、c2(+2.3%)、c4(+0.8%)的 TTFT 中位数,均在几毫秒以内。

  • 评测结果eval_results_all/agg_eval_all.json,3/3 个默认 gsm8k 评测,infrastructure_success: true,各 n=1319):TP4/EP1 c32 em_strict 0.9545(flexible 0.9545);TP4/EP4 DPA c128 em_strict 0.9545(flexible 0.9553);TP4/EP4 DPA c256 em_strict 0.9538(flexible 0.9553)。已发布基线评测 N/A(该配置没有 2026-08-06 的 gsm8k 记录);本镜像的冒烟评测得分为 0.9515(c1)与 0.9560(c64)。
  • 结果: 最终验证在精确的 PR head 上通过。这是配方族扫描的最终验证,不是 PR 的全局批准;CODEOWNER 审阅与合并仍由人工完成。
  • 下一步: 运行 finish 验证矩阵/结果覆盖并将 PR 标记为可审阅。

@github-actions

Copy link
Copy Markdown
Contributor

@Klaud-Cold
Klaud-Cold marked this pull request as ready for review September 11, 2026 04:34
@Klaud-Cold
Klaud-Cold requested a review from a team September 11, 2026 04:34
@Klaud-Cold

Copy link
Copy Markdown
Collaborator Author

Klaud Cold: validated. All owned runs are terminal. Repairs: 0. Runs: 34548446011, 34557151441, 34557177019.

The full sweep is verified; this PR remains ready for review.


Klaud Cold:validated。所有自有运行均已结束。修复次数:0。运行:34548446011, 34557151441, 34557177019

完整 sweep 已通过验证;PR 保持就绪,等待审查。

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good, straightforward image-version bump. Reviewed the config diff (single image field change in the dsr1-fp4-b200-sglang recipe, no other fields touched) and the changelog diff (new entry appended strictly at the tail with no edits to prior bytes, matching the append-only convention).

Extended reasoning...

Overview

The PR touches exactly two files: configs/nvidia-master.yaml, where the dsr1-fp4-b200-sglang recipe's image field is bumped from lmsysorg/sglang:v0.5.16-cu130 to lmsysorg/sglang:v0.5.19-cu130 with no other field changes, and perf-changelog.yaml, which gains one new entry appended at the physical tail of the file describing the bump (old/new build commits, CUDA/FlashInfer/sgl-kernel versions, and a digest for the new image) plus an assertion that the benchmark script's model/backend/dtype/concurrency-grid configuration is unchanged.

Security risks

None identified. This is a data-only YAML change (an image tag string and a changelog text entry) with no code execution, no credentials, no auth/permissions logic, and no injection surface.

Level of scrutiny

Low. This matches the repository's documented "Klaud Cold" automated image-bump pattern: an exact-version image tag bump paired with a changelog entry citing build commits and a digest. I verified via git diff --stat that the config change is a single line and the changelog change is purely additive at the tail (no historical bytes edited), which satisfies the append-only, byte-sensitive requirement documented in AGENTS.md/CONTRIBUTING.md for this changelog file. The recipe is single-node (no multinode/model.container field), so the multi-node model.container == image coupling rule noted in the task context does not apply here.

Other factors

No CODEOWNERS-restricted paths appear to be implicated beyond standard config/changelog ownership, and the PR conversation timeline shows no outstanding CHANGES_REQUESTED or unaddressed third-party objections. No bugs were reported by the bug-hunting system, and my own review of the diff found nothing beyond the mechanical, well-formed change described.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

1 participant