[Klaud Cold] Update qwen3.5-fp4-b200-sglang SGLang image to v0.5.19-cu130 / 将 qwen3.5-fp4-b200-sglang 的 SGLang 镜像更新至 v0.5.19-cu130 - #2954
Conversation
|
Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase For PR verification, add the PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs 感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 如需进行 PR 验证,请为此 PR 添加 PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档 |
Initial attempt — passed
Smoke benchmark vs published 2026-08-10 baseline (same topology, concurrency and 8k/1k dataset; new values from
Smoke eval (
Upstream source comparison (the
Next step: append the 首次尝试 — 通过
冒烟基准 vs 2026-08-10 已发布基线(相同拓扑、并发与 8k/1k 数据集;数值见上方英文表格):TP4/EP1 并发 4 的每 GPU 吞吐 +23.6%,平均 TPOT −16.2%,平均 TTFT −57.3%;TP2/EP1 并发 4 的每 GPU 吞吐 +13.8%,平均 TPOT −10.8%,平均 TTFT −34.3%。两行均为
冒烟评测(
上游源码对比(
下一步:追加 |
Update the qwen3.5-fp4-b200-sglang master image from lmsysorg/sglang:v0.5.14-cu130 to lmsysorg/sglang:v0.5.19-cu130 (digest sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9). Model, TP4/EP1 and TP2/EP1 topologies, 8k/1k workload and the launch script are unchanged. 将 qwen3.5-fp4-b200-sglang 的主配置镜像从 lmsysorg/sglang:v0.5.14-cu130 更新至 lmsysorg/sglang:v0.5.19-cu130(摘要 sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9)。 模型、TP4/EP1 与 TP2/EP1 拓扑、8k/1k 工作负载以及启动脚本均保持不变。 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Append the perf-changelog.yaml entry for the qwen3.5-fp4-b200-sglang SGLang image update from lmsysorg/sglang:v0.5.14-cu130 to lmsysorg/sglang:v0.5.19-cu130. 为 qwen3.5-fp4-b200-sglang 的 SGLang 镜像更新(lmsysorg/sglang:v0.5.14-cu130 → lmsysorg/sglang:v0.5.19-cu130)追加 perf-changelog.yaml 条目。 Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
0606e60 to
baed662
Compare
Final full sweep (attempt 1) — failed before benchmarks
Next step: Repair 1/5 — rebase the branch onto current 最终完整扫描(第 1 次)— 在基准测试前失败
下一步:修复 1/5 — 将分支变基到当前 |
Repair 1/5 — rebase onto current main (done)
Outcome: capacity recheck passed, 修复 1/5 — 变基到当前 main(已完成)
结果:容量复查通过,已添加 |
Final full sweep (attempt 2) — passed
Full-curve results vs published 2026-08-10 baseline (
Default evals (
Outcome: targeted smoke passed (Initial attempt), exact-head final validation passed (this run), 1 of 5 repairs used (branch rebase only; no image or recipe repair). 最终完整扫描(第 2 次)— 通过
完整曲线 vs 2026-08-10 已发布基线(
默认评测(
结论:定向冒烟通过(首次尝试),精确 head 的最终验证通过(本次运行),5 次修复预算使用 1 次(仅分支变基,无镜像或配方修复)。 |
|
see unofficial run visualizer at https://inferencex.semianalysis.com/inference?unofficialRun=34452191302 |
|
Klaud Cold: validated. All owned runs are terminal. Repairs: 1. Runs: 34446067754, 34451877926, 34451925055, 34452056954, 34452105636, 34452191302. The full sweep is verified; this PR remains ready for review. Klaud Cold:validated。所有自有运行均已结束。修复次数:1。运行:34446067754, 34451877926, 34451925055, 34452056954, 34452105636, 34452191302。 完整 sweep 已通过验证;PR 保持就绪,等待审查。 |
|
Claude finished @Klaud-Cold's task in 1m 5s —— View job Review
LGTM - no blocking issues found. The diff against 审阅LGTM - 未发现阻塞问题。 与 |
There was a problem hiding this comment.
Looks good, straightforward version bump.
What was reviewed: the diff touches only the image field of the qwen3.5-fp4-b200-sglang recipe (model, topologies, scenarios untouched), and appends a single new tail entry to perf-changelog.yaml with no prior bytes altered. I also checked whether leaving the paired qwen3.5-fp4-b200-sglang-mtp recipe on the older v0.5.14-cu130 breaks an established pairing convention (as one candidate issue suggested) — history shows the non-mtp and mtp variants have been bumped independently before (e.g. PR #1018 bumped only non-mtp, with the mtp variant picked up in a separate later PR), so this is not a strict rule that's being violated here.
Extended reasoning...
Overview
This PR touches exactly two files: configs/nvidia-master.yaml, where it bumps the SGLang container image for the single-node qwen3.5-fp4-b200-sglang recipe from v0.5.14-cu130 to v0.5.19-cu130 (verified via git diff — only the image: line changed, all other fields in the recipe block are byte-identical), and perf-changelog.yaml, where it appends one new entry at the tail documenting the bump, the unchanged model/topologies/workload, and linking PR #2954. Both changes are minimal, mechanical, and scoped exactly as described in the PR.
Security risks
None. This is a benchmark configuration repo; the change is a version-tag bump for a container image referenced by an internal benchmark runner, with no code execution, credential, or auth-path changes involved.
Level of scrutiny
Low-to-moderate scrutiny is appropriate: it's a single-field version bump plus a changelog append, matching the repository's own stated conventions (append-only changelog, kebab-case fields). I independently re-derived one candidate concern from the ruled-out list — that the paired qwen3.5-fp4-b200-sglang-mtp recipe is left on the older image, seemingly breaking a "bump both together" pattern seen in PRs #1474 and #2056. Walking the full perf-changelog.yaml history for this config family, I found a counter-example (PR #1018) where only the non-mtp variant was bumped, with the mtp variant updated independently in a later, separate PR. AGENTS.md also contains no written rule requiring synchronized bumps for this family (only for multi-node srt-slurm model.container/image parity, which doesn't apply here since this recipe is single-node with no model.container field). So this is not a hard convention violation, just an optional/best-effort pattern that has been deviated from before without issue.
Other factors
The perf-changelog.yaml diff is a pure append (git diff shows no deleted/modified lines before the new block), consistent with the byte-sensitive, append-only requirement called out in AGENTS.md/CONTRIBUTING.md. No other paired recipes/config files reference this image in a way that would require synchronized changes beyond what was checked. Given the narrow, well-precedented nature of the change, I'm confident a human reviewer does not need to inspect this further.
|
/reuse-sweep-run 34452191302 |
将 main 合并到 PR #2954,保留已验证的配方并复用完整扫描结果。
Bump the
qwen3.5-fp4-b200-sglangmaster image fromlmsysorg/sglang:v0.5.14-cu130tolmsysorg/sglang:v0.5.19-cu130(digestsha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9). Model (nvidia/Qwen3.5-397B-A17B-NVFP4-V2), TP4/EP1 and TP2/EP1 topologies, the 8k/1k fixed-seq-len workload and the launch script are unchanged.Baseline
lmsysorg/sglang:v0.5.14-cu130(digestsha256:5027e95bf6ec536856b1b52a91d1f35ff5c564ab83e8a94758a169ff09bb8df3)cluster:b200-nscale), SGLang, FP4, no speculative decoding, TP4/EP1 at concurrency 4 and TP2/EP1 at concurrency 4-12899e3a3fdaee4f2822ff43442509915f9a3a1bae0; logical curve snapshotcurve_workflow_run_id2287 is not the producer)GET /api/v1/benchmarks?model=Qwen-3.5-397B-A17B&date=2026-08-10&exact=true,GET /api/v1/workflow-info?date=2026-08-10,GET /api/v1/evaluations?model=Qwen-3.5-397B-A17Bv0.5.14= sgl-project/sglang@49e384cv0.5.19= sgl-project/sglang@0bcd822evaluationsfeed has no gsm8k row for this identity on that date; its only rows for B200 SGLang FP4 non-MTP are from 2026-04-08 on an older image and are not comparable).将
qwen3.5-fp4-b200-sglang主配置镜像从lmsysorg/sglang:v0.5.14-cu130更新至lmsysorg/sglang:v0.5.19-cu130(摘要sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9)。模型(nvidia/Qwen3.5-397B-A17B-NVFP4-V2)、TP4/EP1 与 TP2/EP1 拓扑、8k/1k 固定序列长度工作负载以及启动脚本均保持不变。基线
lmsysorg/sglang:v0.5.14-cu130(摘要sha256:5027e95bf6ec536856b1b52a91d1f35ff5c564ab83e8a94758a169ff09bb8df3)cluster:b200-nscale),SGLang,FP4,无投机解码,TP4/EP1 并发 4 与 TP2/EP1 并发 4-12899e3a3fdaee4f2822ff43442509915f9a3a1bae0;逻辑曲线快照curve_workflow_run_id2287 不是生产运行)GET /api/v1/benchmarks?model=Qwen-3.5-397B-A17B&date=2026-08-10&exact=true、GET /api/v1/workflow-info?date=2026-08-10、GET /api/v1/evaluations?model=Qwen-3.5-397B-A17Bv0.5.14= sgl-project/sglang@49e384cv0.5.19= sgl-project/sglang@0bcd822基线数值见上方英文表格(tput/GPU、output tput/GPU、平均 TPOT、平均 TTFT)。
evaluations数据中该身份在该日期没有 gsm8k 记录;B200 SGLang FP4 非 MTP 仅有 2026-04-08 旧镜像的记录,不可比较)。🤖 Generated with Claude Code
Note
Low Risk
Config-only container image pin and changelog; no application logic or runtime behavior changes in-repo.
Overview
Bumps the
qwen3.5-fp4-b200-sglangbenchmark config fromlmsysorg/sglang:v0.5.14-cu130tolmsysorg/sglang:v0.5.19-cu130innvidia-master.yaml.Adds a matching
perf-changelog.yamlentry (PR #2954) noting the digest update and that model, TP/EP topologies, the 8k/1k fixed-seq-len workload, and launch script are unchanged.Reviewed by Cursor Bugbot for commit 5dd10a2. Bugbot is set up for automated code reviews on this repo. Configure here.