Skip to content

[Klaud Cold] Update qwen3.5-fp4-b200-sglang SGLang image to v0.5.19-cu130 / 将 qwen3.5-fp4-b200-sglang 的 SGLang 镜像更新至 v0.5.19-cu130 - #2954

Merged
adibarra merged 3 commits into
mainfrom
klaud/auto-1fa5c7d01fd1a2fd-3d9c146c5e9ad2d6
Sep 10, 2026
Merged

[Klaud Cold] Update qwen3.5-fp4-b200-sglang SGLang image to v0.5.19-cu130 / 将 qwen3.5-fp4-b200-sglang 的 SGLang 镜像更新至 v0.5.19-cu130#2954
adibarra merged 3 commits into
mainfrom
klaud/auto-1fa5c7d01fd1a2fd-3d9c146c5e9ad2d6

Conversation

@Klaud-Cold

@Klaud-Cold Klaud-Cold commented Sep 10, 2026

Copy link
Copy Markdown
Collaborator

Bump the qwen3.5-fp4-b200-sglang master image from lmsysorg/sglang:v0.5.14-cu130 to lmsysorg/sglang:v0.5.19-cu130 (digest sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9). Model (nvidia/Qwen3.5-397B-A17B-NVFP4-V2), TP4/EP1 and TP2/EP1 topologies, the 8k/1k fixed-seq-len workload and the launch script are unchanged.

Baseline

  • Published date: 2026-08-10, image lmsysorg/sglang:v0.5.14-cu130 (digest sha256:5027e95bf6ec536856b1b52a91d1f35ff5c564ab83e8a94758a169ff09bb8df3)
  • Workload: single_turn fixed-seq-len ISL 8192 / OSL 1024, B200 single node (cluster:b200-nscale), SGLang, FP4, no speculative decoding, TP4/EP1 at concurrency 4 and TP2/EP1 at concurrency 4-128
  • Producer run: https://github.com/SemiAnalysisAI/InferenceX/actions/runs/30506346629/attempts/1 (head 99e3a3fdaee4f2822ff43442509915f9a3a1bae0; logical curve snapshot curve_workflow_run_id 2287 is not the producer)
  • API queries: GET /api/v1/benchmarks?model=Qwen-3.5-397B-A17B&date=2026-08-10&exact=true, GET /api/v1/workflow-info?date=2026-08-10, GET /api/v1/evaluations?model=Qwen-3.5-397B-A17B
  • Old engine source: sgl-project/sglang tag v0.5.14 = sgl-project/sglang@49e384c
  • New engine source: sgl-project/sglang tag v0.5.19 = sgl-project/sglang@0bcd822
TP/EP Conc tput/GPU (tok/s) output tput/GPU (tok/s) mean TPOT (ms) mean TTFT (ms)
TP4/EP1 4 1385.5 154.1 5.88 439.6
TP2/EP1 4 2491.3 277.1 6.66 376.4
TP2/EP1 8 3949.4 444.0 8.24 584.2
TP2/EP1 16 5621.2 622.9 11.82 647.0
TP2/EP1 32 7784.4 870.5 16.93 973.4
TP2/EP1 64 10511.9 1166.2 25.34 1536.9
TP2/EP1 128 13655.1 1512.5 39.25 2281.5
  • Published eval: N/A for the 2026-08-10 run (the public evaluations feed has no gsm8k row for this identity on that date; its only rows for B200 SGLang FP4 non-MTP are from 2026-04-08 on an older image and are not comparable).

qwen3.5-fp4-b200-sglang 主配置镜像从 lmsysorg/sglang:v0.5.14-cu130 更新至 lmsysorg/sglang:v0.5.19-cu130(摘要 sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9)。模型(nvidia/Qwen3.5-397B-A17B-NVFP4-V2)、TP4/EP1 与 TP2/EP1 拓扑、8k/1k 固定序列长度工作负载以及启动脚本均保持不变。

基线

  • 发布日期:2026-08-10,镜像 lmsysorg/sglang:v0.5.14-cu130(摘要 sha256:5027e95bf6ec536856b1b52a91d1f35ff5c564ab83e8a94758a169ff09bb8df3
  • 工作负载:single_turn 固定序列长度 ISL 8192 / OSL 1024,B200 单节点(cluster:b200-nscale),SGLang,FP4,无投机解码,TP4/EP1 并发 4 与 TP2/EP1 并发 4-128
  • 生产运行:https://github.com/SemiAnalysisAI/InferenceX/actions/runs/30506346629/attempts/1(head 99e3a3fdaee4f2822ff43442509915f9a3a1bae0;逻辑曲线快照 curve_workflow_run_id 2287 不是生产运行)
  • API 查询:GET /api/v1/benchmarks?model=Qwen-3.5-397B-A17B&date=2026-08-10&exact=trueGET /api/v1/workflow-info?date=2026-08-10GET /api/v1/evaluations?model=Qwen-3.5-397B-A17B
  • 旧引擎源码:sgl-project/sglang 标签 v0.5.14 = sgl-project/sglang@49e384c
  • 新引擎源码:sgl-project/sglang 标签 v0.5.19 = sgl-project/sglang@0bcd822

基线数值见上方英文表格(tput/GPU、output tput/GPU、平均 TPOT、平均 TTFT)。

  • 已发布评测:2026-08-10 运行为 N/A(公开 evaluations 数据中该身份在该日期没有 gsm8k 记录;B200 SGLang FP4 非 MTP 仅有 2026-04-08 旧镜像的记录,不可比较)。

🤖 Generated with Claude Code


Note

Low Risk
Config-only container image pin and changelog; no application logic or runtime behavior changes in-repo.

Overview
Bumps the qwen3.5-fp4-b200-sglang benchmark config from lmsysorg/sglang:v0.5.14-cu130 to lmsysorg/sglang:v0.5.19-cu130 in nvidia-master.yaml.

Adds a matching perf-changelog.yaml entry (PR #2954) noting the digest update and that model, TP/EP topologies, the 8k/1k fixed-seq-len workload, and launch script are unchanged.

Reviewed by Cursor Bugbot for commit 5dd10a2. Bugbot is set up for automated code reviews on this repo. Configure here.

@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase As a PR reviewer and CODEOWNER, I have reviewed this and have.

For PR verification, add the full-sweep-fail-fast label (strongly recommended) to this PR — the benchmark sweep only runs on labeled PRs. Use full-sweep-enabled only if you need matrix jobs to keep running past a failure.

PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs


感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 As a PR reviewer and CODEOWNER, I have reviewed this and have

如需进行 PR 验证,请为此 PR 添加 full-sweep-fail-fast 标签(强烈推荐)— 基准测试 sweep 仅在带有标签的 PR 上运行。仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled

PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档

@Klaud-Cold

Klaud-Cold commented Sep 10, 2026

Copy link
Copy Markdown
Collaborator Author

Initial attempt — passed

  • Image: lmsysorg/sglang:v0.5.14-cu130lmsysorg/sglang:v0.5.19-cu130 (digest sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9, Docker Hub tag pushed 2026-09-04), head 09540c6f7e31c879d82abb4c1a9cb3290234d0de
  • Targeted run (trimmed smoke, lowest concurrency per deployment shape plus its default gsm8k eval): https://github.com/SemiAnalysisAI/InferenceX/actions/runs/34446067754 — completed, success (2 benchmark jobs, 1 eval job, collectors green)
  • Points: TP4/EP1 conc 4 (benchmark) and TP2/EP1 conc 4 (benchmark + gsm8k eval), 8k/1k, no speculative decoding. This is startup and compatibility evidence, not a full curve.
  • Change: master image only. The launch script, model, topology, workloads and eval selection are untouched. Capacity check on b200-nscale passed before the edit and again before dispatch.

Smoke benchmark vs published 2026-08-10 baseline (same topology, concurrency and 8k/1k dataset; new values from results_bmk/agg_bmk.json on the run, both rows image: lmsysorg/sglang:v0.5.19-cu130, power_valid: 1):

Point tput/GPU (tok/s) output tput/GPU (tok/s) mean TPOT (ms) mean TTFT (ms)
TP4/EP1 conc 4 1385.5 → 1712.3 (+23.6%) 154.1 → 190.5 (+23.6%) 5.88 → 4.93 (−16.2%) 439.6 → 187.7 (−57.3%)
TP2/EP1 conc 4 2491.3 → 2834.4 (+13.8%) 277.1 → 315.3 (+13.8%) 6.66 → 5.94 (−10.8%) 376.4 → 247.3 (−34.3%)
  • Server logs: both servers reached "fired up" with the FP8 KV cache allocated (TP4: 6.42M tokens, TP2: 2.48M tokens), CUDA graphs captured, no tracebacks and no non-2xx responses other than pre-ready /health 503s. The only notices are the pre-existing --cuda-graph-max-bs deprecation alias (already present in v0.5.14) and the SM100 default of FlashInfer GDN decode for bf16 SSM state, which is the same in both tags.

Smoke eval (eval_results_all/agg_eval_all.json and lm-eval results_*.json; meta_env.json confirms sglang, fp4, spec_decoding: none, TP2/EP1, conc 4, model Qwen3.5-397B-A17B-NVFP4-V2):

Point gsm8k em_strict se flexible-extract n_eff
TP2/EP1 conc 4 0.9651 0.0051 0.9553 1319
  • Published baseline eval: N/A (no published gsm8k row for this identity on 2026-08-10).

Upstream source comparison (the v0.5.19-cu130 tag is built by the SGLang release Dockerfile from the v0.5.19 release branch; CUDA base 13.0.1 → 13.0.3):

  • sgl-project/sglang v0.5.14 sgl-project/sglang@49e384cv0.5.19 sgl-project/sglang@0bcd822
  • Coupled dependencies (python/pyproject.toml): sglang-kernel 0.4.4 → 0.4.6.post1, flashinfer 0.6.12 → 0.6.18, torch 2.11.0 → 2.13.0, transformers 5.8.1 → 5.12.1, nvidia-cutlass-dsl 4.5.2 → 4.6.2, sgl-deep-gemm 0.1.3 → 0.1.7, flash-attn-4 4.0.0b15 → ≥4.0.0b18.
  • Server-argument handling moved from the monolithic server_args.py post-init into srt/arg_groups/* hooks. Every ServerArgs field the script sets exists in both tags: enable_symm_mem, disable_radix_cache, quantization, kv_cache_dtype, mamba_ssm_dtype, attention_backend, moe_runner_backend, max_prefill_tokens, chunked_prefill_size, mem_fraction_static, stream_interval, scheduler_recv_interval, tokenizer_worker_num, tokenizer_path, context_length, tp_size, dp_size, ep_size.
  • --cuda-graph-max-bs is already a deprecated alias of --cuda-graph-max-bs-decode in v0.5.14 and remains one in v0.5.19 (same dest), so the recipe's per-point graph sizing is unchanged.
  • --disable-radix-cache makes the hybrid-mamba radix-cache resolution return early in both tags (overrides.py in v0.5.19, _handle_mamba_radix_cache in v0.5.14), so the no_buffer check that rejects trtllm_mha is not reached.
  • trtllm_mha with --kv-cache-dtype fp8_e4m3 stays valid on SM100 (the v0.5.19 rejection applies only to SM120). flashinfer_trtllm still accepts modelopt_fp4. The SM100 + bf16 SSM default of FlashInfer GDN decode is the same in both tags.
  • Changed defaults that do not affect this recipe because it pins them: mem_fraction_static auto-sizing, chunked_prefill_size auto-sizing.
  • Potentially useful new options left out of scope (tuning, not compatibility): --linear-attn-prefill-backend flashinfer, --mamba-full-memory-ratio (used by the FP8 B200 sibling in WIP - SGL B200 FP8 8k1k  #2866).
  • The same image already runs on this cluster: qwen3.5-fp8-b200-sglang published 2026-09-08 on lmsysorg/sglang:v0.5.19-cu130 with the same trtllm_mha + flashinfer_trtllm + fp8_e4m3 KV path, and the B300 FP4 MTP sibling validated it in [Klaud Cold] Update qwen3.5-fp4-b300-sglang-mtp SGLang image to v0.5.19-cu130 / 将 qwen3.5-fp4-b300-sglang-mtp 的 SGLang 镜像更新至 v0.5.19-cu130 #2950.
  • Result: no engine patch or script change is needed; the image is used as shipped.

Next step: append the perf-changelog.yaml entry, recheck capacity, and start the final full sweep (draft PR, full-sweep-enabled).


首次尝试 — 通过

  • 镜像:lmsysorg/sglang:v0.5.14-cu130lmsysorg/sglang:v0.5.19-cu130(摘要 sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9,Docker Hub 标签推送于 2026-09-04),head 09540c6f7e31c879d82abb4c1a9cb3290234d0de
  • 定向运行(裁剪后的冒烟测试,每种部署形态的最低并发及其默认 gsm8k 评测):https://github.com/SemiAnalysisAI/InferenceX/actions/runs/34446067754 — 已完成,成功(2 个基准任务、1 个评测任务,收集器均通过)
  • 测试点:TP4/EP1 并发 4(基准)与 TP2/EP1 并发 4(基准 + gsm8k 评测),8k/1k,无投机解码。这只是启动与兼容性证据,不是完整曲线。
  • 变更:仅主配置镜像。启动脚本、模型、拓扑、工作负载与评测选择均未改动。编辑前与派发前对 b200-nscale 的容量检查均通过。

冒烟基准 vs 2026-08-10 已发布基线(相同拓扑、并发与 8k/1k 数据集;数值见上方英文表格):TP4/EP1 并发 4 的每 GPU 吞吐 +23.6%,平均 TPOT −16.2%,平均 TTFT −57.3%;TP2/EP1 并发 4 的每 GPU 吞吐 +13.8%,平均 TPOT −10.8%,平均 TTFT −34.3%。两行均为 image: lmsysorg/sglang:v0.5.19-cu130power_valid: 1

  • 服务日志:两个服务均正常启动并分配 FP8 KV cache(TP4:642 万 token,TP2:248 万 token),CUDA graph 捕获完成,无 traceback,除就绪前的 /health 503 外无非 2xx 响应。仅有的提示是 v0.5.14 中已存在的 --cuda-graph-max-bs 弃用别名,以及 SM100 上 bf16 SSM 状态默认使用 FlashInfer GDN decode(两个标签一致)。

冒烟评测eval_results_all/agg_eval_all.json 与 lm-eval results_*.jsonmeta_env.json 确认 sglangfp4spec_decoding: none、TP2/EP1、并发 4、模型 Qwen3.5-397B-A17B-NVFP4-V2):TP2/EP1 并发 4 的 gsm8k em_strict 0.9651(se 0.0051,flexible-extract 0.9553,n_eff 1319)。

  • 已发布基线评测:N/A(2026-08-10 该身份没有已发布的 gsm8k 记录)。

上游源码对比v0.5.19-cu130 标签由 SGLang 发布 Dockerfile 基于 v0.5.19 发布分支构建;CUDA 基础镜像 13.0.1 → 13.0.3):

  • sgl-project/sglang v0.5.14(49e384ce)→ v0.5.19(0bcd8223),链接见上文英文部分。
  • 耦合依赖(python/pyproject.toml):sglang-kernel 0.4.4 → 0.4.6.post1,flashinfer 0.6.12 → 0.6.18,torch 2.11.0 → 2.13.0,transformers 5.8.1 → 5.12.1,nvidia-cutlass-dsl 4.5.2 → 4.6.2,sgl-deep-gemm 0.1.3 → 0.1.7,flash-attn-4 4.0.0b15 → ≥4.0.0b18。
  • 服务参数处理从单一 server_args.py 的 post-init 拆分为 srt/arg_groups/* 钩子。脚本设置的所有 ServerArgs 字段在两个标签中均存在(列表见英文部分)。
  • --cuda-graph-max-bs 在 v0.5.14 中已是 --cuda-graph-max-bs-decode 的弃用别名,在 v0.5.19 中保持不变(相同 dest),因此各点的 CUDA graph 尺寸不受影响。
  • 由于传入 --disable-radix-cache,两个标签中的混合 mamba radix cache 解析都会提前返回,因此拒绝 trtllm_mhano_buffer 校验不会触发。
  • trtllm_mha 搭配 --kv-cache-dtype fp8_e4m3 在 SM100 上仍然有效(v0.5.19 的拒绝仅针对 SM120)。flashinfer_trtllm 仍接受 modelopt_fp4。SM100 + bf16 SSM 下默认使用 FlashInfer GDN decode 的逻辑在两个标签中一致。
  • 不影响本配方的默认值变化(脚本已显式指定):mem_fraction_static 自动推导、chunked_prefill_size 自动推导。
  • 暂不纳入范围的新选项(属于调优而非兼容性):--linear-attn-prefill-backend flashinfer--mamba-full-memory-ratio(FP8 B200 同级配方在 WIP - SGL B200 FP8 8k1k  #2866 中使用)。
  • 同一镜像已在本集群运行:qwen3.5-fp8-b200-sglang 于 2026-09-08 基于 lmsysorg/sglang:v0.5.19-cu130 发布,使用相同的 trtllm_mha + flashinfer_trtllm + fp8_e4m3 KV 路径;B300 FP4 MTP 同级配方已在 [Klaud Cold] Update qwen3.5-fp4-b300-sglang-mtp SGLang image to v0.5.19-cu130 / 将 qwen3.5-fp4-b300-sglang-mtp 的 SGLang 镜像更新至 v0.5.19-cu130 #2950 中验证该镜像。
  • 结论:无需引擎补丁或脚本修改,镜像按原样使用。

下一步:追加 perf-changelog.yaml 条目,重新检查容量,并启动最终完整扫描(保持草稿 PR,添加 full-sweep-enabled)。

Klaud-Cold and others added 2 commits September 10, 2026 07:51
Update the qwen3.5-fp4-b200-sglang master image from lmsysorg/sglang:v0.5.14-cu130
to lmsysorg/sglang:v0.5.19-cu130 (digest sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9).
Model, TP4/EP1 and TP2/EP1 topologies, 8k/1k workload and the launch script are unchanged.

将 qwen3.5-fp4-b200-sglang 的主配置镜像从 lmsysorg/sglang:v0.5.14-cu130 更新至
lmsysorg/sglang:v0.5.19-cu130(摘要 sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9)。
模型、TP4/EP1 与 TP2/EP1 拓扑、8k/1k 工作负载以及启动脚本均保持不变。

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Append the perf-changelog.yaml entry for the qwen3.5-fp4-b200-sglang SGLang image
update from lmsysorg/sglang:v0.5.14-cu130 to lmsysorg/sglang:v0.5.19-cu130.

为 qwen3.5-fp4-b200-sglang 的 SGLang 镜像更新(lmsysorg/sglang:v0.5.14-cu130 →
lmsysorg/sglang:v0.5.19-cu130)追加 perf-changelog.yaml 条目。

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@Klaud-Cold
Klaud-Cold force-pushed the klaud/auto-1fa5c7d01fd1a2fd-3d9c146c5e9ad2d6 branch from 0606e60 to baed662 Compare September 10, 2026 07:51
@Klaud-Cold

Copy link
Copy Markdown
Collaborator Author

Final full sweep (attempt 1) — failed before benchmarks

  • Head: 0606e60cb0cc56b6f904250de970637e7e9d3742 (image bump + perf-changelog.yaml entry for qwen3.5-fp4-b200-sglang, no scenario/eval-selection/append-only modifiers), PR kept draft, label full-sweep-enabled
  • Runs: https://github.com/SemiAnalysisAI/InferenceX/actions/runs/34451925055 (label event) and https://github.com/SemiAnalysisAI/InferenceX/actions/runs/34451877926 (synchronize event on the same head) — both completed, failure in the check-changelog job; every benchmark, eval and collector job was skipped, so no GPU time was used.
  • Diagnosis: check-changelog runs validate_perf_changelog.py --base-ref origin/main --head-ref <PR head>. main advanced past the candidate base 01cea4db with commit c26c9510 (perf(agentx): sweep Kimi K3 with LMCache dev114 / 使用 LMCache dev114 扫描 Kimi K3 #2889), which appended its own perf-changelog.yaml entry. Diffing my branch against the newer origin/main therefore showed that entry as a deletion (Deletions are not allowed in perf-changelog.yaml ... Found deleted line: - config-keys:). The same validator passed locally against the original base. This is branch staleness, not an image or recipe problem.
  • Benchmark/eval results: N/A (no benchmark jobs ran).

Next step: Repair 1/5 — rebase the branch onto current origin/main so the changelog diff is purely additive, re-run the validator against origin/main, recheck capacity and relabel.


最终完整扫描(第 1 次)— 在基准测试前失败

下一步:修复 1/5 — 将分支变基到当前 origin/main,使变更日志差异为纯追加,针对 origin/main 重新运行校验器,重新检查容量并重新添加标签。

@Klaud-Cold

Klaud-Cold commented Sep 10, 2026

Copy link
Copy Markdown
Collaborator Author

Repair 1/5 — rebase onto current main (done)

  • Cause being repaired: the check-changelog gate in the first final sweep (see the "Final full sweep (attempt 1)" comment above) rejected the diff against origin/main because main had gained a newer perf-changelog.yaml entry after the candidate base.
  • Change: no recipe or image change. The two existing commits (image bump, changelog entry) were rebased onto origin/main c26c95106136d6f9fdea3c0a480c17a92aceb11c; the label had been removed before the push. New head baed6629770e3235ed43b23decaa5300e5f67930.
  • Verification before relabeling: git diff origin/main..HEAD -- perf-changelog.yaml has zero deleted lines and every byte of origin/main's changelog is preserved as a prefix; validate_perf_changelog.py --base-ref origin/main --head-ref baed6629 passes; the generated qwen3.5-fp4-b200-sglang matrix (7 points, image lmsysorg/sglang:v0.5.19-cu130) is identical to the smoke-tested head's matrix; configs/nvidia-master.yaml still differs from main only in the one image line.
  • The smoke evidence from the Initial attempt (head 09540c6f) remains valid because the recipe content is unchanged.

Outcome: capacity recheck passed, full-sweep-enabled applied (PR stays draft), and the check-changelog gate passed on head baed6629 (synchronize run https://github.com/SemiAnalysisAI/InferenceX/actions/runs/34452105636). The final full sweep is https://github.com/SemiAnalysisAI/InferenceX/actions/runs/34452191302; results are tracked in the "Final full sweep (attempt 2)" comment.


修复 1/5 — 变基到当前 main(已完成)

  • 修复原因:第一次最终扫描的 check-changelog 门禁(见上方"最终完整扫描(第 1 次)"评论)拒绝了与 origin/main 的差异,因为 main 在候选基线之后新增了 perf-changelog.yaml 条目。
  • 变更:无配方或镜像变更。将现有两个提交(镜像更新、变更日志条目)变基到 origin/main c26c95106136d6f9fdea3c0a480c17a92aceb11c;推送前已移除标签。新 head baed6629770e3235ed43b23decaa5300e5f67930
  • 重新添加标签前的验证:git diff origin/main..HEAD -- perf-changelog.yaml 无删除行,origin/main 的变更日志每个字节都作为前缀保留;validate_perf_changelog.py --base-ref origin/main --head-ref baed6629 通过;生成的 qwen3.5-fp4-b200-sglang 矩阵(7 个点,镜像 lmsysorg/sglang:v0.5.19-cu130)与冒烟测试 head 的矩阵完全一致;configs/nvidia-master.yamlmain 仍仅相差一行镜像。
  • 首次尝试(head 09540c6f)的冒烟证据仍然有效,因为配方内容未变。

结果:容量复查通过,已添加 full-sweep-enabled(PR 保持草稿),head baed6629 上的 check-changelog 门禁已通过(synchronize 运行 https://github.com/SemiAnalysisAI/InferenceX/actions/runs/34452105636)。最终完整扫描为 https://github.com/SemiAnalysisAI/InferenceX/actions/runs/34452191302,结果在"最终完整扫描(第 2 次)"评论中跟踪。

@Klaud-Cold

Klaud-Cold commented Sep 10, 2026

Copy link
Copy Markdown
Collaborator Author

Final full sweep (attempt 2) — passed

  • Head: baed6629770e3235ed43b23decaa5300e5f67930 (image bump + perf-changelog.yaml entry for qwen3.5-fp4-b200-sglang, rebased onto main c26c9510 in Repair 1/5; no scenario/eval-selection/append-only modifiers), PR kept draft, label full-sweep-enabled applied after a passing b200-nscale capacity check
  • Run: https://github.com/SemiAnalysisAI/InferenceX/actions/runs/34452191302 (run-sweep.yml, label event) — completed, success. klaud-sweep-manifest records head baed6629770e3235ed43b23decaa5300e5f67930, run 34452191302 attempt 1, full-sweep: true. The synchronize-triggered run on the same head, https://github.com/SemiAnalysisAI/InferenceX/actions/runs/34452105636, passed check-changelog and did nothing further.
  • Scope executed: the complete family matrix — 7 single-node 8k1k points (TP4/EP1 conc 4; TP2/EP1 conc 4, 8, 16, 32, 64, 128) plus the 2 default gsm8k evals (TP2/EP1 conc 64 and 128). All 17 jobs succeeded (setup, 7 benchmarks, 2 evals, collectors, compare, success-rate, changelog metadata, visualizer comment); nothing trimmed. results_bmk/agg_bmk.json has 7 rows, all image: lmsysorg/sglang:v0.5.19-cu130, spec_decoding: none, power_valid: 1.

Full-curve results vs published 2026-08-10 baseline (lmsysorg/sglang:v0.5.14-cu130, producer run 30506346629; matched per point on topology, concurrency and the 8k/1k dataset):

Point tput/GPU (tok/s) output tput/GPU (tok/s) mean TPOT (ms) mean TTFT (ms)
TP4/EP1 conc 4 1385.5 → 1705.0 (+23.1%) 154.1 → 189.7 (+23.1%) 5.88 → 4.96 (−15.7%) 439.6 → 179.4 (−59.2%)
TP2/EP1 conc 4 2491.3 → 2830.4 (+13.6%) 277.1 → 314.9 (+13.6%) 6.66 → 5.94 (−10.8%) 376.4 → 253.6 (−32.6%)
TP2/EP1 conc 8 3949.4 → 4431.6 (+12.2%) 444.0 → 498.2 (+12.2%) 8.24 → 7.47 (−9.3%) 584.2 → 400.3 (−31.5%)
TP2/EP1 conc 16 5621.2 → 6442.3 (+14.6%) 622.9 → 713.9 (+14.6%) 11.82 → 10.36 (−12.3%) 647.0 → 508.1 (−21.5%)
TP2/EP1 conc 32 7784.4 → 8753.2 (+12.4%) 870.5 → 978.8 (+12.4%) 16.93 → 15.14 (−10.6%) 973.4 → 784.5 (−19.4%)
TP2/EP1 conc 64 10511.9 → 11509.4 (+9.5%) 1166.2 → 1276.9 (+9.5%) 25.34 → 23.28 (−8.1%) 1536.9 → 1269.8 (−17.4%)
TP2/EP1 conc 128 13655.1 → 14443.1 (+5.8%) 1512.5 → 1599.8 (+5.8%) 39.25 → 37.18 (−5.3%) 2281.5 → 2088.4 (−8.5%)
  • No regressions: every point improved throughput per GPU and lowered mean TPOT and TTFT. All 7 points matched a baseline point; none excluded.

Default evals (eval_results_all/agg_eval_all.json and lm-eval results_*.json; meta_env.json confirms sglang, fp4, spec_decoding: none, TP2/EP1, model Qwen3.5-397B-A17B-NVFP4-V2, new image):

Point gsm8k em_strict se flexible-extract n_eff
TP2/EP1 conc 64 0.9682 0.0048 0.9598 1319
TP2/EP1 conc 128 0.9651 0.0051 0.9538 1319
  • Published baseline eval: N/A (no published gsm8k row for this identity on 2026-08-10). The scores are in line with the smoke eval (0.9651 at conc 4) and with the older published 2026-04-08 values (0.967-0.970) for this family.

Outcome: targeted smoke passed (Initial attempt), exact-head final validation passed (this run), 1 of 5 repairs used (branch rebase only; no image or recipe repair). finish verifies coverage and marks the PR ready; CODEOWNER review and merge remain the maintainers' decision.


最终完整扫描(第 2 次)— 通过

  • Head:baed6629770e3235ed43b23decaa5300e5f67930(镜像更新 + qwen3.5-fp4-b200-sglangperf-changelog.yaml 条目,已在修复 1/5 中变基到 main c26c9510;无 scenario / eval-selection / append-only 修饰符),PR 保持草稿,在 b200-nscale 容量检查通过后添加 full-sweep-enabled
  • 运行:https://github.com/SemiAnalysisAI/InferenceX/actions/runs/34452191302(`run-sweep.yml`,标签事件)— 已完成,成功。klaud-sweep-manifest 记录 head baed6629770e3235ed43b23decaa5300e5f67930、运行 34452191302 第 1 次尝试、full-sweep: true。同一 head 上由 synchronize 触发的运行 https://github.com/SemiAnalysisAI/InferenceX/actions/runs/34452105636 通过了 check-changelog,未执行其他任务。
  • 执行范围:完整家族矩阵 — 7 个单节点 8k1k 点(TP4/EP1 并发 4;TP2/EP1 并发 4、8、16、32、64、128)以及 2 个默认 gsm8k 评测(TP2/EP1 并发 64 与 128)。全部 17 个任务成功(setup、7 个基准、2 个评测、收集器、对比、成功率、变更日志元数据、可视化评论);未裁剪任何内容。results_bmk/agg_bmk.json 共 7 行,均为 image: lmsysorg/sglang:v0.5.19-cu130spec_decoding: nonepower_valid: 1

完整曲线 vs 2026-08-10 已发布基线lmsysorg/sglang:v0.5.14-cu130,生产运行 30506346629;按拓扑、并发与 8k/1k 数据集逐点匹配;数值见上方英文表格):TP4/EP1 并发 4 的每 GPU 吞吐 +23.1%,平均 TPOT −15.7%,平均 TTFT −59.2%;TP2/EP1 各点每 GPU 吞吐 +5.8% 至 +14.6%,平均 TPOT −5.3% 至 −12.3%,平均 TTFT −8.5% 至 −32.6%。

  • 无回退:所有点的每 GPU 吞吐均提升,平均 TPOT 与 TTFT 均下降。7 个点全部与基线点匹配,无排除项。

默认评测eval_results_all/agg_eval_all.json 与 lm-eval results_*.jsonmeta_env.json 确认 sglangfp4spec_decoding: none、TP2/EP1、模型 Qwen3.5-397B-A17B-NVFP4-V2、新镜像):TP2/EP1 并发 64 的 gsm8k em_strict 0.9682(se 0.0048,flexible-extract 0.9598);并发 128 为 0.9651(se 0.0051,flexible-extract 0.9538);n_eff 均为 1319。

  • 已发布基线评测:N/A(2026-08-10 该身份没有已发布的 gsm8k 记录)。分数与冒烟评测(并发 4 为 0.9651)以及该家族 2026-04-08 的旧发布值(0.967-0.970)一致。

结论:定向冒烟通过(首次尝试),精确 head 的最终验证通过(本次运行),5 次修复预算使用 1 次(仅分支变基,无镜像或配方修复)。finish 校验覆盖率并将 PR 标记为可评审;CODEOWNER 评审与合并仍由维护者决定。

@github-actions

Copy link
Copy Markdown
Contributor

@Klaud-Cold
Klaud-Cold marked this pull request as ready for review September 10, 2026 08:41
@Klaud-Cold
Klaud-Cold requested a review from a team September 10, 2026 08:41
@Klaud-Cold

Copy link
Copy Markdown
Collaborator Author

Klaud Cold: validated. All owned runs are terminal. Repairs: 1. Runs: 34446067754, 34451877926, 34451925055, 34452056954, 34452105636, 34452191302.

The full sweep is verified; this PR remains ready for review.


Klaud Cold:validated。所有自有运行均已结束。修复次数:1。运行:34446067754, 34451877926, 34451925055, 34452056954, 34452105636, 34452191302

完整 sweep 已通过验证;PR 保持就绪,等待审查。

@claude

claude Bot commented Sep 10, 2026

Copy link
Copy Markdown
Contributor

Claude finished @Klaud-Cold's task in 1m 5s —— View job


Review

  • Gather context (PR body, sweep comments, changed files)
  • Check configs/nvidia-master.yaml image change (public registry, only the intended line)
  • Check perf-changelog.yaml entry (present, appended at end, correct config key, no append-only)
  • Post review summary

LGTM - no blocking issues found.

The diff against origin/main is exactly the two intended changes. The image line for qwen3.5-fp4-b200-sglang moves to lmsysorg/sglang:v0.5.19-cu130, a public Docker Hub tag, with no other config edits. The perf-changelog.yaml entry is appended at the very end of the file, names the correct config key, links this PR, and carries no append-only or eval modifiers. The check-changelog gate and the full 7-point sweep plus both gsm8k evals passed on this exact head (run 34452191302), with no regressions against the 2026-08-10 baseline. I did not run the changelog validator locally because the command was not approved in this session, but the passing CI gate covers it.


审阅

LGTM - 未发现阻塞问题。

origin/main 的差异仅包含预期的两处改动:qwen3.5-fp4-b200-sglang 的镜像行更新为公开的 Docker Hub 标签 lmsysorg/sglang:v0.5.19-cu130,无其他配置修改。perf-changelog.yaml 条目追加在文件末尾,config key 正确,链接指向本 PR,且没有 append-only 或评测修饰符。check-changelog 门禁与完整 7 点扫描及两项 gsm8k 评测已在该 head 上通过(运行 34452191302),相对 2026-08-10 基线无回退。

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Looks good, straightforward version bump.

What was reviewed: the diff touches only the image field of the qwen3.5-fp4-b200-sglang recipe (model, topologies, scenarios untouched), and appends a single new tail entry to perf-changelog.yaml with no prior bytes altered. I also checked whether leaving the paired qwen3.5-fp4-b200-sglang-mtp recipe on the older v0.5.14-cu130 breaks an established pairing convention (as one candidate issue suggested) — history shows the non-mtp and mtp variants have been bumped independently before (e.g. PR #1018 bumped only non-mtp, with the mtp variant picked up in a separate later PR), so this is not a strict rule that's being violated here.

Extended reasoning...

Overview

This PR touches exactly two files: configs/nvidia-master.yaml, where it bumps the SGLang container image for the single-node qwen3.5-fp4-b200-sglang recipe from v0.5.14-cu130 to v0.5.19-cu130 (verified via git diff — only the image: line changed, all other fields in the recipe block are byte-identical), and perf-changelog.yaml, where it appends one new entry at the tail documenting the bump, the unchanged model/topologies/workload, and linking PR #2954. Both changes are minimal, mechanical, and scoped exactly as described in the PR.

Security risks

None. This is a benchmark configuration repo; the change is a version-tag bump for a container image referenced by an internal benchmark runner, with no code execution, credential, or auth-path changes involved.

Level of scrutiny

Low-to-moderate scrutiny is appropriate: it's a single-field version bump plus a changelog append, matching the repository's own stated conventions (append-only changelog, kebab-case fields). I independently re-derived one candidate concern from the ruled-out list — that the paired qwen3.5-fp4-b200-sglang-mtp recipe is left on the older image, seemingly breaking a "bump both together" pattern seen in PRs #1474 and #2056. Walking the full perf-changelog.yaml history for this config family, I found a counter-example (PR #1018) where only the non-mtp variant was bumped, with the mtp variant updated independently in a later, separate PR. AGENTS.md also contains no written rule requiring synchronized bumps for this family (only for multi-node srt-slurm model.container/image parity, which doesn't apply here since this recipe is single-node with no model.container field). So this is not a hard convention violation, just an optional/best-effort pattern that has been deviated from before without issue.

Other factors

The perf-changelog.yaml diff is a pure append (git diff shows no deleted/modified lines before the new block), consistent with the byte-sensitive, append-only requirement called out in AGENTS.md/CONTRIBUTING.md. No other paired recipes/config files reference this image in a way that would require synchronized changes beyond what was checked. Given the narrow, well-precedented nature of the change, I'm confident a human reviewer does not need to inspect this further.

@adibarra

adibarra commented Sep 10, 2026

Copy link
Copy Markdown
Collaborator

/reuse-sweep-run 34452191302

将 main 合并到 PR #2954,保留已验证的配方并复用完整扫描结果。
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

Development

Successfully merging this pull request may close these issues.

2 participants