Skip to content

[Klaud Cold] Update minimaxm3-fp8-h200-vllm-agentic-mtp vLLM image to v0.29.0 (digest-pinned) / 将 minimaxm3-fp8-h200-vllm-agentic-mtp 的 vLLM 镜像更新至 v0.29.0(按摘要固定) - #2953

Closed
Klaud-Cold wants to merge 2 commits into
mainfrom
klaud/auto-b3fff31b6f9a6956-b3fe5685ff662693
Closed

[Klaud Cold] Update minimaxm3-fp8-h200-vllm-agentic-mtp vLLM image to v0.29.0 (digest-pinned) / 将 minimaxm3-fp8-h200-vllm-agentic-mtp 的 vLLM 镜像更新至 v0.29.0(按摘要固定)#2953
Klaud-Cold wants to merge 2 commits into
mainfrom
klaud/auto-b3fff31b6f9a6956-b3fe5685ff662693

Conversation

@Klaud-Cold

Copy link
Copy Markdown
Collaborator

Refresh minimaxm3-fp8-h200-vllm-agentic-mtp from the 2026-09-07 upstream nightly vllm/vllm-openai:nightly-d9105ea8001e0a6d77a96327d17515bb5791fb36 to the vLLM release vllm/vllm-openai:v0.29.0, pinned by manifest digest sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1. Only the master image string changes; benchmarks/single_node/agentic/minimaxm3_fp8_h200_mtp.sh, the TP8 EAGLE3 topology, the Mooncake DRAM-offload arm and the c1-c14 grid are unchanged.

Baseline (published 2026-09-08)

  • Old image: vllm/vllm-openai:nightly-d9105ea8001e0a6d77a96327d17515bb5791fb36 (vLLM main commit d9105ea8, 2026-09-07; Docker Hub digest sha256:254eebf919e8b7b0d530d97fccc606c36f380ff6724d64bc190951bec1aee838). Identity verified on all 8 rows: model MiniMax-M3 (MiniMaxAI/MiniMax-M3-MXFP8), hardware h200, framework vllm, precision fp8, spec_method mtp (EAGLE3 draft Inferact/MiniMax-M3-EAGLE3-GQA, 3 speculative tokens, golden synthetic AL 2.78), disagg false, benchmark_type agentic_traces, TP8 on cluster:h200-dgxc, 1-hour AgentX coding-trace replay.
  • New image source: vLLM tag v0.29.0 = commit 98dff2a8 (release published 2026-09-09; Docker Hub tag pushed 2026-09-09T06:06Z, CUDA 13.0 build). The release branch forked from main at f5c3cc24; the old nightly carries 300 main commits the release lacks and the release carries 14 release-branch commits the nightly lacks (compare).
  • API queries: GET /api/v1/workflow-info?date=2026-09-08&benchmarkType=agentic_traces (changelog entry for this key, PR [Klaud Cold] Update minimaxm3-fp8-h200-vllm-agentic-mtp vLLM image to nightly-d9105ea8001e0a6d77a96327d17515bb5791fb36 #2875, head aac901a1) and GET /api/v1/benchmarks?model=MiniMax-M3&date=2026-09-08&exact=true&sequence=agentic-traces (no view; 14 rows, 8 match this identity). GET /api/v1/evaluations has no row for this single-node identity; eval values below come from the producer run's eval_results_all artifact.
  • Producer of all 8 points and 8 evals: https://github.com/SemiAnalysisAI/InferenceX/actions/runs/34174440759/attempts/1 (PR [Klaud Cold] Update minimaxm3-fp8-h200-vllm-agentic-mtp vLLM image to nightly-d9105ea8001e0a6d77a96327d17515bb5791fb36 #2875 sweep, head SHA ebf645526a9178e1131a99593071e6fa01d04107, reused at merge commit aac901a16a0968702e978d2ab2eb0de725958559). Curve snapshot IDs are not used as producers. Recipe fingerprints: 6d8ae56106e80615a97eb6a60fc098c79431a94e (resident) and b49c7028af3a565848740cbb80701b79d6981242 (Mooncake DRAM offload).
  • Published throughput (tokens/s per GPU; latency in seconds; interactivity in tokens/s per user) and the producer run's eval per point:
conc KV offload total tput/GPU output tput/GPU mean TTFT mean TPOT median interactivity gsm8k em_strict (n=1319)
1 none 1078.8 8.27 2.089 0.011 102.0 0.9719
2 none 1290.3 10.57 0.901 0.010 115.7 0.9719
4 none 1542.8 12.82 0.992 0.018 54.9 0.9735
6 none 2218.9 17.66 0.955 0.023 39.1 0.9735
8 none 3417.9 25.40 0.938 0.024 39.0 0.9682
10 none 3996.4 31.80 1.276 0.028 33.4 0.9666
12 dram (Mooncake 0.3.11.post1) 4263.4 33.91 1.164 0.029 32.5 0.9697
14 dram (Mooncake 0.3.11.post1) 4476.4 37.05 1.299 0.030 31.4 0.9697
  • Eval mismatch: the baseline sweep ran lm-eval gsm8k; the current default eval for this family is the MiniMax vendor minimax_m3_smoke tool-call suite, so eval scores from this PR are N/A against the baseline (different task) and are reported on their own.

minimaxm3-fp8-h200-vllm-agentic-mtp 的镜像从 2026-09-07 上游 nightly vllm/vllm-openai:nightly-d9105ea8001e0a6d77a96327d17515bb5791fb36 刷新为 vLLM 正式版 vllm/vllm-openai:v0.29.0,并按清单摘要 sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1 固定。仅修改主配置中的镜像字符串;benchmarks/single_node/agentic/minimaxm3_fp8_h200_mtp.sh、TP8 EAGLE3 拓扑、Mooncake DRAM 卸载分支以及 c1-c14 网格均保持不变。

基线(发布于 2026-09-08)

  • 旧镜像:vllm/vllm-openai:nightly-d9105ea8001e0a6d77a96327d17515bb5791fb36(vLLM main 提交 d9105ea8,2026-09-07;Docker Hub 摘要 sha256:254eebf919e8b7b0d530d97fccc606c36f380ff6724d64bc190951bec1aee838)。全部 8 行均已核对身份:模型 MiniMax-M3(MiniMaxAI/MiniMax-M3-MXFP8)、硬件 h200、框架 vllm、精度 fp8、spec_method mtp(EAGLE3 草稿模型 Inferact/MiniMax-M3-EAGLE3-GQA,3 个推测 token,黄金合成接受长度 2.78)、disagg false、benchmark_type agentic_traces、cluster:h200-dgxc 上 TP8、1 小时 AgentX 编码轨迹回放。
  • 新镜像来源:vLLM 标签 v0.29.0 = 提交 98dff2a8(正式版发布于 2026-09-09;Docker Hub 标签推送于 2026-09-09T06:06Z,CUDA 13.0 构建)。发布分支从 main 的 f5c3cc24 分叉;旧 nightly 包含 300 个正式版没有的 main 提交,正式版包含 14 个 nightly 没有的发布分支提交(对比)。
  • API 查询:GET /api/v1/workflow-info?date=2026-09-08&benchmarkType=agentic_traces(该键的变更日志条目,PR [Klaud Cold] Update minimaxm3-fp8-h200-vllm-agentic-mtp vLLM image to nightly-d9105ea8001e0a6d77a96327d17515bb5791fb36 #2875,head aac901a1)与 GET /api/v1/benchmarks?model=MiniMax-M3&date=2026-09-08&exact=true&sequence=agentic-traces(不带 view;14 行,其中 8 行匹配该身份)。GET /api/v1/evaluations 没有该单节点身份的记录;下表评测值来自生产运行的 eval_results_all 产物。
  • 全部 8 个点及 8 个评测的生产运行:https://github.com/SemiAnalysisAI/InferenceX/actions/runs/34174440759/attempts/1(PR [Klaud Cold] Update minimaxm3-fp8-h200-vllm-agentic-mtp vLLM image to nightly-d9105ea8001e0a6d77a96327d17515bb5791fb36 #2875 的 sweep,head SHA ebf645526a9178e1131a99593071e6fa01d04107,在合并提交 aac901a16a0968702e978d2ab2eb0de725958559 时复用)。未将曲线快照 ID 用作生产者。配方指纹:6d8ae56106e80615a97eb6a60fc098c79431a94e(常驻)与 b49c7028af3a565848740cbb80701b79d6981242(Mooncake DRAM 卸载)。
  • 已发布吞吐(每 GPU tokens/s;延迟单位为秒;交互性为每用户 tokens/s)及生产运行的各点评测:
并发 KV 卸载 每 GPU 总吞吐 每 GPU 输出吞吐 平均 TTFT 平均 TPOT 中位交互性 gsm8k em_strict(n=1319)
1 1078.8 8.27 2.089 0.011 102.0 0.9719
2 1290.3 10.57 0.901 0.010 115.7 0.9719
4 1542.8 12.82 0.992 0.018 54.9 0.9735
6 2218.9 17.66 0.955 0.023 39.1 0.9735
8 3417.9 25.40 0.938 0.024 39.0 0.9682
10 3996.4 31.80 1.276 0.028 33.4 0.9666
12 dram(Mooncake 0.3.11.post1) 4263.4 33.91 1.164 0.029 32.5 0.9697
14 dram(Mooncake 0.3.11.post1) 4476.4 37.05 1.299 0.030 31.4 0.9697
  • 评测不匹配:基线 sweep 运行的是 lm-eval gsm8k;该家族当前默认评测为 MiniMax 官方 minimax_m3_smoke 工具调用套件,因此本 PR 的评测分数相对基线为 N/A(任务不同),单独报告。

🤖 Generated with Claude Code

Update minimaxm3-fp8-h200-vllm-agentic-mtp from
vllm/vllm-openai:nightly-d9105ea8001e0a6d77a96327d17515bb5791fb36 to
vllm/vllm-openai:v0.29.0 pinned by manifest digest
sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1.
The recipe script, model, precision, TP8 topology, EAGLE3 settings,
workloads and the c1-c14 grid are unchanged.

将 minimaxm3-fp8-h200-vllm-agentic-mtp 的 vLLM 镜像从
vllm/vllm-openai:nightly-d9105ea8001e0a6d77a96327d17515bb5791fb36
更新为按清单摘要固定的 vllm/vllm-openai:v0.29.0
(sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1)。
配方脚本、模型、精度、TP8 拓扑、EAGLE3 设置、工作负载与 c1-c14 网格均未改动。

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution! Please reach out to respective companies' CODEOWNER to fill in the latest PR_REVIEW_CHECKLIST.md before pinging core maintainer on Slack for review. In order for the signoff PR check bot to trigger, you must follow the PR_REVIEW_CHECKLIST.md template correctly, including the phrase As a PR reviewer and CODEOWNER, I have reviewed this and have.

For PR verification, add the full-sweep-fail-fast label (strongly recommended) to this PR — the benchmark sweep only runs on labeled PRs. Use full-sweep-enabled only if you need matrix jobs to keep running past a failure.

PR authors are responsible for ensuring that after merging, all GitHub Action jobs fully pass. A lot of the time, failures are just flakes and simply re-running the failed jobs will fix it. See GitHub's docs on re-running failed jobs


感谢你的贡献!请联系相应公司的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,然后再在 Slack 上联系核心维护者进行审阅。为了触发 signoff PR 检查机器人,你必须正确遵循 PR_REVIEW_CHECKLIST.md 模板,包括保留英文语句 As a PR reviewer and CODEOWNER, I have reviewed this and have

如需进行 PR 验证,请为此 PR 添加 full-sweep-fail-fast 标签(强烈推荐)— 基准测试 sweep 仅在带有标签的 PR 上运行。仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled

PR 作者有责任确保合并后所有 GitHub Action 任务完全通过。 很多时候失败只是偶发抖动(flake),重新运行失败的任务即可解决。参见 GitHub 关于重新运行失败任务的文档

@Klaud-Cold

Klaud-Cold commented Sep 10, 2026

Copy link
Copy Markdown
Collaborator Author

Initial attempt

Status: completed, success (run finished 02:27Z, all 8 non-skipped jobs green: c1 and c12 throughput, c1 and c12 default evals, collect-evals, collect-results, calc-success-rate). Targeted smoke passed; this is startup/compatibility evidence, not full validation.

  • Image / commit: vllm/vllm-openai:v0.29.0@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1 (vLLM tag v0.29.0 = commit 98dff2a8; Docker Hub tag pushed 2026-09-09T06:06Z, manifest-list digest above, linux/amd64 image sha256:082ca6f035279109041ffd3fe0695cb568b29bc580b35c4f297a66a08b216c1b; docker/Dockerfile default CUDA_VERSION=13.0.3, arch list includes 9.0a, identical to the old nightly's Dockerfile). PR head 937d423e844b55f9c2ad6889ce99c5f758f3d166.
  • Change: only the master image string. The generated matrix is identical to the base SHA except the image field (8 points, TP8, nodes:1, run-eval on every point). The trimmed smoke keeps c1 (resident) and c12 (Mooncake DRAM offload) with the default minimax-vendor / minimax_m3_smoke eval. utils/matrix_logic tests: 297 passed.
  • Source comparison (release 98dff2a8 vs old nightly d9105ea8, compare: diverged, release branch forked at f5c3cc24, 300 main commits missing from the release, 14 release-branch commits missing from the nightly):
  • Run: https://github.com/SemiAnalysisAI/InferenceX/actions/runs/34422803079 (dispatched 2026-09-10T00:47:50Z from main, inputs.ref=937d423e844b55f9c2ad6889ce99c5f758f3d166, test-config --config-files configs/nvidia-master.yaml --config-keys minimaxm3-fp8-h200-vllm-agentic-mtp --trim-conc, fail-fast=true, klaud-run=true). Capacity gate for telemetry cluster h200 exited 0 at 00:43Z (before edits) and 00:47Z (before dispatch).
  • Eval results (minimax_m3_smoke, MiniMax provider verifier, 1 sample, exact_match strict): c1 = 1.0, c12 (Mooncake DRAM offload) = 1.0; tool-call schema validation errors 0, tool_calls_match_rate 1.0 on both (eval_results_all/agg_eval_all.json). N/A against the baseline's lm-eval gsm8k scores (different task). Server log: Initializing a V1 LLM engine (v0.29.0), Using V2 Model Runner, target attention TRITON_ATTN, MxFp8 MoE backend MARLIN, available KV cache 69.67 GiB = 2,476,416 tokens (identical to the baseline's c1 KV pool).
  • Throughput results (results_bmk/agg_bmk.json, one-hour AgentX replay per point, matched to the frozen baseline by concurrency, TP8 topology, KV-offload arm and dataset; single runs, no repetition):
conc KV offload total tput/GPU new vs base output tput/GPU new vs base mean TTFT (s) new vs base mean TPOT (s) new vs base median interactivity (tok/s) new vs base profiled requests new vs base
1 none 946.7 vs 1078.8 (-12.3%) 7.30 vs 8.27 (-11.7%) 2.491 vs 2.089 (+19.2%) 0.0136 vs 0.0109 (+24.4%) 82.2 vs 102.0 (-19.5%) 159 vs 170
12 dram (Mooncake) 3694.0 vs 4263.4 (-13.4%) 28.48 vs 33.91 (-16.0%) 1.127 vs 1.164 (-3.2%) 0.0363 vs 0.0291 (+24.6%) 25.6 vs 32.5 (-21.2%) 882 vs 979
  • Regression: both smoke points are slower on v0.29.0 than the published nightly, driven by a consistent +24% to +27% mean/p50 TPOT (decode step time) at unchanged GPU prefix-cache hit rate (0.949 vs 0.954 at c1; 0.821 vs 0.819 at c12), unchanged KV pool and 2.7% to 3.5% lower average GPU power. Median TTFT is within ±5%. The in-scope explanation consistent with a decode-only slowdown is the draft attention backend: v0.29.0 predates [Bugfix][Spec Decode] Honour the draft's attention_backend on Model Runner V2 vllm-project/vllm#54826, so on Model Runner V2 the EAGLE3 draft does not receive the recipe's FLASH_ATTN override (the nightly honours it). No recipe change was made; this is reported, not fixed.
  • Mooncake at c12: 2 requests failed with Failing 1 request(s) due to KV load failure after Failed to get 5210/5211 Mooncake keys from sub-batch (error -707) at 01:22Z; 3 of 1014 records were dropped as errors (2 InvalidInferenceResultError, 1 client ClientOSError), below the AIPerf 10% abort threshold. The baseline nightly's own server logs show the same pattern (c12: 2 KV load failures, c14: 8), so this is pre-existing behaviour of the recipe's load_async Mooncake path, not a v0.29.0 regression. The MooncakeStoreScheduler block-table assertion that crashed the GB300 sibling did not fire.
  • Next step: append the changelog entry, keep the PR draft, recheck capacity and apply full-sweep-enabled for the complete 8-point sweep with default evals.

初次尝试

状态: 已完成,成功(运行于 02:27Z 结束,8 个非跳过任务全部通过:c1 与 c12 吞吐、c1 与 c12 默认评测、collect-evals、collect-results、calc-success-rate)。定向冒烟通过;这是启动/兼容性证据,不是完整验证。

  • 镜像 / 提交:vllm/vllm-openai:v0.29.0@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1(vLLM 标签 v0.29.0 = 提交 98dff2a8;Docker Hub 标签推送于 2026-09-09T06:06Z,清单摘要如上,linux/amd64 镜像 sha256:082ca6f035279109041ffd3fe0695cb568b29bc580b35c4f297a66a08b216c1bdocker/Dockerfile 默认 CUDA_VERSION=13.0.3,架构列表包含 9.0a,与旧 nightly 的 Dockerfile 完全相同)。PR head 937d423e844b55f9c2ad6889ce99c5f758f3d166
  • 变更:仅主配置中的 image 字符串。生成的矩阵除镜像字段外与基准 SHA 完全一致(8 个点,TP8,nodes:1,每个点均 run-eval)。裁剪后的冒烟保留 c1(常驻)与 c12(Mooncake DRAM 卸载),使用默认 minimax-vendor / minimax_m3_smoke 评测。utils/matrix_logic 测试:297 通过。
  • 源码对比(正式版 98dff2a8 vs 旧 nightly d9105ea8对比:已分叉,发布分支自 f5c3cc24 分出,正式版缺少 300 个 main 提交,nightly 缺少 14 个发布分支提交):
  • 运行:https://github.com/SemiAnalysisAI/InferenceX/actions/runs/34422803079(2026-09-10T00:47:50Zmain 分派,inputs.ref=937d423e844b55f9c2ad6889ce99c5f758f3d166test-config --config-files configs/nvidia-master.yaml --config-keys minimaxm3-fp8-h200-vllm-agentic-mtp --trim-concfail-fast=trueklaud-run=true)。遥测集群 h200 的容量门禁在 00:43Z(编辑前)与 00:47Z(分派前)均返回 0。
  • 评测结果(minimax_m3_smoke,MiniMax provider verifier,1 个样本,exact_match strict):c1 = 1.0,c12(Mooncake DRAM 卸载)= 1.0;两者工具调用 schema 校验错误 0,tool_calls_match_rate 1.0(eval_results_all/agg_eval_all.json)。相对基线的 lm-eval gsm8k 分数为 N/A(任务不同)。服务端日志:Initializing a V1 LLM engine (v0.29.0)Using V2 Model Runner、目标模型注意力 TRITON_ATTN、MxFp8 MoE 后端 MARLIN、可用 KV 缓存 69.67 GiB = 2,476,416 token(与基线 c1 的 KV 池完全一致)。
  • 吞吐结果(results_bmk/agg_bmk.json,每点一小时 AgentX 回放,按并发、TP8 拓扑、KV 卸载分支与数据集匹配冻结基线;单次运行,无重复):
并发 KV 卸载 每 GPU 总吞吐 新 vs 基线 每 GPU 输出吞吐 新 vs 基线 平均 TTFT(秒)新 vs 基线 平均 TPOT(秒)新 vs 基线 中位交互性(tok/s)新 vs 基线 剖析请求数 新 vs 基线
1 946.7 vs 1078.8(-12.3%) 7.30 vs 8.27(-11.7%) 2.491 vs 2.089(+19.2%) 0.0136 vs 0.0109(+24.4%) 82.2 vs 102.0(-19.5%) 159 vs 170
12 dram(Mooncake) 3694.0 vs 4263.4(-13.4%) 28.48 vs 33.91(-16.0%) 1.127 vs 1.164(-3.2%) 0.0363 vs 0.0291(+24.6%) 25.6 vs 32.5(-21.2%) 882 vs 979
  • 回退:两个冒烟点在 v0.29.0 上都比已发布的 nightly 慢,主要来自一致的 +24% 至 +27% 平均/p50 TPOT(解码步时间),而 GPU 前缀缓存命中率不变(c1 为 0.949 vs 0.954;c12 为 0.821 vs 0.819)、KV 池不变、平均 GPU 功耗低 2.7% 至 3.5%。中位 TTFT 在 ±5% 以内。与仅解码变慢相符的范围内解释是草稿模型注意力后端:v0.29.0 早于 [Bugfix][Spec Decode] Honour the draft's attention_backend on Model Runner V2 vllm-project/vllm#54826,因此在 Model Runner V2 上 EAGLE3 草稿模型不会收到配方的 FLASH_ATTN 覆盖(nightly 会遵循)。未改动配方;此处仅报告,未修复。
  • c12 的 Mooncake:01:22Z 有 2 个请求在 Failed to get 5210/5211 Mooncake keys from sub-batch(错误 -707)后以 Failing 1 request(s) due to KV load failure 失败;1014 条记录中 3 条作为错误丢弃(2 个 InvalidInferenceResultError,1 个客户端 ClientOSError),低于 AIPerf 10% 的中止阈值。基线 nightly 自身的服务端日志呈现同样模式(c12:2 次 KV 加载失败,c14:8 次),因此这是配方 load_async Mooncake 路径的既有行为,而非 v0.29.0 的回退。导致 GB300 同族崩溃的 MooncakeStoreScheduler 块表断言未触发。
  • 下一步:追加变更日志条目,保持 PR 草稿状态,重新检查容量并添加 full-sweep-enabled 以运行完整的 8 点 sweep 及默认评测。

Append the perf-changelog entry for minimaxm3-fp8-h200-vllm-agentic-mtp
(vllm/vllm-openai:v0.29.0, digest-pinned) linking PR #2953.

为 minimaxm3-fp8-h200-vllm-agentic-mtp 追加 perf-changelog 条目
(vllm/vllm-openai:v0.29.0,按摘要固定),关联 PR #2953。

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@Klaud-Cold

Klaud-Cold commented Sep 10, 2026

Copy link
Copy Markdown
Collaborator Author

Final full sweep

Status: sweep green, final validation not certified; session stopped with outcome failed (04:30Z). run-sweep.yml run 34429897054 on the exact PR head 7568b8f8 finished green (check-changelog, setup, all 8 throughput points, all 8 evals, collect-evals, upload-changelog-metadata; klaud-sweep-manifest records full-sweep: true, 8 agentic points, 8 agentic evals). finish then rejected result coverage: the manifest's eval rows declare eval-framework: minimax-vendor / eval-suite: minimax_m3_smoke for every point, but every eval job ran lm-eval gsm8k (EVAL_FRAMEWORK: lm-eval, empty EVAL_SUITE, artifacts eval_*_lm-eval__1), so the verifier found 8 expected minimax_m3_smoke identities missing and 8 unexpected gsm8k identities. All 8 benchmark identities (recipe fingerprint, concurrency, image) matched.

  • Image / head: vllm/vllm-openai:v0.29.0@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1, PR head 7568b8f8529b843ba13dcfa45dfe1c43d49ae23f.
  • Changes since the initial attempt: one appended perf-changelog.yaml entry for minimaxm3-fp8-h200-vllm-agentic-mtp (no scenario, eval-selection or append-only modifiers; all prior bytes preserved, utils/validate_perf_changelog.py against origin/main passed). No recipe or config change.
  • Run: https://github.com/SemiAnalysisAI/InferenceX/actions/runs/34429897054 (run-sweep.yml, pull_request labeled event on head 7568b8f8, label full-sweep-enabled applied 2026-09-10T02:33:10Z with the PR kept in draft; capacity gate for h200 exited 0 at 02:33:09Z). Expected matrix: the complete 8-point family (TP8 c1/c2/c4/c6/c8/c10 resident, c12/c14 Mooncake DRAM offload) with the default minimax_m3_smoke eval on every point. The earlier synchronize run https://github.com/SemiAnalysisAI/InferenceX/actions/runs/34429854496 on the same head ran before the label and executed no benchmark jobs.
  • Results (per-point bmk_agentic_* aggregates, one-hour AgentX replay per point, power_valid = 1 on all 8; matched to the frozen baseline by concurrency, TP8 topology and KV-offload arm; single runs, no repetition). The full sweep's default eval is lm-eval gsm8k (em_strict, n = 1319), so unlike the smoke it is directly comparable to the baseline:
conc KV offload total tput/GPU new vs base output tput/GPU new vs base mean TTFT (s) new vs base mean TPOT (s) new vs base median interactivity (tok/s) new vs base profiled requests new vs base gsm8k em_strict
1 none 947.0 vs 1078.8 (-12.2%) 7.30 vs 8.27 (-11.7%) 2.482 vs 2.089 (+18.8%) 0.0137 vs 0.0109 (+25.4%) 81.9 vs 102.0 (-19.7%) 159/170 vs 170 0.9696739954510993
2 none 1104.8 vs 1290.3 (-14.4%) 8.74 vs 10.57 (-17.3%) 1.094 vs 0.901 (+21.4%) 0.0132 vs 0.0103 (+27.8%) 86.7 vs 115.7 (-25.1%) 219/241 vs 309 0.9711902956785443
4 none 1458.9 vs 1542.8 (-5.4%) 10.94 vs 12.82 (-14.7%) 1.086 vs 0.992 (+9.5%) 0.0239 vs 0.0176 (+36.2%) 34.4 vs 54.9 (-37.4%) 299/343 vs 310 0.9711902956785443
6 none 1980.6 vs 2218.9 (-10.7%) 15.44 vs 17.66 (-12.6%) 1.027 vs 0.955 (+7.5%) 0.0269 vs 0.0234 (+15.1%) 34.4 vs 39.1 (-12.0%) 408/473 vs 485 0.9711902956785443
8 none 3107.8 vs 3417.9 (-9.1%) 22.11 vs 25.40 (-13.0%) 0.961 vs 0.938 (+2.5%) 0.0327 vs 0.0240 (+36.0%) 27.3 vs 39.0 (-30.1%) 810/897 vs 856 0.9727065959059894
10 none 3384.1 vs 3996.4 (-15.3%) 27.00 vs 31.80 (-15.1%) 1.284 vs 1.276 (+0.7%) 0.0369 vs 0.0284 (+30.0%) 25.1 vs 33.4 (-24.9%) 861/970 vs 962 0.9727065959059894
12 dram (Mooncake) 3697.4 vs 4263.4 (-13.3%) 28.49 vs 33.91 (-16.0%) 1.172 vs 1.164 (+0.7%) 0.0366 vs 0.0291 (+25.6%) 25.7 vs 32.5 (-20.9%) 883/1015 vs 979 0.974981046247157
14 dram (Mooncake) 3852.1 vs 4476.4 (-13.9%) 30.99 vs 37.05 (-16.4%) 1.431 vs 1.299 (+10.1%) 0.0378 vs 0.0301 (+25.5%) 24.8 vs 31.4 (-21.1%) 899/1054 vs 1018 0.9742228961334344
  • Evals: gsm8k em_strict 0.9697 to 0.9750 across the 8 points versus 0.9666 to 0.9735 in the baseline sweep (within ±1 standard error, infrastructure_success true everywhere).
  • Regression: every point is slower on v0.29.0 than the published 2026-09-07 nightly: total throughput per GPU -5.4% (c4) to -15.3% (c10), driven by +15% to +36% mean TPOT at unchanged GPU prefix-cache hit rates (0.93 to 0.95 resident; 0.75 to 0.83 with Mooncake), unchanged KV pool and 2% to 9% lower average GPU power. Mean TTFT moves +0.7% to +21% (tail-driven at low concurrency; median TTFT within a few percent). The in-scope explanation consistent with a decode-only slowdown remains the EAGLE3 draft attention backend: v0.29.0 predates [Bugfix][Spec Decode] Honour the draft's attention_backend on Model Runner V2 vllm-project/vllm#54826, so on Model Runner V2 the draft ignores the recipe's FLASH_ATTN override while the nightly honours it. No recipe change was made; this is reported for reviewers, not fixed, and the 300 main commits absent from the release (compare in the initial attempt) make other contributors possible.
  • Mooncake points: c12 dropped 3 of 1015 records (2 InvalidInferenceResultError from KV load failure after Failed to get ... Mooncake keys, 1 client ClientOSError) and c14 dropped 8 of 1054 (all InvalidInferenceResultError), matching the baseline nightly's own logs (c12: 2, c14: 8 KV load failures). Pre-existing load_async behaviour, not a v0.29.0 regression; all points stayed far below the AIPerf 10% abort threshold.
  • Repairs used: 0 of 5. Owned runs: targeted https://github.com/SemiAnalysisAI/InferenceX/actions/runs/34422803079 (success), final https://github.com/SemiAnalysisAI/InferenceX/actions/runs/34429897054 (success), and the pre-label synchronize run https://github.com/SemiAnalysisAI/InferenceX/actions/runs/34429854496 (success, check-changelog only) and the post-cleanup unlabeled run https://github.com/SemiAnalysisAI/InferenceX/actions/runs/34436781249 (success, no benchmark jobs). All terminal.
  • Diagnosis: .github/workflows/run-sweep.yml sweep-agentic-evals calls benchmark-tmpl.yml without the eval-framework / eval-suite inputs (the file contains no eval-framework at all), so the template's default lm-eval applies; e2e-tests.yml forwards matrix.config['eval-framework'], which is why the targeted smoke ran minimax_m3_smoke. The generator has emitted minimax-vendor for MiniMax agentic evals since Add lightweight tool-use verifier evaluations / 添加轻量级工具调用验证评估 #2634 (2026-08-31), and utils/klaud/validation.py expected_evals takes the suite from the matrix, so the sweep and the verifier disagree for this family regardless of the image. The published baseline sweep ([Klaud Cold] Update minimaxm3-fp8-h200-vllm-agentic-mtp vLLM image to nightly-d9105ea8001e0a6d77a96327d17515bb5791fb36 #2875, 2026-09-08) also ran gsm8k under run-sweep.yml. Both files are shared workflow/tooling code outside this PR's edit scope, so no in-scope repair exists and no repair budget was spent.
  • Disposition: outcome failed, phase final-sweep, repairs 0 of 5. The PR is closed and the candidate branch klaud/auto-b3fff31b6f9a6956-b3fe5685ff662693 is retained so the same candidate is not re-run automatically into the same verifier gap (a retry would spend another targeted run plus full sweep). The green sweep artifacts remain on run 34429897054 for maintainer reuse. This is not evidence of image incompatibility: v0.29.0 completed every point and eval.
  • Next step (manual, outside Klaud Cold scope): forward eval-framework / eval-suite from the matrix in run-sweep.yml's eval jobs (matching e2e-tests.yml), or make the Klaud verifier expect the sweep's actual default eval; then reopen this PR or let the autosweep reselect the family. Reviewers should weigh the 5% to 15% throughput regression reported above before adopting v0.29.0 for this family.

最终完整 sweep

状态: sweep 通过,但最终验证未获认证;会话以结果 failed 停止(04:30Z)。精确 PR head 7568b8f8 上的 run-sweep.yml 运行 34429897054 全部通过(check-changelogsetup、全部 8 个吞吐点、全部 8 个评测、collect-evalsupload-changelog-metadataklaud-sweep-manifest 记录 full-sweep: true、8 个 agentic 点、8 个 agentic 评测)。随后 finish 拒绝了结果覆盖校验:清单中每个点的评测行声明 eval-framework: minimax-vendor / eval-suite: minimax_m3_smoke,但每个评测任务实际运行的是 lm-eval gsm8kEVAL_FRAMEWORK: lm-evalEVAL_SUITE 为空、产物为 eval_*_lm-eval__1),因此校验器发现 8 个预期的 minimax_m3_smoke 身份缺失、8 个非预期的 gsm8k 身份。全部 8 个基准身份(配方指纹、并发、镜像)均匹配。

  • 镜像 / head:vllm/vllm-openai:v0.29.0@sha256:c2914767605584b6d8f45686b82de173ecc99e781897aa3d0a66dacd72c51ae1,PR head 7568b8f8529b843ba13dcfa45dfe1c43d49ae23f
  • 相对初次尝试的变更:为 minimaxm3-fp8-h200-vllm-agentic-mtp 追加一条 perf-changelog.yaml 条目(无场景、评测选择或 append-only 修饰符;保留全部既有字节,utils/validate_perf_changelog.pyorigin/main 校验通过)。配方与配置无改动。
  • 运行:https://github.com/SemiAnalysisAI/InferenceX/actions/runs/34429897054(`run-sweep.yml`,head 7568b8f8 上的 pull_request labeled 事件,full-sweep-enabled 标签于 2026-09-10T02:33:10Z 添加且 PR 保持草稿;h200 容量门禁于 02:33:09Z 返回 0)。预期矩阵:完整的 8 点家族(TP8 c1/c2/c4/c6/c8/c10 常驻,c12/c14 Mooncake DRAM 卸载),每个点均运行默认 minimax_m3_smoke 评测。同一 head 上更早的 synchronize 运行 https://github.com/SemiAnalysisAI/InferenceX/actions/runs/34429854496 发生在添加标签之前,未执行任何基准任务。
  • 结果(每点 bmk_agentic_* 聚合,每点一小时 AgentX 回放,8 个点 power_valid 均为 1;按并发、TP8 拓扑与 KV 卸载分支匹配冻结基线;单次运行,无重复)。完整 sweep 的默认评测为 lm-eval gsm8k(em_strict,n = 1319),因此与冒烟不同,可与基线直接对比:
并发 KV 卸载 每 GPU 总吞吐 新 vs 基线 每 GPU 输出吞吐 新 vs 基线 平均 TTFT(秒)新 vs 基线 平均 TPOT(秒)新 vs 基线 中位交互性(tok/s)新 vs 基线 剖析请求数 新 vs 基线 gsm8k em_strict
1 947.0 vs 1078.8 (-12.2%) 7.30 vs 8.27 (-11.7%) 2.482 vs 2.089 (+18.8%) 0.0137 vs 0.0109 (+25.4%) 81.9 vs 102.0 (-19.7%) 159/170 vs 170 0.9696739954510993
2 1104.8 vs 1290.3 (-14.4%) 8.74 vs 10.57 (-17.3%) 1.094 vs 0.901 (+21.4%) 0.0132 vs 0.0103 (+27.8%) 86.7 vs 115.7 (-25.1%) 219/241 vs 309 0.9711902956785443
4 1458.9 vs 1542.8 (-5.4%) 10.94 vs 12.82 (-14.7%) 1.086 vs 0.992 (+9.5%) 0.0239 vs 0.0176 (+36.2%) 34.4 vs 54.9 (-37.4%) 299/343 vs 310 0.9711902956785443
6 1980.6 vs 2218.9 (-10.7%) 15.44 vs 17.66 (-12.6%) 1.027 vs 0.955 (+7.5%) 0.0269 vs 0.0234 (+15.1%) 34.4 vs 39.1 (-12.0%) 408/473 vs 485 0.9711902956785443
8 3107.8 vs 3417.9 (-9.1%) 22.11 vs 25.40 (-13.0%) 0.961 vs 0.938 (+2.5%) 0.0327 vs 0.0240 (+36.0%) 27.3 vs 39.0 (-30.1%) 810/897 vs 856 0.9727065959059894
10 3384.1 vs 3996.4 (-15.3%) 27.00 vs 31.80 (-15.1%) 1.284 vs 1.276 (+0.7%) 0.0369 vs 0.0284 (+30.0%) 25.1 vs 33.4 (-24.9%) 861/970 vs 962 0.9727065959059894
12 dram(Mooncake) 3697.4 vs 4263.4 (-13.3%) 28.49 vs 33.91 (-16.0%) 1.172 vs 1.164 (+0.7%) 0.0366 vs 0.0291 (+25.6%) 25.7 vs 32.5 (-20.9%) 883/1015 vs 979 0.974981046247157
14 dram(Mooncake) 3852.1 vs 4476.4 (-13.9%) 30.99 vs 37.05 (-16.4%) 1.431 vs 1.299 (+10.1%) 0.0378 vs 0.0301 (+25.5%) 24.8 vs 31.4 (-21.1%) 899/1054 vs 1018 0.9742228961334344
  • 评测:8 个点的 gsm8k em_strict 为 0.9697 至 0.9750,基线 sweep 为 0.9666 至 0.9735(在 ±1 个标准误内,infrastructure_success 全部为 true)。
  • 回退:所有点在 v0.29.0 上都比已发布的 2026-09-07 nightly 慢:每 GPU 总吞吐 -5.4%(c4)至 -15.3%(c10),主要来自 +15% 至 +36% 的平均 TPOT,而 GPU 前缀缓存命中率不变(常驻 0.93 至 0.95;Mooncake 0.75 至 0.83)、KV 池不变、平均 GPU 功耗低 2% 至 9%。平均 TTFT 变化 +0.7% 至 +21%(低并发下由尾部驱动;中位 TTFT 在几个百分点以内)。与仅解码变慢相符的范围内解释仍是 EAGLE3 草稿模型注意力后端:v0.29.0 早于 [Bugfix][Spec Decode] Honour the draft's attention_backend on Model Runner V2 vllm-project/vllm#54826,因此在 Model Runner V2 上草稿模型忽略配方的 FLASH_ATTN 覆盖,而 nightly 会遵循。未改动配方;此处仅向评审者报告而未修复,且正式版缺少的 300 个 main 提交(见初次尝试中的对比)意味着也可能有其他因素。
  • Mooncake 点:c12 丢弃 1015 条记录中的 3 条(2 个由 Failed to get ... Mooncake keys 后的 KV load failure 导致的 InvalidInferenceResultError,1 个客户端 ClientOSError),c14 丢弃 1054 条中的 8 条(全部为 InvalidInferenceResultError),与基线 nightly 自身日志一致(c12:2 次,c14:8 次 KV 加载失败)。这是 load_async 的既有行为,而非 v0.29.0 的回退;所有点都远低于 AIPerf 10% 的中止阈值。
  • 已用修复次数:0/5。自有运行:定向 https://github.com/SemiAnalysisAI/InferenceX/actions/runs/34422803079(成功)、最终 https://github.com/SemiAnalysisAI/InferenceX/actions/runs/34429897054(成功)、以及添加标签前的 synchronize 运行 https://github.com/SemiAnalysisAI/InferenceX/actions/runs/34429854496(成功,仅 check-changelog)以及清理后取消标签触发的运行 https://github.com/SemiAnalysisAI/InferenceX/actions/runs/34436781249(成功,无基准任务)。全部已终止。
  • 诊断:.github/workflows/run-sweep.ymlsweep-agentic-evals 调用 benchmark-tmpl.yml 时未传入 eval-framework / eval-suite(该文件中完全没有 eval-framework),因此模板默认的 lm-eval 生效;e2e-tests.yml 会转发 matrix.config['eval-framework'],这就是定向冒烟能运行 minimax_m3_smoke 的原因。生成器自 Add lightweight tool-use verifier evaluations / 添加轻量级工具调用验证评估 #2634(2026-08-31)起为 MiniMax agentic 评测输出 minimax-vendor,而 utils/klaud/validation.pyexpected_evals 从矩阵读取评测套件,因此无论镜像如何,该家族的 sweep 与校验器都不一致。已发布基线 sweep([Klaud Cold] Update minimaxm3-fp8-h200-vllm-agentic-mtp vLLM image to nightly-d9105ea8001e0a6d77a96327d17515bb5791fb36 #2875,2026-09-08)在 run-sweep.yml 下同样运行的是 gsm8k。这两个文件都是共享工作流/工具代码,超出本 PR 的编辑范围,因此不存在范围内的修复,也未消耗修复预算。
  • 处置:结果 failed,阶段 final-sweep,修复 0/5。PR 关闭,候选分支 klaud/auto-b3fff31b6f9a6956-b3fe5685ff662693 予以保留,以免同一候选被自动重跑并再次撞上同一校验器缺口(重试将再消耗一次定向运行加完整 sweep)。通过的 sweep 产物保留在运行 34429897054 上供维护者复用。这不是镜像不兼容的证据:v0.29.0 完成了每一个点与评测。
  • 下一步(人工,超出 Klaud Cold 范围):在 run-sweep.yml 的评测任务中从矩阵转发 eval-framework / eval-suite(与 e2e-tests.yml 一致),或让 Klaud 校验器接受 sweep 的实际默认评测;然后重新打开本 PR 或让自动 sweep 重新选择该家族。评审者在为该家族采用 v0.29.0 前应权衡上文报告的 5% 至 15% 吞吐回退。

@github-actions

Copy link
Copy Markdown
Contributor

@Klaud-Cold

Copy link
Copy Markdown
Collaborator Author

Stop — final validation not certified. The exact-head full sweep is green, but finish rejected eval coverage because run-sweep.yml ran lm-eval gsm8k while the dispatched matrix declares minimax-vendor / minimax_m3_smoke (shared-workflow gap, out of scope). Outcome failed, repairs 0 of 5; the PR closes and the candidate branch is retained for manual review. Details, per-point results and the reported 5% to 15% throughput regression are in the Final full sweep comment.


停止 — 最终验证未获认证。 精确 head 的完整 sweep 已通过,但 finish 拒绝了评测覆盖校验:run-sweep.yml 运行的是 lm-eval gsm8k,而分派矩阵声明的是 minimax-vendor / minimax_m3_smoke(共享工作流缺口,超出范围)。结果 failed,修复 0/5;PR 关闭,候选分支保留以供人工评审。详情、各点结果及所报告的 5% 至 15% 吞吐回退见最终完整 sweep 评论。

@Klaud-Cold

Copy link
Copy Markdown
Collaborator Author

Klaud Cold: failed. Finishing cleanup; owned child runs will be stopped and checked before closure.


Klaud Cold:failed。正在完成清理;将先停止并确认自有子运行的状态,再关闭 PR。

@Klaud-Cold

Copy link
Copy Markdown
Collaborator Author

Klaud Cold: failed. All owned runs are terminal. Repairs: 0. Runs: 34422803079, 34429854496, 34429897054, 34436781249.

PR closed; the exact-candidate branch is retained for manual review. The interruption does not prove image incompatibility.


Klaud Cold:failed。所有自有运行均已结束。修复次数:0。运行:34422803079, 34429854496, 34429897054, 34436781249

PR 已关闭;保留该候选的分支,等待人工审查。运行中断不能证明镜像不兼容。

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Development

Successfully merging this pull request may close these issues.

1 participant