diff --git a/benchmarks/single_node/agentic/minimaxm3_fp8_mi300x_mtp.sh b/benchmarks/single_node/agentic/minimaxm3_fp8_mi300x_mtp.sh index 4234a66490..32a77f2ee8 100755 --- a/benchmarks/single_node/agentic/minimaxm3_fp8_mi300x_mtp.sh +++ b/benchmarks/single_node/agentic/minimaxm3_fp8_mi300x_mtp.sh @@ -162,6 +162,12 @@ VLLM_CMD=( --trust-remote-code --block-size 128 --gpu-memory-utilization 0.90 + # The upstream nightly does not torch-compile MiniMaxM3SparseForConditionalGeneration, + # and with VLLM_USE_BREAKABLE_CUDAGRAPH=0 the default FULL_AND_PIECEWISE graph + # mode aborts at init ("piecewise CUDA graphs unavailable ... Set + # VLLM_USE_BREAKABLE_CUDAGRAPH=1 or cudagraph_mode=NONE/FULL"; run 34174124043). + # Capture full decode-only graphs, as the MI355X sibling does on its nightly (#2825). + --compilation-config '{"cudagraph_mode":"FULL_DECODE_ONLY"}' --enable-chunked-prefill --max-num-batched-tokens 16384 --language-model-only diff --git a/configs/amd-master.yaml b/configs/amd-master.yaml index cd10677cd7..2510dc3b8d 100644 --- a/configs/amd-master.yaml +++ b/configs/amd-master.yaml @@ -1533,7 +1533,7 @@ dsv4-fp4-mi355x-atom-disagg: - "DECODE_NODES=1" # 1P1D TP8 minimaxm3-fp8-mi300x-vllm-agentic-mtp: - image: vllm/vllm-openai-rocm:v0.27.1 + image: vllm/vllm-openai-rocm:nightly-d9105ea8001e0a6d77a96327d17515bb5791fb36 model: MiniMaxAI/MiniMax-M3-MXFP8 model-prefix: minimaxm3 runner: cluster:mi300x-amd diff --git a/perf-changelog.yaml b/perf-changelog.yaml index 2d3102f695..e0ef448195 100644 --- a/perf-changelog.yaml +++ b/perf-changelog.yaml @@ -6938,3 +6938,12 @@ description: - "Update vLLM image from vllm/vllm-openai:v0.27.1 (v0.27.1 release) to vllm/vllm-openai:nightly-d9105ea8001e0a6d77a96327d17515bb5791fb36 (2026-09-07 upstream nightly, digest sha256:254eebf919e8b7b0d530d97fccc606c36f380ff6724d64bc190951bec1aee838, tag commit vllm-project/vllm@d9105ea8; Docker Hub last pushed 2026-09-07T06:16:01Z), the same tag the B200 MiniMax-M3 AgentX recipe moved to in #2860 and the ROCm counterpart of the MI325X/MI300X bumps in #2872/#2873. benchmarks/single_node/agentic/minimaxm3_fp8_h200_mtp.sh is unchanged: TRITON_ATTN attention, fp8 KV, EAGLE3 with the Inferact MiniMax-M3 EAGLE3-GQA draft pinned to FLASH_ATTN and the committed golden synthetic acceptance length 2.78, Mooncake 0.3.11.post1 DRAM offload on the host-tier arm; the resident TP8 c1/c2/c4/c6/c8/c10 and Mooncake DRAM offload c12/c14 grid is unchanged." pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/2875 + +- config-keys: + - minimaxm3-fp8-mi300x-vllm-agentic-mtp + scenario-type: + - agentic-coding + description: + - "Update vLLM ROCm image from vllm/vllm-openai-rocm:v0.27.1 (v0.27.1 release) to vllm/vllm-openai-rocm:nightly-d9105ea8001e0a6d77a96327d17515bb5791fb36 (2026-09-07 upstream ROCm nightly, digest sha256:74d4a95f3ae672ecddf9acb7296917def82d9eca51687fa2862ae72b03ff1907, tag commit vllm-project/vllm@d9105ea8; Docker Hub last pushed 2026-09-07T05:26:48Z). benchmarks/single_node/agentic/minimaxm3_fp8_mi300x_mtp.sh is unchanged: TRITON_ATTN attention, fp8 KV, block-size 128, EAGLE3 speculative decoding with the Inferact MiniMax-M3 EAGLE3-GQA draft and the committed golden synthetic acceptance length, minimax_m3 tool-call and reasoning parsers; the search space is unchanged. The upstream commit-pinned ROCm nightly is the same vllm commit the B200 MiniMax-M3 AgentX recipe moved to in #2860. Note that vllm-openai-rocm commit-nightly tags have expired from Docker Hub within days in the past; node squash caches keep merged configs running, but a re-pin to a durable tag may be needed later." + - "Add --compilation-config cudagraph_mode=FULL_DECODE_ONLY to the serve command. The upstream nightly does not torch-compile MiniMaxM3SparseForConditionalGeneration, so with VLLM_USE_BREAKABLE_CUDAGRAPH=0 the default FULL_AND_PIECEWISE mode aborts at engine init with piecewise CUDA graphs unavailable (first sweep, run 34174124043, eval cell); full decode-only graphs are what the MI355X MiniMax-M3 sibling runs on its nightly (#2825)." + pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/2873