From 465389fb48657d5f02ebe60b87b5ead59365b6f9 Mon Sep 17 00:00:00 2001 From: Klaud-Cold Date: Thu, 10 Sep 2026 12:36:30 +0000 Subject: [PATCH 1/2] config: update qwen3.5-fp8-h200-sglang-agentic-hicache-mtp SGLang image to v0.5.19-cu130 MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Move the H200 Qwen3.5 FP8 SGLang AgentX HiCache MTP recipe from the lmsysorg/sglang:nightly-dev-cu13-20260907-30705c00 dev nightly to the v0.5.19 release image lmsysorg/sglang:v0.5.19-cu130 (digest sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9, build commit sgl-project/sglang@0bcd822377da7b5718e674eaf9c870d349424dd1). Model, TP8/EP1 topology, EAGLE MTP settings, DRAM HiCache offload, concurrency list and the launch script are unchanged. 将 H200 Qwen3.5 FP8 SGLang AgentX HiCache MTP 配方的镜像从开发 nightly lmsysorg/sglang:nightly-dev-cu13-20260907-30705c00 切换到 v0.5.19 发布版 lmsysorg/sglang:v0.5.19-cu130(摘要 sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9, 构建提交 sgl-project/sglang@0bcd822377da7b5718e674eaf9c870d349424dd1)。 模型、TP8/EP1 拓扑、EAGLE MTP 设置、DRAM HiCache 卸载、并发列表和启动脚本均保持不变。 Co-Authored-By: Claude Fable 5.1 --- configs/nvidia-master.yaml | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/configs/nvidia-master.yaml b/configs/nvidia-master.yaml index 0b8993e83..5df5ba41b 100644 --- a/configs/nvidia-master.yaml +++ b/configs/nvidia-master.yaml @@ -7544,7 +7544,7 @@ qwen3.8next-fp8-h200-sglang-agentic-mtp: # H200 AgentX MTP frontier with DRAM HiCache. This is intentionally an MTP-only # submission; the model's non-speculative AgentX arm is not included. qwen3.5-fp8-h200-sglang-agentic-hicache-mtp: - image: lmsysorg/sglang:nightly-dev-cu13-20260907-30705c00 + image: lmsysorg/sglang:v0.5.19-cu130 model: Qwen/Qwen3.5-397B-A17B-FP8 model-prefix: qwen3.5 runner: cluster:h200-dgxc From 8604d93ed9b21e0fb382fb4e5cb5a5fe61181016 Mon Sep 17 00:00:00 2001 From: Klaud-Cold Date: Thu, 10 Sep 2026 16:51:44 +0000 Subject: [PATCH 2/2] changelog: record qwen3.5-fp8-h200-sglang-agentic-hicache-mtp image update to v0.5.19-cu130 MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Append the perf-changelog entry for moving the H200 Qwen3.5 FP8 SGLang AgentX HiCache MTP recipe to lmsysorg/sglang:v0.5.19-cu130 (PR #2966), rebased onto current main after other changelog entries landed. The entry selects the whole family without scenario, eval-selection or append-only modifiers. 为将 H200 Qwen3.5 FP8 SGLang AgentX HiCache MTP 配方切换到 lmsysorg/sglang:v0.5.19-cu130(PR #2966)追加 perf-changelog 条目;在其他 changelog 条目合入后已重新基于当前 main。该条目选择整个配方族,不带 scenario、评测选择或 append-only 修饰符。 Co-Authored-By: Claude Fable 5.1 --- perf-changelog.yaml | 7 +++++++ 1 file changed, 7 insertions(+) diff --git a/perf-changelog.yaml b/perf-changelog.yaml index c3d266eae..6385eb468 100644 --- a/perf-changelog.yaml +++ b/perf-changelog.yaml @@ -7128,3 +7128,10 @@ - "Update the SGLang image from nightly-dev-cu13-20260829-89816a21 to the digest-pinned v0.5.19-cu130 release (lmsysorg/sglang:v0.5.19-cu130@sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9)." - "Move the TP8 and TP4 low-latency recipes and the master router metadata from Dynamo 1.5.0.dev20260902 to 1.5.0.dev20260910, the nightly that carries the ServerArgs.get_model_config()/use_mla_backend() compatibility fix for SGLang #36972 and pins sglang==0.5.19." pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/2969 + +- config-keys: + - qwen3.5-fp8-h200-sglang-agentic-hicache-mtp + description: + - "Update SGLang image from lmsysorg/sglang:nightly-dev-cu13-20260907-30705c00 (2026-09-07 cu13 dev nightly, build commit sgl-project/sglang@30705c00) to the v0.5.19 release image lmsysorg/sglang:v0.5.19-cu130 (digest sha256:d6e7288627be8b02be88e4bba38e73f6d50e2826869f753c13a4c4385ab3eda9, build commit sgl-project/sglang@0bcd822377da7b5718e674eaf9c870d349424dd1, Docker Hub last pushed 2026-09-04T22:50:19Z)." + - "The release image ships the same CUDA 13.0.3, FlashInfer 0.6.18 and sgl-kernel 0.4.6.post1 as the nightly. benchmarks/single_node/agentic/qwen3.5_fp8_h200_mtp.sh is unchanged: SGLANG_ENABLE_SPEC_V2 EAGLE MTP at 3 steps, golden acceptance length 3.39, flashinfer attention with allreduce fusion, fp8 quantization and fp8_e4m3 KV, HiCache kernel IO / page_first layout. TP8/EP1 DRAM HiCache concurrency 2 through 24 unchanged." + pr-link: https://github.com/SemiAnalysisAI/InferenceX/pull/2966