Repository navigation
docs: use Grouped GEMM for the agent launch walkthrough - #4
Merged
Merged
Conversation
YiyanZhai
force-pushed
the
codex/grouped-gemm-mlsys
branch
from
September 29, 2026 18:08
86abbdc to
e29baea
Compare
jinhongyii
reviewed
Sep 29, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The launch walkthrough currently uses KDA. Use Grouped GEMM for the introductory task, while retaining the recorded KDA examples in the following diagnostic and profiling chapters.
grouped_gemm_fp8task and unified setup/evaluation commands, covering four configurations with DeepGEMM as the correctness reference and performance baseline.Launching the Agentchapter.Validation: strict HTML build (
make html, Sphinx warnings treated as errors) andgit diff --checkpassed. GPU end-to-end validation remains pending.Dependency: TIRx-harness #1 migrates the four-configuration task to the new Harness repository and adds the rebuilt DeepGEMM wheels to its benchmark/server dependencies. Merge that PR before using the documented default-branch setup.
Harness now has the packaging changes in its linked PR, with wheels for Python 3.12 and 3.13. The chapter states the wheel runtime requirements: Linux x86_64, glibc >= 2.38, and CUDA Toolkit 13.2.