From 61d9f8401856f1fa63e76a7623e55bea514918a3 Mon Sep 17 00:00:00 2001 From: YiyanZhai Date: Tue, 29 Sep 2026 12:57:40 -0400 Subject: [PATCH 1/7] docs: use Grouped GEMM for the agent launch walkthrough --- README.md | 8 +-- chapters/launching-the-agent.md | 94 +++++++++++++++++++++------------ 2 files changed, 64 insertions(+), 38 deletions(-) diff --git a/README.md b/README.md index 9efbdb7..50686e6 100644 --- a/README.md +++ b/README.md @@ -3,7 +3,8 @@ A book on the components of agentic GPU programming, their implementation in TIRx Harness, and their use in a kernel optimization task. General concepts appear in Part I, the concrete system in Part II, and worked examples in Part -III. Kimi-Delta Attention forward supplies the worked task. +III. Grouped GEMM supplies the workflow tutorial; recorded Kimi Delta +Attention runs supply the diagnostic and review cases. ## Build and preview @@ -68,8 +69,9 @@ including under a URL subdirectory. No application backend is needed. contents. Edit HTML directly; no generation step is required. - `conf.py` owns the theme settings and repository source-link templates. - `chapters/launching-the-agent.md` opens Part III with environment preparation, - launcher tabs, and result inspection. It uses the task prompt and workspace - generated by the harness setup script. `chapters/self-improvement.md` + launcher tabs, and result inspection for M-grouped contiguous FP8 GEMM. + The harness task setup and benchmark entry points prepare and evaluate the run. + `chapters/self-improvement.md` covers review and harness improvement. `chapters/advanced-tips.md` closes the book with practices for a long search. - `_static/book-tabs.js` handles the launcher tabs using the book's normal diff --git a/chapters/launching-the-agent.md b/chapters/launching-the-agent.md index 753aa7a..b079c1a 100644 --- a/chapters/launching-the-agent.md +++ b/chapters/launching-the-agent.md @@ -1,10 +1,19 @@ # Launching the Agent -Part I introduced the compiler harness and the agent workflows +Part I introduced the compiler harness and the [agent workflows](agent-workflows.md) that organize an optimization search; Part II described how TIRx Harness provides the programming and evaluation tools. We now bring them together to optimize the -forward pass of Kimi Delta Attention on a B200 GPU, for one sequence of -8,192 tokens with 96 heads. +M-grouped contiguous FP8 GEMM on a B200 GPU. The agent implements and +optimizes the kernel in TIRx, using DeepGEMM as the correctness reference +and performance baseline. + +We use Grouped GEMM here to help you get started quickly: a focused workload +with a straightforward computation lets you run the workflow and see results +before working through a more involved kernel. Starting with +the next chapter's Synccheck and Racecheck examples, we switch to a recorded +Kimi Delta Attention (KDA) run to examine the agent's use of analysis and +performance feedback in detail. The workflow introduced here carries over +to that task. In this chapter, we will prepare the task workspace and GPU evaluation service, check the baseline, and launch the optimization run using either @@ -21,17 +30,25 @@ as shown below. During the run, the agent retrieves relevant implementations, writes candidate kernels, and submits them for analysis and GPU evaluation. The returned -feedback guides its next experiment. We begin by preparing the environment -in which this work will take place. +feedback guides its next experiment. We first define the task, then prepare +the environment in which this work will take place. + +## The Task: Grouped GEMM + +Grouped GEMM computes $D_g = A_g B_g^\mathsf{T}$ for groups with different +row counts, as in mixture-of-experts models. We use the M-grouped contiguous +layout with FP8 operands and BF16 output. The task covers four configurations +with 4 or 8 groups and different matrix dimensions. + +The agent implements and optimizes this operation in TIRx. The harness fixes +the inputs and correctness checks and measures performance against DeepGEMM. ## Prepare the Environment An agentic run needs a place for the agent to work and a GPU for its -experiments, both ready before the agent starts. This example optimizes -Kimi Delta Attention forward on a B200 with `B=1`, -`T=8192`, `H=96`, and `K=V=128`. We use two machines: one runs the coding -agent, and the other is a GPU server running KCoral to execute kernels and -collect measurements. The agent and KCoral can also run on the same machine. +experiments, both ready before the agent starts. We use two machines: one +runs the coding agent, and the other is a GPU server running KCoral to +execute kernels and collect measurements. The agent and KCoral can also run on the same machine. First, we will clone the harness and start KCoral on the GPU server. On the agent machine, `evolution/setup.py` then creates a task worktree and Python @@ -63,7 +80,7 @@ python -m pip install -r evolution/preparation/requirements.txt ### Start KCoral on the GPU server In the GPU server's terminal, from the harness checkout, install the benchmark -and server dependencies and start KCoral on GPU 0: +and server dependencies: ```{warning} KCoral executes arbitrary code. Allow only trusted clients on an isolated @@ -90,10 +107,10 @@ export KCORAL_URL="http://gpu-server:8000" curl --fail "$KCORAL_URL/health" ``` -Then prepare the fixed-shape task: +Then prepare the task: ```bash -python evolution/setup.py --task kda_forward_b1_t8192_h96 \ +python evolution/setup.py --task grouped_gemm_fp8 \ --remote "$KCORAL_URL" ``` @@ -101,7 +118,7 @@ The run's `.venv` uses the Python interpreter that launched setup and includes the installed `tirx-harness` wheel. The run directory contains `PROMPT.md`, `manifest.json`, `flowverse.yaml`, and `worktree/`; candidate kernels will live under -`candidates/kda/forward_b1_t8192_h96/` in that worktree. +`candidates/grouped_gemm/fp8/` in that worktree. Set `run_dir` to the absolute path printed by setup, enter the worktree, and activate its environment: @@ -116,8 +133,8 @@ cat "$run_dir/PROMPT.md" `PROMPT.md` combines the task definition with this run's worktree path, service address, benchmark command, authoring rules, and candidate-recording instructions. The prompt asks the agent to make the kernel as fast as possible -while preserving correctness. Chunking, fusion, layouts, warp roles, and -pipelining remain the agent's choices. +while preserving correctness. Tile sizes, group scheduling, layouts, warp roles, +and pipelining remain the agent's choices. ### Check the baseline @@ -127,24 +144,28 @@ it has this form: ```bash python evolution/remote/kcoral_remote.py \ - candidates/kda/forward_b1_t8192_h96 baseline \ + candidates/grouped_gemm/fp8 baseline \ --remote "$KCORAL_URL" --timeout 600 ``` -Confirm that the output includes the timed row `kda-fwd-h96-fixed`, seven -stress probes, and two fp64 probes, all at the required H=96, T=8192 shape. -The final summary must report `passed: 10/10`; individual rows also print -`passed: 1/1`. The private holdout result is reported separately. For this -first call, the candidate is FlashKDA itself, the performance baseline; later -calls time each passing candidate against it in the same call. +Confirm that the summary reports `passed: 4/4`. Every group's valid output +rows must pass the correctness check for every configuration; alignment +padding is excluded. For this first call, the candidate is DeepGEMM itself, +the performance baseline. Later calls compare each passing TIRx candidate +against DeepGEMM on the same quantized inputs, with preparation and compilation +outside timing. Each workload reports `DeepGEMM time / candidate time`; +the summary includes the geometric mean across all four workloads. + +To evaluate a candidate, replace `baseline` with its directory relative to +the workload, such as `scratch/first`. That directory contains `solution.py`, +which exports `setup(data, G, M, N, K) -> callable` as specified in `PROMPT.md`. ## Kick Off Agentic Runs An optimization run lasts many turns, so the agent needs a launcher that keeps it working toward the goal instead of stopping after one answer. Start the run -in the prepared worktree with its environment active. All launch -options use the same generated task, correctness checks, benchmark, and -candidate directory. +in `$run_dir/worktree` with the harness environment active. All launch +options use the same task prompt, correctness checks, and benchmark. ::::{container} launch-tabs ```{raw} html @@ -162,11 +183,11 @@ session, enter the following prompt: ```text /goal Read ../PROMPT.md in full and carry out the optimization task. -Reach at least 3.0x speedup over the prescribed performance baseline -with a reproducibly passing kernel and retain it in the frontier. +Optimize against the prescribed DeepGEMM performance baseline and retain +reproducibly passing candidates with their measurements in the frontier. ``` -The 3× speedup target is an example; choose a target that fits your task. +Set a time budget or performance target appropriate to this workload. ::: @@ -202,15 +223,18 @@ recent experiments, current optimization direction, and any issues needing attention, with evidence from those records. In the generated worktree, open `frontier/index.json` under -`candidates/kda/forward_b1_t8192_h96/` for the retained candidates and their -benchmark results. +`candidates/grouped_gemm/fp8/` for the retained +candidates and their benchmark results. ## Continuing Through Part III -The next two chapters follow analysis and performance feedback from a -recorded Kimi Delta Attention run to show how the agent diagnoses errors and -improves its kernels. The remaining chapters turn to reviewing results, -improving the harness, and keeping future searches productive: +With the Grouped GEMM workflow in place, we now turn to the recorded KDA +case study. Beginning with Synccheck and Racecheck in the next chapter, we +follow concrete diagnostics, the agent's responses, and the measurements +that guide further optimization. These examples come from the KDA run, not +the Grouped GEMM run you just prepared. The remaining chapters turn to +reviewing results, improving the harness, and keeping future searches +productive: - [Compiler Analysis Deep Dive](compiler-analysis-deepdive.md) follows how the agent calls the analyses, interprets their findings, and repairs From bab85b27f2d0097f331ed1f172cf347b983e89ab Mon Sep 17 00:00:00 2001 From: YiyanZhai Date: Tue, 29 Sep 2026 12:59:59 -0400 Subject: [PATCH 2/7] docs: simplify DeepGEMM environment setup --- chapters/launching-the-agent.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/chapters/launching-the-agent.md b/chapters/launching-the-agent.md index b079c1a..a7c02cc 100644 --- a/chapters/launching-the-agent.md +++ b/chapters/launching-the-agent.md @@ -80,7 +80,7 @@ python -m pip install -r evolution/preparation/requirements.txt ### Start KCoral on the GPU server In the GPU server's terminal, from the harness checkout, install the benchmark -and server dependencies: +and server dependencies, including DeepGEMM, and start KCoral on GPU 0: ```{warning} KCoral executes arbitrary code. Allow only trusted clients on an isolated From d49851ac27e99dbd7217c84263825bdbde61608c Mon Sep 17 00:00:00 2001 From: YiyanZhai Date: Tue, 29 Sep 2026 13:01:14 -0400 Subject: [PATCH 3/7] docs: align book index with Grouped GEMM tutorial and new repository --- index.md | 5 +++-- 1 file changed, 3 insertions(+), 2 deletions(-) diff --git a/index.md b/index.md index 8a47fa0..5b6cd24 100644 --- a/index.md +++ b/index.md @@ -29,7 +29,7 @@ and we plan to integrate it into the Mellon University. This book is open source. Contributions, corrections, and examples are welcome -through the [GitHub repository](https://github.com/mlc-ai/agentic-gpu-programming). +through the [GitHub repository](https://github.com/mlc-ai/agentic-gpu-programming-for-mlsys). ## How This Book Is Organized @@ -48,7 +48,8 @@ through the [GitHub repository](https://github.com/mlc-ai/agentic-gpu-programmin programming out end to end. It looks closely at the feedback the harness returns over the course of a run and at the interaction patterns between the agent and each element of the harness, and it closes with a few advanced - tips. Kimi Delta Attention supplies the worked example. + tips. Grouped GEMM introduces the workflow; recorded Kimi Delta Attention + runs supply the diagnostic and review examples. ```{toctree} :caption: Part I, Elements of Agentic GPU Programming From e29baea1bcc48487c3d0c4e62445de5a9dbdc0e8 Mon Sep 17 00:00:00 2001 From: YiyanZhai Date: Tue, 29 Sep 2026 13:44:24 -0400 Subject: [PATCH 4/7] docs: state the packaged DeepGEMM runtime requirements --- chapters/launching-the-agent.md | 7 ++++--- 1 file changed, 4 insertions(+), 3 deletions(-) diff --git a/chapters/launching-the-agent.md b/chapters/launching-the-agent.md index a7c02cc..1b7beb0 100644 --- a/chapters/launching-the-agent.md +++ b/chapters/launching-the-agent.md @@ -57,9 +57,10 @@ baseline through KCoral before launching the agent in that worktree. ### Prepare the harness checkout -Use Linux x86_64, Python 3.12 or 3.13, and pip 25.1+ on both machines. -The agent machine needs `uv`, Rust 1.89+ with Cargo, C/C++ build tools, and -Python development headers; the GPU server needs CUDA and a compatible driver. +Use Linux x86_64 (glibc 2.38 or newer), Python 3.12 or 3.13, and pip 25.1+ +on both machines. The agent machine needs `uv`, Rust 1.89+ with Cargo, +C/C++ build tools, and Python development headers; the GPU server needs +CUDA Toolkit 13.2 and a compatible driver. For profiling, install Nsight Compute on both machines. Clone the harness where the agent runs and on the GPU server. If they share From e8b61d5a55764e22a1479b3817784ac202f816eb Mon Sep 17 00:00:00 2001 From: YiyanZhai Date: Tue, 29 Sep 2026 14:25:40 -0400 Subject: [PATCH 5/7] docs: explain the progression from Grouped GEMM to KDA --- chapters/launching-the-agent.md | 15 ++++++++------- 1 file changed, 8 insertions(+), 7 deletions(-) diff --git a/chapters/launching-the-agent.md b/chapters/launching-the-agent.md index 1b7beb0..5bfcf00 100644 --- a/chapters/launching-the-agent.md +++ b/chapters/launching-the-agent.md @@ -7,13 +7,14 @@ M-grouped contiguous FP8 GEMM on a B200 GPU. The agent implements and optimizes the kernel in TIRx, using DeepGEMM as the correctness reference and performance baseline. -We use Grouped GEMM here to help you get started quickly: a focused workload -with a straightforward computation lets you run the workflow and see results -before working through a more involved kernel. Starting with -the next chapter's Synccheck and Racecheck examples, we switch to a recorded -Kimi Delta Attention (KDA) run to examine the agent's use of analysis and -performance feedback in detail. The workflow introduced here carries over -to that task. +We start with Grouped GEMM, using DeepGEMM as the baseline, because its simpler +computation makes it easier to get started and quickly see performance improve +through successive optimization iterations. Starting with the next chapter's +Synccheck and Racecheck examples, we switch to a recorded Kimi Delta Attention +(KDA) run. KDA's greater complexity lets us demonstrate more of the harness's +capabilities, including correctness checks, synchronization analysis, and +profiling feedback, and how they help the agent diagnose problems and improve +performance. The workflow introduced here carries over to that task. In this chapter, we will prepare the task workspace and GPU evaluation service, check the baseline, and launch the optimization run using either From 73449c152055a118f4cd26a0b870c84d8df21d5f Mon Sep 17 00:00:00 2001 From: YiyanZhai Date: Tue, 29 Sep 2026 14:27:08 -0400 Subject: [PATCH 6/7] docs: address launch walkthrough review comments --- chapters/launching-the-agent.md | 12 ++++-------- 1 file changed, 4 insertions(+), 8 deletions(-) diff --git a/chapters/launching-the-agent.md b/chapters/launching-the-agent.md index 5bfcf00..d089828 100644 --- a/chapters/launching-the-agent.md +++ b/chapters/launching-the-agent.md @@ -82,7 +82,7 @@ python -m pip install -r evolution/preparation/requirements.txt ### Start KCoral on the GPU server In the GPU server's terminal, from the harness checkout, install the benchmark -and server dependencies, including DeepGEMM, and start KCoral on GPU 0: +and server dependencies and start KCoral on GPU 0: ```{warning} KCoral executes arbitrary code. Allow only trusted clients on an isolated @@ -158,10 +158,6 @@ against DeepGEMM on the same quantized inputs, with preparation and compilation outside timing. Each workload reports `DeepGEMM time / candidate time`; the summary includes the geometric mean across all four workloads. -To evaluate a candidate, replace `baseline` with its directory relative to -the workload, such as `scratch/first`. That directory contains `solution.py`, -which exports `setup(data, G, M, N, K) -> callable` as specified in `PROMPT.md`. - ## Kick Off Agentic Runs An optimization run lasts many turns, so the agent needs a launcher that keeps @@ -185,11 +181,11 @@ session, enter the following prompt: ```text /goal Read ../PROMPT.md in full and carry out the optimization task. -Optimize against the prescribed DeepGEMM performance baseline and retain -reproducibly passing candidates with their measurements in the frontier. +Reach at least 3.0x speedup over the prescribed performance baseline +with a reproducibly passing kernel and retain it in the frontier. ``` -Set a time budget or performance target appropriate to this workload. +The 3× speedup target is an example; choose a target that fits your task. ::: From f57479150f53952720248d3810610a88964bc70a Mon Sep 17 00:00:00 2001 From: YiyanZhai Date: Tue, 29 Sep 2026 14:32:08 -0400 Subject: [PATCH 7/7] docs: introduce the KDA case study at the analysis chapter opening --- chapters/compiler-analysis-deepdive.md | 8 +++++--- 1 file changed, 5 insertions(+), 3 deletions(-) diff --git a/chapters/compiler-analysis-deepdive.md b/chapters/compiler-analysis-deepdive.md index fb11610..5859a59 100644 --- a/chapters/compiler-analysis-deepdive.md +++ b/chapters/compiler-analysis-deepdive.md @@ -1,8 +1,10 @@ # Compiler Analysis Deep Dive -The last chapter showed how to carry out agentic GPU programming with the -TIRx Harness. In this chapter, we will take a deep dive and review how an -agent interacts with Synccheck and Racecheck over one specific run. +The last chapter introduced the agent workflow through Grouped GEMM, using +DeepGEMM as the baseline. Starting here, we switch to a recorded Kimi Delta +Attention (KDA) run: its more complex kernel lets us explore the harness's +analysis tools in greater depth. This chapter follows how the agent uses +Synccheck and Racecheck to diagnose and fix synchronization errors and data races. A hang or a race can survive many GPU runs before it shows up, so the agent needs feedback that does not depend on timing. Synccheck and Racecheck supply