From 62293029e7e0901ea28733f6709aee0071f0ef88 Mon Sep 17 00:00:00 2001 From: Hongyi Jin Date: Tue, 29 Sep 2026 12:26:19 -0400 Subject: [PATCH 1/7] docs: align tutorial usage with current TIRx Harness --- chapters/benchmark-server-deepdive.md | 26 +++++++++++++------ chapters/launching-the-agent.md | 36 +++++++++++++++++---------- 2 files changed, 42 insertions(+), 20 deletions(-) diff --git a/chapters/benchmark-server-deepdive.md b/chapters/benchmark-server-deepdive.md index 81891a0..1311e2c 100644 --- a/chapters/benchmark-server-deepdive.md +++ b/chapters/benchmark-server-deepdive.md @@ -165,16 +165,25 @@ efficiently, changes the kernel, and benchmarks the result. A second profile helps it check whether the change had the intended effect. The kernel is a single fused kernel with one CTA per head, in which eight -preprocessing warps prepare each chunk for the MMA and state warps. `$WORKLOAD` stands for the -task's workload directory. Output excerpts keep only the key lines; `...` -marks omitted text. +preprocessing warps prepare each chunk for the MMA and state warps. The commands +below use paths relative to the generated worktree root. Set the workload path +there before running them: + +```bash +WORKLOAD=candidates/kda/forward_b1_t8192_h96 +``` + +Capture and analysis scripts such as `capture_iket.py` and `iket_analyze.py` +were written by the agent during this run. Output excerpts keep only the key +lines; `...` marks omitted text. At one point, the best kernel runs at 0.497283 ms, 2.1162× over the baseline. The agent captures an IKET timeline of it through KCoral and summarizes the time each preprocessing warp spends in each stage: ```bash -python kcoral_iket.py --remote http://10.0.2.2:8901 --send scratch/product-gamma \ +python evolution/remote/kcoral_iket.py --remote http://10.0.2.2:8901 \ + --send "$WORKLOAD/scratch/product-gamma" \ --output-dir iket_out -- profile --postprocess json -- python capture_iket.py python iket_analyze.py iket_out/iket_pid_0x8e.trace.json 13 ``` @@ -238,7 +247,8 @@ moves the diagonal inversions into the Gram tail: The CPU checks pass, and the official benchmark measures a new best: ```bash -python kcoral_remote.py $WORKLOAD scratch/overlap-diag-inverse-gram-tail --remote http://10.0.2.2:8901 --timeout 600 +python evolution/remote/kcoral_remote.py "$WORKLOAD" scratch/overlap-diag-inverse-gram-tail \ + --remote http://10.0.2.2:8901 --timeout 600 ``` ```text @@ -261,7 +271,8 @@ the agent profiles one launch of it with NCU through KCoral and reads the report's source counters locally: ```bash -python kcoral_ncu.py --remote http://10.0.2.2:8901 --send scratch/profile-qacc-ring \ +python evolution/remote/kcoral_ncu.py --remote http://10.0.2.2:8901 \ + --send "$WORKLOAD/scratch/profile-qacc-ring" \ -o kda-full.ncu-rep --set full --launch-count 1 --kernel-name kda_fwd_kernel -- python capture_ncu.py ncu --import kda-full.ncu-rep --page details --section SourceCounters ``` @@ -326,7 +337,8 @@ picks out its own head in shared memory: The CPU checks pass, and the official benchmark measures a new best: ```bash -python kcoral_remote.py $WORKLOAD scratch/beta-tma --remote http://10.0.2.2:8901 --timeout 600 +python evolution/remote/kcoral_remote.py "$WORKLOAD" scratch/beta-tma \ + --remote http://10.0.2.2:8901 --timeout 600 ``` ```text diff --git a/chapters/launching-the-agent.md b/chapters/launching-the-agent.md index 4ec7963..1f97e68 100644 --- a/chapters/launching-the-agent.md +++ b/chapters/launching-the-agent.md @@ -33,10 +33,17 @@ Kimi Delta Attention forward on a B200 with `B=1`, agent, and the other is a GPU server running KCoral to execute kernels and collect measurements. The agent and KCoral can also run on the same machine. -### Install the harness - -The environment running the agent needs Python, Rust, and `uv`. The GPU server -needs CUDA and, for NCU profiling, Nsight Compute. +### Prepare the harness checkout + +This example follows the harness's +{harness}`registered-workload setup `, which builds +the harness in each run's environment. Use Linux x86_64 with Python 3.12 or +3.13 and pip 25.1 or later. The agent machine also needs `uv`, Rust 1.89.0 or +later with Cargo, and the C/C++ build tools and Python development headers +listed in the {harness}`source-build prerequisites `. +Install the CUDA Toolkit wherever kernels are compiled, and a compatible +NVIDIA driver on the GPU server. NCU captures need Nsight Compute on the +server; reading the returned reports also needs `ncu` on the agent machine. Clone the harness where the agent runs and on the GPU server. If they share one machine, a single checkout is enough: @@ -47,11 +54,11 @@ cd TIRx-harness git submodule update --init thirdparty/tvm-rust-ext ``` -On the agent machine, install the harness in your existing Python environment -from the repository root: +On the agent machine, install the setup dependencies in the Python environment +you will use to launch setup, from the repository root: ```bash -python -m pip install . +python -m pip install -r evolution/preparation/requirements.txt ``` ### Start KCoral on the GPU server @@ -65,8 +72,8 @@ network; never expose the server to the public internet. ``` ```bash -python -m pip install --group benchmark 'kcoral[server]' -python -m kcoral --gpus 0 --host 0.0.0.0 --port 8000 +python -m pip install --group server +python -m kcoral server --gpus 0 --host 0.0.0.0 --port 8000 ``` Leave this process running. If both roles use the same machine, open another @@ -91,9 +98,11 @@ python evolution/setup.py --task kda_forward_b1_t8192_h96 \ --remote "$KCORAL_URL" ``` -Setup creates a worktree with the required environment, skills, and task -prompt. The run directory contains `PROMPT.md`, `manifest.json`, and -`worktree/`; candidate kernels will live under +Setup creates a worktree and `.venv` using the Python interpreter that launched +it, then installs the harness and benchmark dependencies from `uv.lock` and +prepares the skills, references, and task prompt. This also applies when the +GPU is remote. The run directory contains `PROMPT.md`, `manifest.json`, +`flowverse.yaml`, and `worktree/`; candidate kernels will live under `candidates/kda/forward_b1_t8192_h96/` in that worktree. Set `run_dir` to the absolute path printed by setup, enter the worktree, and @@ -166,7 +175,8 @@ The 3× speedup target is an example; choose a target that fits your task. :::{container} launch-panel :name: flame-chase -The harness environment includes Humanize. Run +Install Humanize separately using its +[installation instructions](https://github.com/humanfia/humanize#install), then run [Flame Chase](https://docs.humanfia.ai/humanize/flows/flame-chase) in the prepared terminal: From e305ed94bb7fe3e5c34c3ecdee47292d685e560b Mon Sep 17 00:00:00 2001 From: Hongyi Jin Date: Tue, 29 Sep 2026 13:32:51 -0400 Subject: [PATCH 2/7] docs: preserve the recorded benchmark trace excerpts --- chapters/benchmark-server-deepdive.md | 26 +++++++------------------- 1 file changed, 7 insertions(+), 19 deletions(-) diff --git a/chapters/benchmark-server-deepdive.md b/chapters/benchmark-server-deepdive.md index 1311e2c..81891a0 100644 --- a/chapters/benchmark-server-deepdive.md +++ b/chapters/benchmark-server-deepdive.md @@ -165,25 +165,16 @@ efficiently, changes the kernel, and benchmarks the result. A second profile helps it check whether the change had the intended effect. The kernel is a single fused kernel with one CTA per head, in which eight -preprocessing warps prepare each chunk for the MMA and state warps. The commands -below use paths relative to the generated worktree root. Set the workload path -there before running them: - -```bash -WORKLOAD=candidates/kda/forward_b1_t8192_h96 -``` - -Capture and analysis scripts such as `capture_iket.py` and `iket_analyze.py` -were written by the agent during this run. Output excerpts keep only the key -lines; `...` marks omitted text. +preprocessing warps prepare each chunk for the MMA and state warps. `$WORKLOAD` stands for the +task's workload directory. Output excerpts keep only the key lines; `...` +marks omitted text. At one point, the best kernel runs at 0.497283 ms, 2.1162× over the baseline. The agent captures an IKET timeline of it through KCoral and summarizes the time each preprocessing warp spends in each stage: ```bash -python evolution/remote/kcoral_iket.py --remote http://10.0.2.2:8901 \ - --send "$WORKLOAD/scratch/product-gamma" \ +python kcoral_iket.py --remote http://10.0.2.2:8901 --send scratch/product-gamma \ --output-dir iket_out -- profile --postprocess json -- python capture_iket.py python iket_analyze.py iket_out/iket_pid_0x8e.trace.json 13 ``` @@ -247,8 +238,7 @@ moves the diagonal inversions into the Gram tail: The CPU checks pass, and the official benchmark measures a new best: ```bash -python evolution/remote/kcoral_remote.py "$WORKLOAD" scratch/overlap-diag-inverse-gram-tail \ - --remote http://10.0.2.2:8901 --timeout 600 +python kcoral_remote.py $WORKLOAD scratch/overlap-diag-inverse-gram-tail --remote http://10.0.2.2:8901 --timeout 600 ``` ```text @@ -271,8 +261,7 @@ the agent profiles one launch of it with NCU through KCoral and reads the report's source counters locally: ```bash -python evolution/remote/kcoral_ncu.py --remote http://10.0.2.2:8901 \ - --send "$WORKLOAD/scratch/profile-qacc-ring" \ +python kcoral_ncu.py --remote http://10.0.2.2:8901 --send scratch/profile-qacc-ring \ -o kda-full.ncu-rep --set full --launch-count 1 --kernel-name kda_fwd_kernel -- python capture_ncu.py ncu --import kda-full.ncu-rep --page details --section SourceCounters ``` @@ -337,8 +326,7 @@ picks out its own head in shared memory: The CPU checks pass, and the official benchmark measures a new best: ```bash -python evolution/remote/kcoral_remote.py "$WORKLOAD" scratch/beta-tma \ - --remote http://10.0.2.2:8901 --timeout 600 +python kcoral_remote.py $WORKLOAD scratch/beta-tma --remote http://10.0.2.2:8901 --timeout 600 ``` ```text From 7f35592cb6a3e1070f953a342804e3212ed1d9e0 Mon Sep 17 00:00:00 2001 From: Hongyi Jin Date: Tue, 29 Sep 2026 13:33:53 -0400 Subject: [PATCH 3/7] docs: shorten launch prerequisites --- chapters/launching-the-agent.md | 13 ++++--------- 1 file changed, 4 insertions(+), 9 deletions(-) diff --git a/chapters/launching-the-agent.md b/chapters/launching-the-agent.md index 1f97e68..bfc9eaa 100644 --- a/chapters/launching-the-agent.md +++ b/chapters/launching-the-agent.md @@ -35,15 +35,10 @@ collect measurements. The agent and KCoral can also run on the same machine. ### Prepare the harness checkout -This example follows the harness's -{harness}`registered-workload setup `, which builds -the harness in each run's environment. Use Linux x86_64 with Python 3.12 or -3.13 and pip 25.1 or later. The agent machine also needs `uv`, Rust 1.89.0 or -later with Cargo, and the C/C++ build tools and Python development headers -listed in the {harness}`source-build prerequisites `. -Install the CUDA Toolkit wherever kernels are compiled, and a compatible -NVIDIA driver on the GPU server. NCU captures need Nsight Compute on the -server; reading the returned reports also needs `ncu` on the agent machine. +Use Linux x86_64, Python 3.12 or 3.13, and pip 25.1+ on both machines. +The agent machine needs `uv`, Rust 1.89+ with Cargo, C/C++ build tools, and +Python development headers; the GPU server needs CUDA and a compatible driver. +For profiling, install Nsight Compute on both machines. Clone the harness where the agent runs and on the GPU server. If they share one machine, a single checkout is enough: From 6d7a36c52acfc947e5556babf1108dd043fb47a2 Mon Sep 17 00:00:00 2001 From: Hongyi Jin Date: Tue, 29 Sep 2026 13:34:43 -0400 Subject: [PATCH 4/7] docs: explain run preparation before setup commands --- chapters/launching-the-agent.md | 16 ++++++++++------ 1 file changed, 10 insertions(+), 6 deletions(-) diff --git a/chapters/launching-the-agent.md b/chapters/launching-the-agent.md index bfc9eaa..9c91e0e 100644 --- a/chapters/launching-the-agent.md +++ b/chapters/launching-the-agent.md @@ -33,6 +33,11 @@ Kimi Delta Attention forward on a B200 with `B=1`, agent, and the other is a GPU server running KCoral to execute kernels and collect measurements. The agent and KCoral can also run on the same machine. +First, we will clone the harness and start KCoral on the GPU server. On the +agent machine, `evolution/setup.py` then creates a task worktree and Python +environment with the required packages, skills, and prompt. We will check the +baseline through KCoral before launching the agent in that worktree. + ### Prepare the harness checkout Use Linux x86_64, Python 3.12 or 3.13, and pip 25.1+ on both machines. @@ -49,8 +54,8 @@ cd TIRx-harness git submodule update --init thirdparty/tvm-rust-ext ``` -On the agent machine, install the setup dependencies in the Python environment -you will use to launch setup, from the repository root: +From the checkout root on the agent machine, install the dependencies for +`evolution/setup.py`: ```bash python -m pip install -r evolution/preparation/requirements.txt @@ -93,10 +98,9 @@ python evolution/setup.py --task kda_forward_b1_t8192_h96 \ --remote "$KCORAL_URL" ``` -Setup creates a worktree and `.venv` using the Python interpreter that launched -it, then installs the harness and benchmark dependencies from `uv.lock` and -prepares the skills, references, and task prompt. This also applies when the -GPU is remote. The run directory contains `PROMPT.md`, `manifest.json`, +The run's `.venv` uses the Python interpreter that launched setup, with harness +and benchmark dependencies installed from `uv.lock`. The run directory +contains `PROMPT.md`, `manifest.json`, `flowverse.yaml`, and `worktree/`; candidate kernels will live under `candidates/kda/forward_b1_t8192_h96/` in that worktree. From 8c15c78ad615dff316a95d14ecf78d7b1558223d Mon Sep 17 00:00:00 2001 From: Hongyi Jin Date: Tue, 29 Sep 2026 13:36:02 -0400 Subject: [PATCH 5/7] docs: omit dependency lock detail from launch guide --- chapters/launching-the-agent.md | 7 +++---- 1 file changed, 3 insertions(+), 4 deletions(-) diff --git a/chapters/launching-the-agent.md b/chapters/launching-the-agent.md index 9c91e0e..cb0faff 100644 --- a/chapters/launching-the-agent.md +++ b/chapters/launching-the-agent.md @@ -98,10 +98,9 @@ python evolution/setup.py --task kda_forward_b1_t8192_h96 \ --remote "$KCORAL_URL" ``` -The run's `.venv` uses the Python interpreter that launched setup, with harness -and benchmark dependencies installed from `uv.lock`. The run directory -contains `PROMPT.md`, `manifest.json`, -`flowverse.yaml`, and `worktree/`; candidate kernels will live under +The run's `.venv` uses the Python interpreter that launched setup. The run +directory contains `PROMPT.md`, `manifest.json`, `flowverse.yaml`, and +`worktree/`; candidate kernels will live under `candidates/kda/forward_b1_t8192_h96/` in that worktree. Set `run_dir` to the absolute path printed by setup, enter the worktree, and From b5962c0e5d0b558f890a9df27c5671508b88626d Mon Sep 17 00:00:00 2001 From: Hongyi Jin Date: Tue, 29 Sep 2026 13:37:31 -0400 Subject: [PATCH 6/7] docs: initialize submodules during clone --- chapters/launching-the-agent.md | 3 +-- 1 file changed, 1 insertion(+), 2 deletions(-) diff --git a/chapters/launching-the-agent.md b/chapters/launching-the-agent.md index cb0faff..bc9e5e0 100644 --- a/chapters/launching-the-agent.md +++ b/chapters/launching-the-agent.md @@ -49,9 +49,8 @@ Clone the harness where the agent runs and on the GPU server. If they share one machine, a single checkout is enough: ```bash -git clone https://github.com/mlc-ai/TIRx-harness.git +git clone --recursive https://github.com/mlc-ai/TIRx-harness.git cd TIRx-harness -git submodule update --init thirdparty/tvm-rust-ext ``` From the checkout root on the agent machine, install the dependencies for From 0f9eb88ab0de79be363dc89a1d198c75b168797b Mon Sep 17 00:00:00 2001 From: Hongyi Jin Date: Tue, 29 Sep 2026 13:38:11 -0400 Subject: [PATCH 7/7] docs: mention the installed harness wheel in the run environment --- chapters/launching-the-agent.md | 5 +++-- 1 file changed, 3 insertions(+), 2 deletions(-) diff --git a/chapters/launching-the-agent.md b/chapters/launching-the-agent.md index bc9e5e0..753aa7a 100644 --- a/chapters/launching-the-agent.md +++ b/chapters/launching-the-agent.md @@ -97,8 +97,9 @@ python evolution/setup.py --task kda_forward_b1_t8192_h96 \ --remote "$KCORAL_URL" ``` -The run's `.venv` uses the Python interpreter that launched setup. The run -directory contains `PROMPT.md`, `manifest.json`, `flowverse.yaml`, and +The run's `.venv` uses the Python interpreter that launched setup and includes +the installed `tirx-harness` wheel. The run directory contains `PROMPT.md`, +`manifest.json`, `flowverse.yaml`, and `worktree/`; candidate kernels will live under `candidates/kda/forward_b1_t8192_h96/` in that worktree.