Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -68,7 +68,7 @@ implementation of the same op on that workload.
| Device time | Compare `device_busy_ms`, never wall-clock span |
| Two questions | `Ratio`: is another kernel faster. `SOL`: how much faster the hardware allows anyone to go, its binding resource (`mem`/`comp`/`lat`) in a `Bound` column. Import the SOL arithmetic and thresholds from the checkout's roofline tool (M5); never re-derive them here |
| Which page an op lands on | The manifest entry's `family:`, through `_MANIFEST_FAMILY` — an op TileOPs adds needs no change here. One the manifest does not declare falls back to its package, then to keywords |
| Page order | `DATA_PAGES`: Elementwise, RoPE, Reduction, Normalization, Conv & Pool, GEMM, Quantization & Dequantization, Attention, MoE, Sampling, Linear Attention, SSM, Other. One page per family except `Conv & Pool` (two) and `Other` (FFT, mHC, Engram, the rest). `TopkSelectorFwdOp` declares `family: attention`, so its row is on Attention, while the API Reference documents it on the Sampling page. The API Reference nav follows it, with FFT, mHC and Engram after SSM and Top-k on the Sampling page. `_BENCH_ORDER` in `hooks.py` repeats it — change all three together |
| Page order | `DATA_PAGES`: Elementwise, RoPE, Reduction, Normalization, Conv & Pool, GEMM, Quantization & Dequantization, Attention, MoE, Sampling, Linear Attention, SSM, Other. One page per family except `Conv & Pool` (two) and `Other` (FFT, mHC, Engram, the rest). `TopKSelectFwdOp` declares `family: attention`, so its row is on Attention, while the API Reference documents it on the Sampling page. The API Reference nav follows it, with FFT, mHC and Engram after SSM and Top-k on the Sampling page. `_BENCH_ORDER` in `hooks.py` repeats it — change all three together |
| Op order within a page | The order `docs/api/` names them, read by `api_op_order()`. An op no API page names comes last, ranked by verdict |
| Rows follow the manifest | One row group per manifest label, one row per dtype under it in a `dtype` column. Labels keep the snapshot's order, which is the manifest's; the key above the table repeats it. A row no manifest describes takes its id, trailing dtype names split off, as its label |
| Workload shapes | The snapshot names a workload but carries no shapes. `scripts/workload_shape.py` reads them from the spec manifest at the commit the benchmark ran on, joined by the `<label>-<dtype>` the benchmark id is built from. A workload the manifest does not declare keeps its id and gets no shapes — never a guessed one |
Expand Down
4 changes: 2 additions & 2 deletions docs/api/attention.md
Original file line number Diff line number Diff line change
Expand Up @@ -55,13 +55,13 @@ what runs when you call `op(...)`.

## Native sparse attention

::: tileops.attention.NSACmpVarlenFwdOp
::: tileops.attention.NSACompressedVarlenFwdOp
options:
show_root_heading: true
heading_level: 3
members: ["__init__", "forward"]

::: tileops.attention.NSATopkVarlenFwdOp
::: tileops.attention.NSATopKVarlenFwdOp
options:
show_root_heading: true
heading_level: 3
Expand Down
6 changes: 3 additions & 3 deletions docs/api/elementwise.md
Original file line number Diff line number Diff line change
Expand Up @@ -61,7 +61,7 @@ what runs when you call `op(...)`.
heading_level: 3
members: ["__init__", "forward"]

::: tileops.elementwise.LerpFwdOp
::: tileops.elementwise.LerpScalarFwdOp
options:
show_root_heading: true
heading_level: 3
Expand Down Expand Up @@ -397,7 +397,7 @@ what runs when you call `op(...)`.
heading_level: 3
members: ["__init__", "forward"]

::: tileops.elementwise.ClampFwdOp
::: tileops.elementwise.ClampTensorFwdOp
options:
show_root_heading: true
heading_level: 3
Expand All @@ -409,7 +409,7 @@ what runs when you call `op(...)`.
heading_level: 3
members: ["__init__", "forward"]

::: tileops.elementwise.MaskedFillFwdOp
::: tileops.elementwise.MaskedFillTensorFwdOp
options:
show_root_heading: true
heading_level: 3
Expand Down
4 changes: 2 additions & 2 deletions docs/api/linear-algebra.md
Original file line number Diff line number Diff line change
Expand Up @@ -15,7 +15,7 @@ what runs when you call `op(...)`.
merge_init_into_class: false
show_signature_annotations: false

::: tileops.gemm.GemmFp8FwdOp
::: tileops.gemm.GemmFP8FwdOp
options:
show_root_heading: true
heading_level: 3
Expand All @@ -41,7 +41,7 @@ what runs when you call `op(...)`.
merge_init_into_class: false
show_signature_annotations: false

::: tileops.gemm.BmmFp8FwdOp
::: tileops.gemm.BmmFP8FwdOp
options:
show_root_heading: true
heading_level: 3
Expand Down
18 changes: 6 additions & 12 deletions docs/api/linear-attention.md
Original file line number Diff line number Diff line change
Expand Up @@ -7,25 +7,19 @@ what runs when you call `op(...)`.

## DeltaNet

::: tileops.linear_attention.DeltaNetAutogradFwdOp
::: tileops.linear_attention.DeltaNetChunkFwdOp
options:
show_root_heading: true
heading_level: 3
members: ["__init__", "forward"]

::: tileops.linear_attention.DeltaNetFwdOp
::: tileops.linear_attention.DeltaNetChunkBwdOp
options:
show_root_heading: true
heading_level: 3
members: ["__init__", "forward"]

::: tileops.linear_attention.DeltaNetBwdOp
options:
show_root_heading: true
heading_level: 3
members: ["__init__", "forward"]

::: tileops.linear_attention.DeltaNetDecodeFwdOp
::: tileops.linear_attention.DeltaNetRecurrentFwdOp
options:
show_root_heading: true
heading_level: 3
Expand All @@ -47,19 +41,19 @@ what runs when you call `op(...)`.

## Gated linear attention

::: tileops.linear_attention.GLAFwdOp
::: tileops.linear_attention.GLAChunkFwdOp
options:
show_root_heading: true
heading_level: 3
members: ["__init__", "forward"]

::: tileops.linear_attention.GLABwdOp
::: tileops.linear_attention.GLAChunkBwdOp
options:
show_root_heading: true
heading_level: 3
members: ["__init__", "forward"]

::: tileops.linear_attention.GLADecodeFwdOp
::: tileops.linear_attention.GLARecurrentFwdOp
options:
show_root_heading: true
heading_level: 3
Expand Down
6 changes: 3 additions & 3 deletions docs/api/mamba.md
Original file line number Diff line number Diff line change
Expand Up @@ -15,13 +15,13 @@ what runs when you call `op(...)`.

## SSD stages

::: tileops.mamba.DaCumsumFwdOp
::: tileops.mamba.SSDChunkCumsumFwdOp
options:
show_root_heading: true
heading_level: 3
members: ["__init__", "forward"]

::: tileops.mamba.CBProducerFwdOp
::: tileops.mamba.SSDChunkCouplingFwdOp
options:
show_root_heading: true
heading_level: 3
Expand All @@ -47,7 +47,7 @@ what runs when you call `op(...)`.

## Decode

::: tileops.mamba.SSDDecodeFwdOp
::: tileops.mamba.SSDRecurrentFwdOp
options:
show_root_heading: true
heading_level: 3
Expand Down
22 changes: 11 additions & 11 deletions docs/api/moe.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,24 +5,24 @@ constructor takes what the kernel is compiled with; the call takes the tensors.
Both are documented under each op — `__init__` and `forward`, where `forward` is
what runs when you call `op(...)`.

A routed mixture-of-experts layer is available two ways here. `FusedMoeFwdOp` runs
the whole FFN, and `FusedMoeSharedExpertFwdOp` adds a shared expert beside the routed ones. The rest are its stages, callable on their own: the op that picks
A routed mixture-of-experts layer is available two ways here. `FusedMoEFwdOp` runs
the whole FFN, and `FusedMoESharedExpertFwdOp` adds a shared expert beside the routed ones. The rest are its stages, callable on their own: the op that picks
each token's experts, the ops that move tokens into an expert-contiguous layout and
back, and the expert GEMMs that run on it. `MoeGroupedGemmFwdOp` is one grouped
GEMM; `MoeExpertMLPFwdOp` is the pair of them with the gated activation fused into
back, and the expert GEMMs that run on it. `MoEGroupedGemmFwdOp` is one grouped
GEMM; `MoEExpertMLPFwdOp` is the pair of them with the gated activation fused into
the first; `FusedMoEExpertsFwdOp` is that MLP with the permutes around it, on the
tight (no-pad) layout, and `IndexedExpertMLPFwdOp` is the backend it picks instead when
the routes are few enough to read the weights once per route rather than once per expert. `SharedExpertMLPFwdOp` is the dense gated MLP of the shared expert. The routing has to produce the layout the GEMM expects.

## Fused forward

::: tileops.moe.FusedMoeFwdOp
::: tileops.moe.FusedMoEFwdOp
options:
show_root_heading: true
heading_level: 3
members: ["__init__", "forward"]

::: tileops.moe.FusedMoeSharedExpertFwdOp
::: tileops.moe.FusedMoESharedExpertFwdOp
options:
show_root_heading: true
heading_level: 3
Expand All @@ -36,33 +36,33 @@ the routes are few enough to read the weights once per route rather than once pe
heading_level: 3
members: ["__init__", "forward"]

::: tileops.moe.MoePrePermuteFwdOp
::: tileops.moe.MoEPrePermuteFwdOp
options:
show_root_heading: true
heading_level: 3
members: ["__init__", "forward"]

::: tileops.moe.MoePermuteAlignFwdOp
::: tileops.moe.MoEPermuteAlignFwdOp
options:
show_root_heading: true
heading_level: 3
members: ["__init__", "forward"]

::: tileops.moe.MoePostPermuteFwdOp
::: tileops.moe.MoEPostPermuteFwdOp
options:
show_root_heading: true
heading_level: 3
members: ["__init__", "forward"]

## Expert GEMMs

::: tileops.moe.MoeGroupedGemmFwdOp
::: tileops.moe.MoEGroupedGemmFwdOp
options:
show_root_heading: true
heading_level: 3
members: ["__init__", "forward"]

::: tileops.moe.MoeExpertMLPFwdOp
::: tileops.moe.MoEExpertMLPFwdOp
options:
show_root_heading: true
heading_level: 3
Expand Down
15 changes: 3 additions & 12 deletions docs/api/reduction.md
Original file line number Diff line number Diff line change
Expand Up @@ -93,19 +93,10 @@ what runs when you call `op(...)`.

## Vector norms

::: tileops.reduction.L1NormFwdOp
options:
show_root_heading: true
heading_level: 3
members: ["__init__", "forward"]

::: tileops.reduction.L2NormFwdOp
options:
show_root_heading: true
heading_level: 3
members: ["__init__", "forward"]
One op serves every order `torch.linalg.vector_norm` is called at here: pass
`ord` as 1, 2 or `inf`.

::: tileops.reduction.InfNormFwdOp
::: tileops.reduction.VectorNormFwdOp
options:
show_root_heading: true
heading_level: 3
Expand Down
16 changes: 6 additions & 10 deletions docs/api/rope.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,23 +5,19 @@ constructor takes what the kernel is compiled with; the call takes the tensors.
Both are documented under each op — `__init__` and `forward`, where `forward` is
what runs when you call `op(...)`.

## NeoX layout
## Base frequencies

::: tileops.rope.RopeNeoxFwdOp
options:
show_root_heading: true
heading_level: 3
members: ["__init__", "forward"]
One op serves both rotation conventions: pass `rope_layout` as `"neox"` to
rotate the two halves of a head against each other, or `"interleaved"` to rotate
each adjacent pair.

::: tileops.rope.RopeNeoxPositionIdsFwdOp
::: tileops.rope.RopeFwdOp
options:
show_root_heading: true
heading_level: 3
members: ["__init__", "forward"]

## Interleaved layout

::: tileops.rope.RopeNonNeoxFwdOp
::: tileops.rope.RopeNeoxPositionIdsFwdOp
options:
show_root_heading: true
heading_level: 3
Expand Down
2 changes: 1 addition & 1 deletion docs/api/sampling.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,7 +12,7 @@ and offset as tensors so a draw is reproducible.

## Top-k selection

::: tileops.attention.TopkSelectorFwdOp
::: tileops.attention.TopKSelectFwdOp
options:
show_root_heading: true
heading_level: 3
Expand Down
17 changes: 17 additions & 0 deletions docs/api/transform.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,17 @@
# Transform Operators

An orthogonal transform along one axis. A transform rotates a tensor before it is
quantized; the quantization that follows one is on the
[Quantization](quantization.md) page.

## Walsh-Hadamard

`HadamardTransformFwdOp` is specified and not yet implemented, so it has no
constructor to document here. The signature it will serve is in
`src/tileops/manifest/spec/transform.yaml`: a fast Walsh-Hadamard transform along
the last axis scaled by $1/\sqrt{n}$, where $n$ is `base_order` times a power of
two, in float16 or bfloat16 with the butterfly accumulated in float32. It is the
rotation step of rotation-based quantization, which spreads outliers across
channels so a lower bit width stays accurate.

This page gains the op's `__init__` and `forward` when a kernel serves it.
6 changes: 3 additions & 3 deletions docs/user-guide/manifest/examples.md
Original file line number Diff line number Diff line change
Expand Up @@ -48,13 +48,13 @@ GemmFwdOp:
| 4 | workload rows give the values of parameters and indices; the first row expands into two cases, with ids `square-1k-float16` and `square-1k-bfloat16` | [Spec fields § 8](writing.md#workloads) |
| 5 | `roofline` gives only `flops`; `bytes` is derived from the signature and equals `(M*K + K*N + M*N)` times the bytes per element | [Spec fields § 9](writing.md#roofline) |

## 2. MoePrePermuteFwdOp: ADT, let and generator {#moe}
## 2. MoEPrePermuteFwdOp: ADT, let and generator {#moe}

The type of the `layout` parameter is the ADT `MGroupedLayout` defined in `spec/types.yaml`; its definition is in [Extensions § 3](extensions.md#adt).

```yaml
# src/tileops/manifest/spec/moe.yaml
MoePrePermuteFwdOp:
MoEPrePermuteFwdOp:
family: moe
status: implemented
signature:
Expand Down Expand Up @@ -90,7 +90,7 @@ MoePrePermuteFwdOp:
flops: "0"
```

**Table 2** The forms used in `MoePrePermuteFwdOp`
**Table 2** The forms used in `MoEPrePermuteFwdOp`

| No. | Form in the spec | Described in |
| --- | --- | --- |
Expand Down
6 changes: 3 additions & 3 deletions docs/user-guide/manifest/examples.zh.md
Original file line number Diff line number Diff line change
Expand Up @@ -48,13 +48,13 @@ GemmFwdOp:
| 4 | workload 行给出参数与 index 的取值;第一行展开为两个 case,id 分别是 `square-1k-float16` 与 `square-1k-bfloat16` | [写一个 spec 8](writing.md#workloads) |
| 5 | `roofline` 只给出 `flops`;`bytes` 由签名推导,等于 `(M*K + K*N + M*N)` 乘以每个元素的字节数 | [写一个 spec 9](writing.md#roofline) |

## 2. MoePrePermuteFwdOp:ADT、let 与 generator {#moe}
## 2. MoEPrePermuteFwdOp:ADT、let 与 generator {#moe}

`layout` 参数的类型是 `spec/types.yaml` 中定义的 ADT `MGroupedLayout`,定义见[扩展写法 3](extensions.md#adt)。

```yaml
# src/tileops/manifest/spec/moe.yaml
MoePrePermuteFwdOp:
MoEPrePermuteFwdOp:
family: moe
status: implemented
signature:
Expand Down Expand Up @@ -90,7 +90,7 @@ MoePrePermuteFwdOp:
flops: "0"
```

**表 2** `MoePrePermuteFwdOp` 中各处写法的说明
**表 2** `MoEPrePermuteFwdOp` 中各处写法的说明

| No. | spec 中的写法 | 说明所在 |
| --- | --- | --- |
Expand Down
1 change: 1 addition & 0 deletions mkdocs.yml
Original file line number Diff line number Diff line change
Expand Up @@ -77,6 +77,7 @@ nav:
- Linear Attention: api/linear-attention.md
- Mamba: api/mamba.md
- FFT: api/fft.md
- Transform: api/transform.md
- mHC: api/mhc.md
- Engram: api/engram.md
- Trace: api/trace.md
Expand Down
3 changes: 3 additions & 0 deletions scripts/gen_bench_pages.py
Original file line number Diff line number Diff line change
Expand Up @@ -196,6 +196,9 @@ def api_op_order(api_dir: str = API_DIR,
"quantization": "quantization", "sampling": "sampling", "mamba": "ssm",
"rope": "positional", "gemm": "linear_algebra",
"convolution": "convolution", "fft": "fft",
# One op so far, and a rotation is not a quantization step: it publishes on
# Other rather than carrying a page of its own.
"transform": "other",
}
# For an op the manifest does not declare: its package, then the keywords above.
# Every package that maps to a family belongs here — `linear_attention` left out
Expand Down
Loading