Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
34 changes: 27 additions & 7 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -68,7 +68,7 @@ implementation of the same op on that workload.
| Device time | Compare `device_busy_ms`, never wall-clock span |
| Two questions | `Ratio`: is another kernel faster. `SOL`: how much faster the hardware allows anyone to go, its binding resource (`mem`/`comp`/`lat`) in a `Bound` column. Import the SOL arithmetic and thresholds from the checkout's roofline tool (M5); never re-derive them here |
| Which page an op lands on | The manifest entry's `family:`, through `_MANIFEST_FAMILY` — an op TileOPs adds needs no change here. One the manifest does not declare falls back to its package, then to keywords |
| Page order | `DATA_PAGES` follows the API nav's order over the same families: one page per family except `Conv & Pool` (two) and `Other` (Top-k, FFT, mHC, Engram, the rest). `_BENCH_ORDER` in `hooks.py` repeats it — change one, change the other |
| Page order | `DATA_PAGES`: Elementwise, RoPE, Reduction, Normalization, Conv & Pool, GEMM, Quantization & Dequantization, Attention, MoE, Sampling, Linear Attention, SSM, Other. One page per family except `Conv & Pool` (two) and `Other` (FFT, mHC, Engram, the rest). `TopkSelectorFwdOp` declares `family: attention`, so its row is on Attention, while the API Reference documents it on the Sampling page. The API Reference nav follows it, with FFT, mHC and Engram after SSM and Top-k on the Sampling page. `_BENCH_ORDER` in `hooks.py` repeats it — change all three together |
| Op order within a page | The order `docs/api/` names them, read by `api_op_order()`. An op no API page names comes last, ranked by verdict |
| Rows follow the manifest | One row group per manifest label, one row per dtype under it in a `dtype` column. Labels keep the snapshot's order, which is the manifest's; the key above the table repeats it. A row no manifest describes takes its id, trailing dtype names split off, as its label |
| Workload shapes | The snapshot names a workload but carries no shapes. `scripts/workload_shape.py` reads them from the spec manifest at the commit the benchmark ran on, joined by the `<label>-<dtype>` the benchmark id is built from. A workload the manifest does not declare keeps its id and gets no shapes — never a guessed one |
Expand All @@ -78,14 +78,16 @@ implementation of the same op on that workload.
English at the site root, Chinese under `/zh/`. A Chinese page is a
`<name>.zh.md` beside the English `<name>.md`, full prose, never an
`include-markdown` shell. `backends.md`, `torch-compile.md`, everything under
`performance-guides/memory-bound/` and the two guides under `user-guide/manifest/`
and `user-guide/dispatch/` were authored in Chinese: edit the `.zh.md` first, then
bring the English page in line. Everything else goes the other way.
`performance-guides/memory-bound/`, `blog/`, `user-guide/development.md` and the
two guides under `user-guide/manifest/` and `user-guide/dispatch/` were authored
in Chinese: edit the `.zh.md` first, then bring the English page in line.
Everything else goes the other way.

| Rule | Detail |
|------|--------|
| Coverage | Whichever pages have a `.zh.md` — `ls docs/**/*.zh.md` |
| Never translate | `api/` and `benchmarks/`, both generated; `design/`, mirrored English |
| Mirrored, translated | `performance-guides/trace-timeline.md` mirrors TileOPs `docs/perf/trace-timeline.md`; its `.zh.md` is a full translation. When upstream changes that file, update the translation in step |
| Missing translation | Falls back to English at the same URL, and `hooks.py` prepends a "本页暂无中文版" notice. The fallback runs zh → en only: a page that exists only as `.zh.md` leaves its `nav` entry on a missing file and the English sidebar renders a dead link |
| Figures | A figure with text needs one SVG per language: `img/<name>.svg` for English and `img/<name>.zh.svg` beside it, which the `zh` build picks up for the same reference. Translate the `<text>` nodes and the `aria-label`, keep the geometry. English runs longer than Chinese — grow the `viewBox` rather than let text overflow. The user-guide figures are drawn from sources under `figures/user-guide/` (`<name>.zh.puml`, `<name>.en.puml`, `manifest/overview.py`): edit the source, then run `figures/user-guide/render.sh` |
| Nav labels | `nav_translations` in the `i18n` plugin block; keep an entry for every `nav` title |
Expand All @@ -99,23 +101,41 @@ bring the English page in line. Everything else goes the other way.

## Nav

`nav` in `mkdocs.yml` is the page list; its six sections run in the reader's
order, Design last as contributor-facing.
`nav` in `mkdocs.yml` is the page list: Home, then Blog, then the remaining
sections in the reader's order, Design last as contributor-facing.

- Add a new page to `nav`, and its label to `nav_translations`.
- Put a user-facing topic under User Guide.
- Put a user-facing topic under User Guide, in the group its index lists it in,
and keep the nav in the index's order. The index is a plain list per group;
`hooks.py` marks it and `extra.css` draws each item as a card.
- Keep a label short enough for one line in the sidebar; the page's H1 carries
the full title.
- Merging a page into another: delete it, and add its old URL to `_REDIRECTS`
in `hooks.py` so published links redirect.
- List a section's `<dir>/index.md` as a bare path with no title. Given a title
it is promoted anyway and its sidebar row disappears.
- Leave `toc.integrate` off: the page TOC renders in the right column, and it is
incompatible with `navigation.indexes`.

## Blog

`docs/blog/`: plain pages, not Material's `blog` plugin. Why: the plugin
renders no post under `mkdocs-static-i18n` and warns on its archive pages.

- One post is `blog/<slug>.zh.md` and `blog/<slug>.md`, listed in `nav` under
Blog and as one link on `blog/index.md`, newest first.
- A post opens with a short H1 and a one-line subtitle paragraph. `hooks.py`
gives that paragraph the `post-subtitle` class that `extra.css` styles; keep
classes and attribute lists out of the post's Markdown.
- No date line, no in-page TOC: the right column carries the TOC.

## Conventions

- Measure every number and state its conditions. Say when a count will drift.
- Admonitions (`!!! note`, `!!! warning`) for callouts; relative Markdown links
for internal cross-references.
- Write links as plain Markdown. `extra.css` gives every link in running text an
arrow, east within the site and north-east off it, and a teal wash on hover.
- Link to the TileOPs repo rather than duplicating it. A page authored here that
mirrors upstream content will drift.
- Gitignored: `site/`, `__pycache__/`, `.cache/`, `TileOPs/`, and
Expand Down
25 changes: 12 additions & 13 deletions docs/api/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,30 +13,29 @@ op = GemmFwdOp() # construct once, reuse
d = op(a, b) # the specialized kernel is built on first call
```

The pages are ordered by how much an op composes: the pointwise transforms first, then
the axis reductions and the normalizations built on them, then the matmul and the
expert routing over it, then the windowed and spectral transforms, then the
sequence-model kernels built on all of the above. It is the order `tileops` declares
its op families in. The exception is Top-k, whose one op is exported from
`tileops.attention`.
The pages run in the order a reader meets the ops in a model: the pointwise
transforms and the positional rotation, then the axis reductions and the
normalizations built on them, the windowed kernels, the matmul and the
quantization around it, then attention, the expert routing after it and sampling,
then the sequence-model kernels. FFT, mHC and Engram follow, and Trace,
a tool rather than an op, comes last. The Benchmarks pages use the same order.

| Page | What it covers |
| --- | --- |
| [Elementwise](elementwise.md) | unary and binary maps, activations, dropout, and the in-place forms |
| [RoPE](rope.md) | rotary position embedding — NeoX and interleaved layouts, Llama 3.1, YaRN, LongRoPE |
| [Reduction](reduction.md) | sums, extrema, arg-reductions, cumulative scans, softmax |
| [Normalization](normalization.md) | RMSNorm, LayerNorm, GroupNorm, BatchNorm and the fused variants |
| [Quantization](quantization.md) | INT8, FP8 and INT4 quantization, and INT8 dequantization |
| [Top-k](topk.md) | top-k selection |
| [GEMM](linear-algebra.md) | dense matmul — plain, batched, and the fp8 variants |
| [Pooling](pool.md) | average, max and adaptive pooling, with and without indices, plus the chunked sequence mean |
| [Convolution](convolution.md) | forward convolution over 1D, 2D and 3D inputs |
| [FFT](fft.md) | the discrete transform |
| [MoE](moe.md) | the routed mixture-of-experts FFN and its separately callable stages |
| [Sampling](sampling.md) | logits masks (top-k, top-p, min-p) and sampling, including chain speculative sampling |
| [RoPE](rope.md) | rotary position embedding — NeoX and interleaved layouts, Llama 3.1, YaRN, LongRoPE |
| [GEMM](linear-algebra.md) | dense matmul — plain, batched, and the fp8 variants |
| [Quantization & Dequantization](quantization.md) | INT8, FP8 and INT4 quantization, and INT8 dequantization |
| [Attention](attention.md) | forward and backward attention, including the paged and decode kernels |
| [MoE](moe.md) | the routed mixture-of-experts FFN and its separately callable stages |
| [Top-k & Sampling](sampling.md) | top-k selection, logits masks (top-k, top-p, min-p) and sampling, including chain speculative sampling |
| [Linear Attention](linear-attention.md) | DeltaNet, Gated DeltaNet and gated linear attention |
| [Mamba](mamba.md) | the SSD scan, its decode step, and the chunked forms |
| [FFT](fft.md) | the discrete transform |
| [mHC](mhc.md) | Manifold-Constrained Hyper-Connections — the pre/post pair around a layer |
| [Engram](engram.md) | the Engram GateConv pair and its decode step |
| [Trace](trace.md) | the in-kernel timeline tracer, a tool rather than an op |
Expand Down
2 changes: 1 addition & 1 deletion docs/api/quantization.md
Original file line number Diff line number Diff line change
@@ -1,4 +1,4 @@
# Quantization Operators
# Quantization and Dequantization Operators

Every op on this page is used the same way: construct it once, then call it. The
constructor takes what the kernel is compiled with; the call takes the tensors.
Expand Down
17 changes: 13 additions & 4 deletions docs/api/sampling.md
Original file line number Diff line number Diff line change
@@ -1,13 +1,22 @@
# Sampling Operators
# Top-k and Sampling Operators

Every op on this page is used the same way: construct it once, then call it. The
constructor takes what the kernel is compiled with; the call takes the tensors.
Both are documented under each op — `__init__` and `forward`, where `forward` is
what runs when you call `op(...)`.

The masks set every logit a filter drops to `-inf`, so a softmax over the result
renormalizes over what is kept. The samplers draw from a distribution, and take the
random seed and offset as tensors so a draw is reproducible.
The top-k selector returns the indices of the largest scores in each row. The masks
set every logit a filter drops to `-inf`, so a softmax over the result renormalizes
over what is kept. The samplers draw from a distribution, and take the random seed
and offset as tensors so a draw is reproducible.

## Top-k selection

::: tileops.attention.TopkSelectorFwdOp
options:
show_root_heading: true
heading_level: 3
members: ["__init__", "forward"]

## Logits masks

Expand Down
14 changes: 0 additions & 14 deletions docs/api/topk.md

This file was deleted.

Loading
Loading