Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
12 changes: 7 additions & 5 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -77,21 +77,23 @@ implementation of the same op on that workload.

English at the site root, Chinese under `/zh/`. A Chinese page is a
`<name>.zh.md` beside the English `<name>.md`, full prose, never an
`include-markdown` shell. `backends.md`, `torch-compile.md` and everything under
`performance-guides/memory-bound/` were authored in Chinese: edit the `.zh.md`
first, then bring the English page in line. Everything else goes the other way.
`include-markdown` shell. `backends.md`, `torch-compile.md`, everything under
`performance-guides/memory-bound/` and the two guides under `user-guide/manifest/`
and `user-guide/dispatch/` were authored in Chinese: edit the `.zh.md` first, then
bring the English page in line. Everything else goes the other way.

| Rule | Detail |
|------|--------|
| Coverage | Whichever pages have a `.zh.md` — `ls docs/**/*.zh.md` |
| Never translate | `api/` and `benchmarks/`, both generated; `design/`, mirrored English |
| Missing translation | Falls back to English at the same URL, and `hooks.py` prepends a "本页暂无中文版" notice. The fallback runs zh → en only: a page that exists only as `.zh.md` leaves its `nav` entry on a missing file and the English sidebar renders a dead link |
| Figures | A figure with text needs one SVG per language: translate the `<text>` nodes and the `aria-label`, keep the geometry. English runs longer than Chinese — grow the `viewBox` rather than let text overflow |
| Figures | A figure with text needs one SVG per language: `img/<name>.svg` for English and `img/<name>.zh.svg` beside it, which the `zh` build picks up for the same reference. Translate the `<text>` nodes and the `aria-label`, keep the geometry. English runs longer than Chinese — grow the `viewBox` rather than let text overflow. The user-guide figures are drawn from sources under `figures/user-guide/` (`<name>.zh.puml`, `<name>.en.puml`, `manifest/overview.py`): edit the source, then run `figures/user-guide/render.sh` |
| Nav labels | `nav_translations` in the `i18n` plugin block; keep an entry for every `nav` title |
| Chinese search | Requires `jieba` |
| Punctuation | Full-width in Chinese prose: `,。:;()`. Latin quotes and brackets stay half-width inside code spans |
| Latin in Chinese | A space either side of a Latin token: `由 spec 驱动`, `形状和 dtype`. Not inside code spans |
| Keep in English | kernel, spec, agent, dtype, roofline, GEMM, target, and every op name. Why: translating them loses the link to the API |
| Keep in English | kernel, spec, agent, dtype, roofline, GEMM, target, op, family, backend, and every op name. Why: translating them loses the link to the API |
| Terms | 「kernel 接口」 for kernel interface, 「in-tree 实现」 for an in-tree implementation, 「build identity」 and 「构建函数」 for what `entry_for` returns |
| Inline code | Real identifiers only (`GemmOp`, `eval_roofline`, paths, flags). A concept mentioned in prose is not code |
| Type | `extra.css` gives `html[lang="zh"]` looser leading and headings at 700, not 800 — at 800 a CJK fallback face closes up the strokes. Scoped away from fallback pages, whose body text is English. No CJK webfont: Han glyphs come from the platform UI face (`--tf-cjk`) |

Expand Down
2 changes: 1 addition & 1 deletion docs/api/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -45,7 +45,7 @@ Two things this reference does not carry:

- **What each op is allowed to receive.** The authoritative dtype domains, shape rules
and measured workloads are in the op's spec; see [Writing a
Spec](../manifest.md).
Spec](../user-guide/manifest/index.md).
- **How fast it is.** Device time against the fastest alternative on each workload is on
the [Benchmarks](../benchmarks/index.md) pages.

Expand Down
531 changes: 241 additions & 290 deletions docs/backends.md

Large diffs are not rendered by default.

369 changes: 173 additions & 196 deletions docs/backends.zh.md

Large diffs are not rendered by default.

48 changes: 26 additions & 22 deletions docs/index.md
Original file line number Diff line number Diff line change
@@ -1,20 +1,24 @@
# TileOPs

TileOPs is an operator library for large-model inference, built on
[TileLang](https://github.com/tile-ai/tilelang), where one set of operator
interfaces can be implemented by different backends on different hardware.

What sets it apart from a hand-written library is how it is organised: every
operator is declared as a spec first, and an agent then derives the
implementation from that spec. The spec is the only input to generation and the
standard the result is judged by — correctness against the reference the spec
names, performance against the bound the roofline model gives, neither of them a
judgement call. An implementation can therefore be regenerated from its spec,
while the reverse does not hold.

To a caller it is simply a set of operators: shapes and dtype come from the call,
the specialized kernel is built and cached on first use and works under CUDA
graphs afterwards, and each op declares whether it supports
[TileLang](https://github.com/tile-ai/tilelang). One set of op interfaces can be
implemented by different backends on different hardware.

TileOPs differs from a hand-written operator library in how it is organised: every
op is first declared as a spec, and an agent then generates the implementation from
that spec. The spec is the only input to code generation and the standard the result
is accepted against:

- correctness is judged against the reference implementation the spec names;
- performance is judged against the bound the roofline model gives.

Neither check depends on human judgement. An implementation can therefore be
regenerated from its spec at any time, while a spec cannot be derived from an
implementation.

To a caller, TileOPs is a set of ops that can be called directly. Shapes and dtype
are fixed at call time; the specialized kernel is built and cached on first use and
can then be used with CUDA graphs. Each op declares whether it supports
`torch.compile(fullgraph=True)`.

## Installation
Expand All @@ -25,8 +29,8 @@ pip install tileops

## Quick Start

An op commits to nothing at construction. Shapes and dtype come from the inputs
of the call, and the specialized kernel is built and cached on first use.
An op binds no shape at construction. Shapes and dtype come from the tensors passed
to the call, and the specialized kernel is compiled and cached on the first call.

```python
import torch
Expand All @@ -42,12 +46,12 @@ flops, nbytes = op.eval_roofline() # what the call had to do and move

## Where to go next

- [User Guide](user-guide/index.md) — writing a spec, bringing an op into
`torch.compile`, how a benchmark is timed, and adding a hardware backend
- [API Reference](api/index.md) — constructor parameters and call signatures, one
page per op family
- [Benchmarks](benchmarks/index.md) — measured nightly on an H200 against the
alternatives
- [User Guide](user-guide/index.md): reading and writing the manifest, bringing an op
into `torch.compile`, how a benchmark is timed, and adding a hardware backend.
- [API Reference](api/index.md): the constructor parameters and call signatures of
each op family.
- [Benchmarks](benchmarks/index.md): measured nightly on an H200, each workload
against other implementations.

## Links

Expand Down
21 changes: 13 additions & 8 deletions docs/index.zh.md
Original file line number Diff line number Diff line change
@@ -1,10 +1,15 @@
# TileOPs

TileOPs 是一个面向大模型推理的算子库,构建在 [TileLang](https://github.com/tile-ai/tilelang) 之上,同一套算子接口可以由不同后端在不同硬件上实现。
TileOPs 是一个面向大模型推理的算子库,构建在 [TileLang](https://github.com/tile-ai/tilelang) 之上。同一套 op 接口可以由不同 backend 在不同硬件上实现。

它与手写算子库的不同之处在于组织方式:每个算子先以一份 spec 声明,再由 agent 依据这份 spec 生成实现。spec 既是代码生成的唯一依据,也是验收的标准 —— 正确性对照 spec 指定的参考实现,性能对照 roofline 模型给出的上界,两项都不依赖人的判断。因此一个实现可以随时从 spec 重新生成,而反过来做不到。
TileOPs 与手写算子库的区别在于组织方式:每个 op 先以一份 spec 声明,再由 agent 依据这份 spec 生成实现。spec 是代码生成的唯一依据,也是验收的标准:

对使用者而言,它就是一批可以直接调用的算子:形状与 dtype 在调用时确定,特化后的 kernel 在首次使用时构造并缓存,随后可以与 CUDA graph 配合使用;每个算子各自声明是否支持 `torch.compile(fullgraph=True)`。
- 正确性以 spec 指定的参考实现为准;
- 性能以 roofline 模型给出的上界为准。

两项验收都不依赖人的判断。因此一个实现可以随时从 spec 重新生成,spec 却不能从实现反推。

对使用者而言,TileOPs 提供一批可以直接调用的 op。形状与 dtype 在调用时确定;特化后的 kernel 在首次使用时构造并缓存,之后可以与 CUDA graph 配合使用。每个 op 各自声明是否支持 `torch.compile(fullgraph=True)`。

## 安装

Expand All @@ -14,7 +19,7 @@ pip install tileops

## 快速开始

算子在构造时不绑定任何形状。形状和 dtype 取自调用传入的张量,特化后的 kernel 于首次调用时编译并缓存。
op 在构造时不绑定任何形状。形状和 dtype 取自调用时传入的张量,特化后的 kernel 在首次调用时编译并缓存。

```python
import torch
Expand All @@ -28,11 +33,11 @@ d = op(a, b) # -> [M, N]
flops, nbytes = op.eval_roofline() # 本次调用所需的计算量与访存量
```

## 从这里继续
## 后续阅读

- [使用指南](user-guide/index.md) —— 读写 manifest、接入 `torch.compile`、benchmark 怎么计时、接入新硬件后端
- [API 参考](api/index.md) —— 各算子族的构造参数与调用方式
- [性能数据](benchmarks/index.md) —— 每晚在 H200 上实测,逐个 workload 与其他实现对比
- [使用指南](user-guide/index.md):读写 manifest、接入 `torch.compile`、benchmark 的计时方法、接入新硬件 backend。
- [API 参考](api/index.md):各 op family 的构造参数与调用方式。
- [性能数据](benchmarks/index.md):每晚在 H200 上实测,逐个 workload 与其他实现对比。

## 相关链接

Expand Down
Loading
Loading