Complexity Framework is the PyTorch research and training stack behind the released TR-HASH MoE 200M language-model lineage and separate experimental multimodal systems.
- TR-HASH MoE 200M release — architecture, checkpoints, metrics, evaluation protocol, and limitations.
- Getting started — install, construct, load, and test a model.
- Architecture and naming — TR-GQA, TR-MHA, and TR-MoE.
- TR-Hash execution engine — routes, widths, backends, and CUDA Graph constraints.
- Training — base pretraining, refinement, full SFT, evaluation, export, and resume boundaries.
- GPU and dispatch paths — PyTorch fallback, Triton/CGGR, Liger, ROCm, and reporting.
- API reference — public Python and CLI surfaces.
- Released clean SFT v2 — audited 300K mixture, reusable 32K token shards, completed training, and epoch metrics.
All non-Vision release recipes follow pretraining -> same-corpus refinement -> SFT. Vision is the documented exception because its clean-image
refinement is already integrated into the detector recipe.
130B replay-scheduled base pretraining
↓ fresh optimizer, weights only
32.07B unique-token full-parameter refinement (stopped at step 8,156)
↓ full checkpoint weights
3 epochs audited 300K full-parameter instruction SFT v2
↓ PIQA selection
epoch 3 / step 5,982 copied to the root F32 SafeTensors release
↓
TR-Hash-i64 OpenAI-compatible serving
The source-token lineage is approximately 162.07B exposures. SFT v2 then uses 202,948,693 tokenized training tokens per epoch, for 608,846,079 token exposures over three epochs. The SFT is full parameter, not LoRA. See the release page before quoting any metric.
| Area | Status | Entry point |
|---|---|---|
| 200M release and metrics | current | Release |
| Architecture and TR-Hash runtime | current | Architecture, engine |
| 200M pretraining, refinement, full SFT | current | Training, streaming data |
| 200M clean SFT v2 | released; epoch 3 at repository root | SFT v2 |
| CUDA, Triton, Liger, fallback | current | GPU and dispatch |
| Public Python surface | current | API reference |
| TokenRoutedMLP conversion | compatibility only | Migration |
| o200k Dense/TR comparison plans | historical, non-runnable | Run configurations, B200 runbook |
| Multimodal and generative modules | experimental | Multimodal index |
| Vision detector | separate released lineage | Object detection |
Historical documents preserve evidence provenance. Their commands must not be treated as supported entry points unless a current page explicitly says so.
| Public name | Attention | Feed-forward path |
|---|---|---|
| TR-GQA | grouped-query attention | TR-MoE |
| TR-MHA | multi-head attention | TR-MoE |
| TR-MoE | attention-independent | shared SwiGLU plus fixed token-ID experts |
The registry values tr_mha and tr_mha_v2 are experimental token-routed
residual adapters inside attention. They are not synonyms for the released
GQA + TR-MoE architecture.
- Custom models and registries
- Efficient training
- Run configurations and planners
- Historical TokenRoutedMLP migration
- Hugging Face organization card
- Use of generative AI tools
- Multimodal prototypes
- TR-Hash image editor
- TR-Hash image-text-to-text
- TR-Hash text-to-image
- TR-Hash object detection and serving
- TR-Hash Vision ONNX deployment
- Vision v8 COCO accuracy gates
- Vision v8 ONNX quantization
- Detector specialization and ablations
- TR-Hash sensor fusion
- Vision dependency stack
A configuration or implementation is not a completed experiment. Documents use these labels:
- current: matches the released code and artifacts;
- compatibility: supported only for migration or conversion;
- historical: preserves an earlier experiment but is not the default path;
- experimental: implemented without release-level evidence;
- planned: a protocol or launch shape without completed metrics.
Every numerical claim should identify the checkpoint, data or token exposure, evaluation split, runtime, and artifact that supports it.