Skip to content

Latest commit

 

History

History
112 lines (91 loc) · 5.12 KB

File metadata and controls

112 lines (91 loc) · 5.12 KB

Complexity Framework documentation

Complexity Framework is the PyTorch research and training stack behind the released TR-HASH MoE 200M language-model lineage and separate experimental multimodal systems.

Start here

  1. TR-HASH MoE 200M release — architecture, checkpoints, metrics, evaluation protocol, and limitations.
  2. Getting started — install, construct, load, and test a model.
  3. Architecture and naming — TR-GQA, TR-MHA, and TR-MoE.
  4. TR-Hash execution engine — routes, widths, backends, and CUDA Graph constraints.
  5. Training — base pretraining, refinement, full SFT, evaluation, export, and resume boundaries.
  6. GPU and dispatch paths — PyTorch fallback, Triton/CGGR, Liger, ROCm, and reporting.
  7. API reference — public Python and CLI surfaces.
  8. Released clean SFT v2 — audited 300K mixture, reusable 32K token shards, completed training, and epoch metrics.

All non-Vision release recipes follow pretraining -> same-corpus refinement -> SFT. Vision is the documented exception because its clean-image refinement is already integrated into the detector recipe.

Current 200M release path

130B replay-scheduled base pretraining
        ↓ fresh optimizer, weights only
32.07B unique-token full-parameter refinement (stopped at step 8,156)
        ↓ full checkpoint weights
3 epochs audited 300K full-parameter instruction SFT v2
        ↓ PIQA selection
epoch 3 / step 5,982 copied to the root F32 SafeTensors release
        ↓
TR-Hash-i64 OpenAI-compatible serving

The source-token lineage is approximately 162.07B exposures. SFT v2 then uses 202,948,693 tokenized training tokens per epoch, for 608,846,079 token exposures over three epochs. The SFT is full parameter, not LoRA. See the release page before quoting any metric.

Document status

Area Status Entry point
200M release and metrics current Release
Architecture and TR-Hash runtime current Architecture, engine
200M pretraining, refinement, full SFT current Training, streaming data
200M clean SFT v2 released; epoch 3 at repository root SFT v2
CUDA, Triton, Liger, fallback current GPU and dispatch
Public Python surface current API reference
TokenRoutedMLP conversion compatibility only Migration
o200k Dense/TR comparison plans historical, non-runnable Run configurations, B200 runbook
Multimodal and generative modules experimental Multimodal index
Vision detector separate released lineage Object detection

Historical documents preserve evidence provenance. Their commands must not be treated as supported entry points unless a current page explicitly says so.

Architecture vocabulary

Public name Attention Feed-forward path
TR-GQA grouped-query attention TR-MoE
TR-MHA multi-head attention TR-MoE
TR-MoE attention-independent shared SwiGLU plus fixed token-ID experts

The registry values tr_mha and tr_mha_v2 are experimental token-routed residual adapters inside attention. They are not synonyms for the released GQA + TR-MoE architecture.

Supporting guides

Multimodal and vision guides

Evidence policy

A configuration or implementation is not a completed experiment. Documents use these labels:

  • current: matches the released code and artifacts;
  • compatibility: supported only for migration or conversion;
  • historical: preserves an earlier experiment but is not the default path;
  • experimental: implemented without release-level evidence;
  • planned: a protocol or launch shape without completed metrics.

Every numerical claim should identify the checkpoint, data or token exposure, evaluation split, runtime, and artifact that supports it.