Open, efficient AI research
Research and open tooling for compact language and vision models,
deterministic token routing, multimodal generation, and accessible training.
|
The current 201.2M-parameter language-model release: deterministic multi-hash top-2 routing, a completed 130B-token base run, an interrupted 32.07B-token full-parameter refinement, and a promoted full-parameter SFT assistant. The released assistant is full-parameter SFT, not LoRA. Epoch 2 was promoted at 68.82% PIQA accuracy and 69.31% normalized accuracy, and is served by TR-Hash-i64. Try the live 200M chat → |
A compact hierarchical token-routed vision detector with native multi-scale features, shifted-window attention, and NMS and NMS-free detection heads. Status: Trained on COCO 2017. Official val2017 mAP50-95: 0.20. |
Deterministic multi-hash routing supports long-horizon training in a compact language model
Boris Peyriguere · Research Square · 2026
DOI: 10.21203/rs.3.rs-10788774/v1
@article{peyriguere2026deterministic,
title={Deterministic multi-hash routing supports long-horizon training in a compact language model},
author={Boris Peyriguere},
year={2026},
publisher={Research Square},
doi={10.21203/rs.3.rs-10788774/v1},
url={https://doi.org/10.21203/rs.3.rs-10788774/v1}
}TR-HASH replaces learned routing with stable identity-based selection. Token or spatial identity chooses a small parameter subspace while shared computation continues to process the complete contextual hidden state.
identity ──► fixed layer-specific routing ──► selected experts
│ │
└──────── contextual hidden state ──────────┴──► output
|
The research and training layer: model definitions, distributed training, Triton kernels, exact resume, evaluation, ablations, multimodal generation, and Vision v8. |
The inference server for TR-HASH language models: exact hash-table routing, paged KV cache, CUDA-graph and MPS-graph decode, quantization, and HTTP serving. |
|
The product-facing SDK for TR-HASH Vision: prediction, validation, fine-tuning, export, benchmarking, and HTTP serving without carrying the research framework. |
Open TR-HASH checkpoints, model cards, demos, and progressively published training artifacts. |
|
Open text, image, and image-edit datasets with provenance-oriented releases. |
| Hierarchical tower P2 · P3 · P4 · P5 native features |
Attention Shifted-window |
Detection NMS and NMS-free heads |
The released 80-class COCO detector has 2.53M parameters at 640 px input, trained end-to-end on COCO 2017. Accuracy claims and YOLO comparisons are published only against the same-protocol val2017 evaluation.
|
Text instruction and chat SFT corpus. |
336K provenance-aware image-text pairs. |
336K instruction-guided editing triplets. |
Framework scope
- deterministic TR-HASH MoE with separate expert learning rates;
- language models with GQA/MHA and shared plus routed feed-forward paths;
- hierarchical detection, classification, segmentation, depth, pose, and OBB models;
- image generation/editing, speech, and video research components;
- single-device, DDP, FSDP, CUDA, MPS, and CPU execution;
- exact resumable checkpoints with optimizer, scheduler, cursor, and distributed RNG state.
We separate implemented architecture, active training, and validated results. Claims are tied to realized checkpoints and explicit evaluation protocols. Parameters, compute, latency, memory, and accuracy are reported together whenever possible. Planned runs are never presented as completed benchmarks.
Community contributions, replications, and critical evaluations are welcome.





