This page covers the stable research-facing surface. The source remains the authority for experimental fields.
For the released 200M model, prefer
scripts.train_tr_hash_200m_200b.make_config() over manually reconstructing
the shape. See TR-HASH MoE 200M release.
from complexity import ComplexityModel, ModelConfigThe package also exports:
- component registries and registration decorators;
- GQA, MHA, and MQA classes;
- the canonical TR-Hash MLP and experimental registered components;
- normalization and position-embedding components.
| Field | Meaning |
|---|---|
hidden_size |
residual width |
num_hidden_layers |
decoder depth |
intermediate_size |
routed expert pool width or dense FFN width |
vocab_size |
tokenizer vocabulary |
max_position_embeddings |
configured context bound |
| Field | Meaning |
|---|---|
attention_type |
registry key such as gqa, mha, tr_mha_v2 |
num_attention_heads |
query head count |
num_key_value_heads |
K/V head count |
use_qk_norm |
Q/K RMS normalization |
use_sdpa |
PyTorch SDPA path |
sliding_window |
optional local-attention window |
| Field | Meaning |
|---|---|
mlp_type |
set tr_hash_engine for TR-MoE |
num_experts |
routed expert count |
intermediate_size |
total stored routed width; divided by num_experts in TRHashEngineMLP |
routing_strategy |
canonical values: modulo_cyclic, token_id_balanced_hash, token_id_multi_hash, token_id_hierarchical_hash |
route_hash_count |
independent rendezvous channels for token_id_multi_hash (2–8) |
shared_expert |
enable dense shared SwiGLU |
shared_intermediate_size |
shared branch width |
shared_expert_chunk_tokens |
chunk shared computation over tokens |
top_k |
number of deterministic expert routes |
top_k_primary_weight |
blend assigned to primary route |
use_shared_routed_gates |
learn shared/routed scalar gates |
collect_moe_telemetry |
collect route/RMS diagnostics |
use_custom_kernels |
custom-kernel policy |
use_cggr |
grouped-GEMM policy |
Configuration validates shape, routing, and range invariants in
ModelConfig.__post_init__.
model = ComplexityModel(config)
model = ComplexityModel.from_config("config.yaml")
model = ComplexityModel.from_pretrained("checkpoint-directory")result = model(
input_ids,
attention_mask=None,
past_key_values=None,
use_cache=False,
return_hidden_states=False,
return_logits=True,
)Return mapping:
| Key | Value |
|---|---|
logits |
[batch, sequence, vocabulary], or None |
last_hidden_state |
final normalized hidden states |
past_key_values |
optional per-layer cache/state list |
hidden_states |
optional embedding and layer states |
Set return_logits=False for fused or chunked tied-head loss paths.
model.save_pretrained("checkpoint")
restored = ComplexityModel.from_pretrained("checkpoint")For distributed DTensor/FSDP saves, every rank must enter
save_pretrained because full-tensor gathering is collective.
model.generate() intentionally raises RuntimeError. Use the external
serving client.
from complexity.inference import (
ExternalGenerationConfig,
OpenAICompatibleBackend,
create_external_backend,
)create_external_backend accepts "vllm" or "sglang" as client labels and
calls /v1/completions or /v1/chat/completions. These labels select an HTTP
adapter; they do not teach an upstream runtime the TR-MoE architecture. The
released model is served by
TR-Hash-i64.
backend = create_external_backend(
"vllm",
base_url="http://localhost:8000",
model="tr-gqa",
)
answer = backend.chat(
[{"role": "user", "content": "Summarize the experiment."}],
ExternalGenerationConfig(max_tokens=128),
)The client is synchronous and non-streaming in the current implementation.
from complexity.core.registry import (
ATTENTION_REGISTRY,
MLP_REGISTRY,
NORMALIZATION_REGISTRY,
POSITION_REGISTRY,
register_attention,
register_mlp,
)Principal attention keys:
gqa, mha, mqa
tr_mha, tr_mha_v2
lexical_gqa, lexical_key_gqa
causal_conv, causal_state_conv, causal_fast_weight_conv
Principal MLP keys:
tr_hash_engine, tr_hash_moe
dense_deterministic
lexical_modulated, lexical_channel_modulated, lexical_object_micro_expert
swiglu/gelu/geglu/standard/mixtral/token_routed were removed —
constructing a config with any of them raises a clear error pointing at the
replacement (see token-routed.md for token_routed
checkpoints specifically).
Several aliases exist for checkpoint compatibility. New documentation should use the principal key.
from complexity.core.mlp import MLPConfig, TRHashEngineMLP
layer = TRHashEngineMLP(
MLPConfig(
hidden_size=896,
intermediate_size=256, # total routed width: 4 experts × 64
vocab_size=32_000,
num_experts=4,
shared_expert=True,
shared_intermediate_size=3072,
routing_strategy="token_id_multi_hash",
route_hash_count=2,
top_k=2,
top_k_primary_weight=0.5,
routed_output_scale=2.0,
)
)
output = layer(hidden_states, token_ids=input_ids)Useful diagnostics:
layer.engine.last_backend
layer.capability_summary()
layer.training_telemetry()Install:
pip install -e ".[tools]"Imports:
from complexity.mcp import (
MCPTool,
MCPToolResult,
OfficialMCPStdioClient,
OfficialMCPStdioConfig,
)The wrapper launches and calls an MCP server through the official Python SDK stdio transport. It does not reimplement tools.
complexity
cf-plan-run
cf-plan-cluster
cf-image-train
cf-image-edit-train
cf-detector-train
cf-vision-pretrain
cf-detector-serve
cf-sensor-fusion-train
cf-sensor-fusion-submit
cf-o200k-pretrain and cf-check-pipeline were removed along with the
o200k training pipeline (see training.md). cf-plan-run
and cf-plan-cluster remain for token-budget and cluster-sizing arithmetic.
Some older complexity subcommands remain experimental.