Aligning Semantic Cognition with World Action Modeling via Event-Driven Interfaces
π Project Page Β Β·Β π Paper Β Β·Β π€ Model Β Β·Β βοΈ License
Robot policies increasingly incorporate semantic reasoning and future-world prediction, yet combining these capabilities does not guarantee that local predictions and actions remain aligned with task progress. We introduce CogWAM, a cognition-guided world-action model that establishes an explicit semantic interface between task reasoning and world-action learning through a persistent Semantic State, which stores completed task events and the active subtask. CogWAM updates this state only when observations indicate semantic transitions, allowing task-level context to persist across multiple action chunks. To bridge semantic context with physical prediction and control, CogWAM employs progress-conditioned WORLD and ACTION queries that selectively extract task-relevant information for future-world prediction and action generation. During training, the Semantic State provides shared task-progress context for both branches, while inference removes the future-prediction branch and directly generates actions from observations and the maintained state. Without additional robot-action pre-training, CogWAM reaches 15.56 / 11.70 Score/SR on RoboDojo.
Scope of this repository. This is the reproduction release for the RoboDojo experiment: one frozen recipe, end to end β data contract, training, checkpoint release, policy serving, and evaluation. The BiCoord and real-world experiments in the paper are not part of this release.
An event-triggered KEEP/UPDATE mechanism maintains the Semantic State during
closed-loop execution. WORLD and ACTION queries condition parallel branches for
future multi-view DINO feature prediction and continuous action generation.
- π§ Event-Triggered Semantic State. A persistent
(completed events, active subtask)pair that the model rewrites only when it emits<UPDATE>. A<KEEP>costs one forward pass and no autoregressive generation, so semantic context survives across many action chunks instead of being regenerated on a clock. - π― Progress-Conditioned WorldβAction Interface. Two sets of learnable query tokens appended after the observation, instruction and Semantic State. Their hidden states route to different objectives β WORLD to future-latent prediction, ACTION to control β so both branches share task-level semantics while keeping objective-specific representations.
- βοΈ Boundary-Aware Semantic Sampling. Semantic transitions are sparse (6.4% of
decision points). Each semantic batch is 2
UPDATE+ 2 hardKEEP(the decision points adjacent to a transition) + 2 randomKEEP, which forces the model to localise transitions rather than learn the prior.
CogWAM/
βββ cogwam/
β βββ models/
β β βββ cogwam.py # the framework: planner, queries, KEEP/UPDATE objective
β β βββ world_action_mot.py # causal DINO MoT: world + action streams
β β βββ vlm_interface.py # RynnBrain1.1 / Qwen3.5 backbone wrapper
β β βββ dino_v3.py # frozen DINOv3 encoder, multi-layer fusion
β β βββ base.py # framework base class and build_model
β βββ data/
β β βββ dataset.py # RoboDojo LeRobot v2.1 loader
β β βββ event_memory.py # Semantic State labels, index, boundary-aware sampler
β β βββ composite.py # tri-view composite (shared by trainer and server)
β β βββ robodojo.py # the single data configuration
β β βββ lerobot/ # pruned LeRobot v2.1 reader and transforms
β βββ training/ # entry point, dual-stream trainer, optimizer, checkpoints
β βββ serve/ # stateless WebSocket policy server
β βββ eval/ # RoboDojo adapter, rollout, Table-1 aggregation
β βββ recipe.py # the frozen reproduction contract + fingerprint
βββ configs/
β βββ cogwam_robodojo_h25_eventmem_dino_multilayer_50k.yaml
β βββ accelerate/zero2.yaml deepspeed/zero2.json robodojo_deploy.yml
βββ scripts/ # train_multi_node.sh, serve_policy.sh, eval_robodojo.sh
βββ tools/ # convert_checkpoint.py, verify_checkpoint.py, label_subtasks.py
βββ UPSTREAM_SOURCES.json # per-file provenance for every ported file
Two environments are required and cannot be merged: training pins
transformers 5.x on Python 3.12, while the RoboDojo simulator pins Isaac Sim on
Python 3.11. They talk over a WebSocket, so they never share an interpreter and
may sit on different machines.
conda create -n cogwam python=3.12 -y && conda activate cogwam
pip install torch==2.6.0 torchvision==0.21.0 --index-url https://download.pytorch.org/whl/cu124
pip install -r requirements.txt
pip install flash-attn==2.7.4.post1 --no-build-isolation # optional; the recipe uses SDPA
pip install -e .| Component | Version | Why pinned |
|---|---|---|
| Python | 3.12 (3.10+ works) | reference image |
| torch / torchvision | 2.6.0+cu124 / 0.21.0 | see Known gaps |
| transformers | >= 5.2.0 | first release exposing Qwen3_5ForConditionalGeneration; 4.x cannot load the backbone at all |
| flash-linear-attention | 0.3.2 | not optional β RynnBrain1.1 interleaves three linear_attention layers per full-attention layer; later releases change numerics |
| causal_conv1d | 1.5.0.post8 | required by the same layers |
| triton | 3.2.0 | matches those two kernels |
| deepspeed / accelerate | 0.16.9 / 1.5.2 | ZeRO-2, bf16 |
| numpy / av / decord / pyarrow | 1.26.4 / 12.3.0 / 0.6.0 / 14.0.1 | data path |
git clone --recurse-submodules https://github.com/RoboDojo-Benchmark/RoboDojo.git
cd RoboDojo && git checkout 9b4cc885e8f530ed3ab14a30a312ae70242771c4
git submodule update --init --recursive
export ROBODOJO_ROOT="$PWD"
pip install -r /path/to/CogWAM/requirements-eval.txt| Component | Pin |
|---|---|
| RoboDojo | 0.2.0 @ 9b4cc885e8f530ed3ab14a30a312ae70242771c4 |
XPolicyLab |
fe71eb54675cef495fea817a637386a4f4529153 |
third_party/IsaacLab |
afca7b09d60d8beb9c1cb28b43066499940b969b |
third_party/curobo |
d17b54ce32cba095c0b000c4c58777075d11de0e (isaacsim-warp18) |
| Isaac Sim / Python / CUDA | 5.1 / 3.11 / 12.8 |
export COGWAM_BASE_VLM=/models/rynnbrain1.1-2B
export COGWAM_DINO_MODEL=/models/dinov3-vitb16
export COGWAM_DATA_ROOT=/datasets/RoboDojo
export COGWAM_RUN_ROOT=/outputs/cogwam
# 8 nodes x 8 GPUs; set the rendezvous vars on each node
export COGWAM_MAIN_PROCESS_IP=<rank-0 address> COGWAM_NUM_MACHINES=8 COGWAM_MACHINE_RANK=<0..7>
bash scripts/train_multi_node.sh configs/cogwam_robodojo_h25_eventmem_dino_multilayer_50k.yamlSingle-GPU smoke test (this is deliberately not a reproduction, so the recipe contract must be waived):
COGWAM_SKIP_RECIPE_VALIDATION=1 python -m cogwam.training.train \
--config_yaml configs/cogwam_robodojo_h25_eventmem_dino_multilayer_50k.yaml \
--trainer.max_train_steps=10 --datasets.vla_data.per_device_batch_size=1| VLM | RynnBrain1.1-2B (Qwen3.5), hidden 2048, SDPA, BF16 |
| Queries | 16 WORLD + 25 ACTION (must equal the action horizon) |
| Physical model | causal DINO MoT, 30 layers, 24 heads Γ 128; world stream 512/2048, action stream 1024/4096 |
| Visual latents | frozen DINOv3 ViT-B/16, layers [2, 5, 8, 11] mean-fused, 24Γ20 patches pooled to a 12Γ10 / 120-token grid, future frame at t + 16 |
| Action | 14-D absolute joint positions, horizon 25, flow matching, 20 Euler steps at inference |
| Semantic State | event_driven_memory_ntp, decision every 10 steps, loss weight 0.005 |
| Optimisation | ZeRO-2 bf16, global batch 768, 50k steps, 2k warmup, cosine to 5e-7 |
The multi-layer DINO fusion applies to both the clean current prefix and the
future denoising target β they share one world_input projection, so a target
drawn from one feature space and a prefix from another would make the objective
incoherent. Every weight shape is identical to the single-layer variant.
cogwam/recipe.py compares 162 dotted config paths against a frozen contract
and raises before the backbones load, then emits a SHA-256 fingerprint written
next to the run. Storage locations are deliberately not frozen β upstream
enforced absolute paths, which made the recipe unrunnable elsewhere.
Training uses the RoboDojo demonstration corpus in LeRobot v2.1 format
(3,500 episodes / 1,859,602 frames). Three cameras (cam_high,
cam_left_wrist, cam_right_wrist) are stitched into one tri_view_composite
image; actions are 14-D absolute joint positions.
Two extra per-frame columns supervise the entire semantic objective:
| Column | Meaning |
|---|---|
subtask_text |
the subtask in progress at this frame |
complete_text |
the events already finished, joined as ". ".join(items) + ".", or None. |
Both are piecewise constant over right-open spans and flip on the same frame;
neither may ever be empty. Everything else is derived at load time: at a decision
frame t, the loader reads t and t-10 and emits semantic_memory,
cached_current_subtask, semantic_decision, memory_add and
semantic_cache_valid. "Changed" is judged under N1 normalisation (collapse
whitespace, casefold, strip trailing .;,:!?), so re-punctuation does not
manufacture a spurious UPDATE.
The guard that matters: only 6.4% of decision points are UPDATEs, and
that ratio is a contract, not an observation. The first launch builds
semantic_index_v1.npz and refuses to continue outside 0.064 Β± 0.005 β below
the band the semantic head learns to always answer KEEP; above it, the
event-driven premise is gone. Widening the tolerance to make a run start is how
the objective silently becomes a different objective.
To validate a dataset, or to annotate your own:
python tools/label_subtasks.py --dataset-root "$COGWAM_DATA_ROOT" --check--check makes no model calls. Labelling needs an OpenAI-compatible endpoint via
COGWAM_LABEL_BASE_URL / COGWAM_LABEL_MODEL / COGWAM_LABEL_API_KEY, plus a
checklist file giving the ordered subtask strings per task β the model only
decides where the boundaries fall, which keeps the subtask vocabulary stable
across episodes. The checklists used for the released annotations are not part of
this release; re-labelling will produce your own segmentation, and the ratio
guard is how you tell whether it is in family.
Evaluation splits across two processes: a stateless policy server holding the model, and the Isaac Sim client holding all per-environment Semantic State. That split is the contract β the server cannot leak state between episodes because it has none.
# machine A (GPUs): N servers on ports 7777.., blocks until every handshake, prints its address
COGWAM_ARTIFACT_DIR=/ckpts/cogwam-robodojo-h25-50k NUM_SERVERS=8 bash scripts/serve_policy.sh
# machine B (Isaac Sim)
export ROBODOJO_ROOT=/path/to/RoboDojo ROBODOJO_PYTHON=/opt/robodojo-env/bin/python
export COGWAM_CKPT_PATH=/ckpts/cogwam-robodojo-h25-50k COGWAM_POLICY_HOST=<address from A>
bash scripts/eval_robodojo.sh
python -m cogwam.eval.summarize --eval-root <rollout dir> \
--checkpoint /ckpts/cogwam-robodojo-h25-50k --ckpt-name cogwam-robodojo-h25-50k --seed 0Protocol: 42 tasks in 5 dimensions (Generalization 12, Precision 8,
Long-Horizon 8, Memory 6, Open 8), 54 eval configs (the 12 Generalization
tasks run twice, <task> and <task>_random, 25 episodes each; the other 30 run
50), 2100 episodes per seed. Aggregation is equal-weight at every level, so
Memory's 6 tasks weigh as much as Generalization's 12. Episodes flagged
unstable leave the denominator. The public leaderboard protocol uses 3 seeds
(6300 episodes); a single-seed number is one third of that sample size.
Control semantics: the model predicts a 25-step chunk, of which only the
first 10 are executed before replanning; the 15-step tail is RTC overlap and
is discarded (RTC_ENABLED is false in the released configuration). A semantic
KEEP/UPDATE decision fires on exactly the same tick β the client aborts
unless the server advertises event_replan_interval == 10 and
event_semantic_offset == -10.
Results as reported in the paper. CogWAM uses no prior embodied robot-data pre-training, so the comparable group is the lower block:
| Method | Generalization | Precision | Long-Horizon | Memory | Open | Average |
|---|---|---|---|---|---|---|
| Fast-WAM | 2.34 / 1.00 | 1.96 / 0.00 | 9.14 / 5.17 | 3.55 / 3.44 | 0.42 / 0.42 | 3.48 / 2.03 |
| AHA-WAM | 5.79 / 3.00 | 5.86 / 2.42 | 8.61 / 2.67 | 2.97 / 2.78 | 0.88 / 0.83 | 4.82 / 2.39 |
| StarVLA-Ξ± | 3.94 / 2.50 | 9.90 / 4.33 | 14.15 / 6.50 | 3.34 / 2.44 | 0.68 / 0.58 | 6.40 / 3.24 |
| X-WAM | 7.39 / 3.00 | 6.72 / 1.83 | 17.47 / 9.08 | 6.32 / 4.67 | 0.57 / 0.25 | 7.69 / 3.83 |
| Fast-WAM + CogWAM Interface | 12.28 / 9.00 | 19.61 / 13.50 | 23.69 / 14.75 | 9.17 / 8.00 | 1.93 / 1.75 | 13.33 / 9.40 |
| CogWAM (ours) | 15.53 / 12.17 | 24.45 / 19.00 | 28.14 / 19.00 | 7.65 / 6.33 | 2.05 / 2.00 | 15.56 / 11.70 |
Each cell is Score / Success Rate (%). A full single-seed sweep is roughly 30 hours with 8 GPUs per side.
The paper also evaluates closed-loop dual-arm manipulation on a PiPER dual-arm rig with three RealSense D435 cameras, across basic and generalization settings (spatial location, object appearance, distractors, novel objects) β reporting 16.4Γ fewer Semantic State regenerations than step-wise updating.
Those experiments are not part of this release. The code here targets RoboDojo; the real-robot stack, its data and its checkpoints are separate.
@article{wang2026cogwam,
title={CogWAM: Aligning Semantic Cognition with World Action Modeling via Event-Driven Interfaces},
author={Wang, Sen and Liu, Liu and Wang, Xinjiang and Chen, Zequn and Jiang, Haoyi and Ding, Taojun and Xiao, Tingyang and Su, Zhizhong and Wang, Jie and Zhou, Sanping},
journal={arXiv preprint arXiv:2609.37721},
year={2026}
}CogWAM stands on work we did not do.
- StarVLA β this codebase is derived from
the StarVLA research platform; its modular framework/dataloader/trainer separation
is what made extracting a single recipe into a standalone repository tractable.
Per-file provenance, with the upstream SHA-256 each file was ported from and every
deviation we introduced, is in
UPSTREAM_SOURCES.json. - RoboDojo β the benchmark, its task suite and evaluation protocol, and XPolicyLab for the policy interface.
- RynnBrain1.1 (Alibaba DAMO Academy) β the vision-language backbone.
- DINOv3 (Meta AI) β the frozen visual encoder behind both the current prefix and the future target.
- NVIDIA Isaac Lab / Isaac Sim and cuRobo, which RoboDojo builds on, and the LeRobot dataset format with its GR00T-lineage reader.
MIT, with portions derived from StarVLA (MIT) β see LICENSE. The two backbone models carry their own terms and are not redistributed here.

