Review MOTUS - Heungwoo/research GitHub Wiki
Model: Motus — unified latent-action world model (WAM+VLA hybrid) · Tsinghua University (THU-ML) & collaborators (Jun Zhu / Hang Su group; incl. Horizon Robotics) Paper: arXiv 2512.13030 (v1 Dec 15 2025 · v2 Dec 25 2025) · open weights + code (GitHub thu-ml/Motus, HF motus-robotics/Motus, project page) Status: arXiv preprint (not yet peer-reviewed) — filed under Latest Papers. Authors: Hongzhe Bi, Hengkai Tan, Shenghao Xie, Zeyuan Wang, … Zhizhong Su, Lei Ma, Hang Su, Jun Zhu.
One of the concrete instances of the WAM+VLA hybrid NVIDIA's blog predicts (NVIDIA WAM thesis §2.6) — a single Mixture-of-Transformers that can be a world model, a VLA, an inverse-dynamics model, or a video generator on demand. Companion: VLA Architectures §4.2b · World Models · DYNA-2 · Being-H0.7.
- One model, five modes. Motus is a Mixture-of-Transformers (MoT) integrating three experts — understanding, video generation, action — with a UniDiffuser-style scheduler that flexibly switches among world model · VLA · inverse-dynamics model · video generation · video-action joint prediction. Unifying all of these in one pretraining is the contribution.
- Latent actions from optical flow. It learns latent actions by extracting a pixel-level "delta action" from optical flow, which lets it pretrain action representations at scale from action-free video (human + multi-robot).
- 8B params across four towers: video-generation model 5.00B, VLM 2.13B, action expert 641.5M, understanding expert 253.5M.
- SOTA on RoboTwin 2.0. On the RoboTwin 2.0 randomized multi-task setting: 88.66% vs X-VLA 72.80% and π0.5 42.98% (= +15% over X-VLA, +45% over π0.5). Real-world (two platforms, 9 tasks): +11–48% over π0.5.
- It operationalizes the "unified" corner of the WAM+VLA hybrid. Where DYNA-2 drops its video tower at inference (reactive) and π0.7 bolts a world model onto a VLA modularly, Motus keeps all modes live in one MoT and schedules which to run — the most literal reading of "collapse WAM and VLA into one pretraining."
- Optical-flow latent actions are a cheap, general action-supervision bridge. Like DYNA-2's hand-pose pseudo-actions, this turns action-free video into trainable action signal — but via a task-agnostic motion cue (flow) rather than hand tracking, so it extends to non-hand, multi-robot footage.
- Open weights + code. Unlike DYNA-2 (closed) or Cosmos 3 (partial), Motus is fully released, so its unified-MoT recipe is independently checkable.

Mixture-of-Transformers, four towers, one scheduler. Each modality is a specialized expert; the UniDiffuser-style scheduler sets per-modality noise levels so the same weights realize different models:
flowchart LR
IN[obs · language] --> UND[Understanding expert · 253M]
IN --> VGM[Video-generation expert · 5.0B]
IN --> ACT[Action expert · 641M]
VLM[VLM · 2.13B] --- UND
UND <-. shared self-attention .-> VGM
VGM <-. shared self-attention .-> ACT
UND <-. shared self-attention .-> ACT
SCHED[UniDiffuser scheduler<br/>sets per-modality noise level] -. selects mode .-> VGM
SCHED -. selects mode .-> ACT
ACT ==> OUT[action]
| Scheduler mode | What is conditioned on what | Equivalent model |
|---|---|---|
| clean obs → predict future frames | video from observation | world model |
| clean obs + language → action | action from obs+text | VLA |
| clean obs + clean future → action | action from a transition | inverse-dynamics model |
| noise → frames | unconditional video | video generator |
| joint denoise frames + action | co-generate both | video-action joint prediction |
Latent actions via optical flow. Motus derives a pixel-level "delta action" from optical flow between frames, giving an embodiment-agnostic latent action label learnable from any video — the substrate for large-scale action pretraining.
Three-phase training over a six-layer data pyramid.
- Phase 1 — Learning Visual Dynamics: adapt the video-generation model on multi-robot and human videos.
- Phase 2 — Learning Action Representations: pretrain the unified model with the optical-flow latent actions.
- Phase 3 — Specializing for the Target Robot: fine-tune on target-robot data.
- Data pyramid (quantity ↓, quality ↑): web data → egocentric human video → synthetic → task-agnostic → multi-robot trajectories → target-robot task data.

| Setting | Motus | X-VLA | π0.5 |
|---|---|---|---|
| RoboTwin 2.0 (randomized multi-task) | 88.66% | 72.80% | 42.98% |
- Real-world, two robot platforms, 9 tasks (fold towel · brew coffee · grind beans · pour water · touch keyboard · grab cube · place cube · get water from dispenser · put bread in oven): +11–48% over π0.5.
- Ablations report that unifying all functionalities/priors in one model is what drives the downstream gains (vs single-mode baselines).
Significance. Motus is the clearest open demonstration that a single scheduled MoT can serve every world/action role — the unified endpoint of Category E (Review-VLA-Architecture §4.2b). The optical-flow latent-action recipe is a general alternative to hand-pose (DYNA-2) or reconstruction-free latents (ω-0).
Limitations.
- arXiv preprint, unreplicated. Numbers are the authors' own; no peer review yet.
- No explicit limitations/failure section. Future work only ("more universal motion priors; latent actions from internet-scale video").
- Optical flow as action proxy is coarse — flow conflates camera and object/effector motion; how well "delta action" grounds fine contact-rich control isn't isolated.
- RoboTwin 2.0 is simulation; real-world evidence is a 9-task, two-platform demo, not a broad hardware study.
- Paper: arXiv 2512.13030 · code/weights: GitHub · HF · project
- VLA Architectures §4.2b (WAM×VLA hybrids) · World Models
- Sibling hybrids: DYNA-2 · ω-0 · Being-H0.7 · Cortex 2.0 · Cosmos 3 / NVIDIA WAM
- Latest Papers
← Back to Latest Papers · Home · Reviews