Review DreamDojo - Heungwoo/research GitHub Wiki
Venue: ICML 2026 (Poster) Β· arXiv: 2602.06949 Β· Backbone: Cosmos-Predict2.5 (+ WAN2.2 tokenizer) Β· Traction: ~42 arXiv citations (Jun 2026) Category: World model / neural simulator (not a policy) Companions: World Models Β· NVIDIA WAM thesis + Cosmos 3 Β· Independent Visual Representation (P-cluster) Β· GR00T Series Β· per-paper stub: DreamDojo
Figures: paper Figure 1 / Figure 3 are embedded from the repo's local assets; additional figures are embedded from the arXiv HTML; the 3-phase pipeline is an original schematic.
DreamDojo is not a VLA β it is a generative robot world model: given the current observation and an action, it predicts the future video (simulates the outcome). Its three contributions answer why robot world models had failed to generalize:
- Unlock 44.7k hours of unlabeled human video for world-model pretraining by using continuous latent actions as a universal action label (the missing-action-label bottleneck).
- Transfer physical/interaction knowledge to dexterous robots out-of-distribution β it beats its own non-human-pretrained Cosmos-Predict2.5 baseline (e.g. 73.5% physics human-preference for the 14B model).
- Make a diffusion world model real-time interactive via distillation β 10.8 FPS with a 12-frame context (vs the teacher's 2.72 FPS), enough to put the simulator in the loop for policy evaluation (r = 0.995 with real success), model-based planning (+17%), and live teleoperation.
The one-sentence read: DreamDojo is the "imagine" half of the world-action-model thesis built into a real-time simulator a VLA plugs into β not a competitor to VLAs but an evaluator / planner / data-engine for them.

Game/driving world models thrive on discrete controls and abundant logged data. Robot manipulation breaks both: action spaces are high-dimensional and continuous, and robot data is scarce and hardware-fragmented (RT-1 900 h, BridgeData V2 130 h, DROID 350 h, AgiBot-World 2.9k h). So prior robot world models stay in-distribution and fail on unseen interactions/objects/scenes. The abundant resource β diverse human video β has the interaction knowledge but no action labels. DreamDojo's bet: a learned latent action can serve as the universal label that bridges human video and robot control.
A 700M spatiotemporal Transformer (24 encoder + 24 decoder blocks) is trained as a VAE with an information bottleneck (Ξ² = 10β»βΆ, 400k steps, batch 256) to compress the transition between consecutive frames into a 32-D continuous latent β a disentangled proxy action that is defined identically whether the video is human or robot. This is what lets unlabeled human video supply "action-conditioned" supervision to the world model.

The world model fine-tunes Cosmos-Predict2.5 (WAN2.2 tokenizer, 4-frame temporal compression). Three design choices tame the continuous robot action space:
- Relative-action rebaselining β robot joint poses are expressed relative to the start of each latent frame, shrinking the effective action range.
- Action chunking β 4 consecutive actions are concatenated per latent frame, respecting causality and avoiding "future-action" leakage.
- Action injection β actions are projected by a lightweight MLP and added to the timestep embedding via adaptive layer-norm (AdaLN) β i.e. conditioning, not extra tokens.
flowchart LR
P1[Phase 1 Β· Human-video pretrain<br/>44.7k h Β· latent-action conditioned] --> P2[Phase 2 Β· Robot post-train<br/>30k-50k steps Β· target robot]
P2 --> P3[Phase 3 Β· Distillation<br/>Self-Forcing Β· real-time student]
- Phase 1 β human-video pretrain. Mixture = In-lab (55 h, Manus-glove + Vive-tracker, retargetable to GR-1) + EgoDex (829 h) + the crowdsourced DreamDojo-HV (43,827 h, 1.135M trajectories, 6,015 skills, ~1.135M scenes) = 44,711 h. That is ~15Γ duration, ~96Γ skills, ~2,000Γ scenes vs the previously largest robot dataset. Objective = flow-matching + a temporal-transition loss (β_final = β_flow + 0.1Β·β_temporal).
- Phase 2 β robot post-train. The action-MLP first layer is re-initialized; full fine-tune for 30kβ50k steps on limited target-robot data.
- Phase 3 β distillation (Self-Forcing). A diffusion teacher (35 steps, 2.72 FPS) is distilled into an autoregressive student (4 steps) that streams latent frames in real time, conditions on multiple context frames (robust to occlusion/camera-shift), and runs ~4Γ faster on one H100.
Latent-action pretraining beats action-free / no-pretrain (Table 2, PSNR):
| Conditioning | In-lab | EgoDex |
|---|---|---|
| w/o pretraining | 20.58 | 19.95 |
| action-free | 20.80 | 19.92 |
| latent action | 20.91 | 20.34 |
| retargeted (oracle) action | 20.96 | β |
Out-of-distribution physical realism β human preference vs the Cosmos-Predict2.5 baseline (Table 4):
| Model | physics win-rate | action-following |
|---|---|---|
| DreamDojo-2B > Cosmos-Predict2.5 | 62.5% | 63.45% |
| DreamDojo-14B > Cosmos-Predict2.5 | 73.5% | 72.55% |
Design ablation (Table 5, PSNR): baseline 16.20 β +relative 16.52 β +chunked 17.63 (GR-1 Val); counterfactual 19.45 β 19.48 β 20.78 β 20.98 (+temporal loss). Data scaling (Table 3): full mixture (DreamDojo-14B) 21.41 PSNR / 0.788 SSIM / 0.208 LPIPS, improving monotonically as human datasets are added.
Distillation β the real-time payoff (Table 6):
| Metric | Teacher | Student |
|---|---|---|
| PSNR | 14.09 | 13.15 |
| FPS | 2.72 | 10.81 |
| Context length | 1 frame | 12 frames |
The student trades only minor fidelity for 4Γ speed and a 12Γ longer context β the systems result that makes "world-model-in-the-loop" practical.
DreamDojo's value is realized through policies/VLAs, not instead of them:
- Policy evaluation (the headline use): ranking policy checkpoints by simulated success correlates with real success at Pearson r = 0.995, Mean Maximum Rank Violation 0.003 β i.e. it can stand in for expensive real-robot evaluation when selecting VLA checkpoints (validated against GR00T N1.5 policies).
- Model-based planning: sampling actions and picking by imagined outcome lifts a high-variance policy's success +17% (~2Γ vs uniform sampling).
- Live teleoperation / data engine: real-time rollouts support VR-controller teleop and synthetic interaction data.
The crucial framing: DreamDojo and a VLA solve inverse problems.
| Axis | Typical VLA (Ο0.5 / OpenVLA / GR00T) | DreamDojo (world model) |
|---|---|---|
| Role | policy: observation + language β action | simulator: observation + action β future observation |
| Output | executable robot commands | predicted future video frames |
| Pretrain data | robot teleop (+ some web/human video) | 44.7k h unlabeled human video |
| Action representation | action tokens / flow-matching head | latent actions (proxy label) + relative-chunked robot actions for conditioning |
| Real-time | 10β200 Hz control | 10.8 FPS rollout (distilled) |
| Generalization lever | training-data diversity | human-video physics transfer |
| Relationship | the model being evaluated / planned for | evaluates, plans for, and supplies data to the VLA |
Two deeper points:
- Same idea, different target as latent-action VLAs. Latent-action pretraining from human video is exactly the lever behind latent-action policies (LAPA, ViPRA, GR00T's latent-action base, XR-1). DreamDojo applies it to the world model instead of the policy β the P-cluster (predictive) entry in Review-Independent-Visual-Representation, on the imagine side.
- It is a world model, not a WAM. Unlike a World-Action Model (e.g. Cosmos 3, DreamZero) that emits actions, DreamDojo only imagines β it is the substrate a WAM/VLA sits on. Tellingly it is built on the same Cosmos-Predict2.5 lineage, but stops at simulation rather than crossing into action generation. So DreamDojo is "pretrained to imagine" without the "fine-tuned to act" half.
- Latent actions are the universal adapter that turns human video into world-model fuel. The bottleneck was never video β it was labels; a 32-D bottleneck latent supplies them for free, and human-video pretraining then transfers OOD physics that robot-only data cannot.
- Distillation is what makes neural world models usable. A 2.72β10.8 FPS, 1β12-frame-context jump moves world models from "offline video generator" to "interactive simulator" β the precondition for policy-eval/planning/teleop.
- The most valuable near-term use is evaluation, not control. r = 0.995 with real success means DreamDojo can rank VLA checkpoints cheaply β a direct accelerant for VLA development, and a cleaner value proposition than replacing the policy.
- It complements, not competes with, the WAM debate. Where Review-WAM-vs-VLA-Robustness asks "should the policy be a video model?", DreamDojo answers a different question β "can a simulator be learned from human video and run in real time?" β and shows yes.
Stated by the paper:
- Uncommon / fast actions (e.g. slapping, fast waving) are simulated poorly.
- Over-optimistic simulator β absolute success in DreamDojo is often higher than real, i.e. it cannot render nuanced failures. (Fine for ranking policies; risky if read as absolute success.)
- No native multi-view simulation β yet SOTA policies often need multi-view input.
- Post-training knowledge retention is not studied in depth.
Added critical reading:
- It does not act. DreamDojo outputs frames, not commands β every downstream use needs a separate policy/planner/teleoperator. It is infrastructure, not an agent.
- The simulator has its own sim-to-real gap. PSNR/SSIM/LPIPS and human-preference measure video plausibility, not task-relevant controllability β the same proxy-metric weakness flagged for latent world models in Review-World-Models Β§6.
- Latent actions are opaque and not directly executable; their disentanglement quality is asserted via downstream PSNR, not directly probed.
- Teacherβstudent fidelity drop (PSNR 14.09 β 13.15) and long-horizon degradation are real, if modest; robustness of the 12-frame context under large viewpoint shift is shown qualitatively.
DreamDojo makes human-video pretraining a scalable axis for world-model-driven robot learning, and β via distillation β turns a diffusion world model into a real-time, in-the-loop simulator. Its strongest practical contribution is cheap, faithful policy evaluation (r = 0.995), which directly serves VLA development. In the 2026 map it sits in Review-VLA-Architecture's Category-E world-model cluster (alongside DUST, VLAW) but is best understood as the imagine-only substrate beneath the world-action-model thesis: same Cosmos lineage, but stopping at simulation rather than emitting actions.
- arXiv: 2602.06949 Β· arXiv HTML (figures): html/2602.06949v1 Β· ICML 2026 poster: icml.cc/virtual/2026/poster/65193
- Companions: Review-World-Models Β· NVIDIA WAM thesis + Cosmos 3 Β· Review-WAM-vs-VLA-Robustness Β· Independent Visual Representation Β· GR00T Series
- Per-paper stub: DreamDojo

