Review DreamDojo - Heungwoo/research GitHub Wiki

In-Depth Review β€” DreamDojo: a real-time robot world model from large-scale human video

Venue: ICML 2026 (Poster) Β· arXiv: 2602.06949 Β· Backbone: Cosmos-Predict2.5 (+ WAN2.2 tokenizer) Β· Traction: ~42 arXiv citations (Jun 2026) Category: World model / neural simulator (not a policy) Companions: World Models Β· NVIDIA WAM thesis + Cosmos 3 Β· Independent Visual Representation (P-cluster) Β· GR00T Series Β· per-paper stub: DreamDojo

Figures: paper Figure 1 / Figure 3 are embedded from the repo's local assets; additional figures are embedded from the arXiv HTML; the 3-phase pipeline is an original schematic.


1. TL;DR

DreamDojo is not a VLA β€” it is a generative robot world model: given the current observation and an action, it predicts the future video (simulates the outcome). Its three contributions answer why robot world models had failed to generalize:

  1. Unlock 44.7k hours of unlabeled human video for world-model pretraining by using continuous latent actions as a universal action label (the missing-action-label bottleneck).
  2. Transfer physical/interaction knowledge to dexterous robots out-of-distribution β€” it beats its own non-human-pretrained Cosmos-Predict2.5 baseline (e.g. 73.5% physics human-preference for the 14B model).
  3. Make a diffusion world model real-time interactive via distillation β€” 10.8 FPS with a 12-frame context (vs the teacher's 2.72 FPS), enough to put the simulator in the loop for policy evaluation (r = 0.995 with real success), model-based planning (+17%), and live teleoperation.

The one-sentence read: DreamDojo is the "imagine" half of the world-action-model thesis built into a real-time simulator a VLA plugs into β€” not a competitor to VLAs but an evaluator / planner / data-engine for them.

Figure 1 β€” DreamDojo overview: latent actions act as unified labels across human and robot data, letting one world model absorb physical knowledge from large-scale human video.


2. Problem β€” why robot world models don't generalize

Game/driving world models thrive on discrete controls and abundant logged data. Robot manipulation breaks both: action spaces are high-dimensional and continuous, and robot data is scarce and hardware-fragmented (RT-1 900 h, BridgeData V2 130 h, DROID 350 h, AgiBot-World 2.9k h). So prior robot world models stay in-distribution and fail on unseen interactions/objects/scenes. The abundant resource β€” diverse human video β€” has the interaction knowledge but no action labels. DreamDojo's bet: a learned latent action can serve as the universal label that bridges human video and robot control.


3. Method

3.1 Latent Action Model β€” an information bottleneck between frames

A 700M spatiotemporal Transformer (24 encoder + 24 decoder blocks) is trained as a VAE with an information bottleneck (Ξ² = 10⁻⁢, 400k steps, batch 256) to compress the transition between consecutive frames into a 32-D continuous latent β€” a disentangled proxy action that is defined identically whether the video is human or robot. This is what lets unlabeled human video supply "action-conditioned" supervision to the world model.

Figure 3 β€” Latent Action Model: the information-bottleneck design produces a continuous latent vector representing the action between frames, usable as a unified label across embodiments.

3.2 World-model architecture (on Cosmos-Predict2.5)

The world model fine-tunes Cosmos-Predict2.5 (WAN2.2 tokenizer, 4-frame temporal compression). Three design choices tame the continuous robot action space:

  • Relative-action rebaselining β€” robot joint poses are expressed relative to the start of each latent frame, shrinking the effective action range.
  • Action chunking β€” 4 consecutive actions are concatenated per latent frame, respecting causality and avoiding "future-action" leakage.
  • Action injection β€” actions are projected by a lightweight MLP and added to the timestep embedding via adaptive layer-norm (AdaLN) β€” i.e. conditioning, not extra tokens.

3.3 Three-phase training

flowchart LR
  P1[Phase 1 Β· Human-video pretrain<br/>44.7k h Β· latent-action conditioned] --> P2[Phase 2 Β· Robot post-train<br/>30k-50k steps Β· target robot]
  P2 --> P3[Phase 3 Β· Distillation<br/>Self-Forcing Β· real-time student]
Loading
  • Phase 1 β€” human-video pretrain. Mixture = In-lab (55 h, Manus-glove + Vive-tracker, retargetable to GR-1) + EgoDex (829 h) + the crowdsourced DreamDojo-HV (43,827 h, 1.135M trajectories, 6,015 skills, ~1.135M scenes) = 44,711 h. That is ~15Γ— duration, ~96Γ— skills, ~2,000Γ— scenes vs the previously largest robot dataset. Objective = flow-matching + a temporal-transition loss (β„’_final = β„’_flow + 0.1Β·β„’_temporal).
  • Phase 2 β€” robot post-train. The action-MLP first layer is re-initialized; full fine-tune for 30k–50k steps on limited target-robot data.
  • Phase 3 β€” distillation (Self-Forcing). A diffusion teacher (35 steps, 2.72 FPS) is distilled into an autoregressive student (4 steps) that streams latent frames in real time, conditions on multiple context frames (robust to occlusion/camera-shift), and runs ~4Γ— faster on one H100.

4. Results

Latent-action pretraining beats action-free / no-pretrain (Table 2, PSNR):

Conditioning In-lab EgoDex
w/o pretraining 20.58 19.95
action-free 20.80 19.92
latent action 20.91 20.34
retargeted (oracle) action 20.96 β€”

Out-of-distribution physical realism β€” human preference vs the Cosmos-Predict2.5 baseline (Table 4):

Model physics win-rate action-following
DreamDojo-2B > Cosmos-Predict2.5 62.5% 63.45%
DreamDojo-14B > Cosmos-Predict2.5 73.5% 72.55%

Design ablation (Table 5, PSNR): baseline 16.20 β†’ +relative 16.52 β†’ +chunked 17.63 (GR-1 Val); counterfactual 19.45 β†’ 19.48 β†’ 20.78 β†’ 20.98 (+temporal loss). Data scaling (Table 3): full mixture (DreamDojo-14B) 21.41 PSNR / 0.788 SSIM / 0.208 LPIPS, improving monotonically as human datasets are added.

Distillation β€” the real-time payoff (Table 6):

Metric Teacher Student
PSNR 14.09 13.15
FPS 2.72 10.81
Context length 1 frame 12 frames

The student trades only minor fidelity for 4Γ— speed and a 12Γ— longer context β€” the systems result that makes "world-model-in-the-loop" practical.

Figure 4 β€” the six OOD evaluation sets (In-lab, EgoDex, DreamDojo-HV, Counterfactual, + two novel splits) used to test generalization beyond the training distribution.


5. Downstream uses β€” and the connection to VLAs

DreamDojo's value is realized through policies/VLAs, not instead of them:

Figure 5 β€” downstream applications: simulator-based policy evaluation, model-based planning, and live VR teleoperation inside the learned world model.

  • Policy evaluation (the headline use): ranking policy checkpoints by simulated success correlates with real success at Pearson r = 0.995, Mean Maximum Rank Violation 0.003 β€” i.e. it can stand in for expensive real-robot evaluation when selecting VLA checkpoints (validated against GR00T N1.5 policies).
  • Model-based planning: sampling actions and picking by imagined outcome lifts a high-variance policy's success +17% (~2Γ— vs uniform sampling).
  • Live teleoperation / data engine: real-time rollouts support VR-controller teleop and synthetic interaction data.

6. Comparison with existing VLAs

The crucial framing: DreamDojo and a VLA solve inverse problems.

Axis Typical VLA (Ο€0.5 / OpenVLA / GR00T) DreamDojo (world model)
Role policy: observation + language β†’ action simulator: observation + action β†’ future observation
Output executable robot commands predicted future video frames
Pretrain data robot teleop (+ some web/human video) 44.7k h unlabeled human video
Action representation action tokens / flow-matching head latent actions (proxy label) + relative-chunked robot actions for conditioning
Real-time 10–200 Hz control 10.8 FPS rollout (distilled)
Generalization lever training-data diversity human-video physics transfer
Relationship the model being evaluated / planned for evaluates, plans for, and supplies data to the VLA

Two deeper points:

  1. Same idea, different target as latent-action VLAs. Latent-action pretraining from human video is exactly the lever behind latent-action policies (LAPA, ViPRA, GR00T's latent-action base, XR-1). DreamDojo applies it to the world model instead of the policy β€” the P-cluster (predictive) entry in Review-Independent-Visual-Representation, on the imagine side.
  2. It is a world model, not a WAM. Unlike a World-Action Model (e.g. Cosmos 3, DreamZero) that emits actions, DreamDojo only imagines β€” it is the substrate a WAM/VLA sits on. Tellingly it is built on the same Cosmos-Predict2.5 lineage, but stops at simulation rather than crossing into action generation. So DreamDojo is "pretrained to imagine" without the "fine-tuned to act" half.

7. Insight

  • Latent actions are the universal adapter that turns human video into world-model fuel. The bottleneck was never video β€” it was labels; a 32-D bottleneck latent supplies them for free, and human-video pretraining then transfers OOD physics that robot-only data cannot.
  • Distillation is what makes neural world models usable. A 2.72β†’10.8 FPS, 1β†’12-frame-context jump moves world models from "offline video generator" to "interactive simulator" β€” the precondition for policy-eval/planning/teleop.
  • The most valuable near-term use is evaluation, not control. r = 0.995 with real success means DreamDojo can rank VLA checkpoints cheaply β€” a direct accelerant for VLA development, and a cleaner value proposition than replacing the policy.
  • It complements, not competes with, the WAM debate. Where Review-WAM-vs-VLA-Robustness asks "should the policy be a video model?", DreamDojo answers a different question β€” "can a simulator be learned from human video and run in real time?" β€” and shows yes.

8. Limitations

Stated by the paper:

  • Uncommon / fast actions (e.g. slapping, fast waving) are simulated poorly.
  • Over-optimistic simulator β€” absolute success in DreamDojo is often higher than real, i.e. it cannot render nuanced failures. (Fine for ranking policies; risky if read as absolute success.)
  • No native multi-view simulation β€” yet SOTA policies often need multi-view input.
  • Post-training knowledge retention is not studied in depth.

Added critical reading:

  • It does not act. DreamDojo outputs frames, not commands β€” every downstream use needs a separate policy/planner/teleoperator. It is infrastructure, not an agent.
  • The simulator has its own sim-to-real gap. PSNR/SSIM/LPIPS and human-preference measure video plausibility, not task-relevant controllability β€” the same proxy-metric weakness flagged for latent world models in Review-World-Models Β§6.
  • Latent actions are opaque and not directly executable; their disentanglement quality is asserted via downstream PSNR, not directly probed.
  • Teacherβ†’student fidelity drop (PSNR 14.09 β†’ 13.15) and long-horizon degradation are real, if modest; robustness of the 12-frame context under large viewpoint shift is shown qualitatively.

9. Significance & positioning

DreamDojo makes human-video pretraining a scalable axis for world-model-driven robot learning, and β€” via distillation β€” turns a diffusion world model into a real-time, in-the-loop simulator. Its strongest practical contribution is cheap, faithful policy evaluation (r = 0.995), which directly serves VLA development. In the 2026 map it sits in Review-VLA-Architecture's Category-E world-model cluster (alongside DUST, VLAW) but is best understood as the imagine-only substrate beneath the world-action-model thesis: same Cosmos lineage, but stopping at simulation rather than emitting actions.


10. Links

← Back to ICML-2026 Β· Home

⚠️ **GitHub.com Fallback** ⚠️