ICML 2026 DreamDojo - Heungwoo/research GitHub Wiki

DreamDojo โ€” a real-time robot world model from large-scale human videos

Venue: ICML 2026 (Poster) Category: World Model Traction (2026-06): 42 citations (arXiv) ๐Ÿ“– In-depth review (insight ยท VLA comparison ยท limitations ยท figures): Review-DreamDojo

DreamDojo overview: latent actions as unified labels across human and robot data (Figure 1 from the DreamDojo authors, 2026)

Problem

Generative world models could revolutionize generalist agents by simulating action outcomes, but modeling dexterous, contact-rich robot dynamics is hard: robot data has limited coverage due to hardware variability and collection cost, and the high-dimensional continuous action spaces of manipulation lag far behind the discrete controls of game/driving world models. Existing robot world models are largely confined to in-distribution settings and fail to generalize to unseen interactions, objects, and scenes. The core bottleneck is the scarcity of action labels in the abundant, diverse human video that could otherwise supply broad interaction knowledge.

Method

DreamDojo is a foundation world model pretrained on 44k hours of egocentric human video (the largest video corpus to date for world-model pretraining), trained in three phases: human-video pretraining, post-training on small-scale target-robot data, and distillation.

  • Data mixture (44,711 hours, 1.18M trajectories): In-lab (Manus-glove + Vive-tracker tabletop data retargetable to GR-1), the public EgoDex dexterous dataset (829 h), and the in-house crowdsourced DreamDojo-HV (43,827 h, 1.135M trajectories, ~6,015 skills across household/industrial/retail/educational/administrative scenes). The mixture has ~15ร— longer duration, ~96ร— more skills, and ~2,000ร— more scenes than the previously largest dataset.
  • Continuous latent actions serve as unified proxy actions: a 700M spatiotemporal Transformer (24 encoder + 24 decoder blocks, 32-D latent) extracts disentangled latent actions via an information-bottleneck, enabling interaction-knowledge transfer from unlabeled human video.
  • Architecture (on WAN2.2): robot joint poses are rebaselined into relative actions (per 4-timestep latent frame) to shrink the action space; actions are injected as chunks of 4 consecutive actions per latent frame to respect causality and avoid future-action noise.
  • Distillation converts the teacher into an autoregressive student that streams latent frames in real time, conditions on multiple context frames (robust to occlusion/camera shift), and runs ~4ร— faster on a single H100.

Latent action model: information-bottleneck design produces continuous action latents between frames (Figure 3 from the DreamDojo authors, 2026)

Results

The distilled student achieves real-time performance at 10.93 FPS (paper Table 6 reports 10.81 FPS), versus 2.72 FPS for the teacher, with only minor long-horizon degradation (PSNR 13.146 vs. 14.086; context length 12 vs. 1). Latent-action pretraining consistently beats actionless pretraining, and across four OOD evaluations (In-lab, EgoDex, DreamDojo-HV, Counterfactual) a human-video-pretrained teacher outperforms a non-pretrained (Cosmos-Predict2.5) teacher even after distillation (e.g., In-lab PSNR 20.733 vs. 20.304). Scaling data continues to improve all six OOD benchmarks. Applications include live teleoperation, policy evaluation, and model-based planning.

Significance

DreamDojo shows that massive unlabeled human video, unified through continuous latent actions, can pretrain a generalizable robot world model that transfers physical knowledge to dexterous robots โ€” and that distillation can make such models real-time interactive. This positions human-video pretraining as a scalable axis for world-model-driven robot learning in 2026.

Links

โ† Back to ICML-2026