CoRL 2025 DreamGen - Heungwoo/research GitHub Wiki

DreamGen โ€” Unlocking Generalization through Neural Trajectories

Venue: CoRL 2025 ยท Author: NVIDIA (GEAR Lab) ยท arXiv: 2505.12705 Full title: DreamGen: Unlocking Generalization in Robot Learning through Neural Trajectories Category: World Models for Policy Trend tag: World models for data generation

Approach diagram

flowchart LR
  D[Real robot data] --> WM[Fine-tuned video world model]
  WM -- generated videos --> IDM[IDM / latent action model]
  IDM -- pseudo-actions --> NT[Neural trajectories]
  NT -- supervised IL --> POL[Visuomotor policy]
  POL -- deploy --> REAL[Real robot]

Problem

Real-world data is expensive and sim-to-real has a gap. Video world models are now good enough (Sora, Cosmos, VEO-class) that we can imagine visually plausible rollouts โ€” if we could train a policy entirely inside those rollouts, we'd scale beyond teleoperation.

Method

DreamGen is a simple yet effective 4-stage offline data-generation pipeline โ€” not online policy-training-in-imagination. The video world model is not action-conditioned; actions are recovered afterward:

  1. Fine-tune a state-of-the-art image-to-video world model (Cosmos, WAN 2.1, etc.) on target-robot trajectories.
  2. Generate synthetic robot videos from an initial frame + language instruction (familiar or novel tasks/environments).
  3. Recover pseudo-actions for each video using a latent-action model or an inverse-dynamics model (IDM), yielding neural trajectories (synthetic video + action pairs).
  4. Train a visuomotor policy via standard imitation learning on the neural trajectories. Demonstrated with Diffusion Policy, ฯ€โ‚€, and GR00T N1 (GR00T N1 used for the headline generalization runs).

Results

Trained on teleoperation data from a single pick-and-place task in one environment, DreamGen unlocks zero-shot behavior and environment generalization โ€” a GR1 humanoid performs 22 new verbs across 10 new environments.

Setting Baseline + DreamGen
New behaviors, seen env (GR1) ~0% 43.2%
New behaviors, unseen env (GR1) ~0% 28.5%
4 GR1 humanoid tasks (avg) 37% 46.4%
3 Franka tasks (avg) 23% 37%
2 SO-100 tasks (avg) 21% 45.5%

Scaling neural trajectories shows a log-linear improvement in policy success (0 โ†’ 240k trajectories in RoboCasa). DreamGen Bench (evaluating Cosmos, WAN 2.1, Hunyuan, CogVideoX) finds video-generation quality positively correlates with downstream policy success.

Significance

DreamGen is the production-scale signal that video world models move into the training pipeline โ€” as a synthetic-data engine rather than an online imagination simulator. It reframes the world model as a way to expand effective robot data (~10ร— over real teleoperation) and is integrated into NVIDIA's GR00T stack (released as GR00T-Dreams). Threads into ICLR 2026 via Ctrl-World, Cosmos Policy, WorldGym, and VLA-RFT (RL-in-world-model).

Links

Related pages

โ† Back to CoRL-2025