NeurIPS 2025 DreamVLA - Heungwoo/research GitHub Wiki
Venue: NeurIPS 2025 ยท Authors: SJTU / EIT / Tsinghua / Galbot / PKU / UIUC / USTC ยท arXiv: 2507.04447 Category: World Models / VLA Architecture
flowchart LR
V[Obs + language] --> VLM[VLA backbone Seer-based]
VLM --> BA[Block-wise structured attention<br/>masks cross-cue attention]
BA --> DR[Dynamic region]
BA --> D[Depth / spatial]
BA --> S[Semantic features<br/>DINOv2 + SAM]
DR & D & S --> IDM[Inverse-dynamics loop]
IDM --> DiT[Diffusion transformer action head]
DiT --> A[Actions]
Predicting just the next action gives a weak supervision signal โ the model doesn't explicitly learn "what's about to change in the scene." Prior world-model-augmented VLAs (DreamGen, WorldVLA) predict future RGB frames, but RGB is expensive and noisy. What the policy actually needs is structured world knowledge: what moves, how far, what's behind it.
Rather than generating full future RGB frames, DreamVLA introduces world embeddings that forecast three compact world-knowledge cues โ dynamic, spatial, and semantic โ as auxiliary supervision alongside action prediction:
- Dynamic region โ where in the scene will motion happen? (dynamic-region-guided prediction is the organizing cue.)
- Depth โ spatial/geometric structure: how far?
- High-level semantic features โ distilled from DINOv2 and SAM foundation-model features (not raw segmentation labels).
To stop these heterogeneous cues from interfering during training, DreamVLA uses a block-wise structured attention mechanism that masks mutual attention between the dynamic, spatial, and semantic representations, keeping each one clean and disentangled โ a core contribution.
Each forecast is conditioned on current observation + language; the model then closes a perception-prediction-action loop via inverse-dynamics modeling โ given current state and forecast future state, infer the action. Actions are produced by a diffusion-based transformer that disentangles action representations from the shared latent features. The implementation builds on the Seer codebase.
- 76.7% real-robot success on manipulation tasks (Franka), vs. baselines such as Diffusion Policy, Octo-Base, and OpenVLA.
- 4.44 average task length on CALVIN ABC-D โ above prior methods reported in the paper (e.g., VPP 4.29, Seer 4.28).
- 92.6% average on LIBERO (Spatial / Object / Goal / Long).
- Generalizes across held-out scenes better than RGB-only world-model baselines.
Extends the DreamGen / video-world-model line into structured forecasting. The multi-modal forecast idea is adopted in VideoVLA and โ in a different form โ in Cosmos Policy and Ctrl-World. Listed as Category E (World-Model / VAM backbones) in Review-VLA-Architecture.
- arXiv: https://arxiv.org/abs/2507.04447
- NeurIPS virtual page: https://neurips.cc/virtual/2025/poster/118226
- GitHub: https://github.com/Zhangwenyao1/DreamVLA
- DreamGen (CoRL 2025 ancestor)
- VideoVLA (joint action + video sibling)
- Cosmos Policy ยท Ctrl-World (ICLR 2026 descendants)
- Review: VLA Architectures โ ยง5.E
โ Back to NeurIPS-2025