NeurIPS 2025 DreamVLA - Heungwoo/research GitHub Wiki

DreamVLA โ€” A VLA Dreamed with Comprehensive World Knowledge

Venue: NeurIPS 2025 ยท Authors: SJTU / EIT / Tsinghua / Galbot / PKU / UIUC / USTC ยท arXiv: 2507.04447 Category: World Models / VLA Architecture

Approach diagram

flowchart LR
  V[Obs + language] --> VLM[VLA backbone Seer-based]
  VLM --> BA[Block-wise structured attention<br/>masks cross-cue attention]
  BA --> DR[Dynamic region]
  BA --> D[Depth / spatial]
  BA --> S[Semantic features<br/>DINOv2 + SAM]
  DR & D & S --> IDM[Inverse-dynamics loop]
  IDM --> DiT[Diffusion transformer action head]
  DiT --> A[Actions]
Loading

Problem

Predicting just the next action gives a weak supervision signal โ€” the model doesn't explicitly learn "what's about to change in the scene." Prior world-model-augmented VLAs (DreamGen, WorldVLA) predict future RGB frames, but RGB is expensive and noisy. What the policy actually needs is structured world knowledge: what moves, how far, what's behind it.

Method

Rather than generating full future RGB frames, DreamVLA introduces world embeddings that forecast three compact world-knowledge cues โ€” dynamic, spatial, and semantic โ€” as auxiliary supervision alongside action prediction:

  • Dynamic region โ€” where in the scene will motion happen? (dynamic-region-guided prediction is the organizing cue.)
  • Depth โ€” spatial/geometric structure: how far?
  • High-level semantic features โ€” distilled from DINOv2 and SAM foundation-model features (not raw segmentation labels).

To stop these heterogeneous cues from interfering during training, DreamVLA uses a block-wise structured attention mechanism that masks mutual attention between the dynamic, spatial, and semantic representations, keeping each one clean and disentangled โ€” a core contribution.

Each forecast is conditioned on current observation + language; the model then closes a perception-prediction-action loop via inverse-dynamics modeling โ€” given current state and forecast future state, infer the action. Actions are produced by a diffusion-based transformer that disentangles action representations from the shared latent features. The implementation builds on the Seer codebase.

Results

  • 76.7% real-robot success on manipulation tasks (Franka), vs. baselines such as Diffusion Policy, Octo-Base, and OpenVLA.
  • 4.44 average task length on CALVIN ABC-D โ€” above prior methods reported in the paper (e.g., VPP 4.29, Seer 4.28).
  • 92.6% average on LIBERO (Spatial / Object / Goal / Long).
  • Generalizes across held-out scenes better than RGB-only world-model baselines.

Significance

Extends the DreamGen / video-world-model line into structured forecasting. The multi-modal forecast idea is adopted in VideoVLA and โ€” in a different form โ€” in Cosmos Policy and Ctrl-World. Listed as Category E (World-Model / VAM backbones) in Review-VLA-Architecture.

Links

Related pages

โ† Back to NeurIPS-2025

โš ๏ธ **GitHub.com Fallback** โš ๏ธ