ICML 2026 DUST - Heungwoo/research GitHub Wiki

DUST: Dual-Stream Diffusion for World-Model Augmented VLA — Resolving the state/action modality conflict

Venue: ICML 2026 (Poster) Category: World Model Traction (2026-06): 13 citations (arXiv)

DUST overview: a VLM backbone conditions a dual-stream diffusion model that jointly predicts action sequences and future observation embeddings (Figure 1 / x1 from Yoon et al., 2026)

Problem

Augmenting VLAs with world-modeling objectives — training a policy to predict future observations alongside actions — has shown promise for grounding robots in physical dynamics. But jointly predicting next-state observations and action sequences is hard because the two modalities are fundamentally different: actions are low-dimensional and temporally smooth, while future image observations are high-dimensional and spatially structured. Forcing them into a shared latent space creates a modality conflict that limits gains.

Method

DUST (DUal-STream diffusion) is built on a central VLM backbone that produces semantic conditioning features Φₜ from the current observation and instruction. These condition a core diffusion model π_θ that ingests the triplet (proprioceptive state, noised action sequence, noised future-observation embedding), processed by a stack of Multi-Modal Diffusion Transformer (MMDiT) blocks.

MMDiT block: action and vision token streams stay separate, concatenating only briefly in a shared cross-modal attention layer (mmdit figure from Yoon et al., 2026)

Three design choices distinguish DUST:

  • Separate modality streams. Within each MMDiT block, action and vision tokens flow through separate pathways, concatenating only temporarily during a shared cross-modal attention layer, then splitting back — enabling cross-modal sharing without a unified latent space.
  • Decoupled flow-matching loss with per-modality noise. Inspired by Diffusion Forcing, actions and image embeddings receive independent noise timesteps (τ_A, τ_o). Asymmetric noising teaches bidirectional causality: a clean future from a noisy action learns the inverse ("what action produced this state?"), and vice versa for the forward dynamics.
  • Asynchronous joint sampling (test-time scaling). At inference, vision tokens (needing more denoising steps) update every fine-grained step while action tokens update only every q steps (N_o = q × N_A), exploiting the decoupled design for inference-time scaling.

Results

On RoboCasa, DUST + GR00T-N1.5 beats both the GR00T baseline and the FLARE world-model baseline at every data scale: 0.501 avg (100 demos), 0.585 (300), 0.663 (1000 demos) — up to ~6% over baselines. On GR-1, it likewise outperforms GR00T-N1.5 and FLARE. Test-time scaling by raising the vision step count N_o lifts RoboCasa-1000 from 0.663 → 0.697 and GR-1-1000 to 0.471, an extra 2–5% boost. Transfer learning: pre-training only the world-modeling term on action-free BridgeV2 video raises RoboCasa from 0.501 → 0.585, showing DUST can absorb cheap video data before policy fine-tuning. On real-world Franka Research 3 tasks, DUST improves success rates by 13% over baselines.

Significance

DUST shows that the state/action modality conflict in world-model-augmented VLAs is best handled by architectural and training-time decoupling rather than a unified latent space. The dual-stream design unlocks bidirectional dynamics learning, asynchronous test-time scaling, and large-scale pre-training on action-free video — a promising recipe for scalable VLA pretraining.

Links

← Back to ICML-2026