ICLR 2026 dVLA - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 (under review; submitted Sep 2025) Category: VLA Architecture + Reasoning Trend tag: Trends 1 + 2 Authors / affil: Junjie Wen, Minjie Zhu, Yichen Zhu, Yi Xu et al. — Midea Group, Peking University, Shanghai Jiao Tong University arXiv: 2509.25681
flowchart TB
In[Image + Instruction] --> Tx[Single discrete-diffusion transformer<br/>built on MMaDA]
Tx --> Out[Iterative masked denoising]
Out --> F[Subgoal image<br/>visual CoT]
Out --> CoT[Text chain-of-thought<br/>subtask description]
Out --> Act[Action chunk<br/>FAST tokens]
Embodied chain-of-thought (ECoT) usually emits reasoning tokens before the action. With an autoregressive decoder, this destroys inference latency — the model must serialize hundreds of reasoning tokens before any action is produced. Many ECoT-trained models drop reasoning at deployment and lose the training benefit.
Built on MMaDA (a discrete-diffusion multimodal LLM), dVLA performs discrete-diffusion co-generation of a multimodal chain-of-thought:
- A subgoal image (visual CoT) — a single predicted future-state frame sampled at ~[0.9C, 1.1C] timesteps ahead (C = action chunk length, 5 for LIBERO, 50 for real-world), encoded with MAGVIT-v2
- Text chain-of-thought reasoning (high-level subtask descriptions), via the LLaDA tokenizer
- Action tokens, discretized with the FAST tokenizer
All three modalities share one expanded vocabulary (~136,704 tokens) and are produced by iterative masked denoising: within each denoising step the model predicts masked tokens across all modalities simultaneously rather than autoregressively token-by-token. Adding reasoning tokens therefore does not cost one forward pass per token.
- LIBERO: 96.4% average success (Spatial 97.4 / Object 97.9 / Goal 98.2 / Long 92.2), ahead of continuous baselines (π₀ 94.2%, GR00T-N1 93.9%) and discrete ones (Discrete Diffusion VLA 96.3%, WorldVLA 81.8%).
- Real-world Franka (4 tasks × 10 trials): 65% average (26/40).
- Ablation: removing the multimodal CoT drops LIBERO from 96.4% → 89.8% and real-world from 65% → 52.5% — the CoT is decisive, not cosmetic.
- Acceleration: a prefix (block-causal) attention mask plus KV caching (dLLM-Cache) gives ~2× speedup with marginal degradation (LIBERO 1.3 → 2.9 Hz; real-world 1.5 → 3 Hz).
Resolves one of the longest-standing objections to ECoT: that the inference penalty makes it impractical at deployment. Because reasoning and action tokens denoise in the same parallel loop, "emit reasoning" and "emit actions fast" are no longer in tension. dVLA's own headline claim is being the first VLA framework built on a diffusion language model (MMaDA). Combined with Hybrid Training (which makes ECoT skippable at deployment), it closes the practical objections to reasoning-augmented VLAs.
- arXiv: https://arxiv.org/abs/2509.25681
- OpenReview (ICLR 2026): https://openreview.net/forum?id=2rxgospB5s
- Discrete Diffusion VLA
- Hybrid Training
- Embodied-R1 (RL for reasoning tokens)
← Back to ICLR-2026