ICLR 2026 Discrete Diffusion VLA - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 Β· arXiv: 2508.20072 Category: VLA Architecture Trend tag: Trend 1
flowchart LR
V[Vision tokens] --> T[Single Transformer<br/>SAME as VLM]
L[Language tokens] --> T
M[Masked action tokens] --> T
T --> P{Prediction<br/>cross-entropy}
P -- unmask high-confidence first --> R1[Round 1 unmasking]
R1 --> RM[Secondary re-masking<br/>of low-confidence]
RM --> Done[Final action chunk]
See Figure 1 of the original paper (link below) for the authors' architecture diagram with adaptive ordering.
Existing non-AR VLA decoders β most notably flow matching in the Ο0 family β attach a separate continuous action expert outside the VLM backbone. This causes:
- Parameter / interface mismatch between the VLM (token-trained) and the action head (vector-field-trained)
- A fixed denoising trajectory: no adaptive decoding order
- No error correction: once a denoising step commits, the model cannot revisit
Single-transformer policy. Built on the OpenVLA architecture (Prismatic-7B VLM with SigLIP+DINOv2 visual encoders and a Llama-2 backbone); the original autoregressive causal-attention backbone is converted into a bidirectional transformer so discrete diffusion over actions runs inside the same VLM. Continuous controls (translation, rotation, gripper) are discretized with a 256-bin quantile-based scheme (1stβ99th percentile) β 7 tokens per timestep β and grouped into chunks. These tokens are generated by masked discrete diffusion inside the same transformer as the VLM, trained with the VLM's own cross-entropy objective. At inference it runs a small fixed number of refinement rounds (12 by default) with a cosine mask schedule. Two tricks:
- Adaptive decoding order β confident tokens unmasked first, hard tokens later
- Secondary re-masking β revisit uncertain earlier predictions across refinement rounds
- LIBERO: 96.3% average success (+0.9% vs OpenVLA-OFT discrete)
- SimplerEnv-Fractal: 71.2% visual matching / 56.9% variant aggregation β 64.1% overall
- SimplerEnv-Bridge: 54.2% overall (+14.7% over Ο0, +6.4% over Ο0-FAST) Beats AR, MLP-head, and continuous-diffusion baselines under matched training conditions, while using fewer NFEs than AR.
Crystallizes the discrete-diffusion VLA design philosophy: unify action generation into the VLM transformer, reuse the VLM training objective, and inherit the entire token-based decoding ecosystem (MaskGIT-style sampling, dLLM tricks). The most compelling 2026 architectural alternative to flow-matching action experts.
- arXiv: https://arxiv.org/abs/2508.20072
- HuggingFace paper page: https://huggingface.co/papers/2508.20072
- OpenReview: https://openreview.net/forum?id=YWeNCMxdhM
For a long-form review with the paper's architecture figure, full LIBERO (96.3%) + SimplerEnv (64.1% Fractal / 54.2% Bridge) numbers, all ablations (adaptive ordering, re-masking, temperature schedule, OOD robustness), and limitations: In-Depth Review of Discrete Diffusion VLA.
For how discrete diffusion compares against AR / flow matching / continuous diffusion / world-model backbones / dual-system / reasoning-augmented / tokenizer-centric / small-efficient / 3D-tactile / hybrid approaches β 10+ categories across ~60 VLA papers: VLA Architectures Review.
- Ο0.6 (the flow-matching baseline this paper challenges)
- Unified Diffusion VLA (joint frame+action denoising)
- dVLA (with multimodal CoT)
- DIVA & Fast-dVLA (latency optimizations)
β Back to ICLR-2026