Review Discrete Diffusion VLA - Heungwoo/research GitHub Wiki
Paper: Discrete Diffusion VLA: Bringing Discrete Diffusion to Action Decoding in Vision-Language-Action Policies Β· Liang, Li, Yang, Wu, Mao, Nian, Pei, Zhou, Yang, Pang, Mu, Luo Β· ICLR 2026 Β· arXiv 2508.20072 Code: https://github.com/Liang-ZX/DiscreteDiffusionVLA Related summary: Discrete Diffusion VLA
This page is the long-form companion to the Discrete Diffusion VLA summary. Everything is sourced from arXiv 2508.20072v3 and the official code release.
Discrete Diffusion VLA (DDVLA) is the flagship ICLR 2026 paper for the "unified token stream" architectural family β the direct counter-proposal to the Ο-series' separate flow-matching action expert. Three design philosophies competing at ICLR 2026:
| Family | Interface | Objective | Exemplar |
|---|---|---|---|
| Same-stack MoE + prefix KV (Cat B) | Action expert as separate weights, same transformer stack, matched head-dim | Flow matching (continuous) | Ο0.6 Β· Ο0.7 |
| Unified token stream (Cat D) | No interface β actions are masked tokens in the VLM's own vocab | Cross-entropy (discrete) | DDVLA Β· Unified Diffusion VLA Β· dVLA Β· DIVA/Fast-dVLA |
| Embedded dual-system (Cat F) | S1 is last N blocks of S2 | Diffusion MSE + AR CE | Fast-in-Slow |
See Review-VLM-Action-Connection Β§4 (Unified token stream) and Review-VLA-Architecture Β§5.D.
DDVLA brings masked discrete diffusion to VLA action decoding inside a single transformer, eliminating the separate flow/diffusion action expert that the Ο-series depends on. Actions are quantized to 256-bin tokens (7 per timestep: 3 translation + 3 rotation + 1 gripper), treated as extra vocabulary in the VLM, and generated by iterative parallel unmasking with adaptive ordering + secondary re-masking. Headline: 96.3% average LIBERO (vs. Ο0 94.2%); 71.2% SimplerEnv-Fractal visual matching (+9.3 over Ο0-FAST); 54.2% SimplerEnv-Bridge (+6.4). Inference: T=12 forward passes vs. AR's 56, giving 4.7Γ fewer NFEs and 2Γ real latency speedup over OpenVLA while matching optimized continuous-diffusion latency. The key architectural insight: discrete diffusion preserves the VLM's discrete-token interface and cross-entropy training signal, so the action head doesn't corrupt VLM features β Knowledge Insulation becomes unnecessary because there's only one loss.
Every VLA has to answer: how does the VLM talk to the action generator? The 2025 field had two answers:
- Autoregressive (AR) discrete tokens β OpenVLA-style. Same transformer, same CE loss, but sequential decoding (56 forward passes for a 7Γ8 chunk). Slow.
- Continuous diffusion / flow-matching head β Ο-series. Separate transformer in the same stack with matched head-dim; parallel decoding in 5 Euler steps (63 ms on H100). Fast, but the two-loss training corrupts VLM features unless you add Knowledge Insulation to block gradient backflow.
DDVLA proposes a third way:
- Discrete like AR (same CE loss, same vocab interface) β VLM knowledge preserved
- Parallel refinement like diffusion (iterative unmasking) β fast decoding
- Adds adaptive ordering (easy-first, hard-later) and secondary re-masking (revisit uncertain tokens) that AR and BERT-style one-shot parallel can't do

Figure 1 of Liang et al. 2025. Four competing decoding paradigms over action chunks: continuous diffusion (left, sequential Euler steps over continuous vectors), AR (sequential token emission), BERT-style one-shot parallel (parallel but no refinement), and Discrete Diffusion with re-masking (right β parallel iterative refinement; uncertain tokens get re-masked across rounds). Included for scholarly review.

Figure 2 of the paper. Single-transformer VLM backbone (SigLIP+DINOv2 ViT for vision, Llama 2 for language) consumes multi-view RGB + language instruction + masked action tokens. The backbone emits all action tokens in parallel; adaptive decoding (bottom left) ranks tokens by confidence (max or gap) and commits high-confidence ones; secondary re-masking (bottom right) applies two consistency checks (threshold + residual-drop) to re-mask uncertain commitments for the next round.
flowchart LR
V[3 camera images] --> E1[SigLIP+DINOv2 ViT]
L[Language instruction] --> E2[Llama tokenizer]
MA[Masked action tokens<br/>all [MASK] at t=T] --> E3[Action token embeds]
E1 & E2 & E3 --> VLM[Single-transformer VLM<br/>bidirectional attn over action tokens]
VLM --> SEL[Adaptive decoding<br/>max-confidence OR confidence-gap]
SEL --> C[Commit high-confidence]
C --> RM[Secondary re-masking<br/>threshold + residual-drop checks]
RM -- uncertain β [MASK] again --> VLM
RM -- stable β keep --> OUT[Action chunk<br/>H=8 Γ 7 tokens]
OUT -- T=12 rounds total --> ACT[Actions]
- 256-bin quantile-based quantization (1stβ99th percentiles per dimension)
- 7 tokens per timestep: 3 translation + 3 rotation + 1 gripper
- Chunk sizes: H=8 for LIBERO / SimplerEnv-Fractal, H=3 for SimplerEnv-Bridge
- Total sequence length: L = H Γ 7 (so 56 tokens for H=8)
- Backbone: Prismatic-7B (OpenVLA's backbone β SigLIP + DINOv2 ViT for vision, Llama 2 for language)
- Key change from OpenVLA: converts causal attention to bidirectional attention over action tokens so every action token sees every other during refinement
- Single transformer β no separate action expert, no MoE, no cross-attention module
Forward process: at corruption level [MASK]. Training loss is masked cross-entropy computed only on masked positions:
where $\tilde{a}t$ is the corrupted action chunk, $c$ is the context (vision+language), and $p\theta = \text{softmax}(W f_\theta(\tilde{a}_t, c))$.
This is the VLM's native cross-entropy objective. Unlike the Ο-series, there's no separate flow-matching loss β so no gradient backflow problem, no Knowledge Insulation needed.
Round t = T, T-1, β¦, 1:
- Forward pass β get per-token distributions over all masked positions
-
Confidence selection:
-
Max-confidence:
$s_{t,i} = \max_k p_\theta(k \mid a_t, c)$ - Confidence-gap: top1 β top2 probability
-
Max-confidence:
- Commit the top-scoring tokens (step-dependent number)
-
Secondary re-masking β two cheap consistency checks:
-
Threshold check: re-mask if confidence < step-dependent threshold
$\eta_t^{\text{abs}}$ - Residual-drop check: re-mask if confidence degrades relative to previous round or falls outside top-Q
-
Threshold check: re-mask if confidence < step-dependent threshold
- Remaining + re-masked tokens β next round
Refinement rounds: T = 12 (efficiency/accuracy knee point). Temperature schedule: linear decay from Ο=1.0 to Ο=0.
| Component | Value |
|---|---|
| Backbone | Prismatic-7B (SigLIP+DINOv2 ViT + Llama 2) |
| Vision/language encoders | Frozen |
| Action head | Bidirectional attention fine-tuned |
| Batch size | 32 |
| GPU | 4Γ NVIDIA A800 |
| LIBERO Spatial/Object | 150k steps |
| LIBERO Goal/Long | 300k steps |
| SimplerEnv | 100k steps |
| Image resolution | 224Γ224 |
| Default T | 12 refinement rounds |
| Temperature | Linear decay 1.0 β 0 |
| Suite | DDVLA | OpenVLA-OFT (discrete, same tokens) | Ο0 (flow matching) | Diffusion Policy |
|---|---|---|---|---|
| Spatial | 97.2% | 96.2% | β | β |
| Object | 98.6% | 98.2% | β | β |
| Goal | 97.4% | 95.6% | β | β |
| Long | 92.0% | 92.0% | β | β |
| Average | 96.3% | 95.5% | 94.2% | 72.4% |
- +0.8% over OpenVLA-OFT (same tokenization, different decoder) β isolates the discrete-diffusion mechanism from tokenization
- +2.1% over Ο0 (flow matching)
- +23.9% over vanilla Diffusion Policy
| Metric | DDVLA | Ο0-FAST | Ο0 |
|---|---|---|---|
| Visual matching | 71.2% | 61.9% | 58.8% |
| Variant aggregation | 56.9% | 59.0% | β |
| Overall average | 64.1% | β | β |
Per-task visual matching: Pick Coke 85.4%, Move Near 67.5%, Drawer 60.6%.
| Metric | DDVLA | Ο0-FAST | Ο0 |
|---|---|---|---|
| Overall average | 54.2% | 48.3% | 40.1% |
- +5.9% over Ο0-FAST, +14.1% over Ο0
- Put-eggplant grasp rate: 20.8% Β· success: 70.8%
DDVLA is evaluated against all four decoder families:
- Autoregressive (OpenVLA) β DDVLA wins on speed + accuracy
- Continuous diffusion (Ο0, Dita) β DDVLA wins on accuracy, ties on latency
- BERT-style one-shot parallel (OpenVLA-OFT) β DDVLA wins by +0.9% average
- Flow matching (GR00T-N1) β DDVLA wins on the same benchmarks
| Strategy | Accuracy |
|---|---|
| Parallel one-shot (BERT-style baseline) | 95.6% |
| + Random order | 95.8% |
| + Confidence-gap selection | 96.6% |
| + Max-confidence selection | 97.0% |
| + Secondary re-masking | 97.4% (+1.8 from baseline) |
Every component matters. Random order helps marginally; confidence-based ordering is the big jump; re-masking adds the final 0.4%.
| Schedule | Accuracy |
|---|---|
| Hard argmax (Ο=0) | 96.2% |
| Fixed Ο=1.0 | 96.4% |
| Linear decay 1.0 β 0 | 97.4% |
Anneal rather than clip β +1.2% over argmax.
- Diminishing returns beyond T=12
- Most accuracy at Tβ[8,12]
- T=12 is the efficiency/accuracy knee point
| Model | Language-aug degradation | Vision-aug degradation |
|---|---|---|
| Parallel one-shot (OpenVLA-OFT) | β8.0% | β22.6% |
| Continuous diffusion | β2.4% | β29.0% |
| DDVLA | β1.4% | β21.0% |
DDVLA's discrete-token interface preserves pretrained VLM priors better under distribution shift β the clearest numerical argument for the unified-token-stream thesis.
| Model | NFEs | Latency (H800) | Throughput |
|---|---|---|---|
| OpenVLA (AR) | 56 (L=HΒ·D) | 136.2 ms/chunk | β |
| DDVLA | 12 | 68.8 ms/chunk | 14.53 Hz |
| OpenVLA-OFT (continuous diffusion, 12 steps) | 12 | 67.1 ms/chunk | β |
| Parallel one-shot (BERT-style) | 1 | 31.1 ms | β |
- 4.7Γ fewer NFEs than AR β 2Γ real latency speedup
- Matches optimized continuous-diffusion latency (67 ms vs. 69 ms)
- Secondary re-masking overhead: <1 ms (negligible)
- Parallel one-shot is 2Γ faster but loses 1.8% accuracy
DDVLA sits on the Pareto frontier: same latency as continuous diffusion, better accuracy, no gradient-insulation headache.
DDVLA's empirical wins rest on four structural properties that AR, BERT-style parallel, and continuous diffusion can't combine:
- Parallel but iterative. AR is sequential. BERT-style is parallel but one-shot. Continuous diffusion is iterative but on vectors. DDVLA is parallel AND iterative over discrete tokens β the only way to get both.
- Adaptive order. AR forces left-to-right. BERT forces all-at-once. DDVLA lets the model decide which tokens to commit first based on its own confidence β instance-wise routing that matches the actual difficulty structure of the action chunk.
- Re-masking = error correction. Once AR or BERT commits a token, it can't be revisited. DDVLA's secondary re-masking lets uncertain commitments be re-examined in later rounds β the only category-D mechanism with error correction.
- Single loss, single interface. The VLM's CE loss trains every parameter. No separate flow-matching objective, no gradient backflow, no Knowledge Insulation needed. This is why DDVLA preserves VLM knowledge better under OOD.
The paper does not have a dedicated limitations section. Concerns worth flagging:
- Fixed chunk length. H must be chosen in advance (H=8 for LIBERO, H=3 for Bridge). Not as flexible for variable-horizon tasks.
- Refinement is required. T=12 forward passes vs. one-shot parallel's 1 β DDVLA is 2Γ slower than BERT-style but 1.8% more accurate. That accuracy/latency trade-off is real.
- OOD gap is not zero. β21% vision-aug degradation is better than baselines but still large in absolute terms.
- Discrete β 256 bins. Fine-grained manipulation (sub-millimeter precision) could hit the quantization ceiling.
- No head-to-head with Ο0.6 or Ο0.7. The paper compares to Ο0 and Ο0-FAST (2024β25 era) but not the Nov 2025 Ο0.6 (Knowledge Insulation + Gemma3-4B) or Apr 2026 Ο0.7. The Ο-series has moved on; the numerical comparison isn't current.
- No real-robot deployment. All results are sim (LIBERO + SimplerEnv). Ο0.6/Ο0.7 claim production-grade real-robot performance; DDVLA hasn't demonstrated that yet.
- Chunk-size hyperparameter leakage. H=8 for LIBERO/Fractal, H=3 for Bridge β this is a per-benchmark tuning that's not always transparent.
- The accuracy wins are small (+0.8% over OpenVLA-OFT with matched tokenization). Most of the gain over Ο0 comes from tokenization + backbone differences, not from the discrete-diffusion mechanism itself.
- Closed-weights competitors (Ο0.6, Ο0.7) have better real-world evidence. Open code β winning in production.
- Fixed denoising schedule. T=12 and linear temperature decay are not adaptive to task difficulty β DexterityGen-style task-adaptive inference is absent.
DDVLA is the strongest empirical case for "unified token stream" (Category D). Together with Unified Diffusion VLA (joint frame+action denoising), dVLA (with text CoT channel), and DIVA & Fast-dVLA (latency optimizations), it defines the ICLR 2026 alternative to the Ο-series flow-matching paradigm.
Three structural contributions:
- Discrete diffusion as an action decoder β not a minor tokenizer tweak, a genuine new decoder family at the mechanism level
- Adaptive ordering + re-masking β the first category-D system with error correction across refinement rounds
- Empirical parity with Ο0 on continuous-control benchmarks while keeping the unified-transformer simplicity
What DDVLA does not prove: that discrete diffusion can match Ο0.6/Ο0.7 / Gemini Robotics on real-robot production benchmarks. That comparison hasn't been run. The 2026 flow-matching vs. discrete-diffusion debate is still open β DDVLA's case rests on LIBERO/SimplerEnv, while the Ο-series case rests on "cleans real homes / folds UR5e laundry zero-shot."
Placement (see Review-VLA-Architecture Β§5.D and Review-VLM-Action-Connection Β§4): DDVLA is the canonical ICLR 2026 Category D exemplar. Any subsequent discrete-diffusion VLA paper that doesn't beat DDVLA on LIBERO is not seriously in the race.
- Summary page: Discrete Diffusion VLA
- Architecture family siblings: Unified Diffusion VLA (+ future-frame co-denoising), dVLA (+ text CoT), DIVA & Fast-dVLA (latency optimizations)
- The competing paradigm: Ο0.6 Β· Ο0.7 β separate flow-matching action expert with Knowledge Insulation
- Per-paper reviews: Review-pi06 Β· Review-pi07 Β· Review-Fast-in-Slow Β· Review-VLM4VLA
- Cross-paper reviews: Review-VLA-Architecture Β§5.D (Unified token stream) Β· Review-VLM-Action-Connection Β§4 (Unified token stream mechanism)
- arXiv: https://arxiv.org/abs/2508.20072 Β· HTML: https://arxiv.org/html/2508.20072
- OpenReview: https://openreview.net/forum?id=YWeNCMxdhM
- Hugging Face paper: https://huggingface.co/papers/2508.20072
- Code: https://github.com/Liang-ZX/DiscreteDiffusionVLA/tree/libero
- Authors: Liang, Li, Yang, Wu, Mao, Nian, Pei, Zhou, Yang, Pang, Mu, Luo Β· Affiliations: HKU, Shanghai AI Laboratory, Shanghai Jiao Tong University, Huawei Cloud
β Back to ICLR-2026-Discrete-Diffusion-VLA Β· ICLR-2026 Β· Home