ICLR 2026 Reflective DDVLA - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 Category: VLA Architecture — Discrete diffusion / Autonomous driving Trend tag: Diffusion · Reflective inference · Safety constraints Affiliations: LiAuto, Tsinghua University (Li, Zheng, Wang, Wang, Zhao, Liu, Zhan, Zhan, Lang)
flowchart LR
Obs[Multi-view cameras<br/>front, front-left, front-right<br/>+ ego state + nav instruction] --> Tok[Trajectory tokenizer<br/>uniform 1D codebook per axis]
Tok --> DLM[Discrete Diffusion LM<br/>LLaDA-V backbone, fine-tuned]
DLM --> Goal[Goal-Conditioned Generation<br/>Top-K' goals, NMS dNMS = 0.9 m, K = 3]
Goal --> Cand[K candidate trajectories]
Cand --> Sg[Global Scorer Sglobal]
Sg --> Best[Best initial trajectory τ*]
Best --> Loop{Safety-Guided<br/>Regeneration Loop<br/>≤ 10 iterations}
Loop -->|Safety Scorer Ssafe| Viol[Identify violating waypoints V]
Viol -->|Local Scorer Slocal| Anchor[Manhattan local search δ ≤ 10]
Anchor --> Inpaint[Diffusion inpainting<br/>around safe anchors]
Inpaint --> Loop
Loop -->|safe| Out[Final trajectory]
End-to-end driving stacks struggle to encode physical / safety rules during training. Existing fixes are either (1) rule-based post-processing that demands human engineering, (2) RL with simulation-only training and unsafe online rollouts, or (3) gradient-based diffusion guidance that is computationally expensive and brittle around safety constraints. The paper targets a gradient-free way to inject safety into a learned VLA-style planner.
Each 2D waypoint (x, y) is quantized independently with a uniform 1D codebook A = {a₁, a₂, …} over spatial range [-M, M] = [-100, 100] m (Table 4) at resolution Δg; the paper defines Δg symbolically but does not report a numeric value for it or the resulting codebook cardinality. A trajectory of N waypoints is flattened into a length-2N token sequence
y = Q(τ) = (y1,x, y1,y, ..., yN,x, yN,y) ∈ A^{2N}
This makes BEV-space search efficient and supports the inpainting story of discrete diffusion.
Forward process masks tokens with a cosine-style noise schedule. Reverse process is trained with the standard masked-token NLL
L(θ) = E[ -Σ_{i: m_i^(s) = 1} log p_θ(y_i | ỹ^(s), c, s) ]
conditioned on context c = three camera views + ego state + language navigation command. Backbone is LLaDA-V (You et al., 2025), a pretrained Diffusion Language Model. Training uses classifier-free guidance.
Training hyperparameters (Table 4):
| Parameter | Value |
|---|---|
| Dataset | NAVSIM training split (Table 4 does not name the split or sample count) |
| Batch size | 16 |
| Gradient accumulation | 1 |
| Learning rate | 1 × 10⁻⁵ |
| Epochs | 3 |
| Max context length | 8192 |
| LR scheduler | Cosine, warmup ratio 0.03 |
| Weight decay | 0.0 |
| Precision | bfloat16 |
Inference config (Table 3): 5 diffusion steps, answer length 32, block length 32, low-confidence remasking, K = 3 goal candidates, NMS distance threshold 0.9 m, max refinement iterations 10.
Three scoring functions (implementation in Appendix C):
- Sglobal(τ): overall trajectory quality including safety + coherence; returns 0 if any critical rule is violated.
- Ssafe(τ): safety oracle that scores each waypoint by worst-violation-in-local-window.
- Slocal(ax, ay): per-token-pair scorer used during local search.
Stage 1 — Goal-Conditioned Generation. The model produces a distribution over the terminal waypoint pθ(yN | c, s). Top-K' goal candidates are extracted, then Non-Maximum Suppression with threshold dNMS = 0.9 m returns K = 3 spatially diverse goals G = {G1, G2, G3}. For each goal Gk, a full trajectory is generated by inpainting the intermediate waypoints conditioned on Gk. The K trajectories are scored with Sglobal; the highest-scoring τ* proceeds to Stage 2.
Stage 2 — Safety-Guided Regeneration. Iterative loop:
- Ssafe identifies violating waypoint indices
V = {t | Ssafe(τ*)t < τ_safe}. - For each
t ∈ T* ⊆ V, search the Manhattan neighborhoodNδ(δ ≤ 10 tokens) for the corrected pair maximizing Slocal — these become safety anchors. - Diffusion inpainting regenerates the surrounding tokens conditioned on the safety anchors.
- Repeat until V is empty or budget = 10 iterations is hit; fall back to highest-Ssafe candidate otherwise.
Empirically most violations are resolved in 1–3 iterations (the paper does not report a numeric average).
PDMS aggregates NC (no-collision), DAC (drivable-area compliance), TTC (time-to-collision), Comfort, EP (ego progress).
| Method | Paradigm | Input | NC↑ | DAC↑ | TTC↑ | Comf.↑ | EP↑ | PDMS↑ |
|---|---|---|---|---|---|---|---|---|
| UniAD | E2E | Cam | 97.8 | 91.9 | 92.9 | 100.0 | 78.8 | 83.4 |
| PARA-Drive | E2E | Cam | 97.9 | 92.4 | 93.0 | 99.8 | 79.3 | 84.0 |
| Transfuser | E2E | C+L | 97.7 | 92.8 | 92.8 | 100.0 | 79.2 | 84.0 |
| Hydra-MDP | Augmented | C+L | 98.3 | 96.0 | 94.6 | 100.0 | 78.7 | 86.5 |
| DiffusionDrive | Diffusion | C+L | 98.2 | 96.2 | 94.7 | 100.0 | 82.2 | 88.1 |
| GoalFlow | Diffusion | C+L | 98.4 | 98.3 | 94.6 | 100.0 | 85.0 | 90.3 |
| AutoVLA (Post-RFT) | Autoregressive VLA | Cam | 98.4 | 95.6 | 98.0 | 99.9 | 81.9 | 89.1 |
| ReflectDrive (w/o R.I.) | Discrete Diffusion | Cam | 96.9 | 95.4 | 92.2 | 100.0 | 79.0 | 84.8 |
| ReflectDrive (Ours) | Discrete Diffusion | Cam | 97.7 | 99.3 | 93.5 | 100.0 | 86.9 | 91.1 |
| ReflectDrive† (GT agents) | Discrete Diffusion | Cam | 99.7 | 99.5 | 99.1 | 99.9 | 88.9 | 94.7 |
| Human | – | – | 100.0 | 100.0 | 100.0 | 99.9 | 87.5 | 94.8 |
Camera-only ReflectDrive surpasses Camera+LiDAR Augmented planners on PDMS, with the largest margin on DAC = 99.3 (drivable-area compliance), where reflection injects hard safety. With ground-truth-agent oracle (†), the system matches human performance within noise.
Reflection improves over the no-RI baseline by +3.9 DAC, +1.3 TTC, +0.8 NC, +7.9 EP → +6.3 PDMS.
| Goal-Cond. | Safety-Guided | NC↑ | DAC↑ | TTC↑ | Comf.↑ | EP↑ | PDMS↑ |
|---|---|---|---|---|---|---|---|
| ✗ | ✗ | 96.9 | 95.4 | 92.2 | 100.0 | 79.0 | 84.8 |
| ✓ | ✗ | 96.6 | 96.5 | 91.5 | 100.0 | 83.8 | 87.4 |
| ✗ | ✓ | 98.1 | 98.9 | 94.8 | 99.9 | 84.1 | 90.3 |
| ✓ | ✓ | 97.7 | 99.3 | 93.5 | 99.9 | 86.9 | 91.1 |
Goal-Conditioned generation primarily lifts EP (progress); Safety-Guided primarily lifts DAC and TTC. They are complementary.
- Generation steps (Fig. 4a). Non-monotonic: PDMS peaks at 5 steps, then declines.
- Goal candidates K and NMS dNMS (Fig. 4b). Multi-modal goal generation gives a wider option set; scaling K beyond 3 yields diminishing returns.
- Exploration steps and max iterations (Fig. 4c). PDMS scales positively with both → "inference-time scaling" exists for discrete-diffusion planners.
- Single-frame inputs. Three-view images of the current frame only; velocity / motion of surrounding agents is unobserved. Future: incorporate history and predict obstacle trajectories jointly.
-
Reflection design.
- Goal-Conditioned Generation reuses the PDM scorer without task-specific tuning of high-level navigation/efficiency objectives.
- Safety-Guided Regeneration: more iterations don't always help. Failure modes:
- Boundary oscillation in narrow drivable space (discrete-token rounding error compounds).
- Navigation correctness not encoded in the reward.
- Goal-point selection is suboptimal in some scenarios; constrained search range can't recover.
- Sample efficiency. Limited engineering optimization so far; the authors note substantial room to improve inference efficiency (no concrete latency figures are reported in the paper).
Provenance note (important): This is the only "discrete-diffusion VLA" actually at ICLR 2026, and it lives in the autonomous driving family — distinct from the manipulation-focused DDVLA paper at NeurIPS 2025 which the community sometimes conflates with it. The contribution is not "discrete diffusion for VLA" per se; it is the gradient-free reflection / inpainting loop that uses the discrete-token structure to perform local search and regeneration around safety anchors.
Compared to neighbors:
- AutoVLA (Post-RFT): same VLA-for-driving paradigm, autoregressive instead of discrete-diffusion, slightly weaker on PDMS (89.1 vs 91.1) and notably weaker on EP (81.9 vs 86.9). Reflection lets ReflectDrive correct without RL fine-tuning.
- Hydra-MDP / DiffusionDrive / GoalFlow: rely on trajectory anchors or rule-based guidance and Camera+LiDAR. ReflectDrive matches/beats them with camera only.
- Diffusion Planner / classifier-guided diffusion: gradient-based guidance during denoising — expensive and parameter-sensitive. ReflectDrive's discrete inpainting is gradient-free.
- Manipulation discrete-diffusion VLAs (DDVLA NeurIPS 2025): shares the discrete-action-codebook idea but optimizes for sampling efficiency, not safety constraint enforcement.
The broader bet is that discrete action spaces + masked diffusion make it natural to compose learned generators with rule-based search — the same recipe that worked for code (FIM, infilling) now applied to safety-critical planning.
- OpenReview: https://openreview.net/forum?id=XJxXSMLDoZ
- PDF: https://openreview.net/pdf?id=XJxXSMLDoZ
← Back to ICLR-2026