ICLR 2026 HybridVLA - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 Category: VLA Architecture — Hybrid action decoder Trend tag: AR ↔ diffusion fusion Affiliations: Peking University (State Key Lab of Multimedia Information Processing), BAAI, CUHK
flowchart LR
V[Multi-view images + language] --> Enc[DINOv2 + SigLIP<br/>or CLIP for 2.7B]
State[Robot state] --> MLP1[Learnable MLP]
Enc --> Tok[Token sequence]
MLP1 --> Tok
Tok --> LLM[Shared LLM backbone<br/>LLaMA-2 7B / Phi-2 2.7B]
LLM --> BOD[BOD marker]
BOD --> Diff[Diffusion tokens<br/>DDIM denoising in LLM]
Diff --> EOD[EOD marker]
EOD --> AR[Autoregressive<br/>discrete action tokens]
Diff --> dAct[Diffusion action a^d]
AR --> arAct[AR action a^ar + confidence c^ar]
dAct --> Ens["Collaborative Action Ensemble<br/>if c^ar > 0.96 → average, else diffusion only"]
arAct --> Ens
Ens --> A[Action]
Autoregressive VLAs (OpenVLA, RT-2, ManipLLM) discretize 7-DoF actions into bins to inherit the LLM's pretrained next-token-prediction paradigm — efficient and scalable, but quantization disrupts the continuity of SE(3) poses and degrades fine-grained control. Diffusion-head VLAs (π0, CogACT, RDT-1B, DiVLA) attach a separate denoising head after the VLM and predict continuous actions — precise, but the head sits outside the LLM, treating the VLM as a frozen feature extractor and forfeiting the pretrained iterative-generation mechanism. HybridVLA asks whether both paradigms can run inside one shared LLM backbone without mutual interference, and whether their predictions reinforce rather than redundantly cover each other.
- Two model sizes: HybridVLA-7B uses LLaMA-2 7B as the LLM with combined DINOv2 (B×Nv×1024) + SigLIP (B×Nv×1152) vision encoders concatenated channel-wise to give B×Nv×2176 vision tokens. HybridVLA-2.7B uses Phi-2 with CLIP only.
- Initialization: Both inherit Prismatic VLM pretrained parameters (Karamcheti et al., 2024).
- Robot state injection: instead of discretizing state into the language prompt (Type 1 ablation, weaker), a learnable MLP maps the robot state directly into the LLM embedding space at fr ∈ R^{B×1×4096}.
- Action representation: end-effector pose, 7-DoF for single-arm (Δx, Δy, Δz, roll, pitch, yaw, gripper ∈ {0,1}); 14-DoF for dual-arm.
The paper compares four orderings (Table 1 of paper). The chosen Type 1 ("Ours") places diffusion tokens before autoregressive tokens, bracketed by special <BOD> (begin-of-diffusion) and <EOD> (end-of-diffusion) markers, with the robot state injected via a learnable MLP rather than as discrete bins in the language prompt (the latter is Type 3). Three justifications:
- Putting AR tokens first (Type 4) risks GT leakage during training because diffusion modeling conditions on all preceding tokens — and AR tokens contain the discrete action GT.
- Diffusion tokens leak no information because they operate on noise.
- Diffusion-token features serve as continuous latent conditions for the subsequent AR prediction, improving AR quality vs. standalone autoregression.
Type 1 outperforms Types 2/3/4 on both the diffusion-only metric (0.66 vs Type2 0.56 / Type3 0.61 / Type4 0.57) and the AR-only metric (0.62 vs 0.54 / 0.59 / 0.60).
Joint loss L_hybrid = L_dif + L_ar where
-
L_dif = E ||ε − ε_π(a^i_t, i, c)||^2— standard diffusion-policy MSE between sampled Gaussian noise and predicted noise. -
L_ar = cross-entropyover discrete action tokens (vocabulary partially overwritten with action bins per OpenVLA).
The two heads target the same normalized action distribution; the discrete bins are simply a quantized representation of that distribution. The paper uses an unweighted sum L_hybrid = L_dif + L_ce; classifier-free guidance is deliberately not used (for stable arm behavior).
- Pretraining: 5 epochs across 35 datasets, 760K trajectories, 33M frames (OXE, DROID, ROBOMIND etc.), using only single 2D observations. Reportedly >10K A800 GPU-hours.
- Fine-tuning: on self-collected simulation/real data, single- or multi-view depending on task. RGB resized to 224×224. 300 epochs on downstream tasks, mixed precision.
- Diffusion path: DDIM with n=4 steps (the paper reduces denoising steps from 30 to 4 with no significant accuracy degradation; Figure 4). A novel diffusion KV cache caches keys/values before the diffusion tokens, forwarding conditional information, the denoising timestep, and pure noise only on the first sampling step. Removing the cache drops 9.4 Hz → 5.0 Hz at the same accuracy.
-
AR path: runs after
<EOD>, conditioned on the just-generated diffusion tokens (this is what makes HybridVLA-ar stronger than vanilla OpenVLA-style AR). -
Ensemble rule: let
c^ar_{t+1}= mean confidence of the AR action tokens. Ifc^ar > θ(threshold θ=0.96), output(a^d + a^ar) / 2; otherwise fall back toa^dalone. Ablation (Table 8) on θ ∈ {0.90, 0.92, 0.94, 0.96, 0.98} → success 0.66 / 0.64 / 0.70 / 0.74 / 0.69, confirming 0.96 is optimal: below 0.94 the AR predictions become unreliable, while at 0.98 too few AR votes pass the gate so the ensemble collapses toward diffusion-only.
RLBench (10-task multi-task, single front-view, Franka, 100 trajectories per task, 20 rollouts × 3 seeds)
| Model | Mean S.R. | Var. | Speed |
|---|---|---|---|
| ManipLLM (7B) | 0.38 | ±0.042 | 2.2 Hz |
| OpenVLA (7B) | 0.41 | ±0.038 | 6.3 Hz |
| π0 (2.6B) | 0.55 | ±0.035 | 13.8 Hz |
| CogACT (7B) | 0.60 | ±0.041 | 9.8 Hz |
| HybridVLA-dif (7B) | 0.66 | ±0.040 | 9.4 Hz |
| HybridVLA (2.7B) | 0.58 | ±0.031 | 12.3 Hz |
| HybridVLA (7B) | 0.74 | ±0.037 | 6.1 Hz |
(HybridVLA-ar (7B) = 0.62 is reported separately in the per-task Table 7, not in the main RLBench Table 2.)
Per-task highlights for HybridVLA-7B (Table 2): Close box 0.85 / Close laptop lid 0.95 / Toilet seat down 1.00 / Sweep to dustpan 0.90 / Close fridge 1.00. Hardest tasks: Phone on base 0.50, Umbrella out 0.50, Wine at rack 0.50, Water plants 0.50, Frame off hanger 0.70.
| Model | Single-arm mean | Dual-arm mean |
|---|---|---|
| π0 (2.6B) | 0.45 | 0.55 |
| CogACT (7B) | 0.61 | – (no multi-view) |
| HybridVLA-dif (7B) | 0.80 | 0.66 |
| HybridVLA (7B) | 0.83 | 0.71 |
CogACT lacks multi-view support, so the dual-arm comparison is against π0 only. Standout single-arm wins: Pour water 0.80 vs π0 0.45 (+35%); Unplug charger 0.95 vs CogACT 0.70; Pick and place 0.90.
The paper reports two tasks × four unseen conditions (object / background / height / lighting). Per-scenario success and drop vs. the original setting:
| Scenario | HybridVLA (single-arm) | CogACT | HybridVLA (dual-arm) | π0 |
|---|---|---|---|---|
| Original | 0.90 | 0.80 | 0.80 | 0.65 |
| Unseen object | 0.60 (−33%) | 0.45 (−43%) | 0.75 (−6%) | 0.60 (−8%) |
| Unseen background | 0.80 (−11%) | 0.50 (−37%) | 0.60 (−25%) | 0.50 (−23%) |
| Unseen height | 0.75 (−17%) | 0.50 (−37%) | 0.60 (−25%) | 0.45 (−31%) |
| Unseen lighting | 0.70 (−22%) | 0.60 (−25%) | 0.75 (−6%) | 0.55 (−15%) |
HybridVLA degrades less than both baselines in nearly every condition; the largest absolute baseline collapses are CogACT under unseen object/background/height (−37% to −43%).
Components: AR / Dif (the two generation paths), LSP (large-scale pretraining on the assembled robotic datasets), RSE (injected robot-state embedding), CTR (collaborative training recipe with the hybrid objective L_hybrid), CAE (collaborative action ensemble).
| Cfg | Description | Mean |
|---|---|---|
| Ex0 | Full HybridVLA (AR + Dif + LSP + RSE + CTR + CAE) | 0.74 |
| Ex1 | HybridVLA-dif (diffusion-only inference) | 0.66 |
| Ex2 | autoregressive-only branch | 0.60 |
| Ex3 | HybridVLA-ar (AR-only inference) | 0.62 |
| Ex4 | diffusion path trained without CTR | 0.57 |
| Ex5 | no large-scale pretraining (no LSP) | 0.22 |
| Ex6 | no robot-state embedding (no RSE) | 0.68 |
Large-scale pretraining is the single largest contributor: removing it (Ex5) collapses success to 0.22 (vs 0.74 full), even though the VLM is still initialized from pretrained Prismatic weights. The collaborative training recipe lets the two paradigms reinforce rather than interfere (compare HybridVLA-dif 0.66 / HybridVLA-ar 0.62 under CTR against their individually trained counterparts), and the full ensemble (0.74) beats diffusion-only (0.66) and AR-only (0.62). Robot-state embedding adds a modest +0.06 (Ex0 0.74 vs Ex6 0.68).
- DDIM steps: reduced from 30 to 4 with no significant degradation (Figure 4); 4 chosen.
- KV-cache for diffusion: 9.4 Hz → 5.0 Hz without it, at the same ~0.66 success — the first integration of a KV cache into an LLM's diffusion-based action generation.
- Confidence threshold (Table 8): 0.90→0.66, 0.92→0.64, 0.94→0.70, 0.96→0.74, 0.98→0.69.
The conclusion explicitly acknowledges one limitation: inference speed is bottlenecked by autoregressive generation, similar to all AR-based VLAs (OpenVLA, RT-2, ManipLLM). The mitigation is that diffusion-only mode (HybridVLA-dif at 9.4 Hz) is competitive with the full ensemble at 6.1 Hz, so for latency-critical deployment the AR vote can be dropped. Appendix D additionally documents three real-world failure categories: (1) rotational prediction deviations (e.g. accumulated/incorrect rotation in Pour water, Place bottle at rack), (2) poses exceeding the robot's DoF / workspace limits or kinematically infeasible configurations, and (3) dual-arm coordination failures where one arm's interaction changes the object state and invalidates the other arm's already-predicted action.
HybridVLA stakes out a position in the AR vs. diffusion architectural debate that has shaped most of ICLR 2026's VLA submissions:
- Against pure-AR + parallel decoding (OpenVLA-OFT, Discrete Diffusion VLA): HybridVLA agrees that AR alone hurts continuity but recovers it by running diffusion inside the same LLM rather than abandoning AR.
- Against flow/diffusion + frozen VLM (π0.6, CogACT, RDT-1B, GR00T N1.5): the paper argues these waste the LLM's pretrained iterative-generation knowledge by relegating the VLM to a feature extractor. HybridVLA-dif's +0.06 over CogACT and +0.11 over π0 on RLBench (0.66 vs 0.60 and 0.55, without the AR vote) is the empirical evidence for this claim — same diffusion target, but with the LLM as the denoiser.
- Versus OneTwoVLA (which unifies thinking and acting as alternating modes) HybridVLA pursues a different unification axis: not reasoning vs. acting but two action-decoding paradigms in the same forward pass.
The collaborative ensemble's confidence-gated rule is also a soft form of self-verification: the AR branch's token confidence acts as a learned veto on diffusion's continuous output, shipping a kind of policy-uncertainty signal almost for free. This is reminiscent of Action-Aware Pruning but on the head side rather than the encoder.
- arXiv: https://arxiv.org/abs/2503.10631 (v3, 23 Jun 2025)
- Project page: https://hybrid-vla.github.io
- OpenReview (ICLR 2026): https://openreview.net/forum?id=H1KDMNOKQn
- PDF: https://openreview.net/pdf?id=H1KDMNOKQn
- Discrete Diffusion VLA — discrete-token denoising as a different unification
- π0.6 — flow-matching baseline this paper contrasts with
- dVLA
- OneTwoVLA — a different "unified" axis: reasoning vs. acting
- FASTER — parallel decoding for AR VLAs
← Back to ICLR-2026