CoRL 2025 DexVLA - Heungwoo/research GitHub Wiki
Venue: CoRL 2025 ยท arXiv: 2502.05855 Category: VLA Architecture Trend tag: Separate action expert (the pattern ICLR 2026 challenges)
flowchart LR
V[Vision] --> VLM[Qwen2-VL 2B]
L[Language] --> VLM
VLM -. features .-> DE[Plug-in ScaleDP<br/>~1B Diffusion<br/>Action Expert]
noise --> DE
DE --> A[Action chunk<br/>cross-embodiment]
VLAs have generally either (a) produced actions autoregressively from the VLM directly (slow at scale) or (b) used small MLP/diffusion heads (no generative capacity). A large, expressive action generator decoupled from the VLM could marry both.
DexVLA attaches a ~1B-parameter diffusion action expert (ScaleDP โ a transformer-based scaling of Diffusion Policy: 32 layers, hidden 1280, 16 heads) as a plug-in module to a pretrained Qwen2-VL 2B backbone. The expert is trained across robot embodiments (single-arm, bimanual, dexterous hands), so a single expert drives different hardware. A smaller 410M ScaleDP variant is offered to save memory.
Training uses a three-stage embodiment curriculum:
- Cross-embodiment pre-training of the diffusion expert alone on diverse robot data (using DistilBERT + ViT for conditioning, no VLM).
- Embodiment-specific alignment โ jointly train VLM, projection layers, and diffusion expert on single-embodiment data, with the VLM's visual encoder frozen.
- Task-specific adaptation on downstream demonstrations, with sub-step language reasoning generated as output.
Strong cross-embodiment performance on long-horizon, dexterous tasks (laundry folding from crumpled clothing, dryer unloading, bin picking, table bussing, drink pouring with a dexterous hand) using only direct language prompting. Reported scores include 0.92 on shirt folding and 0.90 on a novel embodiment with 100 demonstrations, substantially exceeding Octo, OpenVLA, and Diffusion Policy. Establishes that billion-scale action generators are feasible and beneficial.
Canonical example of the "VLM + separate action expert" pattern โ the same architecture ฯ0/ฯ0.5/ฯ0.6/ฯ0.7 use. Exactly this pattern is what ICLR 2026's discrete-diffusion VLAs explicitly challenge by unifying action generation back into the same transformer as the VLM:
- ฯ0.5 (the sibling flow-matching instance)
- Survey (CoRL 2025)
โ Back to CoRL-2025