ICRA 2026 VLA Practicality - Heungwoo/research GitHub Wiki
Venue: ICRA 2026 · Authors: Wenxuan Song, Jiayi Chen, Xiaoquan Sun, Huashuo Lei, Yikai Qin, Wei Zhao, Pengxiang Ding, Han Zhao, Tongxin Wang, Pengxu Hou, Zhide Zhong, Haodong Yan, Donglin Wang, Jun Ma, Haoang Li · arXiv: 2602.22663 Category: Benchmark / robustness for VLA Trend tag: Practicality (latency · data · robustness)
flowchart LR
subgraph CEBench["CEBench benchmark"]
SIM["14.4k sim trajectories<br/>36 tasks"] --> DR["domain randomization<br/>(clutter · lighting · texture · table height)"]
REAL["1.6k real-world trajectories<br/>8 tasks"] --> DR
EMB["embodiments:<br/>single-arm · bimanual · mobile bimanual"] --> DR
end
DR --> EVAL["evaluate VLAs<br/>(OpenVLA, RDT-1B, TinyVLA, ACT, DP, …)"]
subgraph BASE["LLaVA-VLA baseline (0.5B)"]
MV["multi-view images<br/>(1st + 3rd person)"] --> VLM["LLaVA-OneVision-0.5B"]
PROP["proprioception tokenizer"] --> VLM
VLM --> ACT["action chunks (size 5)<br/>hybrid nav + manip action space"]
end
EVAL --> BASE
VLAs are pushed toward ever-larger backbones, costly large-scale pre-training, and single-embodiment scopes — yet these choices are rarely interrogated against practical deployment. The paper organizes its study around three questions: Q1 how much performance actually depends on parameter scale (and which techniques let small models match large ones); Q2 whether pre-training is necessary for small models in a target scenario; and Q3 how to define a unified action space for cross-embodiment manipulation, covering fixed-base and mobile robots. The authors argue existing benchmarks do not jointly stress data efficiency, visual generalization (via domain randomization), and cross-embodiment deployability.
CEBench is a benchmark spanning single-arm, bimanual, and mobile-bimanual embodiments in both simulation and the real world, built with domain randomization (clutter, random lighting, diverse textures, variable table heights) so that seen vs domain-randomized (DR) generalization can be measured directly. It collects 14.4k simulated trajectories across 36 tasks and 1.6k expert-curated real-world trajectories across 8 tasks. Evaluated policies include OpenVLA, RDT-1B, TinyVLA, RoboFlamingo, and generative/diffusion methods alongside ACT and Diffusion Policy baselines.
LLaVA-VLA is the improved baseline: a lightweight VLA built on a pre-trained LLaVA-OneVision-0.5B backbone. It takes multi-view images (first- and third-person, concatenated), adds a proprioception tokenizer, and emits action chunks (chunk size 5). A hybrid action space mixes direction and value tokens so a single policy can switch between navigation and manipulation — making it, per the authors, an end-to-end VLA for mobile manipulation. Training is two-stage: post-training then fine-tuning, with post-training on 8× NVIDIA H100 and fine-tuning on a single NVIDIA 4090 — i.e., consumer-grade-GPU reach.
- CALVIN: first sub-task success 96.2%, comparable to the 97.4% of its 7B counterpart; last sub-task 50.6%. The 0.5B model thus tracks a 7B model on the entry task.
- RoboTwin: 40.3% average success on seen tasks vs 28.6% under domain randomization — quantifying the seen→DR generalization gap.
- Real-world bimanual: 44.2% seen / 30.7% DR average success, on par with or above compared baselines.
(Numbers above are quoted from the paper; latency/throughput figures were not confirmable from the public text and are omitted.)
The paper reframes VLA progress as a practicality problem rather than a scale race: a 0.5B model with proprioception tokenization, action chunking, and a unified nav+manip action space can rival 7B policies on entry tasks while training on a single 4090. CEBench's explicit seen vs domain-randomized split makes it a natural companion to the robustness/benchmark thread — RobustVLA (perturbation robustness), LIBERO-Plus (factor-controlled generalization probing) — and the cross-embodiment design connects to the taxonomy in VLA Architectures review. For practitioners, the headline is that small + cross-embodiment + DR-tested is a viable axis distinct from raw parameter count.
- arXiv: 2602.22663 · HTML · PDF
- RobustVLA (perturbation-robustness benchmark thread)
- LIBERO-Plus (factor-controlled generalization probing)
- VLA Architectures review (cross-embodiment / lightweight-backbone context)
- ICRA 2026 Survey
← Back to ICRA-2026