ICRA 2026 VLA Practicality - Heungwoo/research GitHub Wiki

Rethinking VLA Practicality — Benchmark + Improved Baseline

Venue: ICRA 2026 · Authors: Wenxuan Song, Jiayi Chen, Xiaoquan Sun, Huashuo Lei, Yikai Qin, Wei Zhao, Pengxiang Ding, Han Zhao, Tongxin Wang, Pengxu Hou, Zhide Zhong, Haodong Yan, Donglin Wang, Jun Ma, Haoang Li · arXiv: 2602.22663 Category: Benchmark / robustness for VLA Trend tag: Practicality (latency · data · robustness)

Approach diagram

flowchart LR
  subgraph CEBench["CEBench benchmark"]
    SIM["14.4k sim trajectories<br/>36 tasks"] --> DR["domain randomization<br/>(clutter · lighting · texture · table height)"]
    REAL["1.6k real-world trajectories<br/>8 tasks"] --> DR
    EMB["embodiments:<br/>single-arm · bimanual · mobile bimanual"] --> DR
  end
  DR --> EVAL["evaluate VLAs<br/>(OpenVLA, RDT-1B, TinyVLA, ACT, DP, …)"]
  subgraph BASE["LLaVA-VLA baseline (0.5B)"]
    MV["multi-view images<br/>(1st + 3rd person)"] --> VLM["LLaVA-OneVision-0.5B"]
    PROP["proprioception tokenizer"] --> VLM
    VLM --> ACT["action chunks (size 5)<br/>hybrid nav + manip action space"]
  end
  EVAL --> BASE
Loading

Problem

VLAs are pushed toward ever-larger backbones, costly large-scale pre-training, and single-embodiment scopes — yet these choices are rarely interrogated against practical deployment. The paper organizes its study around three questions: Q1 how much performance actually depends on parameter scale (and which techniques let small models match large ones); Q2 whether pre-training is necessary for small models in a target scenario; and Q3 how to define a unified action space for cross-embodiment manipulation, covering fixed-base and mobile robots. The authors argue existing benchmarks do not jointly stress data efficiency, visual generalization (via domain randomization), and cross-embodiment deployability.

Method

CEBench is a benchmark spanning single-arm, bimanual, and mobile-bimanual embodiments in both simulation and the real world, built with domain randomization (clutter, random lighting, diverse textures, variable table heights) so that seen vs domain-randomized (DR) generalization can be measured directly. It collects 14.4k simulated trajectories across 36 tasks and 1.6k expert-curated real-world trajectories across 8 tasks. Evaluated policies include OpenVLA, RDT-1B, TinyVLA, RoboFlamingo, and generative/diffusion methods alongside ACT and Diffusion Policy baselines.

LLaVA-VLA is the improved baseline: a lightweight VLA built on a pre-trained LLaVA-OneVision-0.5B backbone. It takes multi-view images (first- and third-person, concatenated), adds a proprioception tokenizer, and emits action chunks (chunk size 5). A hybrid action space mixes direction and value tokens so a single policy can switch between navigation and manipulation — making it, per the authors, an end-to-end VLA for mobile manipulation. Training is two-stage: post-training then fine-tuning, with post-training on 8× NVIDIA H100 and fine-tuning on a single NVIDIA 4090 — i.e., consumer-grade-GPU reach.

Results

  • CALVIN: first sub-task success 96.2%, comparable to the 97.4% of its 7B counterpart; last sub-task 50.6%. The 0.5B model thus tracks a 7B model on the entry task.
  • RoboTwin: 40.3% average success on seen tasks vs 28.6% under domain randomization — quantifying the seen→DR generalization gap.
  • Real-world bimanual: 44.2% seen / 30.7% DR average success, on par with or above compared baselines.

(Numbers above are quoted from the paper; latency/throughput figures were not confirmable from the public text and are omitted.)

Significance

The paper reframes VLA progress as a practicality problem rather than a scale race: a 0.5B model with proprioception tokenization, action chunking, and a unified nav+manip action space can rival 7B policies on entry tasks while training on a single 4090. CEBench's explicit seen vs domain-randomized split makes it a natural companion to the robustness/benchmark thread — RobustVLA (perturbation robustness), LIBERO-Plus (factor-controlled generalization probing) — and the cross-embodiment design connects to the taxonomy in VLA Architectures review. For practitioners, the headline is that small + cross-embodiment + DR-tested is a viable axis distinct from raw parameter count.

Links

Related pages

← Back to ICRA-2026

⚠️ **GitHub.com Fallback** ⚠️