ICML 2026 VLA Arena - Heungwoo/research GitHub Wiki

VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models โ€” structured robustness diagnostics for VLA policies

Venue: ICML 2026 (Poster) Category: Benchmark Affiliations: Borong Zhang, Jiahao Li, Jiachen Shen, Yishuai Cai, Yuhao Zhang, Yuanpei Chen, Juntao Dai, Jiaming Ji, Yaodong Yang Traction (2026-06): 17 citations (arXiv)

Problem

Vision-Language-Action (VLA) models report high headline success rates, but those single-number scores conflate genuine generalization with memorization of training distributions. They also obscure whether failures stem from the task structure, the language command, or the visual observation. Existing benchmarks rarely separate these axes or stress-test safety behavior, so it is hard to tell why a model succeeds or fails, or whether one model's ranking advantage is robust or an artifact of a particular difficulty setting.

Method

VLA-Arena is a comprehensive, open-source benchmark built around a structured task design organized along three orthogonal dimensions: task structure, language commands, and visual observations. Rather than a flat list of tasks, the benchmark spans 11 task suites grouped into four diagnostic categories:

  • Safety โ€” whether the policy respects safety constraints,
  • Distractor โ€” robustness to irrelevant objects/clutter,
  • Extrapolation โ€” generalization beyond the training distribution,
  • Long Horizon โ€” multi-step task completion.

In total there are 170 tasks at three difficulty levels (L0โ€“L2). To probe robustness systematically, the framework applies graded language perturbations (W0โ€“W4) and visual perturbations (V0โ€“V4) as diagnostic tools, isolating which input modality drives a failure. The release includes scaled dataset variants (VLA-Arena-S/M/L), an evaluation toolchain, reference models, and a public leaderboard.

flowchart LR
    A[Task suites: 11] --> B{4 categories}
    B --> S[Safety]
    B --> D[Distractor]
    B --> E[Extrapolation]
    B --> L[Long Horizon]
    A --> C[170 tasks ยท L0โ€“L2]
    C --> P1[Language perturb. W0โ€“W4]
    C --> P2[Visual perturb. V0โ€“V4]
    P1 --> R[Robustness diagnostics]
    P2 --> R

Results

Evaluating state-of-the-art VLAs across this structured suite surfaces three consistent failure modes: "memorization over generalization, superficial visual perception, and a neglect of safety constraints." The authors further report that "model rank reversals across L0โ€“L2 validate that each level provides non-redundant insights" โ€” i.e., the relative ordering of models flips between difficulty levels, demonstrating that a single difficulty setting (and a single headline number) would give a misleading picture of comparative model quality.

Significance

VLA-Arena reframes VLA evaluation from a leaderboard chase into a diagnostic exercise. By factoring tasks along structure/language/vision and layering graded perturbations plus an explicit Safety category, it gives the community a shared, reproducible way to attribute failures and to expose brittleness that aggregate success rates hide. The fully open release โ€” datasets, toolchain, models, and leaderboard โ€” lowers the barrier for systematic, comparable robustness reporting across future VLA work.

Links

โ† Back to ICML-2026