ICML 2026 VLA Arena - Heungwoo/research GitHub Wiki
VLA-Arena: An Open-Source Framework for Benchmarking Vision-Language-Action Models โ structured robustness diagnostics for VLA policies
Venue: ICML 2026 (Poster) Category: Benchmark Affiliations: Borong Zhang, Jiahao Li, Jiachen Shen, Yishuai Cai, Yuhao Zhang, Yuanpei Chen, Juntao Dai, Jiaming Ji, Yaodong Yang Traction (2026-06): 17 citations (arXiv)
Problem
Vision-Language-Action (VLA) models report high headline success rates, but those single-number scores conflate genuine generalization with memorization of training distributions. They also obscure whether failures stem from the task structure, the language command, or the visual observation. Existing benchmarks rarely separate these axes or stress-test safety behavior, so it is hard to tell why a model succeeds or fails, or whether one model's ranking advantage is robust or an artifact of a particular difficulty setting.
Method
VLA-Arena is a comprehensive, open-source benchmark built around a structured task design organized along three orthogonal dimensions: task structure, language commands, and visual observations. Rather than a flat list of tasks, the benchmark spans 11 task suites grouped into four diagnostic categories:
- Safety โ whether the policy respects safety constraints,
- Distractor โ robustness to irrelevant objects/clutter,
- Extrapolation โ generalization beyond the training distribution,
- Long Horizon โ multi-step task completion.
In total there are 170 tasks at three difficulty levels (L0โL2). To probe robustness systematically, the framework applies graded language perturbations (W0โW4) and visual perturbations (V0โV4) as diagnostic tools, isolating which input modality drives a failure. The release includes scaled dataset variants (VLA-Arena-S/M/L), an evaluation toolchain, reference models, and a public leaderboard.
flowchart LR
A[Task suites: 11] --> B{4 categories}
B --> S[Safety]
B --> D[Distractor]
B --> E[Extrapolation]
B --> L[Long Horizon]
A --> C[170 tasks ยท L0โL2]
C --> P1[Language perturb. W0โW4]
C --> P2[Visual perturb. V0โV4]
P1 --> R[Robustness diagnostics]
P2 --> R
Results
Evaluating state-of-the-art VLAs across this structured suite surfaces three consistent failure modes: "memorization over generalization, superficial visual perception, and a neglect of safety constraints." The authors further report that "model rank reversals across L0โL2 validate that each level provides non-redundant insights" โ i.e., the relative ordering of models flips between difficulty levels, demonstrating that a single difficulty setting (and a single headline number) would give a misleading picture of comparative model quality.
Significance
VLA-Arena reframes VLA evaluation from a leaderboard chase into a diagnostic exercise. By factoring tasks along structure/language/vision and layering graded perturbations plus an explicit Safety category, it gives the community a shared, reproducible way to attribute failures and to expose brittleness that aggregate success rates hide. The fully open release โ datasets, toolchain, models, and leaderboard โ lowers the barrier for systematic, comparable robustness reporting across future VLA work.
Links
- ICML 2026: https://icml.cc/virtual/2026/poster/61388
โ Back to ICML-2026