ICML 2026 Dismantling the Illusion of Vision Language Action - Heungwoo/research GitHub Wiki
Dismantling the Illusion of Vision-Language-Action Models Competence via Explicit Distributional Shifts — LIBERO-Gen diagnostic benchmark
Venue: ICML 2026 (Poster) Category: Benchmark / Analysis-Insight Authors: Xueyang Zhou, Yangming Xu, Guiyao Tie, Yongchao Chen, Chaoran Hu, Bo Tao, Xingwei Zhao, Xiang Xiang, Pan Zhou, Lichao Sun
Vision-Language-Action (VLA) models post strong numbers on standard manipulation benchmarks, yet generalize poorly in the real world. The authors argue this apparent competence is partly an illusion arising from two complementary pathologies: spurious invariance to semantic changes (the model ignores meaningful changes in the instruction or task) and extreme brittleness to trivial environmental perturbations (small visual changes break it). Conventional aggregate metrics mask both failure modes, because they are driven by intuition-based heuristics rather than explicit assumptions about which distribution shift is being tested.
The paper introduces LIBERO-Gen, a diagnostic benchmark that restructures VLA evaluation around explicit distributional assumptions rather than ad-hoc heuristics. It uses a hierarchical testing protocol that separates evaluation into three tiers:
- In-distribution (ID): standard conditions matching training.
- Compositional generalization (CG): novel recombinations of seen elements — reported along Spatial and Task axes (Spatial-CG, Task-CG).
- Domain generalization (DG): shifts in the underlying domain / environment.
By disentangling these tiers, the benchmark exposes whether a model's success reflects genuine task understanding or memorization of fixed correlations. The authors also study sampling strategies for building more robust policies, validating a structured "Stair" sampling scheme as an effective technique for improving robustness under these explicit shifts.
flowchart TD
A[VLA model] --> B[LIBERO-Gen protocol]
B --> C[In-distribution]
B --> D[Compositional generalization<br/>Spatial-CG / Task-CG]
B --> E[Domain generalization]
C & D & E --> F[Diagnose: spurious invariance<br/>vs. environmental brittleness]
Across the evaluated VLA models, Pi0.5 is the top performer, reaching 64.0% on Spatial-CG and 21.2% on Task-CG — the large gap between the two illustrating how much harder compositional task generalization is than spatial recombination. The analysis attributes the dominant failures to perceptual instability and action binding collapse (the policy losing the correct mapping from grounded perception to action). Structured "Stair" sampling is shown to improve robustness relative to standard sampling under the benchmark's explicit shift conditions.
LIBERO-Gen reframes VLA evaluation from "does it score well on average" to "under exactly which distributional assumption does it succeed or fail." By making the shift type explicit and tiered, it surfaces spurious invariance and brittleness that aggregate benchmarks hide, and provides a diagnostic lens (perceptual instability, action binding collapse) plus a concrete mitigation (Stair sampling). For a field where headline success rates can overstate real competence, this offers a more honest measurement substrate for comparing VLA models.
- ICML 2026: https://icml.cc/virtual/2026/poster/64080
← Back to ICML-2026