CVPR 2026 LIBERO Plus - Heungwoo/research GitHub Wiki
LIBERO-Plus โ Progressive Robustness Benchmark for Visual-Language-Action Models
Venue: CVPR 2026 (Poster #38735) Category: Benchmark / Robustness Trend tag: Trend 6 (robustness backlash) Affiliations: Fudan + Shanghai AI Lab
Approach diagram
flowchart LR
LIB["LIBERO base benchmark"] --> PERT["7-factor perturbation framework"]
PERT --> AX1["light conditions"]
PERT --> AX2["camera viewpoints"]
PERT --> AX3["background textures"]
PERT --> AX4["objects layout"]
PERT --> AX5["language instructions"]
PERT --> AX6["robot initial states"]
PERT --> AX7["sensor noise"]
AX1 --> EVAL["evaluate top VLAs"]
AX2 --> EVAL
AX3 --> EVAL
AX4 --> EVAL
AX5 --> EVAL
AX6 --> EVAL
AX7 --> EVAL
EVAL --> RES["95% โ less than 30%"]
Problem
Standard LIBERO numbers are at saturation โ top VLAs all report 95 %+. But the benchmark conditions are nearly identical to training. The question: how brittle are these models under modest, realistic perturbations?
Method
An automated 7-factor perturbation framework that systematically varies objects layout, camera viewpoints, robot initial states, language instructions, light conditions, background textures, and sensor noise in LIBERO. Tasks are stratified into five difficulty levels, producing a "progressive robustness" suite of 10,030 tasks.
Results
Across nine models (OpenVLA and its OFT variants, ฯโ and ฯโ-fast, Nora, WorldVLA, UniVLA, RIPT-VLA), top VLAs drop from 95 %+ to under 30 % under modest perturbations. The most damaging factors are camera viewpoints and robot initial states, not the semantic/visual axes one might expect.
A striking finding: models are largely insensitive to language variations โ further experiments show the models tend to ignore language instructions almost entirely, relying instead on visual/spatial priors. This is arguably the paper's sharpest indictment of current VLAs as "vision-action" rather than true "vision-language-action" models.
Significance
LIBERO-Plus is the single most important benchmark paper at CVPR 2026 for the practical state of VLAs. It calls out benchmark inflation explicitly and provides the tooling to re-evaluate every 2024โ2026 paper. Expect a wave of follow-up work re-reporting numbers on the perturbed version. The clearest 2026 echo of RoboArena-style honest evaluation.
Links
- arXiv: 2510.13626 (submitted Oct 2025, rev. Dec 2025)
- Code: github.com/sylvestf/LIBERO-plus
- Project page: sylvestf.github.io/LIBERO-plus
- CVPR Poster: #38735
Related pages
โ Back to CVPR-2026