ICML 2026 Can VLMs Diagnose and Recover - Heungwoo/research GitHub Wiki
Can VLMs Diagnose and Recover from VLA Manipulation Faults? — VLA-FixBench, FaultEval, and VLM-VLA recovery
Venue: ICML 2026 (Poster) Category: Benchmark Affiliations: Bowen Yan, Jiahao Xiao, Kehui Liu, Jianbo Zhang, Zicheng Zhang, Qi Jia, Zhongjie Jia, Haoming Song, Chunyi Li, Bin Zhao, Guangtao Zhai
VLA models frequently fail in robotic manipulation, but their faults are poorly structured and often require expert diagnosis to interpret and repair. There is no standard way to categorize manipulation failures (perception vs. planning vs. control), nor to measure whether a vision-language model (VLM) can recognize, localize, and recover from them. This paper asks directly whether VLMs can serve as diagnostic and recovery agents for VLA faults.
The work contributes three pieces:
- VLA-FixBench — a fault dataset covering perception, planning, and control failures, annotated with task stages and repair strategies, providing structured ground truth for fault understanding.
- FaultEval — an evaluation framework that benchmarks 20 VLMs across fault-related dimensions (diagnosing and characterizing perception/planning/control failures).
- A VLM-VLA collaboration mechanism that localizes spatiotemporal deviations in a failing rollout and rolls back task execution to the deviation point, enabling targeted recovery rather than restarting the whole task.
flowchart LR
A[VLA executes task] --> B{Fault occurs?}
B -- yes --> C[VLM diagnoses fault type<br/>perception / planning / control]
C --> D[Localize spatiotemporal deviation]
D --> E[Roll back execution to deviation]
E --> F[Targeted recovery / retry]
F --> A
B -- no --> G[Task success]
The authors report that an idealized feedback loop can improve task success rates by 13% on LIBERO and 35% on real-world robotic systems, indicating substantial headroom for VLM-driven fault recovery and that the gains are larger in the messier real-world setting. The FaultEval benchmark over 20 VLMs quantifies how well current models diagnose the three fault categories, establishing baselines for future fault-aware policies.
By turning unstructured VLA failures into a labeled benchmark (VLA-FixBench) with a shared evaluation protocol (FaultEval) and a concrete recovery mechanism, this work reframes robustness as a diagnosable, recoverable property rather than an opaque failure mode. The closed-loop VLM-VLA collaboration — localize the deviation, roll back, recover — offers a practical path to more reliable manipulation, and the large real-world gain (35%) underscores the value of fault diagnosis over blind retries.
- ICML 2026: https://icml.cc/virtual/2026/poster/64203
← Back to ICML-2026