Review VLA Evaluation - Heungwoo/research GitHub Wiki
Cross-Paper Review — Evaluating VLA Policies Credibly
Topic survey (updated Aug 2026) · the evaluation crisis and its 2026 answers. Structure: 📈 trend · ⚖️ approaches · ✅ emerging norms · ⚠️ limitations. Companion pages: PolaRiS · LIBERO-X · LIBERO-Plus · Qwen-RobotManip (the OOD manifesto) · WAM vs VLA Robustness · RoboMME.
1. The indictment — why in-distribution benchmarking broke
Three independent 2026 demonstrations:
- From-scratch ≈ pretrained, in-distribution. RobotManip §6.1: models with no large-scale robot pretraining match or beat π0.5-class models on LIBERO/RoboTwin (StarVLA 98.0, scratch 98.2 vs π0.5 97.6 on LIBERO) — the benchmarks reward memorization, not the pretraining the field spends its money on. The same paper's scaling ablation shows in-distribution scores are flat in pretraining data volume while OOD scores scale.
- Cliff-edge collapses under compound shift. LIBERO-X's L1→L5 pyramid: representative VLAs fall 39.4 → 29.6 → 17.0 → 11.5 → 8.2 as perturbations stack; StarVLA drops 85.7 → 10.6 on RoboTwin Clean→Rand.
- Sim benchmarks don't rank real policies. Existing sim suites correlate weakly with real-world generalist performance — the gap PolaRiS was built to close (and quantifies via 600 real + 93k sim paired rollouts).
2. Trend arc
| Era | Practice | Failure mode |
|---|---|---|
| ≤2024 | Single-suite success rates (LIBERO, Calvin) | Train/test distribution overlap |
| 2025 | Perturbation extensions (LIBERO-Plus/PRO, SafeLIBERO) | Single-axis, independent perturbations; no difficulty progression |
| H1 2026 | Evaluation methodology becomes a publishable contribution class: hierarchical capability-decomposed protocols, real-to-sim reconstruction, statistical rigor, safety-aware sim, benchmark-free world-model evaluation | Fragmentation; each lab ships its own protocol |
3. Approach taxonomy (definitions, pros, cons)
| Approach | Definition | Pros | Cons / limits |
|---|---|---|---|
| Static sim benchmarks | Fixed scenes/tasks, shared splits | Cheap, comparable, reproducible | Saturated; memorization-prone; insensitive to pretraining |
| Capability-decomposed perturbation suites | Progressive, labeled perturbation axes ([LIBERO-X](/Heungwoo/research/wiki/RSS-2026-LIBERO-X)'s spatial→topology→attribute→semantic pyramid; [LIBERO-Plus](/Heungwoo/research/wiki/CVPR-2026-LIBERO-Plus)'s 7 dims; RoboTwin-IF's language-grounding probes; RoboTwin-XE's embodiment swap) | Diagnostic — tells you which capability failed | Still sim; training-side diversity must co-evolve or the gap is unfair |
| Real-to-sim reconstruction | Scan real scenes → neural reconstruction → interactive sim evals ([PolaRiS](/Heungwoo/research/wiki/RSS-2026-PolaRiS)) | Anchored to reality; validated rank-correlation; scalable env creation; shareable hubs | Physics fidelity for contact; validation so far concentrated on π-family policies |
| Statistically rigorous real evaluation | Confidence-aware comparison beyond binary success (RSS #76); betting-based sequential evaluation (RSS #90); offline policy evaluation via discounted liveness (RSS #154) | Correct inferences from few rollouts | Doesn't reduce rollout cost by orders of magnitude |
| Safety-aware evaluation | Damage-aware simulation (OopsieVerse, RSS #98) | Measures what deployment actually risks | Young; no adoption yet |
| World-model evaluation | Roll policies out in a learned WM ([Ctrl-World](/Heungwoo/research/wiki/ICLR-2026-Ctrl-World), WorldGym, [Qwen-RobotWorld](/Heungwoo/research/wiki/Review-Qwen-RobotWorld)'s stated direction; [dWorldEval](/Heungwoo/research/wiki/ICML-2026-Scaling-Real-World-Robot-Policy-Evaluation) makes actions first-class tokens) | Unlimited, benchmark-drift-free | WM fidelity is the evaluator's own confound; language-actioned WMs can't consume continuous policy actions (dWorldEval is the first crack in this) |
| Tiered diagnostic benchmarks (ICML 2026) | [LIBERO-Gen](/Heungwoo/research/wiki/ICML-2026-Dismantling-the-Illusion-of-Vision-Language-Action) splits ID / compositional / domain-generalization tiers; [VLA-Arena](/Heungwoo/research/wiki/ICML-2026-VLA-Arena) (170 tasks, L0–L2, safety/distractor/extrapolation axes) | Exposes spurious invariance standard metrics mask (π0.5: 64.0% Spatial-CG vs 21.2% Task-CG) | Still sim; tier taxonomies not yet standardized across suites |
| Failure diagnosis & adversarial probes (ICML 2026) | [VLA-FixBench](/Heungwoo/research/wiki/ICML-2026-Can-VLMs-Diagnose-and-Recover) benchmarks 20 VLMs on fault diagnosis (idealized recovery loop: +13% LIBERO / +35% real); [TRAP](/Heungwoo/research/wiki/ICML-2026-TRAP) demonstrates the first targeted adversarial-patch attack on VLA chain-of-thought | Opens the recovery and security axes | Both nascent; no defense literature yet |
4. Emerging norms worth adopting
- Report OOD, not just in-distribution — with capability decomposition (spatial / object / instruction axes at minimum).
- Validate your sim against reality — PolaRiS-style paired-rollout correlation, not assumed transfer.
- Fine-tune-then-perturb protocols (train on Clean, evaluate on Rand) expose what generalist pretraining actually bought.
- Instruction-following probes (RoboTwin-IF-class) catch VLA-to-VA degradation — GR00T-N1.7's 16.6% there, despite respectable visual-perturbation scores, shows why this axis must be separate.
- Statistical reporting: trials-per-cell and uncertainty, especially for ≤10-rollout real evaluations.
5. ⚠️ Limitations of the current toolkit
- Contact-rich and deformable-object tasks still lack faithful simulated evaluation anywhere — the shared blind spot of both perturbation suites and real-to-sim.
- Real-correlation evidence (PolaRiS) covers a narrow policy family so far; correlation for AR/tokenized or humanoid policies is unmeasured.
- Humanoid evaluation lags a full generation: intervention-assisted scoring, no OOD protocol, no perturbation suite (Review-Humanoid-VLA).
- No venue-level enforcement: OOD reporting remains voluntary, so the in-distribution leaderboard culture persists in parallel.