Review VLA Evaluation - Heungwoo/research GitHub Wiki

Cross-Paper Review — Evaluating VLA Policies Credibly

Topic survey (updated Aug 2026) · the evaluation crisis and its 2026 answers. Structure: 📈 trend · ⚖️ approaches · ✅ emerging norms · ⚠️ limitations. Companion pages: PolaRiS · LIBERO-X · LIBERO-Plus · Qwen-RobotManip (the OOD manifesto) · WAM vs VLA Robustness · RoboMME.

1. The indictment — why in-distribution benchmarking broke

Three independent 2026 demonstrations:

  1. From-scratch ≈ pretrained, in-distribution. RobotManip §6.1: models with no large-scale robot pretraining match or beat π0.5-class models on LIBERO/RoboTwin (StarVLA 98.0, scratch 98.2 vs π0.5 97.6 on LIBERO) — the benchmarks reward memorization, not the pretraining the field spends its money on. The same paper's scaling ablation shows in-distribution scores are flat in pretraining data volume while OOD scores scale.
  2. Cliff-edge collapses under compound shift. LIBERO-X's L1→L5 pyramid: representative VLAs fall 39.4 → 29.6 → 17.0 → 11.5 → 8.2 as perturbations stack; StarVLA drops 85.7 → 10.6 on RoboTwin Clean→Rand.
  3. Sim benchmarks don't rank real policies. Existing sim suites correlate weakly with real-world generalist performance — the gap PolaRiS was built to close (and quantifies via 600 real + 93k sim paired rollouts).

2. Trend arc

Era Practice Failure mode
≤2024 Single-suite success rates (LIBERO, Calvin) Train/test distribution overlap
2025 Perturbation extensions (LIBERO-Plus/PRO, SafeLIBERO) Single-axis, independent perturbations; no difficulty progression
H1 2026 Evaluation methodology becomes a publishable contribution class: hierarchical capability-decomposed protocols, real-to-sim reconstruction, statistical rigor, safety-aware sim, benchmark-free world-model evaluation Fragmentation; each lab ships its own protocol

3. Approach taxonomy (definitions, pros, cons)

Approach Definition Pros Cons / limits
Static sim benchmarks Fixed scenes/tasks, shared splits Cheap, comparable, reproducible Saturated; memorization-prone; insensitive to pretraining
Capability-decomposed perturbation suites Progressive, labeled perturbation axes ([LIBERO-X](/Heungwoo/research/wiki/RSS-2026-LIBERO-X)'s spatial→topology→attribute→semantic pyramid; [LIBERO-Plus](/Heungwoo/research/wiki/CVPR-2026-LIBERO-Plus)'s 7 dims; RoboTwin-IF's language-grounding probes; RoboTwin-XE's embodiment swap) Diagnostic — tells you which capability failed Still sim; training-side diversity must co-evolve or the gap is unfair
Real-to-sim reconstruction Scan real scenes → neural reconstruction → interactive sim evals ([PolaRiS](/Heungwoo/research/wiki/RSS-2026-PolaRiS)) Anchored to reality; validated rank-correlation; scalable env creation; shareable hubs Physics fidelity for contact; validation so far concentrated on π-family policies
Statistically rigorous real evaluation Confidence-aware comparison beyond binary success (RSS #76); betting-based sequential evaluation (RSS #90); offline policy evaluation via discounted liveness (RSS #154) Correct inferences from few rollouts Doesn't reduce rollout cost by orders of magnitude
Safety-aware evaluation Damage-aware simulation (OopsieVerse, RSS #98) Measures what deployment actually risks Young; no adoption yet
World-model evaluation Roll policies out in a learned WM ([Ctrl-World](/Heungwoo/research/wiki/ICLR-2026-Ctrl-World), WorldGym, [Qwen-RobotWorld](/Heungwoo/research/wiki/Review-Qwen-RobotWorld)'s stated direction; [dWorldEval](/Heungwoo/research/wiki/ICML-2026-Scaling-Real-World-Robot-Policy-Evaluation) makes actions first-class tokens) Unlimited, benchmark-drift-free WM fidelity is the evaluator's own confound; language-actioned WMs can't consume continuous policy actions (dWorldEval is the first crack in this)
Tiered diagnostic benchmarks (ICML 2026) [LIBERO-Gen](/Heungwoo/research/wiki/ICML-2026-Dismantling-the-Illusion-of-Vision-Language-Action) splits ID / compositional / domain-generalization tiers; [VLA-Arena](/Heungwoo/research/wiki/ICML-2026-VLA-Arena) (170 tasks, L0–L2, safety/distractor/extrapolation axes) Exposes spurious invariance standard metrics mask (π0.5: 64.0% Spatial-CG vs 21.2% Task-CG) Still sim; tier taxonomies not yet standardized across suites
Failure diagnosis & adversarial probes (ICML 2026) [VLA-FixBench](/Heungwoo/research/wiki/ICML-2026-Can-VLMs-Diagnose-and-Recover) benchmarks 20 VLMs on fault diagnosis (idealized recovery loop: +13% LIBERO / +35% real); [TRAP](/Heungwoo/research/wiki/ICML-2026-TRAP) demonstrates the first targeted adversarial-patch attack on VLA chain-of-thought Opens the recovery and security axes Both nascent; no defense literature yet

4. Emerging norms worth adopting

  1. Report OOD, not just in-distribution — with capability decomposition (spatial / object / instruction axes at minimum).
  2. Validate your sim against reality — PolaRiS-style paired-rollout correlation, not assumed transfer.
  3. Fine-tune-then-perturb protocols (train on Clean, evaluate on Rand) expose what generalist pretraining actually bought.
  4. Instruction-following probes (RoboTwin-IF-class) catch VLA-to-VA degradation — GR00T-N1.7's 16.6% there, despite respectable visual-perturbation scores, shows why this axis must be separate.
  5. Statistical reporting: trials-per-cell and uncertainty, especially for ≤10-rollout real evaluations.

5. ⚠️ Limitations of the current toolkit

  • Contact-rich and deformable-object tasks still lack faithful simulated evaluation anywhere — the shared blind spot of both perturbation suites and real-to-sim.
  • Real-correlation evidence (PolaRiS) covers a narrow policy family so far; correlation for AR/tokenized or humanoid policies is unmeasured.
  • Humanoid evaluation lags a full generation: intervention-assisted scoring, no OOD protocol, no perturbation suite (Review-Humanoid-VLA).
  • No venue-level enforcement: OOD reporting remains voluntary, so the in-distribution leaderboard culture persists in parallel.

← Back to Home · Reviews