IROS 2026 IMLE VLA - Heungwoo/research GitHub Wiki

IROS 2026 — IMLE-VLA: Fast Single-Step Action Generation for VLA Policies

Venue: IROS 2026 (Pittsburgh) · paper #4624 · Simon Fraser University · University of Pennsylvania (Hosseinkhani, Peng, Shramko, … Jayaraman, Li). The real-time datapoint of IROS 2026 — replace a VLA's iterative diffusion/flow action head with a single-step generator, hitting 55 Hz and LIBERO 98.0%. Companions: IROS 2026 survey · Real-Time Execution · VLA Architectures.

1. Problem

Leading VLAs couple a VLM backbone with a continuous action head trained via diffusion or flow matching, which needs iterative multi-step sampling (e.g. 10 Euler steps in π0.5). This is an inference bottleneck → stop-and-go robot motion and slow task completion.

2. Method

IMLE-VLA replaces the iterative head with a single-step conditional generator trained via conditional Implicit Maximum Likelihood Estimation (cIMLE). The cIMLE objective promotes multimodal action coverage — avoiding the mode collapse of naive regression heads — while eliminating multi-step sampling entirely. It's a drop-in action-head swap on a π0.5-class VLA.

3. Results

  • Applied to π0.5: inference frequency 3.67× (55 Hz vs 15 Hz) → up to 11× higher action throughput.
  • LIBERO (40 tasks): 98.0% avg success — highest among baselines and fastest.
  • LIBERO-Plus perturbations: retains π0.5's robustness while other baselines degrade sharply (single-step ≠ brittle).
  • Real-world Franka (4 tasks): smoother motion, faster completion, beats π0.5 on every task; 3.9×–6.6× less inference time per episode.

4. Why it matters (efficiency lens)

IMLE-VLA is the cleanest IROS 2026 example of the survey §4 efficiency push: the field is attacking the diffusion/flow action-head latency wall (cf. Reflex streaming, Fast-dVLA, Real-Time Execution). The insight is that single-step generation need not sacrifice multimodality or robustness — cIMLE preserves both, unlike naive regression. It sits alongside the other IROS efficiency entries (BFA++ token pruning, Fast-Enough-to-Act token merging) as evidence that making VLAs real-time is a first-class 2026 track, not an afterthought.

Limitations (reviewer): demonstrated on π0.5 + LIBERO/Franka; cIMLE's coverage vs a well-tuned few-step flow head at matched compute isn't isolated; single-step quality on very high-precision contact tasks untested.

5. Links

← Back to IROS 2026 survey · Home