ICML 2026 VLANeXt - Heungwoo/research GitHub Wiki

VLANeXt: Recipes for Building Strong VLA Models — a unified 12-finding design-space study

Venue: ICML 2026 (Poster) Category: Analysis-Insight Affiliations: S-Lab, Nanyang Technological University; Sun Yat-sen University (authors: Xiao-Ming Wu, Bin Fan, Kang Liao, Jian-jian Jiang, Runze Yang, Yihang Luo, Zhonghua Wu, Wei-Shi Zheng, Chen Change Loy) Traction (2026-06): 5 citations (arXiv)

Performance comparison on the LIBERO and LIBERO-plus benchmarks; despite a smaller model size, VLANeXt achieves higher success rates than prior methods on both standard task performance and robustness/generalization (Figure 1 from Wu et al., 2026)

Problem

Following the rise of large foundation models, Vision–Language–Action models (VLAs) emerged, leveraging strong visual and language understanding for general-purpose policy learning. Yet the current VLA landscape remains fragmented and exploratory: although many groups have proposed their own VLA models, inconsistencies in training protocols and evaluation settings make it difficult to identify which design choices truly matter. This paper brings structure to the space by reexamining the VLA design space under a single unified framework and evaluation setup.

Method

Starting from a simple VLA baseline similar to RT-2 and OpenVLA, the authors systematically dissect design choices along three dimensions and distill 12 key findings into a practical recipe:

  • Foundational components — backbone, VLM choice, and overall pipeline.
  • Perception essentials — multi-view visual inputs, and how proprioception is conditioned.
  • Action modelling perspectives — the policy head and auxiliary objectives. They compare regression, diffusion, flow matching, and a VQ-VAE codebook classifier (codebook size 1024, 3 codes per action); regression achieves the strongest performance, with diffusion the strongest among generative heads. They also augment action prediction with an auxiliary world-modelling objective and a frequency-domain regularizer.

The resulting model, VLANeXt (Figure 7), tokenizes multi-view visual inputs, language instructions, and proprioception, processes them with a multimodal LLM, and uses meta queries to enable soft interaction with a policy module that predicts action chunks via flow matching, further regularized by a frequency-domain objective.

VLANeXt architecture: multi-view visual inputs, language, and proprioception are tokenized and processed by a multimodal LLM, with meta queries enabling soft interaction with the policy module; action chunks are predicted using flow matching (Figure 7 from Wu et al., 2026)

Results

On the LIBERO benchmark (success rate %, Spatial/Object/Goal/Long), VLANeXt reaches 99.0 / 99.2 / 96.6 / 94.6, averaging 97.4 — ahead of OpenVLA-OFT (97.1), FLOWER (96.9), UniVLA (95.2), π0 (86.0), and the original OpenVLA (76.5), despite a smaller model size.

On LIBERO-plus (robustness under camera, robot, language, light, background, noise, and layout perturbations), the baselines it surpasses include OpenVLA-OFT and π0-Fast (total 61.6), π0 (53.6), UniVLA (42.9), and OpenVLA (15.6) — underscoring how fragile many VLAs are to perturbations.

In real-world experiments (success count / 20), VLANeXt leads on all four tasks: single-arm Clean 14/20 and Drawer 11/20, bimanual Clean 11/20 and Lifting 15/20, versus π0 (10/8/10/13) and OpenVLA-OFT (7/7/5/9).

Significance

VLANeXt is a consolidation paper for a field saturated with one-off architectures. By holding training and evaluation fixed and ablating choices end-to-end, it converts scattered intuitions into a reproducible recipe. The 12 findings, an open codebase, and an accompanying "awesome-vla" literature repository make it a likely reference baseline and a low-barrier starting point for fair, controlled VLA comparison in 2026.

Links

← Back to ICML-2026