ICML 2026 VLANeXt - Heungwoo/research GitHub Wiki
VLANeXt: Recipes for Building Strong VLA Models — a unified 12-finding design-space study
Venue: ICML 2026 (Poster) Category: Analysis-Insight Affiliations: S-Lab, Nanyang Technological University; Sun Yat-sen University (authors: Xiao-Ming Wu, Bin Fan, Kang Liao, Jian-jian Jiang, Runze Yang, Yihang Luo, Zhonghua Wu, Wei-Shi Zheng, Chen Change Loy) Traction (2026-06): 5 citations (arXiv)

Problem
Following the rise of large foundation models, Vision–Language–Action models (VLAs) emerged, leveraging strong visual and language understanding for general-purpose policy learning. Yet the current VLA landscape remains fragmented and exploratory: although many groups have proposed their own VLA models, inconsistencies in training protocols and evaluation settings make it difficult to identify which design choices truly matter. This paper brings structure to the space by reexamining the VLA design space under a single unified framework and evaluation setup.
Method
Starting from a simple VLA baseline similar to RT-2 and OpenVLA, the authors systematically dissect design choices along three dimensions and distill 12 key findings into a practical recipe:
- Foundational components — backbone, VLM choice, and overall pipeline.
- Perception essentials — multi-view visual inputs, and how proprioception is conditioned.
- Action modelling perspectives — the policy head and auxiliary objectives. They compare regression, diffusion, flow matching, and a VQ-VAE codebook classifier (codebook size 1024, 3 codes per action); regression achieves the strongest performance, with diffusion the strongest among generative heads. They also augment action prediction with an auxiliary world-modelling objective and a frequency-domain regularizer.
The resulting model, VLANeXt (Figure 7), tokenizes multi-view visual inputs, language instructions, and proprioception, processes them with a multimodal LLM, and uses meta queries to enable soft interaction with a policy module that predicts action chunks via flow matching, further regularized by a frequency-domain objective.

Results
On the LIBERO benchmark (success rate %, Spatial/Object/Goal/Long), VLANeXt reaches 99.0 / 99.2 / 96.6 / 94.6, averaging 97.4 — ahead of OpenVLA-OFT (97.1), FLOWER (96.9), UniVLA (95.2), π0 (86.0), and the original OpenVLA (76.5), despite a smaller model size.
On LIBERO-plus (robustness under camera, robot, language, light, background, noise, and layout perturbations), the baselines it surpasses include OpenVLA-OFT and π0-Fast (total 61.6), π0 (53.6), UniVLA (42.9), and OpenVLA (15.6) — underscoring how fragile many VLAs are to perturbations.
In real-world experiments (success count / 20), VLANeXt leads on all four tasks: single-arm Clean 14/20 and Drawer 11/20, bimanual Clean 11/20 and Lifting 15/20, versus π0 (10/8/10/13) and OpenVLA-OFT (7/7/5/9).
Significance
VLANeXt is a consolidation paper for a field saturated with one-off architectures. By holding training and evaluation fixed and ablating choices end-to-end, it converts scattered intuitions into a reproducible recipe. The 12 findings, an open codebase, and an accompanying "awesome-vla" literature repository make it a likely reference baseline and a low-barrier starting point for fair, controlled VLA comparison in 2026.
Links
- arXiv: 2602.18532
- ICML 2026: https://icml.cc/virtual/2026/poster/64465
- Project: https://dravenalg.github.io/VLANeXt/ · Code: DravenALG/VLANeXt
← Back to ICML-2026