ICML 2026 AIR VLA - Heungwoo/research GitHub Wiki
AIR-VLA: Vision-Language-Action Systems for Aerial Manipulation — The first VLA benchmark for flying manipulators
Venue: ICML 2026 (Poster) Category: Benchmark Traction (2026-06): 1 citation (arXiv)

Problem
Vision-Language-Action (VLA) models have achieved remarkable success in ground-based embodied intelligence, but their application to Aerial Manipulation Systems (AMS) — UAVs equipped with robotic manipulators — remains a largely unexplored frontier. AMS have characteristics that break the assumptions of VLA paradigms designed for static or 2D mobile bases: floating-base dynamics, strong coupling between the UAV and the manipulator, and multi-step, long-horizon operational tasks. Free from the constraints of terrain and altitude, AMS can reach working heights inaccessible to ground robots, but no standardized testbed or data foundation existed to study VLA control in this regime.
Method
AIR-VLA is positioned as the first VLA benchmark tailored for aerial manipulation, delivered as a full-stack platform:
- Simulation environment — a physics-based environment built on NVIDIA Isaac Sim that models the collaborative UAV-plus-manipulator system and its interactive assets.
- Teleoperated dataset — a high-quality multimodal dataset of 3,000 manually teleoperated demonstrations spanning base manipulation, object & spatial understanding, semantic reasoning, and long-horizon planning, with multi-camera observation spaces.
- Task taxonomy — four core task suites targeting manipulation precision, semantic reasoning, visual perception, and planning, emphasizing deep coordination between 3D UAV mobility and manipulator control.
- Multi-dimensional metrics — because of the multi-step nature of aerial tasks, evaluation moves beyond simple task success rates to weighted, sub-metric scores tailored to UAV mobility, manipulator control, and high-level planning. A separate pipeline evaluates VLM high-level planning (planning success rate).

Results
Using the platform, the authors systematically evaluate mainstream VLA models and state-of-the-art VLM models across the four task suites (Table 1, weighted totals). Flow-matching VLA policies such as $\pi_0$ and $\pi_{0.5}$ demonstrate significant advantages, outperforming traditional imitation-learning baselines such as ACT and Diffusion Policy across all metrics. $\pi_{0.5}$ is the best-performing model and is selected for a robustness study (Table 2): under UAV dynamic disturbance and third-person-perception-deprived conditions, scores drop, quantifying the sensitivity of current policies to floating-base perturbation and missing viewpoints. The VLM planning evaluation (Table 3) reports normalized sub-metric scores and planning success rates across scenarios and instruction types, revealing the capability boundaries of current models on long-horizon aerial planning.
Significance
AIR-VLA validates the feasibility of transferring VLA paradigms to aerial systems while exposing their boundaries on UAV mobility, manipulator control, and high-level planning. By releasing a standardized simulation testbed, a 3,000-demonstration multimodal dataset, and aerial-specific multi-dimensional metrics, it establishes a data and evaluation foundation for general-purpose aerial robotics — a domain previously without any VLA benchmark.
Links
- arXiv: 2601.21602
- ICML 2026: https://icml.cc/virtual/2026/poster/64411
← Back to ICML-2026