ICLR 2026 AutoFly - Heungwoo/research GitHub Wiki

AutoFly — VLA model for UAV autonomous navigation in the wild

Venue: ICLR 2026 · Authors: Xiaolou Sun, Wufei Si, Wenhui Ni, Yuntian Li, Dongming Wu, Fei Xie, Runwei Guan, He-Yang Xu, Henghui Ding, Yuan Wu, Yutao Yue, Yongming Huang, Hui Xiong · arXiv: 2602.09657 · Category: Embodied navigation / VLN (aerial) · Trend tag: Vision-Language-Action models for autonomous UAV flight.

Approach diagram

flowchart LR
  RGB[Onboard RGB camera] --> PDE[Pseudo-depth encoder<br/>depth-aware features from RGB]
  RGB --> VE[Visual encoder]
  Inst[Coarse directional guidance] --> VLA
  PDE --> VLA
  VE --> VLA
  subgraph VLA["AutoFly VLA model"]
    Stage1[Stage 1: align vision + depth + language] --> Stage2[Stage 2: action grounding]
  end
  VLA --> Act[UAV action<br/>continuous planning + obstacle avoidance]
Loading

Problem

VLN research for UAVs typically assumes detailed, pre-specified step-by-step instructions. But real outdoor exploration happens in unknown environments where such instructions are unavailable: the drone has only coarse directional guidance and must navigate autonomously through continuous planning and obstacle avoidance. This is the gap AutoFly targets.

Method

AutoFly is an end-to-end Vision-Language-Action (VLA) model for autonomous UAV navigation:

  • Pseudo-depth encoder: derives depth-aware spatial features directly from RGB imagery, improving spatial reasoning without a real depth sensor.
  • Progressive two-stage training: first aligns visual representations, depth information, and language understanding; then grounds these into navigation actions.
  • A new autonomous-navigation dataset emphasizing obstacle avoidance and continuous planning (rather than instruction-following), incorporating real-world data.

Results

AutoFly reports a +3.9% success rate over state-of-the-art VLA baselines, with consistent performance across both simulated and real-world environments. The authors state model, data, and code are released. (Exact simulator and per-task numbers omitted here pending the camera-ready tables.)

Significance

AutoFly reframes aerial VLN away from instruction-following toward autonomous flight in the wild under sparse guidance, and shows that RGB-only pseudo-depth can substitute for true depth sensing in a UAV VLA. It extends the VLA paradigm — dominant in manipulation and ground navigation — into the aerial embodiment.

Links

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️