ICLR 2026 AutoFly - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 · Authors: Xiaolou Sun, Wufei Si, Wenhui Ni, Yuntian Li, Dongming Wu, Fei Xie, Runwei Guan, He-Yang Xu, Henghui Ding, Yuan Wu, Yutao Yue, Yongming Huang, Hui Xiong · arXiv: 2602.09657 · Category: Embodied navigation / VLN (aerial) · Trend tag: Vision-Language-Action models for autonomous UAV flight.
flowchart LR
RGB[Onboard RGB camera] --> PDE[Pseudo-depth encoder<br/>depth-aware features from RGB]
RGB --> VE[Visual encoder]
Inst[Coarse directional guidance] --> VLA
PDE --> VLA
VE --> VLA
subgraph VLA["AutoFly VLA model"]
Stage1[Stage 1: align vision + depth + language] --> Stage2[Stage 2: action grounding]
end
VLA --> Act[UAV action<br/>continuous planning + obstacle avoidance]
VLN research for UAVs typically assumes detailed, pre-specified step-by-step instructions. But real outdoor exploration happens in unknown environments where such instructions are unavailable: the drone has only coarse directional guidance and must navigate autonomously through continuous planning and obstacle avoidance. This is the gap AutoFly targets.
AutoFly is an end-to-end Vision-Language-Action (VLA) model for autonomous UAV navigation:
- Pseudo-depth encoder: derives depth-aware spatial features directly from RGB imagery, improving spatial reasoning without a real depth sensor.
- Progressive two-stage training: first aligns visual representations, depth information, and language understanding; then grounds these into navigation actions.
- A new autonomous-navigation dataset emphasizing obstacle avoidance and continuous planning (rather than instruction-following), incorporating real-world data.
AutoFly reports a +3.9% success rate over state-of-the-art VLA baselines, with consistent performance across both simulated and real-world environments. The authors state model, data, and code are released. (Exact simulator and per-task numbers omitted here pending the camera-ready tables.)
AutoFly reframes aerial VLN away from instruction-following toward autonomous flight in the wild under sparse guidance, and shows that RGB-only pseudo-depth can substitute for true depth sensing in a UAV VLA. It extends the VLA paradigm — dominant in manipulation and ground navigation — into the aerial embodiment.
- arXiv: https://arxiv.org/abs/2602.09657
- OpenReview: https://openreview.net/forum?id=88RKxlFUNY
- ICLR 2026 Survey
- OmniVLA (navigation)
- NavFoM (navigation foundation model, includes drone embodiment)
← Back to ICLR-2026