ICLR 2026 CompassNav - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 · Authors: LinFeng Li, Jian Zhao, Yuan Xie, Xin Tan, Xuelong Li · arXiv: 2510.10154 · Category: Embodied navigation / VLN · Trend tag: Decision-understanding training (SFT+RFT) for navigation LVLMs.
flowchart LR
Expert[Single GT path<br/>R2R-style data] --> Astar[A* geodesic distances<br/>annotate ALL feasible actions]
Astar --> Data[Compass-Data-22k<br/>dense field of correctness]
Data --> SFT[Stage 1: SFT<br/>navigational priors]
SFT --> RFT[Stage 2: RFT<br/>Gap-Aware Hybrid Reward]
RFT --> LVLM[7B open-source LVLM<br/>weighs options, then decides]
LVLM --> Act[Navigation action]
Standard navigation training reduces the task to sequence-to-sequence replication of a single correct path. Datasets like R2R provide only one ground-truth trajectory, forcing rigid path imitation that lacks the counterfactual data needed for robust decision-making — the model memorizes routes instead of learning to weigh options and decide.
CompassNav distills spatial reasoning and decision-making directly into open-source LVLMs:
- Compass-Data-22k: a 22k-trajectory dataset whose RFT subset annotates all feasible actions at each step using A* geodesic distances, producing a dense field of correctness across the decision space rather than a single labeled move.
- Gap-Aware Hybrid Reward: a dynamic reward that adapts to decision certainty — decisive signals for clearly optimal actions, nuanced scores to encourage exploration when options are close.
- Two-stage recipe: preparatory SFT to instill navigational priors, then Reward-based Fine-Tuning (RFT) so the agent learns relative move quality instead of route memorization.
A 7B CompassNav model reaches state-of-the-art on goal-navigation benchmarks, outperforming much larger general-purpose models including GPT-4o and surpassing o1-mini, and the paper demonstrates real-world robot navigation. (Exact SR / SPL figures omitted here pending the camera-ready tables.)
CompassNav reframes navigation training as decision understanding rather than path imitation: by densely supervising the full action space with geodesic correctness and a certainty-aware reward, a compact open LVLM can beat frontier proprietary models on embodied goal navigation.
- ICLR 2026 Survey
- OmniVLA (navigation)
- NavFoM (navigation foundation model)
← Back to ICLR-2026