ICLR 2026 CompassNav - Heungwoo/research GitHub Wiki

CompassNav — from path imitation to decision understanding

Venue: ICLR 2026 · Authors: LinFeng Li, Jian Zhao, Yuan Xie, Xin Tan, Xuelong Li · arXiv: 2510.10154 · Category: Embodied navigation / VLN · Trend tag: Decision-understanding training (SFT+RFT) for navigation LVLMs.

Approach diagram

flowchart LR
  Expert[Single GT path<br/>R2R-style data] --> Astar[A* geodesic distances<br/>annotate ALL feasible actions]
  Astar --> Data[Compass-Data-22k<br/>dense field of correctness]
  Data --> SFT[Stage 1: SFT<br/>navigational priors]
  SFT --> RFT[Stage 2: RFT<br/>Gap-Aware Hybrid Reward]
  RFT --> LVLM[7B open-source LVLM<br/>weighs options, then decides]
  LVLM --> Act[Navigation action]
Loading

Problem

Standard navigation training reduces the task to sequence-to-sequence replication of a single correct path. Datasets like R2R provide only one ground-truth trajectory, forcing rigid path imitation that lacks the counterfactual data needed for robust decision-making — the model memorizes routes instead of learning to weigh options and decide.

Method

CompassNav distills spatial reasoning and decision-making directly into open-source LVLMs:

  • Compass-Data-22k: a 22k-trajectory dataset whose RFT subset annotates all feasible actions at each step using A* geodesic distances, producing a dense field of correctness across the decision space rather than a single labeled move.
  • Gap-Aware Hybrid Reward: a dynamic reward that adapts to decision certainty — decisive signals for clearly optimal actions, nuanced scores to encourage exploration when options are close.
  • Two-stage recipe: preparatory SFT to instill navigational priors, then Reward-based Fine-Tuning (RFT) so the agent learns relative move quality instead of route memorization.

Results

A 7B CompassNav model reaches state-of-the-art on goal-navigation benchmarks, outperforming much larger general-purpose models including GPT-4o and surpassing o1-mini, and the paper demonstrates real-world robot navigation. (Exact SR / SPL figures omitted here pending the camera-ready tables.)

Significance

CompassNav reframes navigation training as decision understanding rather than path imitation: by densely supervising the full action space with geodesic correctness and a certainty-aware reward, a compact open LVLM can beat frontier proprietary models on embodied goal navigation.

Links

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️