ICLR 2026 Embodied Nav FM - Heungwoo/research GitHub Wiki

NavFoM — Embodied Navigation Foundation Model

Venue: ICLR 2026 · OpenReview: kkBOIsrCXh Category: Navigation foundation model — cross-embodiment Trend tag: Foundation models for navigation / multi-embodiment Authors: Jiazhao Zhang, Anqi Li, Yunpeng Qi, Minghan Li (equal); corresponding Zhizheng Zhang, He Wang. PKU + Galbot + USTC + BAAI + Adelaide + Zhejiang U + Differential Robotics.

Approach diagram

flowchart LR
  subgraph Inputs[Egocentric inputs from any rig]
    Cams[Variable camera rigs<br/>quadruped 4-view · drone 4-view<br/>wheeled 1-view · car 6-8 view]
    Inst[Language instruction L]
  end
  Cams --> VE[DINOv2 + SigLIP encoders<br/>concat-channel · 576 patches]
  VE --> GP[Grid Average Pooling<br/>fine 64-d for current obs<br/>coarse 4-d for history]
  GP --> TVI["TVI tokens · E_base + P_time(t) + P_angle(φ)<br/>navigation: time + angle<br/>video QA: time only<br/>image QA: base only"]
  TVI --> BATS["Budget-Aware Temporal Sampling<br/>P(t) = (1-ε)·e^(k(t-T)/T) + ε<br/>token budget B = 1600/2048"]
  Inst --> Tok[Language tokens E_L]
  BATS --> LLM[Qwen2-7B LLM]
  Tok --> LLM
  LLM --> Head[3-layer MLP planner A_θ]
  Head --> Traj["τ = α_task · A_θ E_T^A<br/>M=8 waypoints · normalized to (−1, 1)"]
  Traj --> Embod[VLN / search / track / driving<br/>quadruped · drone · wheeled · car]
Loading

Problem

Navigation models are typically embodiment-specific and task-specific: different stacks for VLN, object search, target tracking, and autonomous driving — each tied to a particular camera rig and history length. Cross-task models (Uni-NaVid, OctoNav, UniGoal) usually fix a single camera config; cross-embodiment ones (One-Ring, ExAug, Embodiment Randomization) usually fix a single task. NavFoM is an early attempt to unify both axes within a single VLA-style model.

Detailed Method

Generalist navigation formulation

Given language instruction L and N-camera image sequence I1:T1:N, predict a trajectory τ = {a1, ..., aM} where a ∈ R4 = (x, y, z, θ); z used only for UAV; θ is yaw. Mapping π(L, I1:T1:N) → τT.

Architecture

Vision encoders. Concatenated DINOv2 (Oquab et al., 2023) and SigLIP (Zhai et al., 2023) along the channel dimension; 576 patches per view. Common Cambrian-style recipe.

Grid pooling. To fight token-count blowup with multi-view streaming video:

  • Fine-grained Vfine ∈ R64×C for the current observation and image-QA frames.
  • Coarse-grained Vcoarse ∈ R4×C for historical and video-QA frames.

Cross-modal projector P(·): 2-layer MLP into the LLM's latent space.

LLM: Qwen2-7B (Yang et al., 2024a) + planner head Aθ = 3-layer MLP.

Temporal-Viewpoint Indicator (TVI) tokens

Inserted between visual tokens to mark viewpoint angle φ and timestep t:

ETVI =

  • Ebase + Ptime(TimePE(t)) + Pangle(AnglePE(φ)) — Navigation
  • Ebase + Ptime(TimePE(t)) — Video QA
  • Ebase — Image QA

AnglePE is sinusoidal applied separately to cos(φ) and sin(φ) (preserves circular continuity 0 ≡ 2π so d(0, ε) < d(0, π)). TimePE is sinusoidal in t. Ptime and Pangle are 2-layer MLPs. Three required attributes: viewpoint-awareness, time-awareness, separability across task types.

Budget-Aware Temporal Sampling (BATS)

A constant token budget Btoken (set to 2048 in simulators, 1600 in real robots) is enforced via an exponential growth sampling probability inspired by the Ebbinghaus forgetting curve:

P(t) = (1 − ε)·ek(t−T)/T + ε, k > 0, ε = 0.1

This keeps recent frames fully sampled and decays older ones. The expected sampled-frame count integrates analytically:

Eframes ≈ (1 − ε)·(1 − e−k)/k · T + εT

Token-budget constraint: ((4+1)·Eframe + (64+1))·N ≤ Btoken; k solved with Brent's method offline. Each timestep contributes 4 coarse + 1 TVI tokens for history and 64 fine + 1 TVI for current obs, multiplied by N cameras.

Compared to Token Merging (Uni-NaVid): no per-step compute overhead. Compared to Uniform Sampling (NaVILA): keeps recent context dense.

Trajectory prediction with task-specific scaling

Raw trajectories range from meters indoors to tens of meters outdoors → distribution divergence. Scaling factor αtask normalizes waypoints to [-1, 1]:

τT = {a1, ..., aM}T = αtask · Aθ(ETA), M = 8

Three different scaling factors for indoor / UAV / cars. For wheeled/car: aidx = (x, y, θ); for UAV: (x, y, z, θ). Loss: MSE on valid action indices. Total loss: L = β·Lnav + LQA, β = 10 to amplify the small navigation MSE relative to next-token cross-entropy.

Training data (12.7M samples total)

Category Count
Navigation samples (total) 8.02 M
↳ VLN (R2R + RxR + OpenUAV) 3.37 M
↳ Object Goal (HM3D ObjNav) 1.02 M
↳ Active Visual Tracking (EVT-Bench) 0.897 M
↳ Autonomous Driving (nuScenes + OpenScene) 0.681 M
↳ Web-video navigation (Sekai, ~182K YouTube videos) 2.03 M
Image QA 3.15 M
Video QA 1.61 M
Total 12.74 M

Embodiments: quadrupeds, drones, wheeled robots, cars. Camera configs: 1, 4, 6, 8 views. QA data sampled at 1 FPS; continuous nav at 2 FPS; discrete habitat actions converted to trajectory waypoints.

Training configuration

  • Hardware: 56× NVIDIA H100 GPUs, ~72 hours, 4,032 GPU-hours total.
  • Vision encoders + LLM: initialized with default pre-trained weights; ViT and Qwen2-7B fine-tuned per VLM training paradigm; single epoch.
  • Loss scaling β=10; cross-entropy for QA in next-token prediction.

Real-world deployment

Remote server with single RTX 4090 + 1600 token budget → at most 0.5 s to generate an 8-waypoint trajectory (paper's stated figure; ≈2 Hz worst case). The server communicates with client robots (quadruped, humanoid, drone, wheeled) over the Internet. (Note: Appendix A.4 / Fig. 15 instead labels the deployment server a GeForce RTX 5090; the main text says RTX 4090. The paper does not report a GPU-memory figure or a 5 Hz / 218 ms latency.)

Comprehensive Results

VLN-CE R2R + RxR (Table 1)

NavFoM uses RGB only (no depth, no odometry). Highlights:

R2R Val-Unseen:

  • Single-view: NavFoM NE 5.01 / SR 56.2 / SPL 51.2 vs StreamVLN-RGB-only (5.10 / 55.7 / 50.9).
  • Multi-view (4 RGB cameras): NavFoM NE 4.61 / SR 61.7 / SPL 55.3 vs HNR* (RGB+depth+odo): NE 4.42 / SR 61.0 / SPL 51.0.

RxR Val-Unseen:

  • Single-view: NavFoM NE 5.51 / SR 57.4 / SPL 49.4 / nDTW 60.2 vs StreamVLN-RGB-only (6.16 / 51.8 / 45.0 / 62.1). +5.6 SR over best previous single-view RGB-only.
  • Multi-view: NavFoM NE 4.74 / SR 64.4 / SPL 56.2 / nDTW 65.8 vs HNR* (RGB+depth+odo): SR 56.3. +8.1 SR using only RGB cameras vs depth+odometry baseline.
  • Single→multi gain: +5.5 SR on R2R, +7.0 SR on RxR.

HM3D-OVON object navigation (Table 2)

NavFoM is zero-shot (not trained on HM3D-OVON):

Method Val Seen SR Val Seen SPL Val Seen Synonyms SR SPL Val Unseen SR SPL
VLFM* (zero-shot) 35.2 18.6 32.4 17.3 35.2 19.6
MTU3D 55.0 23.6 45.0 14.7 40.8 12.1
Uni-NaVid* (zero-shot) 41.3 21.1 43.9 21.8 39.5 19.8
NavFoM Single view* 37.7 25.5 43.3 29.9 43.6 31.3
NavFoM Four views* 40.1 27.1 45.4 32.6 45.2 31.9

Zero-shot 4-view 45.2 SR (Val Unseen) beats the best prior baseline MTU3D (40.8 SR); even NavFoM single-view (43.6) already exceeds it. The SPL gap is even bigger (31.9 vs 12.1).

EVT-Bench active visual tracking (Table 3)

Single Target SR / TR Distracted Target SR / TR
EVT (with SoM+GPT-4o) 32.5 / 49.9 15.7 / 35.7
Uni-NaVid 25.7 / 39.5 11.3 / 27.4
TrackVLA 85.1 / 78.6 57.6 / 63.2
NavFoM Single view 85.0 / 80.5 61.4 / 68.2
NavFoM Four views 88.4 / 80.7 62.0 / 67.9

NAVSIM autonomous driving (Table 4)

Method Camera VLM-Based PDMS ↑
Human - - 94.8
Constant Velocity - - 21.6
Ego Status MLP - - 65.6
UniAD ✓ - 83.4
PARA-Drive ✓ - 84.0
LAW ✓ - 84.6
DrivingGPT ✓ ✓ 82.4
NavFoM (Eight views) ✓ ✓ 84.3

NavFoM on 8-view rig competitive with task-specific autonomous driving baselines despite training on a unified data mix.

Summary across the 7 public benchmarks (Figure 1)

NavFoM SOTA or highly competitive:

  • HM3D-OVON (Val Unseen SR): 45.2 (4-view) / 43.6 (single-view) vs best prior baseline MTU3D 40.8
  • RxR Single-View: 57.4 vs 51.8/49.3
  • RxR Multi-View: 64.4 vs 56.3
  • EVT-Bench DT: 62.0 vs 57.6
  • NavSim: 84.3 vs 84.0/84.6/83.4/82.4

Real-world deployments (Figure 6)

Cross-task and cross-embodiment, including humanoid robots, quadrupeds, drones, and wheeled robots. Tasks include "Follow the man in the white T-shirt and jeans," "Move to the TV," "Track the man," "Go up to the footbridge."

Ablation Studies

History token organization (Table 8 / Figure 8)

Token budget B=2048 vs B=1024 on RxR Val-Unseen:

Method NE ↓ SR ↑ SPL ↑ nDTW ↑
B=1024, Uniform Sampling 5.33 59.7 49.6 57.9
B=1024, Linear Probability Sampling 5.28 61.2 50.9 58.9
B=1024, BATS 4.98 62.5 53.9 64.1
B=2048, Token Merging (Uni-NaVid) 5.01 63.2 54.9 64.4
B=2048, Uniform Sampling 4.90 62.4 54.0 63.9
B=2048, Linear Probability Sampling 4.89 63.0 54.6 64.8
B=2048, BATS 4.74 64.4 56.2 65.8

BATS wins at both budgets; only -1.4 nDTW drop halving from 2048→1024 vs -6.0 / -5.2 drops for baselines.

TVI vs alternatives (Table 8 lower)

TVI variant NE SR SPL nDTW
Viewpoint-history positional embedding (HAMT) 6.27 52.3 46.3 58.7
Individual learned special tokens 5.52 59.1 52.0 59.6
Handcraft Tokens (no Pangle/time) 6.06 53.6 46.1 58.0
TVI Tokens (Eq. 3) 4.74 64.4 56.2 65.8

+12 SR over HAMT-style PE, +5.3 SR over special-tokens — separating viewpoint and temporal identifiers from visual content matters.

Co-tuning across navigation tasks (Figure 7)

Comparing single-task training, +50% other-task data, +100% other-task data:

  • Searching improves from 10.3% to 45.2% (training conditions are single-view + closed-set; multi-task pretraining helps the model generalize to multi-view + open-vocabulary evaluation).
  • Tracking improves from 12.6% to 62.0%.
  • The paper attributes large gains to mismatches between training and evaluation distributions that multi-task data mitigates.

Number of training samples (Figure 5)

Method Samples
NaVid 1.2 M
Uni-NaVid 5.9 M
NavFoM 12.7 M

Limitations

The paper does not have a dedicated "Limitations" section. Implicit/discussed limitations:

  • β scheduling. Constant β=10 to amplify nav loss vs QA loss; the authors flag adaptive β as future work.
  • No DexGraspNet/RLBench/CALVIN evaluation — focus is navigation only; manipulation cross-application untested.
  • Token budget. Real-world deployment uses 1600-token budget on a single 4090; for very long horizons (T ≈ 1120 in 4-camera setup at B=2048), Eq. 5 becomes unsolvable for k — they note this is rare in practice (most timesteps ≈122 in RxR).
  • Failure cases discussed in Appendix F (not detailed in main text).
  • Closed-set training data — VLN-CE / HM3D-OVON / EVT-Bench / nuScenes datasets dominate; the paper notes these are simulation-heavy.

Significance & Positioning

NavFoM pushes navigation toward the single-model, many-embodiments regime that VLAs are pursuing for manipulation.

  • vs NaVid (1.2 M, single embodiment) / Uni-NaVid (5.9 M, single config) / NaVILA (legged-only): NavFoM scales the data 10× and unifies four task families across ground, aerial, and automotive platforms.
  • vs StreamVLN-RGB-only: competitive with previous best RGB-only single-view (+0.5 R2R / +5.6 RxR SR).
  • vs HNR / BEVBert:** beats them on RxR even when those use RGB+depth+odometry; NavFoM uses RGB only.
  • vs TrackVLA: comparable single-view tracking, better multi-view (+3.3 SR on single-target, +0.4 SR on distracted target).
  • vs UniAD / PARA-Drive / LAW (autonomous driving): competitive PDMS (84.3 vs 83.4–84.6) without task-specific autonomous-driving training pipelines.
  • vs UniVLA / X-VLA: the manipulation analogues of NavFoM. NavFoM's key contribution — TVI tokens + BATS — is the navigation-specific answer to "how does one model consume arbitrary camera rigs in real time."

The TVI + BATS recipe is concretely transferable: any cross-embodiment system that needs to encode (camera angle, time, task type) within an LLM context can adopt the same identifier-token + budget-aware sampling design.

Links

Related pages

  • REI-Bench (instruction understanding for embodied agents)
  • UniVLA (cross-embodiment manipulation analogue)
  • X-VLA

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️