ICLR 2026 Embodied Nav FM - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 · OpenReview: kkBOIsrCXh Category: Navigation foundation model — cross-embodiment Trend tag: Foundation models for navigation / multi-embodiment Authors: Jiazhao Zhang, Anqi Li, Yunpeng Qi, Minghan Li (equal); corresponding Zhizheng Zhang, He Wang. PKU + Galbot + USTC + BAAI + Adelaide + Zhejiang U + Differential Robotics.
flowchart LR
subgraph Inputs[Egocentric inputs from any rig]
Cams[Variable camera rigs<br/>quadruped 4-view · drone 4-view<br/>wheeled 1-view · car 6-8 view]
Inst[Language instruction L]
end
Cams --> VE[DINOv2 + SigLIP encoders<br/>concat-channel · 576 patches]
VE --> GP[Grid Average Pooling<br/>fine 64-d for current obs<br/>coarse 4-d for history]
GP --> TVI["TVI tokens · E_base + P_time(t) + P_angle(φ)<br/>navigation: time + angle<br/>video QA: time only<br/>image QA: base only"]
TVI --> BATS["Budget-Aware Temporal Sampling<br/>P(t) = (1-ε)·e^(k(t-T)/T) + ε<br/>token budget B = 1600/2048"]
Inst --> Tok[Language tokens E_L]
BATS --> LLM[Qwen2-7B LLM]
Tok --> LLM
LLM --> Head[3-layer MLP planner A_θ]
Head --> Traj["τ = α_task · A_θ E_T^A<br/>M=8 waypoints · normalized to (−1, 1)"]
Traj --> Embod[VLN / search / track / driving<br/>quadruped · drone · wheeled · car]
Navigation models are typically embodiment-specific and task-specific: different stacks for VLN, object search, target tracking, and autonomous driving — each tied to a particular camera rig and history length. Cross-task models (Uni-NaVid, OctoNav, UniGoal) usually fix a single camera config; cross-embodiment ones (One-Ring, ExAug, Embodiment Randomization) usually fix a single task. NavFoM is an early attempt to unify both axes within a single VLA-style model.
Given language instruction L and N-camera image sequence I1:T1:N, predict a trajectory τ = {a1, ..., aM} where a ∈ R4 = (x, y, z, θ); z used only for UAV; θ is yaw. Mapping π(L, I1:T1:N) → τT.
Vision encoders. Concatenated DINOv2 (Oquab et al., 2023) and SigLIP (Zhai et al., 2023) along the channel dimension; 576 patches per view. Common Cambrian-style recipe.
Grid pooling. To fight token-count blowup with multi-view streaming video:
- Fine-grained Vfine ∈ R64×C for the current observation and image-QA frames.
- Coarse-grained Vcoarse ∈ R4×C for historical and video-QA frames.
Cross-modal projector P(·): 2-layer MLP into the LLM's latent space.
LLM: Qwen2-7B (Yang et al., 2024a) + planner head Aθ = 3-layer MLP.
Inserted between visual tokens to mark viewpoint angle φ and timestep t:
ETVI =
- Ebase + Ptime(TimePE(t)) + Pangle(AnglePE(φ)) — Navigation
- Ebase + Ptime(TimePE(t)) — Video QA
- Ebase — Image QA
AnglePE is sinusoidal applied separately to cos(φ) and sin(φ) (preserves circular continuity 0 ≡ 2π so d(0, ε) < d(0, π)). TimePE is sinusoidal in t. Ptime and Pangle are 2-layer MLPs. Three required attributes: viewpoint-awareness, time-awareness, separability across task types.
A constant token budget Btoken (set to 2048 in simulators, 1600 in real robots) is enforced via an exponential growth sampling probability inspired by the Ebbinghaus forgetting curve:
P(t) = (1 − ε)·ek(t−T)/T + ε, k > 0, ε = 0.1
This keeps recent frames fully sampled and decays older ones. The expected sampled-frame count integrates analytically:
Eframes ≈ (1 − ε)·(1 − e−k)/k · T + εT
Token-budget constraint: ((4+1)·Eframe + (64+1))·N ≤ Btoken; k solved with Brent's method offline. Each timestep contributes 4 coarse + 1 TVI tokens for history and 64 fine + 1 TVI for current obs, multiplied by N cameras.
Compared to Token Merging (Uni-NaVid): no per-step compute overhead. Compared to Uniform Sampling (NaVILA): keeps recent context dense.
Raw trajectories range from meters indoors to tens of meters outdoors → distribution divergence. Scaling factor αtask normalizes waypoints to [-1, 1]:
τT = {a1, ..., aM}T = αtask · Aθ(ETA), M = 8
Three different scaling factors for indoor / UAV / cars. For wheeled/car: aidx = (x, y, θ); for UAV: (x, y, z, θ). Loss: MSE on valid action indices. Total loss: L = β·Lnav + LQA, β = 10 to amplify the small navigation MSE relative to next-token cross-entropy.
| Category | Count |
|---|---|
| Navigation samples (total) | 8.02 M |
| ↳ VLN (R2R + RxR + OpenUAV) | 3.37 M |
| ↳ Object Goal (HM3D ObjNav) | 1.02 M |
| ↳ Active Visual Tracking (EVT-Bench) | 0.897 M |
| ↳ Autonomous Driving (nuScenes + OpenScene) | 0.681 M |
| ↳ Web-video navigation (Sekai, ~182K YouTube videos) | 2.03 M |
| Image QA | 3.15 M |
| Video QA | 1.61 M |
| Total | 12.74 M |
Embodiments: quadrupeds, drones, wheeled robots, cars. Camera configs: 1, 4, 6, 8 views. QA data sampled at 1 FPS; continuous nav at 2 FPS; discrete habitat actions converted to trajectory waypoints.
- Hardware: 56× NVIDIA H100 GPUs, ~72 hours, 4,032 GPU-hours total.
- Vision encoders + LLM: initialized with default pre-trained weights; ViT and Qwen2-7B fine-tuned per VLM training paradigm; single epoch.
- Loss scaling β=10; cross-entropy for QA in next-token prediction.
Remote server with single RTX 4090 + 1600 token budget → at most 0.5 s to generate an 8-waypoint trajectory (paper's stated figure; ≈2 Hz worst case). The server communicates with client robots (quadruped, humanoid, drone, wheeled) over the Internet. (Note: Appendix A.4 / Fig. 15 instead labels the deployment server a GeForce RTX 5090; the main text says RTX 4090. The paper does not report a GPU-memory figure or a 5 Hz / 218 ms latency.)
NavFoM uses RGB only (no depth, no odometry). Highlights:
R2R Val-Unseen:
- Single-view: NavFoM NE 5.01 / SR 56.2 / SPL 51.2 vs StreamVLN-RGB-only (5.10 / 55.7 / 50.9).
- Multi-view (4 RGB cameras): NavFoM NE 4.61 / SR 61.7 / SPL 55.3 vs HNR* (RGB+depth+odo): NE 4.42 / SR 61.0 / SPL 51.0.
RxR Val-Unseen:
- Single-view: NavFoM NE 5.51 / SR 57.4 / SPL 49.4 / nDTW 60.2 vs StreamVLN-RGB-only (6.16 / 51.8 / 45.0 / 62.1). +5.6 SR over best previous single-view RGB-only.
- Multi-view: NavFoM NE 4.74 / SR 64.4 / SPL 56.2 / nDTW 65.8 vs HNR* (RGB+depth+odo): SR 56.3. +8.1 SR using only RGB cameras vs depth+odometry baseline.
- Single→multi gain: +5.5 SR on R2R, +7.0 SR on RxR.
NavFoM is zero-shot (not trained on HM3D-OVON):
| Method | Val Seen SR | Val Seen SPL | Val Seen Synonyms SR | SPL | Val Unseen SR | SPL |
|---|---|---|---|---|---|---|
| VLFM* (zero-shot) | 35.2 | 18.6 | 32.4 | 17.3 | 35.2 | 19.6 |
| MTU3D | 55.0 | 23.6 | 45.0 | 14.7 | 40.8 | 12.1 |
| Uni-NaVid* (zero-shot) | 41.3 | 21.1 | 43.9 | 21.8 | 39.5 | 19.8 |
| NavFoM Single view* | 37.7 | 25.5 | 43.3 | 29.9 | 43.6 | 31.3 |
| NavFoM Four views* | 40.1 | 27.1 | 45.4 | 32.6 | 45.2 | 31.9 |
Zero-shot 4-view 45.2 SR (Val Unseen) beats the best prior baseline MTU3D (40.8 SR); even NavFoM single-view (43.6) already exceeds it. The SPL gap is even bigger (31.9 vs 12.1).
| Single Target SR / TR | Distracted Target SR / TR | |
|---|---|---|
| EVT (with SoM+GPT-4o) | 32.5 / 49.9 | 15.7 / 35.7 |
| Uni-NaVid | 25.7 / 39.5 | 11.3 / 27.4 |
| TrackVLA | 85.1 / 78.6 | 57.6 / 63.2 |
| NavFoM Single view | 85.0 / 80.5 | 61.4 / 68.2 |
| NavFoM Four views | 88.4 / 80.7 | 62.0 / 67.9 |
| Method | Camera | VLM-Based | PDMS ↑ |
|---|---|---|---|
| Human | - | - | 94.8 |
| Constant Velocity | - | - | 21.6 |
| Ego Status MLP | - | - | 65.6 |
| UniAD | ✓ | - | 83.4 |
| PARA-Drive | ✓ | - | 84.0 |
| LAW | ✓ | - | 84.6 |
| DrivingGPT | ✓ | ✓ | 82.4 |
| NavFoM (Eight views) | ✓ | ✓ | 84.3 |
NavFoM on 8-view rig competitive with task-specific autonomous driving baselines despite training on a unified data mix.
NavFoM SOTA or highly competitive:
- HM3D-OVON (Val Unseen SR): 45.2 (4-view) / 43.6 (single-view) vs best prior baseline MTU3D 40.8
- RxR Single-View: 57.4 vs 51.8/49.3
- RxR Multi-View: 64.4 vs 56.3
- EVT-Bench DT: 62.0 vs 57.6
- NavSim: 84.3 vs 84.0/84.6/83.4/82.4
Cross-task and cross-embodiment, including humanoid robots, quadrupeds, drones, and wheeled robots. Tasks include "Follow the man in the white T-shirt and jeans," "Move to the TV," "Track the man," "Go up to the footbridge."
Token budget B=2048 vs B=1024 on RxR Val-Unseen:
| Method | NE ↓ | SR ↑ | SPL ↑ | nDTW ↑ |
|---|---|---|---|---|
| B=1024, Uniform Sampling | 5.33 | 59.7 | 49.6 | 57.9 |
| B=1024, Linear Probability Sampling | 5.28 | 61.2 | 50.9 | 58.9 |
| B=1024, BATS | 4.98 | 62.5 | 53.9 | 64.1 |
| B=2048, Token Merging (Uni-NaVid) | 5.01 | 63.2 | 54.9 | 64.4 |
| B=2048, Uniform Sampling | 4.90 | 62.4 | 54.0 | 63.9 |
| B=2048, Linear Probability Sampling | 4.89 | 63.0 | 54.6 | 64.8 |
| B=2048, BATS | 4.74 | 64.4 | 56.2 | 65.8 |
BATS wins at both budgets; only -1.4 nDTW drop halving from 2048→1024 vs -6.0 / -5.2 drops for baselines.
| TVI variant | NE | SR | SPL | nDTW |
|---|---|---|---|---|
| Viewpoint-history positional embedding (HAMT) | 6.27 | 52.3 | 46.3 | 58.7 |
| Individual learned special tokens | 5.52 | 59.1 | 52.0 | 59.6 |
| Handcraft Tokens (no Pangle/time) | 6.06 | 53.6 | 46.1 | 58.0 |
| TVI Tokens (Eq. 3) | 4.74 | 64.4 | 56.2 | 65.8 |
+12 SR over HAMT-style PE, +5.3 SR over special-tokens — separating viewpoint and temporal identifiers from visual content matters.
Comparing single-task training, +50% other-task data, +100% other-task data:
- Searching improves from 10.3% to 45.2% (training conditions are single-view + closed-set; multi-task pretraining helps the model generalize to multi-view + open-vocabulary evaluation).
- Tracking improves from 12.6% to 62.0%.
- The paper attributes large gains to mismatches between training and evaluation distributions that multi-task data mitigates.
| Method | Samples |
|---|---|
| NaVid | 1.2 M |
| Uni-NaVid | 5.9 M |
| NavFoM | 12.7 M |
The paper does not have a dedicated "Limitations" section. Implicit/discussed limitations:
- β scheduling. Constant β=10 to amplify nav loss vs QA loss; the authors flag adaptive β as future work.
- No DexGraspNet/RLBench/CALVIN evaluation — focus is navigation only; manipulation cross-application untested.
- Token budget. Real-world deployment uses 1600-token budget on a single 4090; for very long horizons (T ≈ 1120 in 4-camera setup at B=2048), Eq. 5 becomes unsolvable for k — they note this is rare in practice (most timesteps ≈122 in RxR).
- Failure cases discussed in Appendix F (not detailed in main text).
- Closed-set training data — VLN-CE / HM3D-OVON / EVT-Bench / nuScenes datasets dominate; the paper notes these are simulation-heavy.
NavFoM pushes navigation toward the single-model, many-embodiments regime that VLAs are pursuing for manipulation.
- vs NaVid (1.2 M, single embodiment) / Uni-NaVid (5.9 M, single config) / NaVILA (legged-only): NavFoM scales the data 10× and unifies four task families across ground, aerial, and automotive platforms.
- vs StreamVLN-RGB-only: competitive with previous best RGB-only single-view (+0.5 R2R / +5.6 RxR SR).
- vs HNR / BEVBert:** beats them on RxR even when those use RGB+depth+odometry; NavFoM uses RGB only.
- vs TrackVLA: comparable single-view tracking, better multi-view (+3.3 SR on single-target, +0.4 SR on distracted target).
- vs UniAD / PARA-Drive / LAW (autonomous driving): competitive PDMS (84.3 vs 83.4–84.6) without task-specific autonomous-driving training pipelines.
- vs UniVLA / X-VLA: the manipulation analogues of NavFoM. NavFoM's key contribution — TVI tokens + BATS — is the navigation-specific answer to "how does one model consume arbitrary camera rigs in real time."
The TVI + BATS recipe is concretely transferable: any cross-embodiment system that needs to encode (camera angle, time, task type) within an LLM context can adopt the same identifier-token + budget-aware sampling design.
- OpenReview: https://openreview.net/forum?id=kkBOIsrCXh
- PDF: https://openreview.net/pdf?id=kkBOIsrCXh
- Project page: https://pku-epic.github.io/NavFoM-Web/
- REI-Bench (instruction understanding for embodied agents)
- UniVLA (cross-embodiment manipulation analogue)
- X-VLA
← Back to ICLR-2026