ICLR 2026 Seeing To Experiencing - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 · Authors: Honglin He, Yukai Ma, Wayne Wu, Bolei Zhou (UCLA) · arXiv: 2507.22028 · Category: Embodied navigation (navigation foundation models) · Trend tag: Video pretraining + RL post-training for interactive navigation.
flowchart LR
Video[Large-scale real-world videos] --> Pre
subgraph Pre["Offline pretraining (Seeing)"]
AGDM[Anchor-Guided Distribution Matching<br/>anchor-based supervision]
end
AGDM --> Base[Navigation foundation model<br/>generalizable from video]
Base --> Post
subgraph Post["RL post-training (Experiencing)"]
RAM[Residual-Attention Module<br/>reactive behaviors w/o erasing priors]
Sim[Simulation interaction]
end
RAM --> Agent[Interactive, safe nav agent]
Sim --> RAM
Agent --> Bench[NavBench-GS<br/>3DGS reconstructions + physics]
Navigation foundation models trained purely on offline video ("seeing") generalize broadly but hit diminishing returns from scale — they never experience the consequences of their actions, so they lack reactive, interactive behavior and safety in real urban environments. The paper asks how to add interaction ("experiencing") without erasing the generalization learned from video.
S2E (Seeing-to-Experiencing) combines video pretraining with RL post-training:
- Anchor-Guided Distribution Matching (offline pretraining): stabilizes learning and models diverse motion patterns through anchor-based supervision over large-scale real-world videos.
- Residual-Attention Module (RL post-training): obtains reactive behaviors from simulation without erasing the model's pretrained knowledge — the residual design protects the video-learned priors while adding interactivity.
- NavBench-GS: an end-to-end evaluation benchmark built on photorealistic 3D Gaussian Splatting reconstructions of real scenes that incorporate physical interactions, used to systematically assess generalizability and safety.
The paper also offers a systematic comparison of RL vs. supervised fine-tuning for scaling navigation models.
S2E maintains the generalizability acquired from large-scale video while improving interactivity and safety via simulation RL, evaluated on NavBench-GS. (Exact metric values omitted here pending the camera-ready tables.)
S2E is a clean statement of the "seeing then experiencing" scaling recipe: video gives broad generalization, RL in physics-aware 3DGS sim adds the reactive, safe behavior video alone cannot. The residual-attention trick — adding interaction without catastrophic forgetting of video priors — is broadly applicable to embodied foundation models.
- ICLR 2026 Survey
- OmniVLA (navigation)
- NavFoM (navigation foundation model)
- SAGE-3D (physics-aware 3DGS for navigation)
← Back to ICLR-2026