ICLR 2026 Seeing To Experiencing - Heungwoo/research GitHub Wiki

S2E — From Seeing to Experiencing: scaling navigation FMs with RL

Venue: ICLR 2026 · Authors: Honglin He, Yukai Ma, Wayne Wu, Bolei Zhou (UCLA) · arXiv: 2507.22028 · Category: Embodied navigation (navigation foundation models) · Trend tag: Video pretraining + RL post-training for interactive navigation.

Approach diagram

flowchart LR
  Video[Large-scale real-world videos] --> Pre
  subgraph Pre["Offline pretraining (Seeing)"]
    AGDM[Anchor-Guided Distribution Matching<br/>anchor-based supervision]
  end
  AGDM --> Base[Navigation foundation model<br/>generalizable from video]
  Base --> Post
  subgraph Post["RL post-training (Experiencing)"]
    RAM[Residual-Attention Module<br/>reactive behaviors w/o erasing priors]
    Sim[Simulation interaction]
  end
  RAM --> Agent[Interactive, safe nav agent]
  Sim --> RAM
  Agent --> Bench[NavBench-GS<br/>3DGS reconstructions + physics]
Loading

Problem

Navigation foundation models trained purely on offline video ("seeing") generalize broadly but hit diminishing returns from scale — they never experience the consequences of their actions, so they lack reactive, interactive behavior and safety in real urban environments. The paper asks how to add interaction ("experiencing") without erasing the generalization learned from video.

Method

S2E (Seeing-to-Experiencing) combines video pretraining with RL post-training:

  • Anchor-Guided Distribution Matching (offline pretraining): stabilizes learning and models diverse motion patterns through anchor-based supervision over large-scale real-world videos.
  • Residual-Attention Module (RL post-training): obtains reactive behaviors from simulation without erasing the model's pretrained knowledge — the residual design protects the video-learned priors while adding interactivity.
  • NavBench-GS: an end-to-end evaluation benchmark built on photorealistic 3D Gaussian Splatting reconstructions of real scenes that incorporate physical interactions, used to systematically assess generalizability and safety.

The paper also offers a systematic comparison of RL vs. supervised fine-tuning for scaling navigation models.

Results

S2E maintains the generalizability acquired from large-scale video while improving interactivity and safety via simulation RL, evaluated on NavBench-GS. (Exact metric values omitted here pending the camera-ready tables.)

Significance

S2E is a clean statement of the "seeing then experiencing" scaling recipe: video gives broad generalization, RL in physics-aware 3DGS sim adds the reactive, safe behavior video alone cannot. The residual-attention trick — adding interaction without catastrophic forgetting of video priors — is broadly applicable to embodied foundation models.

Links

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️