ICLR 2026 Executable 3DGS Nav - Heungwoo/research GitHub Wiki

SAGE-3D — physically executable 3D Gaussian for embodied navigation

Venue: ICLR 2026 (Poster) · Authors: Bingchen Miao, Rong Wei, Zhiqi Ge, Xiaoquan Sun, Shiqi Gao, Jingzhe Zhu, Renhan Wang, Siliang Tang, Jun Xiao, Rui Tang, Juncheng Li · arXiv: 2510.21307 · Category: Embodied navigation / VLN (environments) · Trend tag: Semantically + physically aligned 3D Gaussian Splatting for VLN.

Approach diagram

flowchart LR
  GS[3D Gaussian Splatting scene<br/>photorealistic but inert] --> OCS
  subgraph SAGE["SAGE-3D paradigm"]
    OCS[Object-Centric Semantic Grounding<br/>object-level annotations]
    PAE[Physics-Aware Execution Jointing<br/>collision objects + physical interfaces]
  end
  OCS --> Env[Executable, semantically +<br/>physically aligned environment]
  PAE --> Env
  Env --> Data[InteriorGS: 1K annotated scenes<br/>SAGE-Bench: 2M VLN samples]
  Data --> Agent[VLN agent training / eval]
Loading

Problem

3D Gaussian Splatting (3DGS) renders photorealistically in real time, but as a navigation environment it is inert: it lacks fine-grained object semantics and physical executability (no collision, no interaction). That makes raw 3DGS unusable as a training/evaluation world for Vision-Language Navigation agents.

Method

SAGE-3D (Semantically and physically Aligned Gaussian Environments for 3D navigation) upgrades 3DGS into an executable VLN environment via two components:

  • Object-Centric Semantic Grounding: adds object-level fine-grained annotations to the Gaussian scene so instructions can reference specific objects.
  • Physics-Aware Execution Jointing: embeds collision objects into the 3DGS and constructs physical interfaces, enabling interaction and constraint-respecting motion.

Released artifacts:

  • InteriorGS: ~1K object-annotated 3DGS indoor scenes, spanning diverse indoor/outdoor settings (homes, gyms, concert halls, swimming pools, amusement parks) — over 554k object instances across 755 categories.
  • SAGE-Bench: the first 3DGS-based VLN benchmark, with 2M VLN data points.

Results

The paper reports a +31% improvement on the VLN-CE Unseen task with the SAGE-3D environment, while noting convergence challenges relative to standard methods. (Additional metric values omitted here pending the camera-ready tables.)

Significance

SAGE-3D is the first attempt to make 3D Gaussian Splatting a first-class navigation environment — semantically grounded and physically executable — rather than just a renderer. The InteriorGS scenes and SAGE-Bench give the VLN community a photorealistic-yet-interactive world, complementing physics-aware 3DGS benchmarks like NavBench-GS.

Links

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️