RSS 2026 Learning to Evolve - Heungwoo/research GitHub Wiki

Learning to Evolve: Multi-modal Interactive Fields for Robust Humanoid Navigation in Dynamic Environments

Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Humanoids · paper #23 Authors: Peifeng Jiang, Hong Liu, Jin Jin, Wenshuai Wang, Xia Li arXiv: 2605.21935 · program page

Summary compiled from the arXiv paper (v1); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Multi-modal Interactive Fields overview (Figure 1 of arXiv 2605.21935, © the authors)

The three coupled fields rendered over a reconstructed office scene: (1) an Appearance Field (3DGS semantic heatmap) for dense semantic grounding, (2) a Spatial Field (textual scene graph) that updates when an object is relocated — shown as the humanoid re-planning from the old to the new path — and (3) a Geometry Field that recovers a textured object mesh to verify the terminal interaction pose before manipulation.

Problem

Manipulation-oriented humanoid navigation needs scene memory that survives two coupled failure modes: locomotion-induced semantic-geometric distortion (bipedal camera jitter corrupts the map) and map-reality mismatch when objects are relocated, removed, or added. Existing semantic mapping and scene-graph systems (e.g., HOV-SG, ConceptGraphs) assume stable camera trajectories, static scenes, or coarse object geometry, so stale memory can drive the robot to obsolete coordinates or unsafe interaction poses.

Method

The Multi-modal Interactive Field (MIF) couples three fields in a closed perception-adaptation loop: a confidence-aware semantic 3D Gaussian Splatting Appearance Field whose per-Gaussian reliability gate suppresses gait-corrupted primitives; a topological scene-graph Spatial Field with a multi-modal discrepancy score D (positional + semantic + relational evidence, ROC-selected threshold τ = 0.45) that triggers spatially local memory updates instead of global re-scans; and an on-demand Geometry Field that reconstructs watertight object meshes with a Flow-Matching model from target-centered views to check Interaction Pose Safety (IPS: collision, kinematic reachability, stability). Feature distillation to 32-D semantics cuts VRAM below 4 GB. The system runs on a Unitree G1 humanoid with an RTX 4090 workstation in a ~100 m² office with 100+ object categories.

Results

MIF achieves 92% semantic-grounding success with 0.18 m mean distance error under humanoid locomotion (vs 74%/0.35 m for HOV-SG, 65%/0.42 m for Feature-Splatting) and 94% IPS success with zero observed collisions and 0.12 m terminal error (point-cloud baseline: 62%, 32% collision rate). On scene-change tasks without a global rescan, MIF (Post-Update) reaches 94% / 98% / 86% success for relocated / removed / added objects, versus 12% / 10% / 0% for static HOV-SG memory. Feature distillation reduces semantic memory footprint by 91.4%; the Appearance Field updates incrementally at 22 FPS while mesh generation takes about 6.2 s asynchronously.

Significance

A system-level answer to making humanoid scene memory revisable: the same reliability signal gates rendering, graph construction, and safety checking, and the 12%→94% relocation jump quantifies how much static scene-graph memory costs in dynamic homes and offices. Related wiki threads: Review-Humanoid-VLA · Review-System-0-1-2.

← Back to RSS 2026 survey · RSS-2026-Papers · Home