RSS 2026 HoMMI - Heungwoo/research GitHub Wiki

HoMMI โ€” Learning Whole-Body Mobile Manipulation from Human Demonstrations

Venue: RSS 2026 (Imitation Learning session) ยท Authors: Xiaomeng Xu, Jisang Park, Han Zhang, Eric Cousineau, Aditya Bhat, Jose Barreiros, Dian Wang, Jeannette Bohg, Shuran Song โ€” Stanford ร— Toyota Research Institute Category: Robot-free data collection + whole-body mobile manipulation Trend tag: RSS 2026 thread 2 โ€” human data & cross-embodiment transfer arXiv: 2603.03243 ยท project ยท code

Compiled from the verified RSS 2026 abstract and the paper's Fig. 1.

Key figure

HoMMI overview (Figure 1 of arXiv 2603.03243, ยฉ the authors)

Figure 1 of the paper. (a) Scalable data collection: a human with dual UMI-style handheld grippers plus head-mounted egocentric sensing performs household tasks (towel-to-bin, cart pushing, cushion handling) โ€” no robot in the loop. (b) The challenge: side-by-side human demonstration vs robot deployment showing the two gaps the policy design must bridge โ€” the visual gap (human hands/grippers visible in egocentric views vs robot arms) and the kinematic gap (human head height/reach vs the shorter mobile bimanual robot). (c) Resulting skills: whole-body manipulation with active perception, long-horizon navigation while carrying, and bimanual coordination โ€” all learned from the robot-free demonstrations.

Problem

UMI-style handheld interfaces made robot-free data collection work for tabletop arms โ€” but mobile manipulation needs global context (where the body is going, what the head sees) that a wrist-mounted gripper camera can't provide, and adding egocentric sensing widens the human-to-robot embodiment gap in both observation and action spaces.

Method

  • Interface: UMI augmented with egocentric sensing โ€” portable, robot-free, scalable capture of whole-body mobile-manipulation demonstrations.
  • Cross-embodiment hand-eye policy design to bridge the widened gap: an embodiment-agnostic visual representation, a relaxed head-action representation, and a whole-body controller that realizes commanded hand-eye trajectories through coordinated whole-body motion under robot-specific physical constraints.
  • Full code, data, and hardware design release stated.

Results (as reported)

  • Enables long-horizon mobile manipulation requiring bimanual and whole-body coordination, navigation, and active perception โ€” trained entirely from robot-free human demonstrations.

Significance

The Song/Bohg/TRI lineage (UMI โ†’ this) extending the robot-free data thesis from arms to whole bodies โ€” the practical complement to ฮจโ‚€'s teleop-based humanoid pipeline and EgoHumanoid's (#204) egocentric loco-manipulation co-training. The design insight worth retaining: the fix for a bigger embodiment gap is not more data but interface-level abstraction (hand-eye trajectories + a constraint-aware whole-body controller as the translator) โ€” the mobile-manipulation echo of the canonical-representation moves seen across RSS 2026. See Review-Humanoid-VLA and CoRL-2025-DexUMI for the interface lineage.

โ† RSS 2026 survey ยท Home