IROS 2026 Cross Embodiment WM - Heungwoo/research GitHub Wiki

IROS 2026 — Scaling Cross-Embodiment World Models for Dexterous Manipulation

Venue: IROS 2026 (Pittsburgh) · paper #2451 · Shanghai Jiao Tong University · UC San Diego (Hao Su, Henrik Christensen) · MIT (Yilun Du) (He, Ai, Mu, Liu, Wan, Fu, Du, Christensen, Su). Representative of: world models × cross-embodiment × dexterity — a particle-based world model as a shared geometric interface across human and robot hands. Companions: World Models · Cross-Embodiment · Dexterous-Hand Data Pyramid · IROS 2026 survey.

1. Problem

Cross-embodiment learning wants generalist robots that learn across morphologies, but kinematics and action-space differences block data sharing and control transfer. The question posed: what structure can be shared across embodiments despite these differences?

2. Method

The answer: the physical interactions embodiments induce can be modeled in a shared geometric space, so a world model becomes a common interface for learning and control.

  • Particle representation — human and robot hands are both represented as sets of 3D particles; actions are end-effector particle displacement fields. This abstracts away embodiment-specific joint spaces while preserving interaction-relevant geometry/motion.
  • Graph-based world model trained on random interaction data from diverse simulated robot hands + real human hands.
  • Deployment — integrate the WM with model-predictive control (MPC) on new hardware.

3. Results (three findings)

  • Embodiment diversity → generalization: more training embodiments improves transfer to unseen hands.
  • Sim + real beats either alone: appropriately combining simulated and real interaction data outperforms single-source.
  • One model, distinct kinematics: the same learned WM enables control on robot hands with different kinematics and DoF (rigid and deformable manipulation).

4. Why it matters (WAM × cross-embodiment lens)

This is the IROS 2026 datapoint that makes the world model the cross-embodiment interface itself — a different bet from action-space unification (RT-X), latent-action codes (UniVLA), or soft prompts (X-VLA) in Cross-Embodiment. By modeling shared interaction geometry (particles) rather than shared actions, it sidesteps the joint-space mismatch that blocks cross-hand transfer — directly relevant to the data-pyramid's L4 retargeting bottleneck (here: no retargeting, a shared particle space instead). It also fits the survey §5.1 "WAM as scaffold" trend: the WM is a shared representation + MPC model, not an end-to-end policy.

Limitations (reviewer): particle representation + MPC is a classical-control pipeline (not a learned end-to-end VLA); random-interaction training data may under-cover task-directed behavior; the sim+real particle estimation quality bounds transfer.

5. Links

← Back to IROS 2026 survey · Home