ICML 2026 Cross Embodiment Robot Foundation World Models - Heungwoo/research GitHub Wiki

Cross-Embodiment Robot Foundation World Models with Latent Actions — A unified latent action space that scales positively with embodiment count

Venue: ICML 2026 (Poster) Category: World Model Affiliations: Huang Huang, Sriram Yenamandra, Arjun Majumdar, Elie Aljalbout, Tushar Nagarajan, Jimmy Yang, Akshara Rai, Michael Rabbat, Li Fei-Fei, Jiajun Wu, Tingfan Wu, Franziska Meier

Problem

The diversity of robot embodiments and their action spaces makes it challenging to build robot world models that generalize across different embodiments. A natural baseline is to condition a world model on explicit motion labels (joint commands, end-effector deltas, etc.), but these explicit action spaces differ from robot to robot. The authors show that conditioning on explicit labels creates disjoint action spaces across embodiments, which limits downstream task performance when the model is adapted to a previously unseen robot — and, counterintuitively, makes things worse as more embodiments are added during pretraining.

Method

The paper introduces the Latent Action Conditioned Robot World Model (LAC-WM), which operates within a learned unified latent action space shared across diverse embodiments. Rather than conditioning predictions on each robot's native, explicit action representation, LAC-WM maps actions into a common latent space, so transitions from different robots are described in compatible terms. This unified action space is the key mechanism that lets the world model transfer to new embodiments.

The contrast model is EAC-WM (Explicit Action Conditioned World Model), conditioned on explicit motion labels, which serves as the baseline throughout the study.

flowchart LR
  subgraph EAC-WM (baseline)
    E1[Robot A explicit actions] --> EX[Disjoint action spaces]
    E2[Robot B explicit actions] --> EX
    EX --> ED[Limited transfer; degrades<br/>as #embodiments grows]
  end
  subgraph LAC-WM (ours)
    L1[Robot A actions] --> U[Unified latent action space]
    L2[Robot B actions] --> U
    U --> LD[World model transfers to<br/>unseen robots; scales positively]
  end
Loading

Results

Both models are evaluated on a dexterous manipulation task, with pretraining across multiple embodiments and adaptation to previously unseen robots.

  • LAC-WM achieves up to a 46.7% improvement in performance over EAC-WM when adapting to unseen embodiments.
  • Positive scaling: LAC-WM's downstream performance scales positively with the number of embodiments used during pretraining — more pretraining robots help.
  • Negative scaling for the baseline: EAC-WM's disjoint action space causes downstream performance to decrease as the number of pretraining embodiments increases, the opposite of the desired trend.

These results isolate the action-space representation (latent vs. explicit) as the decisive factor: the same world-model recipe either benefits or suffers from added embodiments depending solely on whether the action space is unified.

Significance

The work pinpoints a clean, important lesson for cross-embodiment robot foundation models: a unified latent action space is necessary for efficient cross-embodiment learning. Without it, naively pooling data from many robots can hurt a world model because the action conditioning becomes fragmented. By making world-model performance scale positively with embodiment diversity, LAC-WM addresses a central obstacle to building robot foundation world models that improve as more robot data is collected.

Links

← Back to ICML-2026

⚠️ **GitHub.com Fallback** ⚠️