ICLR 2026 CrossEmb Offline RL - Heungwoo/research GitHub Wiki

Cross-Embodiment Offline RL — universal control priors from heterogeneous robots

Venue: ICLR 2026 · Authors: Haruki Abe, Takayuki Osa, Yusuke Mukuta, Tatsuya Harada (University of Tokyo / RIKEN AIP) · arXiv:2602.18025 · Category: Data / memory / representation for manipulation · Trend tag: Cross-embodiment / offline RL

Official title: Cross-Embodiment Offline Reinforcement Learning for Heterogeneous Robot Datasets.

Approach diagram

flowchart LR
  Data[Heterogeneous trajectories<br/>16 robot platforms<br/>expert + suboptimal] --> Group[Embodiment grouping<br/>cluster by morphological similarity]
  Group --> OffRL[Offline RL<br/>group-gradient update]
  OffRL --> Prior[Universal control prior]
  Prior --> FT[Downstream fine-tuning]
Loading

Problem

Scalable robot pre-training is bottlenecked by the cost of high-quality demonstrations per platform. Behavior cloning ignores abundant suboptimal data, and naively pooling trajectories from many morphologies causes conflicting gradients across embodiments that impede learning.

Method

The paper unites offline RL (to exploit both expert and plentiful suboptimal data) with cross-embodiment learning (to aggregate heterogeneous trajectories into universal control priors), and analyzes the strengths and limits of this paradigm. Its key fix is an embodiment-based grouping strategy: robots are clustered by morphological similarity and the model is updated with a per-group gradient. This simple, static grouping reduces inter-robot gradient conflict.

Results

Evaluated on a constructed suite of locomotion datasets spanning 16 distinct robot platforms. The combined offline-RL + cross-embodiment approach outperforms pure behavior cloning, especially on data rich in suboptimal trajectories. As the suboptimal fraction and number of robot types grow, gradient conflict worsens; the static morphology-based grouping substantially reduces conflicts and outperforms existing conflict-resolution methods. (Numeric success/reward values are in the paper and omitted here pending confirmation.)

Significance

Shows offline RL — not just imitation — can scale across morphologies if cross-embodiment gradient interference is managed, and that a cheap morphology-similarity grouping beats more elaborate multi-task conflict-resolution schemes. A data-centric recipe for turning messy, mixed-quality, multi-robot datasets into reusable control priors.

Links

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️