ICLR 2026 CrossEmb Offline RL - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 · Authors: Haruki Abe, Takayuki Osa, Yusuke Mukuta, Tatsuya Harada (University of Tokyo / RIKEN AIP) · arXiv:2602.18025 · Category: Data / memory / representation for manipulation · Trend tag: Cross-embodiment / offline RL
Official title: Cross-Embodiment Offline Reinforcement Learning for Heterogeneous Robot Datasets.
flowchart LR
Data[Heterogeneous trajectories<br/>16 robot platforms<br/>expert + suboptimal] --> Group[Embodiment grouping<br/>cluster by morphological similarity]
Group --> OffRL[Offline RL<br/>group-gradient update]
OffRL --> Prior[Universal control prior]
Prior --> FT[Downstream fine-tuning]
Scalable robot pre-training is bottlenecked by the cost of high-quality demonstrations per platform. Behavior cloning ignores abundant suboptimal data, and naively pooling trajectories from many morphologies causes conflicting gradients across embodiments that impede learning.
The paper unites offline RL (to exploit both expert and plentiful suboptimal data) with cross-embodiment learning (to aggregate heterogeneous trajectories into universal control priors), and analyzes the strengths and limits of this paradigm. Its key fix is an embodiment-based grouping strategy: robots are clustered by morphological similarity and the model is updated with a per-group gradient. This simple, static grouping reduces inter-robot gradient conflict.
Evaluated on a constructed suite of locomotion datasets spanning 16 distinct robot platforms. The combined offline-RL + cross-embodiment approach outperforms pure behavior cloning, especially on data rich in suboptimal trajectories. As the suboptimal fraction and number of robot types grow, gradient conflict worsens; the static morphology-based grouping substantially reduces conflicts and outperforms existing conflict-resolution methods. (Numeric success/reward values are in the paper and omitted here pending confirmation.)
Shows offline RL — not just imitation — can scale across morphologies if cross-embodiment gradient interference is managed, and that a cheap morphology-similarity grouping beats more elaborate multi-task conflict-resolution schemes. A data-centric recipe for turning messy, mixed-quality, multi-robot datasets into reusable control priors.
← Back to ICLR-2026