CoRL 2025 X Sim - Heungwoo/research GitHub Wiki

X-Sim โ€” Cross-Embodiment Learning via Real-to-Sim-to-Real

Venue: CoRL 2025 (Oral) ยท Authors: Prithwish Dan, Kushal Kedia et al. (Choudhury lab) โ€” Cornell ยท arXiv: 2505.07096 Category: Cross-Embodiment Trend tag: Embodiment-agnostic supervision

Approach diagram

flowchart LR
  HV[RGBD human video] --> SIM[Real-to-sim:<br/>photorealistic scene +<br/>tracked object trajectories]
  SIM --> RL[RL policy in sim<br/>w/ object-centric rewards]
  RL --> POL["Distill into image-conditioned<br/>diffusion policy<br/>(randomized views/lighting)"]
  POL --> REAL[Real robot<br/>w/ online domain adaptation]
Loading

Problem

Cross-embodiment learning from human video usually tries to retarget human motion to robot motion โ€” brittle, since arm kinematics and end-effectors differ. What's actually transferable across embodiments is what happens to the objects, not the limbs.

Method

X-Sim uses object motion as the transferable signal via a three-stage real-to-sim-to-real pipeline:

  1. Real-to-sim: reconstruct a photorealistic simulation from an RGBD human video and track object trajectories to define object-centric rewards.
  2. Sim training: use those rewards to train a reinforcement-learning policy in simulation, then distill it into an image-conditioned diffusion policy using synthetic rollouts rendered under randomized viewpoints and lighting.
  3. Sim-to-real: an online domain-adaptation step replays real robot actions in sim to align visual encoders between domains, closing the sim-to-real gap without any robot teleoperation data.

The transferable signal is object motion, so no human-to-robot motion retargeting is needed.

Results

CoRL 2025 Oral. Evaluated on 5 manipulation tasks across 2 environments:

  • +30% task progress on average over hand-tracking and sim-to-real baselines.
  • Matches behavior cloning with ~10x less data-collection time (human video substitutes for teleoperation).
  • Generalizes to new camera viewpoints and test-time environment changes.

Significance

A clean, principled cross-embodiment recipe that sidesteps the retargeting trap. Threads into ICLR 2026's generalist multi-embodiment VLAs (X-VLA, XR-1, UniVLA).

Links

๐Ÿ“– In-depth cross-paper review

For a taxonomy of cross-embodiment approaches across ICLR 2026 + CoRL 2025 + ฯ€0.7: Cross-Embodiment Training Review.

Related pages

โ† Back to CoRL-2025

โš ๏ธ **GitHub.com Fallback** โš ๏ธ