ICML 2026 OXE AugE - Heungwoo/research GitHub Wiki

OXE-AugE — Tripling Open X-Embodiment with simulated robot augmentation for cross-embodiment policies

Venue: ICML 2026 (Poster) Category: Benchmark Affiliations: UC Berkeley (EECS); UT Austin (CS); GRASP Laboratory, University of Pennsylvania Traction (2026-06): 5 citations (arXiv)

AugE-Toolkit pipeline: fusing a learned SAM2 mask with a simulation-rendered mask to segment the source robot, inpaint the background, and replay/composite a target embodiment (Figure 1 from Ji et al., 2026)

Problem

Generalist robot policies need large, diverse data, but re-collecting demonstrations for every new arm-and-gripper combination is prohibitively expensive. The Open X-Embodiment (OXE) dataset aggregates demonstrations from over 60 real-world robot datasets, yet it is highly imbalanced: over 85% of real trajectories come from just four robots (Franka, xArm, Kuka iiwa, Google Robot), and most constituent datasets tie a single robot to a fixed scene. This risks overfitting to robot–scene combinations, and policies such as Octo, OpenVLA, GR00T, and π0 still require finetuning on new robots. The authors ask whether existing data can be augmented — swapping in new embodiments — to improve cross-embodiment transfer, generalization, and robustness.

Method

The work builds on cross-painting, a three-stage per-frame pipeline: (i) source-robot segmentation, (ii) background inpainting, and (iii) augmented-robot replay and compositing. Prior learning-based methods (e.g., RoVi-Aug) edit pixels with diffusion models but lack kinematic guarantees and scale poorly; simulation-based methods (e.g., Mirage) render with known camera parameters but need calibration unavailable in large offline datasets.

AugE-Toolkit keeps simulation-based physical accuracy while remaining applicable to uncalibrated data via three components:

  1. Fusion of simulation and learned masks — a SAM2 model fine-tuned on 20 trajectories from each of 16 OXE datasets supplies appearance-aligned masks, fused with geometrically accurate simulation masks through translation alignment (grid-search IoU), distance pruning, and union+morphological smoothing. IoU can also flag/filter bad data.
  2. Automatic base position tuning — iteratively samples ±Δ offsets along (x,y,z), halving the step until max end-effector tracking error falls below 1 cm, discarding unreachable trajectories.
  3. Scalable application across many target embodiments with minimal manual effort.

Applying this at scale yields OXE-AugE: OXE augmented with 9 robot embodiments, over 4.4 million trajectories — more than triple the original OXE.

Results

Simulation study on scaling robot augmentation (Figure 3 from Ji et al., 2026)

A systematic simulation study (100 trials/task, five Robosuite tasks) shows:

  • Robustness: under lighting shifts and occlusions, N-robot augmentation consistently beats 1-robot augmentation and source-only training on the original embodiment.
  • Transfer: training on the augmented target robot reaches 26–65% success across four robots; adding source data improves by 9–19%.
  • Generalization to unseen robots: (N−1)× augmentation substantially outperforms 1× augmentation, often rivaling policies trained directly on the augmented target.

In physical experiments (10 trials/task, 40/embodiment), finetuning on OXE-AugE improved average success by 24% for OpenVLA-OFT and 45% for π0 on unseen robot–gripper combinations across four real-world tasks. On the novel Robotiq++ embodiment, finetuned policies reached 75% (OpenVLA-OFT) and 82% (π0). The simulator-based pipeline also beat diffusion-based RoVi-Aug, which caused a 27–30% drop in final policy performance due to misaligned grippers and geometric inconsistencies.

Significance

OXE-AugE reframes robot augmentation from pairwise transfer into a scalable data pipeline, demonstrating that increasing the number and diversity of augmented embodiments improves generalization to unseen robots and robustness on original robots. The open-source dataset (>4.4M trajectories) and toolkit give the community a practical lever for cross-embodiment scaling without new hardware data collection.

Links

← Back to ICML-2026