RSS 2026 R2RGen - Heungwoo/research GitHub Wiki

R2RGen: Real-to-Real 3D Data Generation for Spatially-generalized Robotic Manipulation

Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Manipulation 3 · paper #121 Authors: Xiuwei Xu, Angyuan Ma, Hankun Li, Bingyao Yu, Zheng Zhu, Jie Zhou, Jiwen Lu arXiv: 2510.08547 · program page

Summary compiled from the arXiv paper (v2; arXiv title "R2RGEN: Real-to-Real 3D Data Generation for Spatially Generalized Manipulation"); all numbers quoted from the paper. Trend context: RSS 2026 survey.

R2RGen simulator-free data generation (Figure 1 of arXiv 2510.08547, © the authors)

From one human-collected demo, R2RGen's three modules (shape processing, group backtrack, camera processing) directly edit pointcloud-action pairs to produce many spatially diverse real-world demos. Bottom/right panels show the covered settings: bimanual and spatially-related tasks, object location and viewpoint changes, and cross-platform camera setups (wrist camera + UMI, exterior camera).

Problem

Spatial generalization — robustness to different placements of objects, environment, and the robot itself — is the main driver of the huge demonstration counts needed for visuomotor imitation learning, and gets worse for mobile manipulators whose base position varies the viewpoint. Prior data-generation approaches rely on simulators/rendering (sim-to-real gap) or are limited to fixed-base, fixed-viewpoint settings.

Method

R2RGen is a simulator- and rendering-free real-to-real pipeline that directly augments pointcloud observation-action pairs in a shared 3D space. Three stages: (1) pre-processing — scene parsing (segmenting K objects, arm, environment) plus lightweight trajectory annotation into skill/motion segments with target and in-hand object IDs; (2) augmentation — objects and robot position are rearranged with a group-wise backtracking strategy that resolves causal conflicts when skills involve multiple related objects; (3) camera-aware 3D post-processing that reshapes generated pointclouds to match the real RGB-D sensor's distribution (visibility/noise) in the deployed camera frame, supporting exterior, wrist (UMI-style), and mixed camera setups. The downstream policy is iDP3, trained purely on generated data.

Results

On 8 real tasks (2 simple, 4 complex, 2 bimanual; single-arm platform plus a MobileAloha-style dual-PiPER mobile base): with one source demo, baseline iDP3 gets 3.4% average success; +DemoGen reaches 15.6-18.8% on the tasks it can handle, while +R2RGen averages 40.3% — comparable to training on 25 human demonstrations (41.0%) and beating 40 demos on several hard tasks. On exterior-wrist and wrist-wrist camera settings R2RGen reaches 28.1-50.0% from one demo, where naive cross-camera deployment fails catastrophically. Success saturates as generated demos grow (attributed to iDP3's lightweight PointNet capacity), and a SAM 3D extension handles non-rigid objects.

Significance

Pushes the "demo multiplication" trend (DemoGen, Real2Render2Real lineage) from fixed tabletops to mobile manipulation and arbitrary camera rigs without any simulator — a practical data-efficiency lever relevant to Review-Cross-Embodiment and 3D-policy threads in Review-Dexterous-Manipulation.

← Back to RSS 2026 survey · RSS-2026-Papers · Home