RSS 2026 Robo3R - Heungwoo/research GitHub Wiki

Robo3R: Enhancing Robotic Manipulation with Accurate Feed-Forward 3D Reconstruction

Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Manipulation 2 · paper #56 Authors: Sizhe Yang, Linning Xu, Hao Li, Juncheng Mu, Jia Zeng, Dahua Lin, Jiangmiao Pang arXiv: 2602.10101 · program page

Summary compiled from the arXiv paper (v2); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Robo3R overview (Figure 1 of arXiv 2602.10101, © the authors)

Figure 1: from RGB frames plus robot state, Robo3R predicts local geometry, relative pose, and a global similarity transformation in a single forward pass (left); its point clouds are visibly cleaner than a depth camera's (middle); and the resulting geometry lifts downstream success rates for imitation learning/sim-to-real, grasp synthesis, and collision-free motion planning (right).

Problem

3D input helps manipulation policies, grasp synthesis, and motion planning, but depth cameras (stereo or ToF) are noisy and fail on transparent, reflective, or tiny objects, while general feed-forward reconstruction models (VGGT, π³, DepthAnything3, MapAnything) lack manipulation-level geometric precision and reliable metric scale. Robo3R (Shanghai AI Lab / CUHK / USTC / Tsinghua) aims to replace depth sensors and calibration with an RGB-only, manipulation-ready reconstruction model.

Method

Robo3R fuses DINOv2 ViT-L image features (1–2 views) with MLP-encoded robot joint states, processed by 18 alternating global/frame-wise attention blocks. It predicts scale-invariant local point maps via a masked point head (separate robot/object/background branches with depth, ray, and mask heads to avoid over-smoothing), a relative pose head (9D rotation orthogonalized by SVD), and similarity-transformation tokens that map registered points into metric-scale geometry in the canonical robot frame. A keypoint head predicts robot-link keypoint heatmaps; solving PnP against forward-kinematics 3D keypoints refines the camera extrinsics. Training is end-to-end on Robo3R-4M, a new Isaac Sim synthetic dataset of 100,000 scenes / 4 million frames built from 16,911 objects, 4,710 textures, and 6,512 environment maps with extensive domain randomization.

Results

On a held-out synthetic benchmark (2,000 scenes / 80,000 frames), Robo3R reaches monocular point error 0.006 and scale error 0.007 — an order of magnitude below π³ (0.061/0.497) and better than MapAnything fine-tuned on the same data (0.010/0.010) — plus relative-pose RTE 0.014 / RRE 0.013 with [email protected] = 0.951. It reconstructs objects as thin as 1.5 mm and handles mirrors and transparent cups that blind a RealSense D455. In real-world downstream tests (Franka Research 3 and bimanual UR5e with XHand, RTX 4090, 10 Hz control): imitation learning with ManiFlow+Robo3R scores 14/16, 15/16, 12/16, 16/16 on Sweep Bean / Insert Screw / Breakfast / BiDex Pour, beating depth-camera, RGB, other feed-forward, and π0 baselines; sim-to-real reaches 16/16 (Push Cube) and 12/16 (Pick Cube) vs 7/16 and 5/16 with depth; grasp synthesis and collision-free motion planning gain most on transparent/reflective/small/thin objects (e.g., 5/5 vs 2/5 on thin obstacles).

Significance

Positions learned feed-forward reconstruction as a drop-in replacement for depth sensors in manipulation stacks — metric-scale, canonical robot frame, calibration-free — and shows the 3D-representation route to sim-to-real consistency. Related wiki threads: Review-Dexterous-Manipulation · Review-VLA-Evaluation.

← Back to RSS 2026 survey · RSS-2026-Papers · Home