RSS 2026 UMI Underwater - Heungwoo/research GitHub Wiki

UMI-Underwater: Learning Underwater Manipulation without Underwater Teleoperation

Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Manipulation 1 · paper #9 Authors: Hao Li, Long Yin Chung, Jack Goler, Ryan Zhang, Xiaochi Xie, Huy Ha, Shuran Song, Mark Cutkosky arXiv: 2603.27012 · program page

Summary compiled from the arXiv paper (v1); all numbers quoted from the paper. Trend context: RSS 2026 survey.

UMI-Underwater overview (Figure 1 of arXiv 2603.27012, © the authors)

Figure 1 shows the two data streams: on the left, task-centric on-land demonstrations collected with the handheld UMI-Aquatic gripper (with depth, goal, and affordance maps for unseen objects); on the right, robot-centric self-supervised grasping data collected in water by the ROV with verify/retry behaviors over diverse objects. The affordance predictor trained on land transfers to the underwater domain zero-shot.

Problem

Underwater grasping suffers from severely degraded, highly variable imagery (attenuation, scattering, turbidity, caustics) and from the expense of collecting diverse underwater demonstrations — most underwater manipulation systems remain teleoperation-centric. End-to-end RGB visuomotor policies are brittle under these distribution shifts. The paper targets both bottlenecks: eliminating underwater teleoperation and bridging the land-to-water perception gap.

Method

The system (all authors at Stanford) has three components. (1) A self-supervised underwater data collection pipeline: a heuristic visuo-servoing PD controller on a tethered QYSEA ROV runs a six-stage grasp routine (yaw alignment, forward approach, depth adjustment, close-range approach with Depth Anything V2 range estimation, grasp, 3-s drag verification), with regrasp and overshoot-recovery behaviors; it collected 536 grasp episodes (~15 h autonomous runtime), of which 233 successes train the policy. (2) UMI-Aquatic, a UMI-derived handheld gripper with an iPhone camera and AprilTags, used to collect 800 on-land demonstrations (6 object types, 8 backgrounds, ~2.2 h); a goal-conditioned, depth-based affordance U-Net (depth 4, base channels 64, 112×112 input) is trained on these and deployed underwater zero-shot via a calibrated plane-at-depth warp into the underwater camera geometry. (3) An affordance- and depth-conditioned diffusion policy (CLIP ViT-B/16 encoder, 1D U-Net denoiser, 16-step action horizon, 6-D actions) trained on the underwater successes; the stack runs at 10 Hz actions with 1 Hz policy inference.

Results

In pool experiments (20 trials per condition, three objects per scene with an on-land goal image specifying the target): in-distribution, DP+Aff+Depth reaches 85% success vs 65% for the RGB-only diffusion policy. Under unseen background wallpapers, the method keeps 80% while DP+RGB and even DP+RGB+Aff+Depth collapse to 0% — adding raw RGB cripples robustness under appearance shift. On novel objects seen only in the on-land data (pitcher, can, power drill), the zero-shot transfer achieves 75% vs 50% for DP+RGB. A recurring failure mode is overshoot-induced target switching.

Significance

A clean demonstration that a depth-based affordance heatmap can serve as a modular, domain-robust perception interface, letting cheap on-land UMI-style data drive underwater deployment without any underwater teleoperation. Connects to the UMI-family data-collection thread and affordance-conditioned policies discussed in Review-Dexterous-Manipulation.

← Back to RSS 2026 survey · RSS-2026-Papers · Home