RSS 2026 Zero Shot Sim to Real Robot Learning A - Heungwoo/research GitHub Wiki
Zero-Shot Sim-to-Real Robot Learning: A Dexterous Manipulation Study on Reactive Catching
Venue: RSS 2026 (Sydney, Jul 13–17) · Session: RL · paper #148 Authors: Kejia Ren, Gaotian Wang, Andrew Morgan, Kaiyu Hang arXiv: 2605.09789 · program page
Summary compiled from the arXiv paper (v1); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Figure 1: a Franka arm reactively catches a human-thrown rubber ball on a flat, low-friction plate with no passive stabilization. The top shows the ball's arc from throw to catch; the bottom row shows sequential impacts during the catching motion (yellow arrows = ball trajectory at impact), ending with the ball stabilized on the plate.
Problem
Dynamic, contact-rich manipulation is acutely sensitive to modeling errors and perception noise, so sim-to-real transfer often fails. Conventional domain randomization samples one physical-parameter instance per episode, giving the policy diverse experience but no structured way to reason about how uncertainty spreads possible outcomes — especially in millisecond-scale tasks like catching on a flat plate, where nothing mechanically stabilizes the object.
Method
Domain-Randomized Instance Set (DRIS) represents N randomized instances (e.g., balls with different restitution/friction) simultaneously; all evolve in parallel under a shared action, and the joint evolution informs the policy update, with theoretical analysis supporting improved robustness. A size-agnostic point-cloud-style autoencoder (conv layers + feature-wise max-pooling) compresses the instance-set state into a fixed latent, and a FiLM module conditions the latent on the plate's tilt before an MLP outputs actions (plate translation δ plus tilt axis/angle α, β). Training uses PPO in 128 parallel environments (encoder pre-trained on 128,000 DRIS state samples in ~10 min; policy ~2 h); ball-ball collisions are disabled so instances stay independent. Deployment converts actions to plate poses via IK and joint-space torque control on a 7-DoF Franka Research 3, zero-shot except controller-gain identification.
Results
In simulation, DRIS policies with as few as 10 instances stay robust as Gaussian observation noise scales 1–4σ, while an E2E single-ball baseline degrades sharply; the same holds under ±0.05 rad execution noise and out-of-distribution restitution (trained on [0.4, 0.7], tested on [0.7, 0.8]). At matched interaction counts, DRIS (128 envs × 50 balls) reaches 0.89 success under 2σ noise vs 0.73 for a 6,400-environment E2E baseline, using 0.71 GB vs 2 GB of simulation VRAM. On the real robot with four ball types (wiffle, rubber, ping-pong, foam) and three ramp release speeds, the DRIS policy (N=200) achieves 41/60 catches (68%) vs 8/60 for E2E and 3/60 for a hand-crafted velocity-tracking baseline; it also catches human-thrown balls, balances a rolling foam ball, and handles irregular objects.
Significance
A simple, compute-cheap reformulation of domain randomization — propagate the uncertainty set, don't just sample it — that yields genuinely zero-shot transfer on one of the most noise-sensitive dynamic manipulation settings; complementary to the robustness discussions in RL and Review-Dexterous-Manipulation.
← Back to RSS 2026 survey · RSS-2026-Papers · Home