ICLR 2026 Emergent Dexterity - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 ยท Authors: Patrick Yin, Tyler Westenbroek, Zhengyu Zhang, Joshua Tran, Ignacio Dagnino, Eeshani Shilamkar, Numfor Mbiziwo-Tiapo, Simran Bagaria, Xinlei Liu, Galen Mullins, Andrey Kolobov, Abhishek Gupta (UW WEIRD Lab et al.) ยท arXiv:2603.15789 ยท Category: Dexterous Manipulation ยท Trend tag: Scaling RL for dexterity
flowchart LR
Resets[OmniReset<br/>programmatic diverse resets] --> Cover[Broad coverage of<br/>robot-object interactions]
Cover --> RL[On-policy RL<br/>single reward, fixed hypers<br/>no curriculum, no demos]
RL --> Compute[More compute -><br/>wider behavioral coverage]
RL --> Distill[Distill to<br/>visuomotor policy]
Distill --> Real[Zero-shot real transfer<br/>robust retrying behavior]
Long-horizon dexterous manipulation is hard for RL primarily because of exploration: an agent rarely stumbles into the diverse robot-object configurations that underlie dexterous skills. Standard fixes โ reward shaping, curricula, human demonstrations, per-task hyperparameter tuning โ are brittle and labor-intensive, and they do not scale gracefully with compute.
The paper argues that the exploration bottleneck can be sidestepped by exploiting a capability simulators already have: resetting to arbitrary states. OmniReset programmatically generates a diverse distribution of reset states โ exposing the RL algorithm systematically to the broad set of contacts, grasps, and object poses that make up a dexterous task โ instead of always starting from a narrow initial distribution.
With this reset machinery, plain on-policy RL solves a broad class of tasks using a single reward function, fixed algorithm hyperparameters, no curriculum, and no human demonstrations. The key scaling property: adding compute (more reset-driven rollouts) converts directly into broader behavioral coverage and continued performance gains, rather than saturating. The resulting state-based policies are then distilled into visuomotor policies for deployment.
- OmniReset scales to long-horizon dexterous tasks beyond the reach of existing approaches, and learns policies robust over significantly wider ranges of initial conditions than baselines.
- Distilled visuomotor policies transfer to the real world zero-shot, exhibiting robust retrying behavior and substantially higher success rates than baselines.
(The paper reports qualitative superiority and broader coverage; specific per-task success numbers are in the full paper / project page.)
Reframes dexterous RL from a reward/curriculum-engineering problem into a reset-design problem, where the simulator's reset primitive becomes the main lever for exploration. This makes the recipe simple and compute-scalable โ a single reward, no demos, no curriculum โ and points toward "just add compute" as a path to emergent dexterous behavior, complementing demonstration-heavy and domain-randomization-heavy alternatives.
โ Back to ICLR-2026