ICLR 2026 Emergent Dexterity - Heungwoo/research GitHub Wiki

OmniReset โ€” Emergent dexterity from diverse resets and large-scale RL

Venue: ICLR 2026 ยท Authors: Patrick Yin, Tyler Westenbroek, Zhengyu Zhang, Joshua Tran, Ignacio Dagnino, Eeshani Shilamkar, Numfor Mbiziwo-Tiapo, Simran Bagaria, Xinlei Liu, Galen Mullins, Andrey Kolobov, Abhishek Gupta (UW WEIRD Lab et al.) ยท arXiv:2603.15789 ยท Category: Dexterous Manipulation ยท Trend tag: Scaling RL for dexterity

Approach diagram

flowchart LR
  Resets[OmniReset<br/>programmatic diverse resets] --> Cover[Broad coverage of<br/>robot-object interactions]
  Cover --> RL[On-policy RL<br/>single reward, fixed hypers<br/>no curriculum, no demos]
  RL --> Compute[More compute -><br/>wider behavioral coverage]
  RL --> Distill[Distill to<br/>visuomotor policy]
  Distill --> Real[Zero-shot real transfer<br/>robust retrying behavior]
Loading

Problem

Long-horizon dexterous manipulation is hard for RL primarily because of exploration: an agent rarely stumbles into the diverse robot-object configurations that underlie dexterous skills. Standard fixes โ€” reward shaping, curricula, human demonstrations, per-task hyperparameter tuning โ€” are brittle and labor-intensive, and they do not scale gracefully with compute.

Method

The paper argues that the exploration bottleneck can be sidestepped by exploiting a capability simulators already have: resetting to arbitrary states. OmniReset programmatically generates a diverse distribution of reset states โ€” exposing the RL algorithm systematically to the broad set of contacts, grasps, and object poses that make up a dexterous task โ€” instead of always starting from a narrow initial distribution.

With this reset machinery, plain on-policy RL solves a broad class of tasks using a single reward function, fixed algorithm hyperparameters, no curriculum, and no human demonstrations. The key scaling property: adding compute (more reset-driven rollouts) converts directly into broader behavioral coverage and continued performance gains, rather than saturating. The resulting state-based policies are then distilled into visuomotor policies for deployment.

Results

  • OmniReset scales to long-horizon dexterous tasks beyond the reach of existing approaches, and learns policies robust over significantly wider ranges of initial conditions than baselines.
  • Distilled visuomotor policies transfer to the real world zero-shot, exhibiting robust retrying behavior and substantially higher success rates than baselines.

(The paper reports qualitative superiority and broader coverage; specific per-task success numbers are in the full paper / project page.)

Significance

Reframes dexterous RL from a reward/curriculum-engineering problem into a reset-design problem, where the simulator's reset primitive becomes the main lever for exploration. This makes the recipe simple and compute-scalable โ€” a single reward, no demos, no curriculum โ€” and points toward "just add compute" as a path to emergent dexterous behavior, complementing demonstration-heavy and domain-randomization-heavy alternatives.

Links

Related pages

โ† Back to ICLR-2026

โš ๏ธ **GitHub.com Fallback** โš ๏ธ