ICML 2026 DexMachina - Heungwoo/research GitHub Wiki

DexMachina: Functional Retargeting for Bimanual Dexterous Manipulation — a decaying virtual-controller curriculum for learning dexterous policies from human demos

Venue: ICML 2026 (Poster) Category: Dexterous Affiliations: Zhao Mandi (Stanford), Yifan Hou (Stanford), Dieter Fox (NVIDIA), Yashraj Narang (NVIDIA), Ajay Mandlekar (NVIDIA), Shuran Song (Stanford) Traction (2026-06): 31 citations (arXiv)

Functional retargeting from one human demonstration to diverse dexterous hands (Figure 1 from Mandi et al., 2026)

Problem

The paper studies functional retargeting: given a human hand-object demonstration, learn dexterous robot policies that manipulate the object to follow the demonstrated trajectory. The focus is long-horizon, bimanual tasks with articulated objects, which are hard because of (1) the large action space of two multi-fingered hands, (2) spatiotemporal discontinuities in contact, and (3) the embodiment gap between human and robot hands. Prior learning-based methods succeed mainly on short, simple tasks and are bottlenecked by manual reward engineering or costly data collection. The key failure mode on long clips is catastrophic early-stage failure: e.g., after lifting a box bimanually, the policy must reposition one hand mid-air to open a lid while the other re-grasps — exploratory attempts drop the object and terminate the episode before any useful learning occurs.

Method

DexMachina is a curriculum-based RL algorithm whose core idea is virtual object controllers with decaying strength.

DexMachina virtual-controller curriculum overview (Figure 2 from Mandi et al., 2026)

  • Virtual object controllers. The demonstration states are treated as control goals; virtual spring-damper constraints drive the object along its target trajectory. Each object gets six virtual 1-DoF joints for its base pose plus a 1-DoF joint for articulation, all actuated by PD controllers using privileged simulation information. Initially the controllers handle most of the object motion, so the policy can learn across the entire sequence without myopic early termination.
  • Auto-curriculum scheduling (Algorithm 1). Training starts with high controller gains at critical damping, then exponentially decays the gains based on the policy's learning progress (tracked via a history of past rewards). As controller influence fades, the policy must assume more control. Because the policy optimizes a weighted sum of task and auxiliary rewards, it learns motion/contact-improving actions while avoiding disruption of the object trajectory.
  • Action formulation & rewards. A hybrid action formulation with restrictive wrist-joint bounds, plus motion and contact auxiliary rewards that act as soft guidance rather than hard targets, giving the policy room to discover hand-specific strategies.

Setup. Data from ARCTIC (5 articulated objects, 7 demonstrations); 6 open-source dexterous hands of varying size/kinematics; Genesis physics simulation; PPO base RL algorithm; state-based observations controlling both hands at once.

Results

  • DexMachina consistently improves on all hands and tasks, with the largest gains on the four long-horizon clips (e.g. "Waffleiron-300", which requires picking up, opening/closing the lid, flipping, and re-opening/closing all mid-air). Task reward alone falls short there; adding auxiliary rewards helps inconsistently; DexMachina significantly outperforms the no-curriculum setting using the same rewards.
  • The re-implemented ObjDex baseline beats its original reported numbers on short-horizon tasks, validating the comparison, yet still trails DexMachina on long horizons. Kinematics-only retargeting visually aligns with human hands but cannot complete tasks (barely lifts objects).
  • Hardware-adaptive strategies emerge. On Notebook-300, the XHand uses one hand to hold and the other to close the cover, while the smaller, less-actuated Inspire Hand learns to use both hands; on Mixer-300, the long-fingered Allegro Hand closes the lid easily whereas the Schunk Hand compensates with extra wrist motion.
  • Ablations show restrictive wrist-motion bounds yield the best overall performance, and the curriculum outperforms ManipTrans-style error-threshold curricula.

Significance

DexMachina turns a brittle long-horizon exploration problem into a tractable one by handing control off gradually from a privileged virtual controller to the learned policy. Beyond the algorithm, the released simulation benchmark (diverse articulated tasks × multiple dexterous hands) enables functional comparison of hardware designs — surfacing how hand morphology shapes achievable strategies — and lowers the barrier for future dexterous-manipulation research.

Is it a "dexterous controller" study? (controller vs policy-learning)

A useful classification note, since the name "virtual controller" can mislead:

  • The "controller" in the method is a virtual object controller, not a hand controller. It is a spring-damper PD on the object (privileged sim info) used as a training scaffold whose gain decays — the contribution is a curriculum learning algorithm, not a control law / gain / impedance design for the hand.
  • What it produces is a learned bimanual multi-finger policy (hybrid action + wrist-joint bounds, state-based, PPO in Genesis, demo-conditioned, sim-only — not deployed on real hardware). So it does yield hand controllers in the broad "policy = controller" sense, and it studies how hand morphology shapes achievable strategies (a controller/embodiment-centric angle + a functional hardware benchmark).
  • It is not a steerable / reusable low-level dexterous controller in the sense of DexterityGen (Family A — an RL-pretrained primitive library you steer). DexMachina is Family C (bimanual) + Family F (human-demo retargeting): a method to acquire per-demo hand policies, not a low-level controller you call.

One line: DexMachina is a retargeting / curriculum policy-learning method (it learns dexterous hand control from human demos), not a hand-controller design — read "dexterous controller" here as learned policy, not as a DexterityGen-style low-level/reusable controller.

Links

← Back to ICML-2026