ICLR 2026 DexNDM - Heungwoo/research GitHub Wiki

DexNDM — Closing the Reality Gap for In-Hand Rotation via Joint-Wise Neural Dynamics

Venue: ICLR 2026 Category: Dexterous Manipulation Trend tag: Trend 6

Approach diagram

flowchart LR
  Exp[Category RL experts<br/>privileged obs, PPO] --> Gen[Distill to single<br/>generalist policy BC]
  Real[Autonomous real data<br/>~7.5k traj, Chaos Box] --> NDM[Joint-wise neural<br/>dynamics model f_psi_i]
  NDM --> Res[Residual policy<br/>a_t + a_res_t]
  Gen --> Res
  Res --> Pol[Action correction at deploy<br/>transfers to real]
Loading

Problem

Sim-to-real transfer for dexterous in-hand manipulation has been dominated by heavy domain randomization — train on a wide distribution and hope the real robot is in it. Works, but brittle and tuning-heavy.

Method

Two-stage pipeline. (1) Specialist-to-generalist policy: train category-specific RL experts (PPO in IsaacGym) with privileged observations, then behavior-clone successful trajectories into a single generalist policy that uses only proprioception history, wrist orientation, and target axis. (2) Joint-wise neural dynamics model: for each joint i, learn a dynamics model q^(t+1)_i = f_ψi(h^i_t) that predicts the next joint state from only that joint's W-step state-action history (not the global hand state). This factorization contracts the high-dimensional reality gap into per-joint low-dimensional terms, making it learnable from limited real data (Claim 3.1, via data-processing inequality on KL between train/test distributions).

The model is not composed into a hybrid simulator to retrain the policy. Instead it supervises a lightweight residual policy (a_t + a^res_t) that corrects the sim-trained base policy's actions at deployment. Real data is collected autonomously via a "Chaos Box" — the hand is placed in soft balls and replays open-loop actions, yielding ~7.5k trajectories of diverse load interactions with no human resets and no object-state estimation.

Results

A single sim-trained policy transfers to a wide range of real objects (size 2–20 cm, aspect ratios up to 5.33:1, regular/small/irregular/animal shapes) across 6 wrist orientations and multi-axis targets.

  • Sim (unseen objects, ±x axis): RotR 144.22±13.91 vs. AnyRotate reimpl. 91.90±11.60; goal-oriented success 88.27±3.21% vs. 64.33±4.70%.
  • Real multi-axis (palm-down z): regular objects 23.82±3.86 rad rotated, 37.50±5.02 s time-to-fall; small objects 9.29±1.63 rad; irregular 8.61±0.76 rad.
  • vs. AnyRotate (shared objects): e.g. tin cylinder 10.79±0.54 rad vs. 2.63±0.75 rad.
  • Data efficiency: ~4,000 autonomous trajectories reach quality that would need ~7.5M task-aware trajectories (power-law scaling). Cross-simulator transfer (IsaacGym→Genesis, MuJoCo) outperforms UAN and ASAP.

Significance

A concrete alternative to domain randomization for sim-to-real. Rather than randomizing dynamics in sim, it measures the residual gap per joint from cheap autonomous data and corrects the policy's actions at deployment. The joint-wise factorization is the key data-efficiency lever — it provably contracts distribution shift — and the Chaos-Box autonomous collection removes the human-in-the-loop bottleneck that limits most real-world dexterous data.

Links

  • arXiv:2510.08556
  • OpenReview
  • Authors: Xueyi Liu (Tsinghua, Shanghai Qi Zhi), He Wang (Peking Univ., Galbot), Li Yi (Tsinghua, Shanghai Qi Zhi)
  • ICLR 2026 listing

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️