ICLR 2026 DexNDM - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 Category: Dexterous Manipulation Trend tag: Trend 6
flowchart LR
Exp[Category RL experts<br/>privileged obs, PPO] --> Gen[Distill to single<br/>generalist policy BC]
Real[Autonomous real data<br/>~7.5k traj, Chaos Box] --> NDM[Joint-wise neural<br/>dynamics model f_psi_i]
NDM --> Res[Residual policy<br/>a_t + a_res_t]
Gen --> Res
Res --> Pol[Action correction at deploy<br/>transfers to real]
Sim-to-real transfer for dexterous in-hand manipulation has been dominated by heavy domain randomization — train on a wide distribution and hope the real robot is in it. Works, but brittle and tuning-heavy.
Two-stage pipeline. (1) Specialist-to-generalist policy: train category-specific RL experts (PPO in IsaacGym) with privileged observations, then behavior-clone successful trajectories into a single generalist policy that uses only proprioception history, wrist orientation, and target axis. (2) Joint-wise neural dynamics model: for each joint i, learn a dynamics model q^(t+1)_i = f_ψi(h^i_t) that predicts the next joint state from only that joint's W-step state-action history (not the global hand state). This factorization contracts the high-dimensional reality gap into per-joint low-dimensional terms, making it learnable from limited real data (Claim 3.1, via data-processing inequality on KL between train/test distributions).
The model is not composed into a hybrid simulator to retrain the policy. Instead it supervises a lightweight residual policy (a_t + a^res_t) that corrects the sim-trained base policy's actions at deployment. Real data is collected autonomously via a "Chaos Box" — the hand is placed in soft balls and replays open-loop actions, yielding ~7.5k trajectories of diverse load interactions with no human resets and no object-state estimation.
A single sim-trained policy transfers to a wide range of real objects (size 2–20 cm, aspect ratios up to 5.33:1, regular/small/irregular/animal shapes) across 6 wrist orientations and multi-axis targets.
- Sim (unseen objects, ±x axis): RotR 144.22±13.91 vs. AnyRotate reimpl. 91.90±11.60; goal-oriented success 88.27±3.21% vs. 64.33±4.70%.
- Real multi-axis (palm-down z): regular objects 23.82±3.86 rad rotated, 37.50±5.02 s time-to-fall; small objects 9.29±1.63 rad; irregular 8.61±0.76 rad.
- vs. AnyRotate (shared objects): e.g. tin cylinder 10.79±0.54 rad vs. 2.63±0.75 rad.
- Data efficiency: ~4,000 autonomous trajectories reach quality that would need ~7.5M task-aware trajectories (power-law scaling). Cross-simulator transfer (IsaacGym→Genesis, MuJoCo) outperforms UAN and ASAP.
A concrete alternative to domain randomization for sim-to-real. Rather than randomizing dynamics in sim, it measures the residual gap per joint from cheap autonomous data and corrects the policy's actions at deployment. The joint-wise factorization is the key data-efficiency lever — it provably contracts distribution shift — and the Chaos-Box autonomous collection removes the human-in-the-loop bottleneck that limits most real-world dexterous data.
- arXiv:2510.08556
- OpenReview
- Authors: Xueyi Liu (Tsinghua, Shanghai Qi Zhi), He Wang (Peking Univ., Galbot), Li Yi (Tsinghua, Shanghai Qi Zhi)
- ICLR 2026 listing
← Back to ICLR-2026