ICLR 2026 DemoGrasp - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 Category: Dexterous Manipulation — Grasping Trend tag: Data efficiency / sim-to-real Authors: Haoqi Yuan, Ziye Huang (equal), Ye Wang, Chuan Mao, Chaoyi Xu; corresponding Zongqing Lu. PKU + BeingBeyond + Renmin U.
flowchart LR
Demo["Single successful grasp trajectory<br/>{q_t*, p_t*ee-obj}_{t=0..T_D}"] --> Editor[Demo Editor π · single-step MDP]
Obs[Initial p_0^ee, p_0^obj, full point cloud c_0^obj<br/>PointNet 128-d → MLP 1024-1024-512-512] --> Editor
Editor --> Edit[T^ee ∈ SE(3) Δxyz∈[-0.05,0.05]m<br/>Δrpy∈[-1.57,1.57]rad<br/>Δq^G ∈ [-1,1]rad]
Edit --> Replay[Open-loop replay in IsaacGym<br/>40 sim steps · motion-planned approach]
Replay --> Reward[r = 1[success] · 1[no collision]<br/>collision randomly disabled in 50% envs]
Reward --> PPO[PPO · 7000 parallel envs · lr=3e-4]
PPO --> Pol[Universal state-based policy]
Pol --> Roll[35K rendered rollouts<br/>RGB + DR + ViT]
Roll --> FM[Flow-matching action head<br/>GR00T-N1.5 architecture]
FM --> Real[Sim-to-real on FR3 + Inspire Hand]
Universal dexterous grasping is a multi-task RL problem with high-dimensional action spaces (typically ~24 DoF for an arm + dexterous hand) and long horizons. Prior work mitigated exploration via complex reward shaping, curriculum learning (UniDexGrasp, UniDexGrasp++), iterative distillation (UniGraspTransformer), or two-stage residual RL (ResDex). These methods (i) often run on floating-wrist hands without arms, (ii) use privileged contact information, (iii) face a trade-off between collision penalties and task rewards, (iv) struggle with small/thin tabletop objects, and (v) require complex pipelines that are hard to extend to new embodiments.
DemoGrasp's reformulation: a single successful grasp trajectory encodes most of the transferable structure of universal grasping (approach center, squeeze, lift). Instead of exploring in the low-level robot action space, explore in the space of edits applied to that trajectory.
The demonstration D is a successful grasp trajectory of one specific object (collected by teleop or hard-coded), expressed in the initial object frame (world frame translated to the object's geometric center at t=0):
D = {(qt*hand, pt*ee-obj)}t=0..TD
where q*hand is hand joint targets and p*ee-obj is the 6D end-effector pose in the initial object frame.
Two parameters: an SE(3) end-effector transformation Tee and delta hand grasp pose ΔqG. Edited targets:
- pt*′ee-obj = Tee pt*ee-obj for t ≤ Tlift; vertical lift Δz afterward.
- qt*′hand linearly interpolates from q0*hand to (qTlift*hand + ΔqG) before Tlift; held thereafter.
Replaying D′ from a motion-planned starting pose works open-loop and already gives ~75% success on training objects (Table 8 row 1).
- State: initial end-effector 6D pose p0ee, initial object 6D pose p0obj, full object point cloud c0obj.
- Action: Tee + ΔqG. End-effector translation Δxyz ∈ [-0.05, 0.05] m; rotation Δrpy ∈ [-1.57, 1.57] rad; hand DoFs Δq ∈ [-1, 1] rad. Rotations are quaternions in observation, Euler in action.
- Transition: replays edited demo for up to T = 40 simulation steps then terminates.
- Reward: r = 1[success] · 1[no collision during execution].
Strict no-collision can prevent grasping flat objects that need fingers to slip under them. The paper randomly disables robot–table collision detection in half of the parallel envs, so:
- Collision-free successful grasps: E[r] = 1
- Successful grasps with table contact: E[r] = 0.5
- Failures: E[r] = 0
This encourages collision-free grasps where possible while permitting minimal contact for hard objects.
- Algorithm: PPO (Schulman et al., 2017).
- Encoder: PointNet on full object point cloud → 128-d feature.
- Actor/critic MLP: [1024, 1024, 512, 512] with ELU activations; tanh output rescaled to action ranges.
- Hyperparameters (Table 12): parallel envs 7,000; learning rate 3e-4; PPO clip ε=0.2; episode length 1; demo-replay execution 40 steps; rollout steps/iteration 1; update epochs/iter 5; minibatches/epoch 4; init Gaussian σ=0.8; gradient clip 1.0.
- Compute: 24 hours on a single NVIDIA RTX 4090.
- 35,000 rendered trajectories collected with the trained RL policy; only successful ones kept.
- Architecture: GR00T-N1.5 (ViT encoder + flow-matching action head). Pretrained ViT fine-tuned; action head trained from scratch (no GR00T-N1.5 weights used). 100K iterations on 4× A800 GPUs, 16 hours.
- Domain randomization: RoboTwin's 300 background images for table textures; 100 FMD material images for object textures; 3 point lights with intensity/ambient ∈ [0.1, 0.8]; camera extrinsics noise [-0.02, 0.02]; table position noise [-0.05, 0.05] m.
- Depth augmentation: depth-range randomization, Gaussian blur+noise, 1% pixel dropout, blending with NYU-Depth-v2 frames at α=0.005 (HERMES recipe).
- Real hardware: 7-DoF Franka Research 3 (FR3) arm + 6-DoF Inspire Hand (6 active + 6 passive); 2× RealSense D435i cameras at diagonal viewpoints.
3,200 training objects; floating 6-DoF wrist + 18-DoF Shadow Hand. Spatial randomization 50 cm × 50 cm (baselines do not randomize position).
| State Train | State Test Seen Cat. | State Test Unseen Cat. | Vision Train | Vision Test Seen Cat. | Vision Test Unseen Cat. | |
|---|---|---|---|---|---|---|
| UniDexGrasp | 79.4 | 74.3 | 70.8 | 73.7 | 68.6 | 65.1 |
| UniDexGrasp++ | 87.9 | 84.3 | 83.1 | 85.4 | 79.6 | 76.7 |
| UniGraspTransformer | 91.2 | 89.2 | 88.3 | 88.9 | 87.3 | 86.8 |
| DemoGrasp | 95.2 | 95.5 | 94.4 | 92.2 | 92.3 | 90.1 |
+5 state / +4 vision over best baseline; only 1% generalization gap between train and unseen — strong universality.
Trained on 175 objects (75 YCB + 100 DexGraspNet). Tested on 5 OOD datasets:
| Method | DGA | EGAD | Omni6DPose | ModelNet40 | Visual Dexterity |
|---|---|---|---|---|---|
| RobustDexGrasp | 64.40 | 93.45 | 73.00 | 75.70 | 92.50 |
| DemoGrasp | 74.40 | 96.75 | 82.24 | 75.58 | 97.80 |
DemoGrasp matches RobustDexGrasp on ModelNet40 and exceeds on the other four datasets.
Six embodiments tested on the same 175-object training set + five test datasets:
- Inspire Hand (FR3+Inspire), Allegro (UR5+Allegro), DClaw (FR3+DClaw), floating Shadow, FR3+Shadow, Schunk SVH (UR5+Schunk), and FR3+Gripper (Panda parallel-jaw).
- All multi-fingered hands achieve >90% on the 175 training objects.
- Average 84.6% across all six unseen test datasets.
- Floating Shadow vs FR3+Shadow: difference of only 1.4% — adding a robot arm doesn't hurt.
- FR3+Gripper underperforms on EGAD/DGA (Panda's limited stroke struggles with wide objects).
- Per-dataset average ranges (across all six embodiments): EGAD 49.4–97.9%, DGA 30.7–79.3%, Omni6DPose 66.6–85.0%, DexGraspNet (Unseen Cat.) 81.1–97.5%, ModelNet40 61.2–80.1%, Visual Dexterity 83.5–99.1%.
110 unseen objects, 5 trials each, 50×50 cm position randomization:
| Shape | Category | Num. | Success rate |
|---|---|---|---|
| Regular | Bottles | 12 | 95.0% |
| Regular | Boxes & Jars | 22 | 93.6% |
| Regular | Balls & Fruit | 12 | 98.3% |
| Regular | Soft Toys | 10 | 96.0% |
| Irregular | (mixed) | 18 | 90.0% |
| Flat & Thin | Tools | 10 | 60.0% |
| Flat & Thin | Others | 14 | 74.3% |
| Small | (mixed) | 12 | 76.7% |
95.3% on normal-sized objects. 68.3% on flat/thin (thickness < 1.5 cm), 76.7% on small (diameter < 3.5 cm). The paper claims this is the first universal sim-to-real policy to grasp small/thin tabletop objects without severe collisions.
| Model | Sim | Real |
|---|---|---|
| Any-DemoGrasp | 83.66% | 82% |
| Instruct-DemoGrasp | 85.33% | 84% |
Language-conditioned slightly outperforms unconditional (the paper attributes this to reduced action uncertainty for imitation learning). Robust to randomized backgrounds + lighting (82% under challenging scene appearance changes). Uses GR00T-N1.5's vision-language model frozen for instruction encoding.
| Method | Success (%) |
|---|---|
| Sampling | 77.56 |
| RL | 96.24 |
Training BC on uniformly sampled successful rollouts gives multimodal/inconsistent data; RL converges to a unimodal optimal policy.
Measures contribution of each editing axis:
| Δxyz | Δrpy | Δq | Train Set | Test Set |
|---|---|---|---|---|
| 75.29 (open-loop replay) | 73.43 | |||
| ✓ | 81.35 | 76.04 | ||
| ✓ | ✓ | 86.40 | 79.68 | |
| ✓ | ✓ | 94.22 | 81.39 | |
| ✓ | ✓ | ✓ | 96.24 | 82.74 |
Adding wrist rotation (Δrpy) is the biggest win (+13 train); hand DoFs add a smaller +2 — using a dexterous hand as a single-DoF gripper already gets most of the way, but Δq enables more robust grasps (vase from the side using thumb+index+ring force closure).
| Train data | DGA | EGAD | Omni6DPose | ModelNet40 | Visual Dexterity |
|---|---|---|---|---|---|
| 175 Objects (YCB+DexGraspNet) | 65.62 | 97.88 | 85.04 | 80.13 | 99.13 |
| Test Sets (train directly) | 71.49 | 99.16 | 88.71 | 81.10 | 99.20 |
Only 2.4% average gap between training on 175 objects and training on the test sets — universal grasping doesn't need millions of training objects with this method.
| Demo | Replay | RL Train | RL Test |
|---|---|---|---|
| small obj. + top | 75.29% | 96.24% | 82.74% |
| small obj. + side | 62.90% | 95.18% | 81.45% |
| big obj. + top | 7.23% | 95.02% | 82.46% |
| big obj. + side | 3.88% | 95.27% | 83.22% |
Even when the demo's open-loop replay is nearly useless (3.88% on a "big-object side grasp" demo), RL recovers >95% — DemoGrasp is robust to demo quality.
| Cameras | YCB Sim | DexGN Sim | Real (5 challenging objects) |
|---|---|---|---|
| Mono-Depth | 80.2% | 95.2% | 5/5, 4/5, 1/5, 0/5, 0/5 |
| Two-Depth | 80.3% | 96.4% | 5/5, 4/5, 1/5, 0/5, 0/5 |
| Mono-RGB | 83.2% | 94.8% | 4/5, 5/5, 4/5, 0/5, 3/5 |
| Two-RGB | 87.0% | 97.3% | 5/5, 5/5, 4/5, 5/5, 5/5 |
RGB beats depth on small/flat objects (sensor noise hides them in depth); two views beat one (richer 3D, less hand-object occlusion).
The paper does not have an explicit "Limitations" section. Implicit/honest gaps surfaced in the experiments:
- Flat/thin objects still hardest — 68.3% real-world, vs 95.3% on normal objects. The collision-free/with-contact reward design helps but the gap remains.
- Single-step open-loop replay depends on a successful motion plan from the home pose to the demo's first frame; cluttered start states could fail.
- Policy is closed-loop only on the vision side — the underlying RL action space is single-step on the editing parameters. For dynamic objects or major mid-grasp disturbances, the open-loop replay window may be short.
- Reset region is 50 cm × 50 cm tabletop — out-of-region or non-tabletop scenarios (deep bins, shelves, occluded objects) are not evaluated.
- No in-hand reorientation, tool use, bimanual — universal "pick and lift" only.
DemoGrasp is the cleanest current example that problem reformulation > scaling for dexterous skills.
- vs UniDexGrasp / UniDexGrasp++ / UniGraspTransformer: all explore in the full robot action space with complex curricula and reward shaping. DemoGrasp's compact 6+J editing space + single-step MDP + binary reward beats them with simpler training.
- vs ResDex / RobustDexGrasp: ResDex uses a two-stage residual RL framework; RobustDexGrasp focuses on robustness via large-scale data. DemoGrasp uses one demonstration and beats RobustDexGrasp on 4/5 OOD datasets with the Allegro Hand.
- vs DextrAH-G / DextrAH-RGB / Singh et al.: these are sim-to-real on a wide variety of objects but still fall short on small/thin tabletop objects. DemoGrasp is, per the paper, the first to grasp previously unseen small/thin objects in tabletop settings without severe collisions.
- vs ClutterDexGrasp: ClutterDexGrasp adds clutter; DemoGrasp adds clutter as an extension and still hits >80%.
- vs DexNDM: DexNDM tackles dexterous policies via a different angle (neural distance manifold). Both papers exemplify the 2026 trend that careful problem framing replaces brute-force scaling.
- vs D-REX: D-REX is a robot-experience scaling effort; DemoGrasp is on the data-efficiency end of the same trend.
The single-step MDP + demo-editing recipe is broadly transferrable to other manipulation skills (pouring, insertion) where a successful trajectory exists and only needs re-aiming to generalize.
- OpenReview: https://openreview.net/forum?id=Bf4FeuW0Mr
- PDF: https://openreview.net/pdf?id=Bf4FeuW0Mr
- Project page: https://research.beingbeyond.com/demograsp
← Back to ICLR-2026