ICLR 2026 DemoGrasp - Heungwoo/research GitHub Wiki

DemoGrasp — Universal Dexterous Grasping from a Single Demonstration

Venue: ICLR 2026 Category: Dexterous Manipulation — Grasping Trend tag: Data efficiency / sim-to-real Authors: Haoqi Yuan, Ziye Huang (equal), Ye Wang, Chuan Mao, Chaoyi Xu; corresponding Zongqing Lu. PKU + BeingBeyond + Renmin U.

Approach diagram

flowchart LR
  Demo["Single successful grasp trajectory<br/>{q_t*, p_t*ee-obj}_{t=0..T_D}"] --> Editor[Demo Editor π · single-step MDP]
  Obs[Initial p_0^ee, p_0^obj, full point cloud c_0^obj<br/>PointNet 128-d → MLP 1024-1024-512-512] --> Editor
  Editor --> Edit[T^ee ∈ SE(3) Δxyz∈[-0.05,0.05]m<br/>Δrpy∈[-1.57,1.57]rad<br/>Δq^G ∈ [-1,1]rad]
  Edit --> Replay[Open-loop replay in IsaacGym<br/>40 sim steps · motion-planned approach]
  Replay --> Reward[r = 1[success] · 1[no collision]<br/>collision randomly disabled in 50% envs]
  Reward --> PPO[PPO · 7000 parallel envs · lr=3e-4]
  PPO --> Pol[Universal state-based policy]
  Pol --> Roll[35K rendered rollouts<br/>RGB + DR + ViT]
  Roll --> FM[Flow-matching action head<br/>GR00T-N1.5 architecture]
  FM --> Real[Sim-to-real on FR3 + Inspire Hand]
Loading

Problem

Universal dexterous grasping is a multi-task RL problem with high-dimensional action spaces (typically ~24 DoF for an arm + dexterous hand) and long horizons. Prior work mitigated exploration via complex reward shaping, curriculum learning (UniDexGrasp, UniDexGrasp++), iterative distillation (UniGraspTransformer), or two-stage residual RL (ResDex). These methods (i) often run on floating-wrist hands without arms, (ii) use privileged contact information, (iii) face a trade-off between collision penalties and task rewards, (iv) struggle with small/thin tabletop objects, and (v) require complex pipelines that are hard to extend to new embodiments.

DemoGrasp's reformulation: a single successful grasp trajectory encodes most of the transferable structure of universal grasping (approach center, squeeze, lift). Instead of exploring in the low-level robot action space, explore in the space of edits applied to that trajectory.

Detailed Method

Demonstration representation

The demonstration D is a successful grasp trajectory of one specific object (collected by teleop or hard-coded), expressed in the initial object frame (world frame translated to the object's geometric center at t=0):

D = {(qt*hand, pt*ee-obj)}t=0..TD

where q*hand is hand joint targets and p*ee-obj is the 6D end-effector pose in the initial object frame.

Demo editing

Two parameters: an SE(3) end-effector transformation Tee and delta hand grasp pose ΔqG. Edited targets:

  • pt*′ee-obj = Tee pt*ee-obj for t ≤ Tlift; vertical lift Δz afterward.
  • qt*′hand linearly interpolates from q0*hand to (qTlift*hand + ΔqG) before Tlift; held thereafter.

Replaying D′ from a motion-planned starting pose works open-loop and already gives ~75% success on training objects (Table 8 row 1).

Single-step MDP

  • State: initial end-effector 6D pose p0ee, initial object 6D pose p0obj, full object point cloud c0obj.
  • Action: Tee + ΔqG. End-effector translation Δxyz ∈ [-0.05, 0.05] m; rotation Δrpy ∈ [-1.57, 1.57] rad; hand DoFs Δq ∈ [-1, 1] rad. Rotations are quaternions in observation, Euler in action.
  • Transition: replays edited demo for up to T = 40 simulation steps then terminates.
  • Reward: r = 1[success] · 1[no collision during execution].

Reward design subtlety — random collision-free disabling

Strict no-collision can prevent grasping flat objects that need fingers to slip under them. The paper randomly disables robot–table collision detection in half of the parallel envs, so:

  • Collision-free successful grasps: E[r] = 1
  • Successful grasps with table contact: E[r] = 0.5
  • Failures: E[r] = 0

This encourages collision-free grasps where possible while permitting minimal contact for hard objects.

RL implementation

  • Algorithm: PPO (Schulman et al., 2017).
  • Encoder: PointNet on full object point cloud → 128-d feature.
  • Actor/critic MLP: [1024, 1024, 512, 512] with ELU activations; tanh output rescaled to action ranges.
  • Hyperparameters (Table 12): parallel envs 7,000; learning rate 3e-4; PPO clip ε=0.2; episode length 1; demo-replay execution 40 steps; rollout steps/iteration 1; update epochs/iter 5; minibatches/epoch 4; init Gaussian σ=0.8; gradient clip 1.0.
  • Compute: 24 hours on a single NVIDIA RTX 4090.

Vision-based sim-to-real

  • 35,000 rendered trajectories collected with the trained RL policy; only successful ones kept.
  • Architecture: GR00T-N1.5 (ViT encoder + flow-matching action head). Pretrained ViT fine-tuned; action head trained from scratch (no GR00T-N1.5 weights used). 100K iterations on 4× A800 GPUs, 16 hours.
  • Domain randomization: RoboTwin's 300 background images for table textures; 100 FMD material images for object textures; 3 point lights with intensity/ambient ∈ [0.1, 0.8]; camera extrinsics noise [-0.02, 0.02]; table position noise [-0.05, 0.05] m.
  • Depth augmentation: depth-range randomization, Gaussian blur+noise, 1% pixel dropout, blending with NYU-Depth-v2 frames at α=0.005 (HERMES recipe).
  • Real hardware: 7-DoF Franka Research 3 (FR3) arm + 6-DoF Inspire Hand (6 active + 6 passive); 2× RealSense D435i cameras at diagonal viewpoints.

Comprehensive Results

DexGraspNet with Shadow Hand (Table 1)

3,200 training objects; floating 6-DoF wrist + 18-DoF Shadow Hand. Spatial randomization 50 cm × 50 cm (baselines do not randomize position).

State Train State Test Seen Cat. State Test Unseen Cat. Vision Train Vision Test Seen Cat. Vision Test Unseen Cat.
UniDexGrasp 79.4 74.3 70.8 73.7 68.6 65.1
UniDexGrasp++ 87.9 84.3 83.1 85.4 79.6 76.7
UniGraspTransformer 91.2 89.2 88.3 88.9 87.3 86.8
DemoGrasp 95.2 95.5 94.4 92.2 92.3 90.1

+5 state / +4 vision over best baseline; only 1% generalization gap between train and unseen — strong universality.

Cross-dataset zero-shot (Allegro on UR5, Table 2)

Trained on 175 objects (75 YCB + 100 DexGraspNet). Tested on 5 OOD datasets:

Method DGA EGAD Omni6DPose ModelNet40 Visual Dexterity
RobustDexGrasp 64.40 93.45 73.00 75.70 92.50
DemoGrasp 74.40 96.75 82.24 75.58 97.80

DemoGrasp matches RobustDexGrasp on ModelNet40 and exceeds on the other four datasets.

Cross-embodiment universality (Figure 3 / Table 10)

Six embodiments tested on the same 175-object training set + five test datasets:

  • Inspire Hand (FR3+Inspire), Allegro (UR5+Allegro), DClaw (FR3+DClaw), floating Shadow, FR3+Shadow, Schunk SVH (UR5+Schunk), and FR3+Gripper (Panda parallel-jaw).
  • All multi-fingered hands achieve >90% on the 175 training objects.
  • Average 84.6% across all six unseen test datasets.
  • Floating Shadow vs FR3+Shadow: difference of only 1.4% — adding a robot arm doesn't hurt.
  • FR3+Gripper underperforms on EGAD/DGA (Panda's limited stroke struggles with wide objects).
  • Per-dataset average ranges (across all six embodiments): EGAD 49.4–97.9%, DGA 30.7–79.3%, Omni6DPose 66.6–85.0%, DexGraspNet (Unseen Cat.) 81.1–97.5%, ModelNet40 61.2–80.1%, Visual Dexterity 83.5–99.1%.

Real-world (Table 3)

110 unseen objects, 5 trials each, 50×50 cm position randomization:

Shape Category Num. Success rate
Regular Bottles 12 95.0%
Regular Boxes & Jars 22 93.6%
Regular Balls & Fruit 12 98.3%
Regular Soft Toys 10 96.0%
Irregular (mixed) 18 90.0%
Flat & Thin Tools 10 60.0%
Flat & Thin Others 14 74.3%
Small (mixed) 12 76.7%

95.3% on normal-sized objects. 68.3% on flat/thin (thickness < 1.5 cm), 76.7% on small (diameter < 3.5 cm). The paper claims this is the first universal sim-to-real policy to grasp small/thin tabletop objects without severe collisions.

Cluttered + language-conditioned grasping (Table 4)

Model Sim Real
Any-DemoGrasp 83.66% 82%
Instruct-DemoGrasp 85.33% 84%

Language-conditioned slightly outperforms unconditional (the paper attributes this to reduced action uncertainty for imitation learning). Robust to randomized backgrounds + lighting (82% under challenging scene appearance changes). Uses GR00T-N1.5's vision-language model frozen for instruction encoding.

Ablation Studies

Sampling vs RL (Table 5)

Method Success (%)
Sampling 77.56
RL 96.24

Training BC on uniformly sampled successful rollouts gives multimodal/inconsistent data; RL converges to a unimodal optimal policy.

Action-space ablation (Table 8)

Measures contribution of each editing axis:

Δxyz Δrpy Δq Train Set Test Set
75.29 (open-loop replay) 73.43
✓ 81.35 76.04
✓ ✓ 86.40 79.68
✓ ✓ 94.22 81.39
✓ ✓ ✓ 96.24 82.74

Adding wrist rotation (Δrpy) is the biggest win (+13 train); hand DoFs add a smaller +2 — using a dexterous hand as a single-DoF gripper already gets most of the way, but Δq enables more robust grasps (vase from the side using thumb+index+ring force closure).

175 objects vs full test sets (Table 7)

Train data DGA EGAD Omni6DPose ModelNet40 Visual Dexterity
175 Objects (YCB+DexGraspNet) 65.62 97.88 85.04 80.13 99.13
Test Sets (train directly) 71.49 99.16 88.71 81.10 99.20

Only 2.4% average gap between training on 175 objects and training on the test sets — universal grasping doesn't need millions of training objects with this method.

Demonstration quality (Table 9)

Demo Replay RL Train RL Test
small obj. + top 75.29% 96.24% 82.74%
small obj. + side 62.90% 95.18% 81.45%
big obj. + top 7.23% 95.02% 82.46%
big obj. + side 3.88% 95.27% 83.22%

Even when the demo's open-loop replay is nearly useless (3.88% on a "big-object side grasp" demo), RL recovers >95% — DemoGrasp is robust to demo quality.

Camera configurations for vision policies (Table 6)

Cameras YCB Sim DexGN Sim Real (5 challenging objects)
Mono-Depth 80.2% 95.2% 5/5, 4/5, 1/5, 0/5, 0/5
Two-Depth 80.3% 96.4% 5/5, 4/5, 1/5, 0/5, 0/5
Mono-RGB 83.2% 94.8% 4/5, 5/5, 4/5, 0/5, 3/5
Two-RGB 87.0% 97.3% 5/5, 5/5, 4/5, 5/5, 5/5

RGB beats depth on small/flat objects (sensor noise hides them in depth); two views beat one (richer 3D, less hand-object occlusion).

Limitations

The paper does not have an explicit "Limitations" section. Implicit/honest gaps surfaced in the experiments:

  • Flat/thin objects still hardest — 68.3% real-world, vs 95.3% on normal objects. The collision-free/with-contact reward design helps but the gap remains.
  • Single-step open-loop replay depends on a successful motion plan from the home pose to the demo's first frame; cluttered start states could fail.
  • Policy is closed-loop only on the vision side — the underlying RL action space is single-step on the editing parameters. For dynamic objects or major mid-grasp disturbances, the open-loop replay window may be short.
  • Reset region is 50 cm × 50 cm tabletop — out-of-region or non-tabletop scenarios (deep bins, shelves, occluded objects) are not evaluated.
  • No in-hand reorientation, tool use, bimanual — universal "pick and lift" only.

Significance & Positioning

DemoGrasp is the cleanest current example that problem reformulation > scaling for dexterous skills.

  • vs UniDexGrasp / UniDexGrasp++ / UniGraspTransformer: all explore in the full robot action space with complex curricula and reward shaping. DemoGrasp's compact 6+J editing space + single-step MDP + binary reward beats them with simpler training.
  • vs ResDex / RobustDexGrasp: ResDex uses a two-stage residual RL framework; RobustDexGrasp focuses on robustness via large-scale data. DemoGrasp uses one demonstration and beats RobustDexGrasp on 4/5 OOD datasets with the Allegro Hand.
  • vs DextrAH-G / DextrAH-RGB / Singh et al.: these are sim-to-real on a wide variety of objects but still fall short on small/thin tabletop objects. DemoGrasp is, per the paper, the first to grasp previously unseen small/thin objects in tabletop settings without severe collisions.
  • vs ClutterDexGrasp: ClutterDexGrasp adds clutter; DemoGrasp adds clutter as an extension and still hits >80%.
  • vs DexNDM: DexNDM tackles dexterous policies via a different angle (neural distance manifold). Both papers exemplify the 2026 trend that careful problem framing replaces brute-force scaling.
  • vs D-REX: D-REX is a robot-experience scaling effort; DemoGrasp is on the data-efficiency end of the same trend.

The single-step MDP + demo-editing recipe is broadly transferrable to other manipulation skills (pouring, insertion) where a successful trajectory exists and only needs re-aiming to generalize.

Links

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️