ICLR 2026 D REX - Heungwoo/research GitHub Wiki

D-REX β€” Differentiable Real-to-Sim-to-Real for Dexterous Grasping

Venue: ICLR 2026 Category: Dexterous Manipulation β€” Sim-to-Real Trend tag: Differentiable simulation / digital twin Authors / Affiliations: Haozhe Lou, Mingtong Zhang, Haoran Geng, Hanyang Zhou, Sicheng He, Zhiyuan Gao, Siheng Zhao, Jiageng Mao, Pieter Abbeel, Jitendra Malik, Daniel Seita, Yue Wang β€” USC (PSI Lab / Viterbi) and UC Berkeley (EECS) Paper: arXiv:2603.01151

Approach diagram

flowchart LR
  Real[Real RGB videos<br/>scene + object] --> GS[2DGS + 3DGS<br/>collision mesh K + visual]
  GS --> MJCF[VLM-generated MJCF<br/>initial physical params]
  MJCF --> Sim[Brax/MuJoCo<br/>differentiable engine]
  Real2[Real robot push trajectory<br/>FoundationPose poses] --> Loss
  Sim --> Loss["L_traj : β€–s_sim(m) βˆ’ s_realβ€–Β²<br/>over T steps"]
  Loss -->|βˆ‡_m via adjoint| Mass[Identified mass m*]
  Hum[Human grasp video] --> HaMeR[HaMeR + MCC-HO<br/>hand+object pose]
  HaMeR --> Retgt[Dex-Retargeting<br/>β†’ Allegro joint angles]
  Mass --> Pol["GraspMLP Ο€_Ο†<br/>predicts (Γ‚, rΜ‚, fΜ‚)<br/>conditioned on K, m"]
  Retgt --> Pol
  Pol --> Deploy[Real-world grasp<br/>force-aware control]
Loading

Problem

Sim-to-real transfer for dexterous grasping is bottlenecked by inaccurate object physics (especially mass) and by the cost of collecting robot-specific demonstrations. Standard real-to-sim pipelines reconstruct geometry but rarely identify dynamics parameters. Position-only policies replicate human grasps but ignore unobserved gravitational forces β€” applying a uniform force across objects of varying masses leads to bounce-off (too heavy a grasp on a light object) or slippage (too light a grasp on a heavy object). The authors set out to (i) recover object mass directly from real-world push interactions and (ii) use the recovered mass to condition a grasping policy that adapts force.

Detailed Method

D-REX is built on top of MuJoCo (forward sim), Brax (differentiable kinematics), and a GradSim-style adjoint for differentiable contact dynamics. The full pipeline has four stages.

1. Visual + geometric reconstruction

For each object the authors capture roughly 306–337 RGB images (Table 4) from a mobile device. A COLMAP SfM pass is followed by 2D Gaussian Splatting (with surface normals, used for collision geometry) and 3D Gaussian Splatting (used for photorealistic rendering). Mesh extraction yields collision meshes ranging from ~13.5 k vertices (Ketchup) to ~67.9 k vertices (Letter A). End-to-end COLMAP+3DGS+2DGS reconstruction takes 30–35 min per object.

A VLM (GPT-4o style) generates an initial MJCF by parsing scene images plus prompts, providing structural priors but unreliable physical parameters.

2. Differentiable mass identification

The dynamics use a Newton-Euler rigid-body model with a compliant penalty contact model:

f_n(s, u, θ) = -n · ( k_e · C(s) + k_d · Ċ(u) )

with stiffness k_e and damping k_d as part of ΞΈ. Integration uses a semi-implicit (symplectic) Euler scheme rather than explicit Euler β€” explicit Euler explodes with stiff contact and small initial-mass guesses; semi-implicit gives a larger stability region.

The trajectory loss matches simulated to real object pose:

L_traj(m) = Ξ£_{t=0..T} β€– s_sim_t(m) βˆ’ s_real_t β€–Β²

s_real_t comes from FoundationPose tracking; s_sim_t from rolling out the same robot control signals in the differentiable engine. Mass is optimized via the discrete adjoint method, with PyTorch backprop computing βˆ‡_m L_traj.

The push interaction is reduced to a planar push-down with a virtual fulcrum to suppress friction. Under that simplification, the trajectory is affine in 1/m (Eq. 23 in the paper), making the identification a well-conditioned least-squares problem in ΞΈ = 1/m.

Adaptive learning strategy. Particle masses initialise at ~0.002 kg/vertex, then the optimizer adapts the per-vertex masses according to the object's overall mass (Appendix A.2.3):

  • Heavy objects (~0.8 kg, e.g. Ketchup): up to 2000 epochs with higher LR.
  • Medium (~0.1 kg): converges in ~100 epochs.
  • Light (~0.05 kg): ~100 epochs with LR decay. The Runtime Analysis (Section 5.4) reports each iteration takes ~1.43–1.68 s with convergence typically within ~200 epochs (~5–20 min). The Table 5 integrator ablation reports per-iteration cost of 1.36–1.43 s for semi-implicit Euler vs 1.17–1.22 s for explicit Euler.

3. Human β†’ robot demo transfer

Each frame I_t of a human grasp video is processed by HaMeR (3D hand transformer) and MCC-HO (hand-held object reconstruction):

h_t ∈ SE(3) Γ— ℝ^{J_h}, o_t ∈ SE(3)

where J_h are finger joints. Dex-Retargeting maps (h_t, o_t) onto the robot hand with J_r DoF, producing target joint angles A_t ∈ ℝ^{J_r}. The authors assume object geometry is shared between human demo and robot execution.

4. GraspMLP policy + force-aware optimization

The policy Ο€_Ο† is a small MLP that consumes positionally-encoded mesh vertices K plus the identified mass m and produces a 19-dim output:

Ο€_Ο†(o) = (Γ‚ ∈ ℝ¹⁢, rΜ‚ ∈ ℝ², fΜ‚ ∈ ℝ¹)

with target grasping force fΜ‚ = mΒ·g / n_active where n_active is the predicted number of active contacts. Contact constraint rΜ‚ has two terms: sustained contact during rollout, and object retention at end of rollout. An indicator I_in_hand(t) flags whether n_active(t) β‰₯ N_min over horizon H.

Two-stage training (Algorithm 2):

  • Stage 1 (Supervised): MSE on action Γ‚ vs retargeted human action; BCE on rΜ‚ with ground-truth label = 1; MSE on fΜ‚ vs supervised target. Total loss L = L_a + L_r + L_f.
  • Stage 2 (Sim refinement): Roll out Γ’ in MuJoCo with the Real2Sim MJCF; observe r_env (whether grasp succeeded) and contact-based force f_env = clip(mΒ·gΒ·n_contacts / f_max, 0, 1). Loss is L = 0.8Β·L_r + 0.3Β·L_f (only the reward + force heads are refined; the action head is kept from Stage 1).

For standard objects ~200 demonstrations per object are used; for objects with higher geometric or dynamic complexity, the dataset scales up to 5000 demonstrations to cover the variance needed for robust policy learning (Appendix A.2.5).

Comprehensive Results

Mass identification across diverse objects (Table 1)

Object Inferred (VLM, g) Identified (g) GT (g) % error
Letter U 500 110 125 12.0%
Letter A 500 145 134 9.0%
Lego 300 53 59 8.6%
Domino 500 117 106 9.3%
Cookie 500 200 210 4.8%
Ketchup 1000 667 726 8.1%

Errors range 4.8 %–12.0 % without object-specific tuning, despite poor VLM priors (e.g. Ketchup was inferred as 1000 g when true mass is 726 g).

Identical geometry, different density (Table 2)

Density Identified (g) GT (g)
ρ₁ 95 82
ρ₂ 129 125
ρ₃ 207 218

All deviations are under 13 g, confirming sensitivity to internal density even when shape is fixed.

Force-aware grasping (Table 3, cross-mass evaluation)

Each row = policy trained on mass ρ_i, each column = evaluated mass ρ_j:

Train\Eval ρ₁ ρ₂ ρ₃
ρ₁ 75 % 30 % 15 %
ρ₂ 40 % 80 % 30 %
ρ₃ 15 % 40 % 95 %

Mass-matched diagonal dominates; off-diagonal failures are physical (under- or over-applied force).

Tabletop grasping vs baselines (Figure 7)

Eight objects of varying geometry and mass; compared against DexGraspNet 2.0 (large-scale sim training) and Human2Sim2Robot (recent real-to-sim-to-real from RGBD human demos). All baselines use D-REX's collision meshes for fairness. The following per-object and average success rates are read directly from the Figure 7 bar chart (legend: DexGraspNet 2.0 / Human2Sim2Robot / DREAM = D-REX):

Object Mass DexGraspNet 2.0 Human2Sim2Robot D-REX
Lightbulb 37 g 95 % 80 % 100 %
Cube 120 g 90 % 45 % 100 %
A 134 g 90 % 80 % 75 %
Cookie 210 g 80 % 10 % 85 %
Spam 365 g 15 % 10 % 70 %
Nutella 414 g 80 % 20 % 85 %
Spray 548 g 90 % 55 % 100 %
Ketchup 726 g 65 % 60 % 75 %
Average β€” 76 % 45 % 86 %

D-REX has the highest average (86 %) with substantially lower variance; baselines drop sharply on the heaviest/hardest objects (Human2Sim2Robot collapses to 10–20 % on Cookie, Spam, and Nutella) because their force command is implicitly tuned to lighter mass, while D-REX maintains 70–100 % across the full mass range. The paper states D-REX "consistently outperforms the baselines across eight objects with diverse geometries and masses, achieving high success rates with substantially lower variance."

Ablation Studies

  • Semi-implicit vs explicit Euler (Table 5, Appendix A.2.4). The paper motivates the semi-implicit (symplectic) Euler scheme on the grounds that explicit Euler explodes with stiff contact and small initial-mass guesses, while semi-implicit Euler has a larger stability region. Table 5 reports the identified mass and per-iteration time under each integrator:

    Object GT (g) Semi-implicit (g) Explicit (g) Semi-implicit (s) Explicit (s)
    Lego 59 51 34 1.38 1.21
    Ketchup 726 667 685 1.43 1.19
    Cookie 210 200 189 1.41 1.18
    Domino 106 117 135 1.36 1.17
    U 125 110 98 1.39 1.20
    A 134 145 120 1.42 1.22

    Semi-implicit Euler yields more accurate mass estimates (errors mostly single-digit grams) at a modest runtime cost (~0.2 s slower per iteration); explicit Euler is slightly faster but consistently less accurate with more frequent numerical instabilities.

  • Identified vs ground-truth mass for the policy (Figure 5). Across Letter A, Cookie, and Ketchup, success rate peaks at the identified mass, sometimes exceeding the ground-truth mass policy. Conditioning on arbitrary mismatched masses causes substantial drops (e.g. Cookie at 50 g or 1000 g near zero success).

  • Mass mismatch failure modes (Figure 6). Light-trained policy on heavy object: object slips out due to insufficient force. Heavy-trained policy on light object: bounce-off due to excessive force. Confirms force conditioning, not just position, is necessary.

Limitations

The authors do not include a dedicated "Limitations" section, but several are acknowledged in the appendix and discussion:

  • Z-axis error from FoundationPose is non-negligible and sometimes requires manual post-processing to align estimated poses with the digital asset.
  • The push-based identification uses a virtual-fulcrum planar assumption to suppress friction; the method does not yet jointly identify friction or contact stiffness alongside mass.
  • Reconstruction is offline (~30–35 min per object) and the authors note that mass-identification scales with mesh vertex count.
  • The retargeting step assumes identical object geometry between human and robot manipulation phases.
  • No demonstration of generalization across object categories with a single shared policy β€” each object is currently trained with its own ~200–5000 demonstrations.
  • Initial MJCF parameters from the VLM are unreliable (e.g. all six masses in Table 1 inferred as 300–1000 g regardless of true mass) and the system depends on real interaction data to correct them.

Significance & Positioning

D-REX couples differentiable rendering (Gaussian Splats) with differentiable physics (Brax + GradSim-style adjoint) for dexterous policy learning. Its niche is the often-overlooked dynamics half of the digital-twin agenda: prior real-to-sim work (URDFormer, Real2Code, Ditto) recovers articulation and geometry, while sim-to-real work (domain randomization, RetinaGAN) blurs the gap statistically β€” D-REX identifies the most consequential dynamics parameter (mass) directly from a robot-object push.

Within the 2026 dexterous-grasping cluster:

  • DemoGrasp addresses the data-efficiency side via a single retargeted demonstration; D-REX provides the missing physical fidelity needed for the policy to act on a heavy object.
  • DexNDM does per-joint dynamics correction on the robot side; D-REX corrects on the object side.
  • X-Sim uses Gaussian-Splat scenes for cross-embodiment policy distillation but without differentiable mass identification.
  • Compared to Human2Sim2Robot and DexGraspNet 2.0, D-REX's force-awareness is what unlocks reliable heavy-object grasping: Figure 7 reports a higher average success rate (86 % vs 76 % / 45 %) with substantially lower variance, with the largest gains on heavy/hard objects (e.g. Spam at 70 % vs 15 % / 10 %) where both baselines, which lack explicit force conditioning, collapse.

D-REX is also a concrete instantiation of the broader trend that VLA / generalist policies have begun to falter on physically demanding tasks β€” and that the route forward is identified-physics rather than larger pretraining corpora. It does not replace VLA pretraining (in the sense of Ο€0.6 / Ο€0.7 / GR00T Series); rather, it provides a complementary recipe for force-aware heads that those policies lack today.

Links

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️