RSS 2026 DexEvolve - Heungwoo/research GitHub Wiki

DexEvolve: Evolutionary Optimization for Robust and Diverse Dexterous Grasp Synthesis

Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Manipulation 2 · paper #59 Authors: René Zurbrügg, Andrei Cramariuc, Marco Hutter arXiv: 2602.15201 · program page

Summary compiled from the arXiv paper (v1); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Evolutionarily refined dexterous grasps (Figure 1 of arXiv 2602.15201, © the authors)

Figure 1: diverse, physically stable XHand grasps produced by evolutionary refinement, shown on two Handles assets (pink handles on the tabletop) and one object asset — candidates are refined directly inside high-fidelity simulation rather than merely filtered by it.

Problem

Data-driven dexterous grasp prediction needs large, diverse datasets, but analytical grasp synthesis makes simplifying assumptions (coarse contact dynamics, friction approximations), so most proposals fail high-fidelity physics verification and get discarded — a sample-inefficient generate-then-filter paradigm. High-fidelity simulators are also non-differentiable, ruling out gradient-based refinement of full grasp configurations.

Method

DexEvolve reinterprets the simulator as a black-box objective: analytical seeds from GraspQP initialize an asynchronous, gradient-free evolutionary algorithm running in Isaac Sim with massively parallel rollouts and early rejection. Grasps G = (wrist pose, joint states, delta joint commands) evolve via density-aware tournament selection (suppressing clustered modes), finger/pose-swapping crossover, Gaussian mutation, and archive-based novelty insertion (candidates within distance τ of a neighbor only replace it on fitness improvement). Fitness combines a disturbance-protocol lifetime score (forces along ±x, ±y, ±z), a contact-distance penalty, and a penetration penalty; contact points and grasp commands are resampled per offspring via FPS + contact-Jacobian solves. Because the objective need not be differentiable, a PointNet++ preference model trained on ~1,000 human pairwise annotations (Bradley–Terry loss) can steer refinement toward natural grasps. The refined distribution is finally distilled into a point-cloud-conditioned diffusion model (DexGraspAnything-style with contact-consistency constraints) for deployment. The paper also introduces a Handles dataset of 90 geometrically distinct, commercially modeled handle/knob assets annotated for the XHand.

Results

On the Handles dataset and a DexGraspNet subset, refinement yields over 120 distinct stable grasps per object — a 1.7–6x improvement over unrefined analytical seeds — while raising success rates and entropy; convergence plateaus by roughly 10k simulator steps. Against diffusion trained on the same seeds, evolutionary refinement achieves ~115 vs 72 unique grasps at 32 seeds (60% better) and 118 vs 81 at 128 seeds (46% better), with higher entropy (~3.0 vs 2.4–2.6). Real-world deployment on a Franka Panda + XHand (RealSense L415-captured point clouds lifted with Depth Anything V3, cuRobo planning) shows successful cabinet-handle grasps across multiple grasp modes.

Significance

Turns high-fidelity simulation from a reject filter into the optimizer itself, and shows quality-diversity machinery (archives, density-aware selection) scaling to full dexterous grasp configurations — a data-generation recipe upstream of the learned-grasping thread in Review-Dexterous-Manipulation.

← Back to RSS 2026 survey · RSS-2026-Papers · Home