RSS 2026 Emerging Extrinsic Dexterity in Cluttered - Heungwoo/research GitHub Wiki

Emerging Extrinsic Dexterity in Cluttered Scenes via Dynamics-aware Policy Learning

Venue: RSS 2026 (Sydney, Jul 13–17) · Session: RL · paper #149 Authors: Yixin Zheng, Jiangran Lyu, Yifan Zhang, Jiayi Chen, Mi Yan, Yuntian Deng, Xuesong Shi, Xiaoguang Zhao, Yizhou Wang, Zhizheng Zhang, He Wang arXiv: 2603.09882 · program page

Summary compiled from the arXiv paper (v2); all numbers quoted from the paper. Trend context: RSS 2026 survey.

DAPL extrinsic dexterity behaviors (Figure 1 of arXiv 2603.09882, © the authors)

Given a current object state (blue outline) and goal state (green), the policy selectively uses environmental contact: (a) avoiding disturbable objects when free-space motion suffices, (b) traversing unavoidable obstacles under contact-rich dynamics, (c) deliberately leveraging a neighbor to flip an object that is hard to flip without contact, and (d) generalizing across diverse real cluttered scenes.

Problem

Extrinsic dexterity — using environmental contact via pushing, sliding, toppling — is needed where grasping fails, but in clutter it requires selectively exploiting contacts among multiple objects with coupled dynamics. Geometry-centric representations (CORN, UniCORN) and planning/primitive pipelines are brittle in dense clutter because they don't model how objects respond once contact occurs.

Method

DAPL (Galbot/PKU/BAAI/CASIA) is a two-stage loop: (1) a physical world model is pretrained to predict point-level object dynamics (dense position/velocity supervision over target-object, scene, and end-effector point clouds carrying 7-D physical features including velocity and mass), yielding a dynamics-aware scene representation; (2) an RL policy conditioned on this representation, proprioception, and a relative goal pose outputs joint commands, with only a sparse success reward plus light contact/goal shaping. A curriculum alternates policy training (~2×10^4 RL steps/stage) with world-model refinement (500k steps) on the policy's own rollouts. Training/benchmarking use Clutter6D, a new 6D rearrangement benchmark with task-oriented scene graphs, ~10K normalized assets, and Sparse (4 objects) / Moderate (8) / Dense (12) tracks.

Results

On held-out Clutter6D scenes, DAPL achieves 71.88 / 51.04 / 44.56% success (Sparse/Moderate/Dense), roughly doubling the best baseline in Dense (CORN 22.22%; UniCORN 5.81%; prehensile GraspGen+CuRobo 3.13%; teleoperation 20%). Curriculum iterations raise success 61.3% → 71.8%. Zero-shot sim-to-real on a Franka Research 3 with three RealSense cameras (mass priors estimated by prompting GPT-5) yields 48% over 10 cluttered scenes vs 52% for human teleoperation, with lower mean execution time (42.6 s vs 55.9 s); a Galbot G1 humanoid grocery-retrieval deployment combines the policy with a planner and grasping module. Sensitivity tests show robustness to 25-50% mass/velocity noise but sharp degradation with goal-pose noise.

Significance

Shows contact-induced dynamics modeling — not more geometry — is the missing ingredient for non-prehensile manipulation in clutter, with emergent avoid/traverse/leverage behaviors; connects the world-model and RL threads (RL, Review-World-Models) to Review-Dexterous-Manipulation.

← Back to RSS 2026 survey · RSS-2026-Papers · Home