RSS 2026 SimToolReal - Heungwoo/research GitHub Wiki
SimToolReal: An Object-Centric Policy for Zero-Shot Dexterous Tool Manipulation
Venue: RSS 2026 (Sydney, Jul 13–17) · Session: RL · paper #151 Authors: Kushal Kedia, Tyler Ga Wei Lum, Jeannette Bohg, Karen Liu arXiv: 2602.16863 · program page
Summary compiled from the arXiv paper (v2); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Top: a single policy deployed zero-shot on novel real tools and tasks — hammer, marker, spatula over a pan, and brushes — on a KUKA arm with a five-fingered hand. Bottom: the characteristic grasp → in-hand rotation → tool-use sequence for sweeping crumpled paper into a dustpan with a brush.
Problem
Tool use is a hard class of dexterity — grasping thin objects lying flat, in-hand reorientation into functional poses, and forceful contact — that parallel-jaw grippers resist and teleoperation captures poorly. Prior sim-to-real RL needed per-task object modeling and reward tuning, limiting generality.
Method
SimToolReal trains a single goal-conditioned, object-centric RL policy in simulation over procedurally generated tool-like primitives, with the universal objective of moving each object through random goal poses (poses represented via D = 4 keypoints); this induces grasping, in-hand rotation, and stable-contact skills without task-specific engineering. Training uses SAPG optimization with an asymmetric critic. At deployment, SAM 3D recovers the object mesh and graspable-region bounding box from RGB-D, and goal-pose trajectories are extracted from a human video, so the sim-trained policy runs zero-shot on real tools. Hardware: 22-DoF Sharpa five-fingered hand on a 7-DoF KUKA iiwa 14 (29-DoF control).
Results
On the authors' DexToolBench (6 tool categories, 12 instances, 24 task trajectories, 120 real rollouts at 5 trials each), the single policy shows strong zero-shot Task Progress; on brush-sweeping variants it scores 98.0% (no rotation) and 82.7% (with 90° rotation) versus 61.0/10.8% for Fixed Grasp and 8.1/0% for Kinematic Retargeting — the claimed 37% average improvement. In simulation it matches specialist per-task RL policies on their own training setups while specialists collapse under object or trajectory changes. Failure modes: pose-tracking loss 43.7%, drops 34.5%, incomplete in-hand rotation 18.2%, grasp failure 3.6%, with consistent re-grasp recovery behavior.
Significance
Shows that "reach random goal poses with random primitives" is a sufficient universal pretext task for dexterous tool use, replacing per-task sim-to-real engineering with one policy steered at test time by human-video trajectories. Related wiki threads: RL · Review-Dexterous-Manipulation.
← Back to RSS 2026 survey · RSS-2026-Papers · Home