RSS 2026 Functional Force Aware Retargeting from Virtual - Heungwoo/research GitHub Wiki
Functional Force-Aware Retargeting from Virtual Human Demos to Soft Robot Policies
Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Imitation learning 3 · paper #202 Authors: Uksang Yoo, Mengjia Zhu, Evan Pezent, Jom Preechayasomboon, Jean Oh, Jeffrey Ichnowski, Amir Memar, Ben Abbatematteo, Homanga Bharadhwaj, Ashish Deshpande, Harsha Prahlad (Robotics Institute, CMU; Meta Reality Labs) arXiv: 2604.01224 · program page
Summary compiled from the arXiv paper (v1); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Figure 1 overview. A human demonstrates in a VR environment (capturing hand kinematics, object motion, and contact forces); SoftAct's two-stage force-aware retargeting (Stage 1 force-balanced finger assignment, Stage 2 contact-informed refinement) maps this onto a non-anthropomorphic pneumatic soft robot hand. A kinematic pose policy feeds a trajectory-tracking controller (forward model + pressure controller) for real-world deployment.
Problem
Soft robot hands are underactuated, compliant, and non-anthropomorphic, so conventional retargeting that assumes finger-to-finger kinematic correspondence fails under extreme embodiment mismatch. The functional intent of a human demonstration — where and how hard the hand pushes — cannot be recovered from joint trajectories alone.
Method
SoftAct explicitly reasons about contact forces. Immersive VR captures rich human demonstrations (hand kinematics, object motion, dense contact patches, and contact forces). A two-stage, force-aware retargeting algorithm follows: Stage 1 attributes demonstrated contact forces to individual human fingers and allocates robot fingers proportionally, giving a force-balanced human→robot mapping; Stage 2 does online retargeting by combining baseline end-effector pose tracking with geodesic-weighted contact refinements (using contact geometry and force magnitude via mesh geodesics) to adjust robot fingertip targets in real time. A diffusion policy is trained to imitate the retargeted trajectories, and a learned MLP inverse model — optimized with gradient descent at execution time — converts desired fingertip displacements into pneumatic chamber pressures on a custom soft hand.
Results
On a custom non-anthropomorphic pneumatic soft hand across six contact-rich tasks (light bulb insertion, light bulb twisting, cup pouring, marker grabbing, bottle unscrewing, box reorienting). The low-level controller reaches 2.28 mm fingertip RMSE vs 5.05 (linear), 6.11 (MLP), 8.74 (KNN) baselines — a 55% RMSE reduction over the strongest baseline (48% mean error; 63%/74% vs learning-based methods), with tracking variance cut by up to 69%. Object-level rotational error drops up to ~80%. At the policy level, SoftAct beats a kinematic-only baseline on all tasks (30 sim rollouts, 20 real per task), raising success by roughly 30–70%; e.g. real-world Paper Cup Pouring 85% vs 35%, Light Bulb Screwing 95% vs 30%, Box Reorienting 70% vs 10%.
Significance
SoftAct shows that modeling contact geometry and force distribution — not just kinematics — is essential for transferring human manipulation skills to compliant soft hands, enabling stable zero-shot real-world deployment across extreme embodiment mismatch. Related: Review-Human-Video-Transfer · Review-Realtime-Execution.
← Back to RSS 2026 survey · RSS-2026-Papers · Home