RSS 2026 Force Policy - Heungwoo/research GitHub Wiki
Force Policy: Learning Hybrid Force-Position Control Policy under Interaction Frame for Contact-Rich Manipulation
Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Manipulation 3 · paper #128 Authors: Hongjie Fang, Shirun Tang, Mingyu Mei, Haoxiang Qin, Zihao He, Jingjing Chen, Feng Ying, Chenxi Wang, Wanxi Liu, Zaixing He, Cewu Lu, Shiquan Wang arXiv: 2602.22088 · program page
Summary compiled from the arXiv paper (v2); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Figure 1 uses EV-charger plugging to illustrate the human-inspired split: vision guides global movement and coarse alignment; on contact, high-frequency force feedback forms and corrects the local contact structure. The right panel shows the resulting architecture — a 5 Hz global vision policy feeding a global feature to a 50 Hz local force policy that outputs structure plus action for hybrid force-position control.
Problem
Learning-based contact-rich policies usually entangle perception, planning, and contact refinement in one monolithic network (force appended to observations, e.g. ForceVLA, TA-VLA), trading global generalization against stable local refinement; control-centric approaches assume a known task structure or learn only controller parameters. Two questions: how to realize a global-local organization, and how to represent the task interaction structure explicitly so it transfers across contact skills.
Method
The paper (Noematrix / Flexiv / SJTU / Zhejiang) formalizes an interaction frame (IF) — an instantaneous local basis derived from the spectral decomposition of environmental stiffness, anchored to intended twist/wrench — and recovers it from non-ideal demonstrations by classifying the dominant power residual (structural vs dissipative, prompted via Gemini 3 Pro on visual context) and orthogonalizing twist against wrench (or vice versa). Interaction patches are classified into Free/Surface/Insertion/Rotation modes that map to a hybrid-control selection mask. Force Policy then splits control: a 5 Hz global vision policy (instantiated with RISE-2; swappable with VLAs) handles free-space motion and provides a global feature; a 50 Hz local force policy (ResNet wrist-image + GRU proprioception/wrench encoders, FiLM-conditioned, MIP diffusion head) predicts the IF, selection mask, reference wrench, and local action chunk for hybrid force-position control. The predicted selection mask doubles as the implicit router, and a dual-policy asynchronous scheduler with DTW-based chunk alignment keeps 50 Hz execution smooth despite inference latency.
Results
On a Flexiv Rizon 4 with wrist + global RealSense D415 cameras and 50 demos/task: across Push-and-Flip, Plug-in-EV-Charger (~160 N insertion), and Scrape-off-Sticker (Easy/Hard, ~35 N pressing), Force Policy tops every stage — e.g., flip 95.0% vs 60.0% (FoAR) and 42.5% (RISE-2); EV plug-in 65.0% vs ≤10% for all baselines (RISE-2 and π0.5 at 0%); Hard sticker full-off 90.0% vs ≤20%. Force regulation: 0.00 cm spurious pushed distance on the heavy-object test and force profiles closely tracking demonstrations while baselines oscillate or under-press. Generalization to unseen objects strongly beats all baselines (e.g., 5/5, 5/5, 4/5 on the first three unseen objects). Ablations: the adaptive IF recovery beats analytic, power-based, and wrench-only labeling (wrench-only labels drop task success from 90% to 50%); semantic classification accuracy is 92–100%; the scheduler improves SPARC smoothness (−2.640 vs −4.515 linear).
Significance
Makes the interaction structure itself a learned, explicit output — reviving hybrid force-position control as the policy's interface rather than a hand-tuned controller — and shows a plug-and-play global/local split where any visuomotor policy or VLA can serve as the global layer. Related wiki threads: Review-Dexterous-Manipulation · Review-Cross-Embodiment.
← Back to RSS 2026 survey · RSS-2026-Papers · Home