CoRL 2025 UniFP - Heungwoo/research GitHub Wiki

UniFP โ€” Unified Policy for Position and Force Control in Legged Loco-Manipulation

Venue: CoRL 2025 (Best Paper) ยท Authors: Peiyuan Zhi, Peiyang Li, Jianqin Yin, Baoxiong Jia, Siyuan Huang โ€” BIGAI (State Key Lab of General AI) / BUPT Paper: arXiv 2505.20829 ยท PMLR 305:652-669 Category: Legged Loco-Manipulation / Whole-Body / Force Control Trend tag: Force as first-class policy output

Approach diagram

flowchart LR
  S[Proprioceptive<br/>state history] --> POL[RL unified policy]
  POL --> EST[Estimate external<br/>force from state]
  EST --> ADJ[Position + velocity<br/>compensation]
  ADJ --> R[Legged robot<br/>loco-manipulation]
Loading

Problem

Contact-rich tasks (pushing doors, opening drawers, carrying loads while walking) require force control, but learned policies almost always output positions. Switching between position-mode and force-mode with hand-engineered logic is brittle, especially on legged platforms where loco and manipulation must be co-scheduled.

Method

A single RL-trained policy that unifies position and force control without relying on force/torque sensors. By simulating diverse combinations of active position and force commands together with external disturbance forces, the policy learns to estimate the external contact force from the robot's historical proprioceptive states and to compensate for it via position and velocity adjustments. A single network thereby supports a spectrum of behaviors โ€” position tracking, force application, force tracking, and compliant interaction โ€” selected by the commanded input rather than hand-engineered mode switching.

Results

CoRL 2025 Best Paper award. Validated on two platforms โ€” a Unitree B2-Z1 quadrupedal manipulator (12-DOF base + 6-DOF Z1 arm) and a Unitree G1 humanoid (29-DOF). When plugged into an imitation-learning pipeline (RGB from end-effector- and base-mounted cameras), the force-aware policy achieves ~39.5% higher success than position-only baselines across four contact-rich tasks (e.g., wiping, cabinet/drawer opening).

Significance

One of CoRL 2025's strongest signals that contact and force are now first-class in robot learning โ€” alongside TA-VLA, DexSkin, and Tactile Beyond Pixels. Previous position-only paradigm was leaving capability on the table for anything contact-rich.

Links

Related pages

โ† Back to CoRL-2025

โš ๏ธ **GitHub.com Fallback** โš ๏ธ