CoRL 2026 CHIP - Heungwoo/research GitHub Wiki

CoRL 2026 โ€” CHIP: Adaptive Compliance for Humanoid Control through Hindsight Perturbation

Venue: CoRL 2026 (Austin, TX, Nov 9โ€“12). Paper: arXiv 2512.14689. Representative of: learned compliance for humanoid manipulation โ€” control not just where to move but how stiff/compliant the arms should be. Companions: Humanoid VLA ยท CoRL 2026 survey.

CHIP training-and-deployment overview: hindsight perturbation during RL yields a controllable-stiffness tracking policy on the Unitree G1 (figure from the authors, arXiv 2512.14689, ยฉ the authors)

1. Problem

Humanoid robots have become impressively agile at locomotion, but they remain weak at forceful manipulation โ€” moving heavy objects, wiping a surface, pushing a cart, opening a door. RL-based motion-tracking controllers optimize for precise tracking of a reference motion, which makes the arms effectively stiff: they resist external contact instead of yielding to it. Classical impedance/compliance control gives that knob but integrates poorly with learned agile tracking. The goal is a controller that lets you dial end-effector stiffness up or down while still tracking dynamic reference motions well.

2. Method

CHIP is a plug-and-play module for keypoint-based motion-tracking policies that adds controllable end-effector stiffness โ€” needing neither data augmentation nor extra reward tuning. The trick is hindsight perturbation: during training a random force f is applied at the end effector, and instead of editing the reference trajectory, CHIP edits the observed tracking goal to a hindsight target g โˆ’ (1/k)ยทf, while the tracking reward stays anchored to the original reference g. The coefficient controls the emulated stiffness k, so at deployment a continuous compliance coefficient sets how much the arm yields under contact. Force is estimated implicitly from proprioceptive history, and the module works with both local and global 3-point tracking policies. Trained on a Unitree G1.

3. Results

Reported on the Unitree G1 (baselines: FALCON force-perturbation and a no-force standard tracker):

  • Agile tracking is preserved โ€” global position tracking error stays around 0.08 m.
  • End-effector displacement scales roughly linearly with the compliance coefficient, i.e. stiffness is genuinely controllable.
  • On multi-robot collaborative grasping/transport, CHIP reaches ~80% success vs. much lower baseline rates.
  • Downstream demos: VR teleoperation with on-the-fly compliance, cart pushing, door opening, and a VLA trained for autonomous wiping (60โ€“80% success).

4. Why it matters

Compliance has been the missing knob for humanoid manipulation: agile trackers are stiff by construction, so contact-rich tasks either fail or fight the environment. CHIP shows you can graft a controllable-stiffness capability onto an existing tracking policy cheaply โ€” no trajectory relabeling, no reward surgery โ€” which makes it an attractive building block for humanoid VLAs and teleoperation.

Limitations (reviewer): stiffness is emulated via the observed-goal offset rather than measured force control, so behavior under large or fast contact forces depends on the quality of the implicit force estimate; results are single-embodiment (G1); reported manipulation success rates leave substantial headroom.

5. Links

โ† Back to CoRL 2026 survey ยท Home