CoRL 2026 CHIP - Heungwoo/research GitHub Wiki
CoRL 2026 โ CHIP: Adaptive Compliance for Humanoid Control through Hindsight Perturbation
Venue: CoRL 2026 (Austin, TX, Nov 9โ12). Paper: arXiv 2512.14689. Representative of: learned compliance for humanoid manipulation โ control not just where to move but how stiff/compliant the arms should be. Companions: Humanoid VLA ยท CoRL 2026 survey.

1. Problem
Humanoid robots have become impressively agile at locomotion, but they remain weak at forceful manipulation โ moving heavy objects, wiping a surface, pushing a cart, opening a door. RL-based motion-tracking controllers optimize for precise tracking of a reference motion, which makes the arms effectively stiff: they resist external contact instead of yielding to it. Classical impedance/compliance control gives that knob but integrates poorly with learned agile tracking. The goal is a controller that lets you dial end-effector stiffness up or down while still tracking dynamic reference motions well.
2. Method
CHIP is a plug-and-play module for keypoint-based motion-tracking policies that adds controllable end-effector stiffness โ needing neither data augmentation nor extra reward tuning. The trick is hindsight perturbation: during training a random force f is applied at the end effector, and instead of editing the reference trajectory, CHIP edits the observed tracking goal to a hindsight target g โ (1/k)ยทf, while the tracking reward stays anchored to the original reference g. The coefficient controls the emulated stiffness k, so at deployment a continuous compliance coefficient sets how much the arm yields under contact. Force is estimated implicitly from proprioceptive history, and the module works with both local and global 3-point tracking policies. Trained on a Unitree G1.
3. Results
Reported on the Unitree G1 (baselines: FALCON force-perturbation and a no-force standard tracker):
- Agile tracking is preserved โ global position tracking error stays around 0.08 m.
- End-effector displacement scales roughly linearly with the compliance coefficient, i.e. stiffness is genuinely controllable.
- On multi-robot collaborative grasping/transport, CHIP reaches ~80% success vs. much lower baseline rates.
- Downstream demos: VR teleoperation with on-the-fly compliance, cart pushing, door opening, and a VLA trained for autonomous wiping (60โ80% success).
4. Why it matters
Compliance has been the missing knob for humanoid manipulation: agile trackers are stiff by construction, so contact-rich tasks either fail or fight the environment. CHIP shows you can graft a controllable-stiffness capability onto an existing tracking policy cheaply โ no trajectory relabeling, no reward surgery โ which makes it an attractive building block for humanoid VLAs and teleoperation.
Limitations (reviewer): stiffness is emulated via the observed-goal offset rather than measured force control, so behavior under large or fast contact forces depends on the quality of the implicit force estimate; results are single-embodiment (G1); reported manipulation success rates leave substantial headroom.
5. Links
- arXiv 2512.14689
- Survey: CoRL 2026 ยท Related: Humanoid VLA
โ Back to CoRL 2026 survey ยท Home