RSS 2026 Perceptive Humanoid Parkour - Heungwoo/research GitHub Wiki

Perceptive Humanoid Parkour: Chaining Dynamic Human Skills via Motion Matching

Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Humanoids · paper #20 Authors: Zhen Wu, Xiaoyu Huang, Lujie Yang, Yuanhang Zhang, Xi Chen, Pieter Abbeel, Rocky Duan, Angjoo Kanazawa, Carmelo Sferrazza, Guanya Shi, Karen Liu arXiv: 2602.15827 · program page

Summary compiled from the arXiv paper (v2); all numbers quoted from the paper. Trend context: RSS 2026 survey.

PHP parkour skills on a Unitree G1 (Figure 1 of arXiv 2602.15827, © the authors)

Real-hardware strobe sequences of the single multi-skill policy: (a) cat-vault then dash-vault chained at ~3 m/s, (b) climbing onto a 1.25 m wall (96% of robot height) and rolling down, (c) speed-vault at ~3 m/s, and (d) a 60-second continuous multi-obstacle course traversal with autonomous skill selection.

Problem

Humanoid locomotion has reached stable walking on varied terrain, but agile parkour requires highly dynamic contact-rich skills (climbing, vaulting, rolling), human-like expressiveness, long-horizon skill composition, and perception-driven decisions about which skill to deploy — capabilities that reward-shaping RL or isolated motion clips alone do not deliver.

Method

PHP (Amazon FAR with UC Berkeley/CMU/Stanford) is a modular pipeline: (1) motion matching — nearest-neighbor search in a feature space over a locomotion database plus retargeted atomic human skill clips (only about two demonstrations per skill, a few seconds each) — composes obstacle-adaptive long-horizon kinematic reference trajectories with smooth transitions; (2) privileged motion-tracking RL expert policies are trained per skill following BeyondMimic-style DeepMimic rewards; (3) experts are distilled into a single depth-based multi-skill student using DAgger combined with PPO, with a warmup curriculum that decays the DAgger weight to 0.1 and relaxes termination thresholds to tolerate mirrored skill modes. The student uses onboard depth and a discrete 2D velocity command to decide whether to step over, climb, vault, or roll off obstacles.

Results

Zero-shot sim-to-real on a Unitree G1: climbing a 1.25 m wall (96% robot height) in 3.63 s, a cat vault clearing a 0.4 m-high obstacle while covering over 2 m (154% of height) at peak 3.41 m/s, drop landings from 1.25 m, and a 48-60 s multi-obstacle course with online adaptation to obstacle displacement. In simulation benchmarks (success rate over obstacle heights/speeds), PHP scores 0.95-1.00 across 36-76 cm obstacles at 1-2 m/s, versus velocity-tracking RL (0.00 beyond 36 cm), uncomposed motion data (≤0.37), and end-to-end depth RL (≤0.95 only on the easiest case). Ablations show DAgger-only distillation collapses on hard skills (e.g., 0.03-0.16 at 1 m/s) while DAgger+RL restores 0.95-1.00.

Significance

A clean recipe — motion matching for composition, teacher-student DAgger+RL for control — that pushes perceptive humanoid agility well past prior reward-shaped locomotion; relevant to Review-Humanoid-VLA and the skill-chaining/system-level discussion in Review-System-0-1-2.

← Back to RSS 2026 survey · RSS-2026-Papers · Home