RSS 2026 Generalizing from References using a - Heungwoo/research GitHub Wiki

Generalizing from References using a Multi-Task Reference and Goal-Driven RL Framework

Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Humanoids · paper #26 Authors: Jiashun Wang, M. Eva Mungai, He Li, Jean Pierre Sleiman, Jessica K. Hodgins, Farbod Farshidian arXiv: 2602.20375 · program page

Summary compiled from the arXiv paper (v1); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Humanoid box-parkour behaviors (Figure 1 of arXiv 2602.20375, © the authors)

Figure 1: hardware filmstrips of the Unitree G1 executing the three learned skills — walk-climb (top), walk-jump (middle), and climb-down (bottom) — on a box-based parkour setup, showing human-like whole-body coordination through contact-rich transitions.

Problem

Reference-tracking humanoid policies (DeepMimic-style, distillation, adversarial imitation) produce natural motion but are brittle outside the demonstration distribution, while purely task-driven RL adapts but loses motion quality and needs heavy reward shaping. The paper (RAI Institute / CMU) asks how to use human reference motion as a behavioral prior rather than a deployment-time constraint.

Method

A single goal-conditioned policy is trained jointly on two tasks sharing observation and action spaces: (i) a reference-guided imitation task, where reference motions define dense tracking rewards and goal conditions but are never policy inputs (no phase variables, no trajectory input, no adversarial discriminator); and (ii) a goal-conditioned generalization task, where 2-D root-position goals are sampled independently of any reference and rewards are sparse task-success terms plus shared regularization/survival terms. A single scalar difficulty variable λ drives an automatic curriculum that simultaneously anneals a virtual assistive wrench at the base and shifts task-sampling probability from imitation toward generalization. Training uses PPO in Isaac Lab with an asymmetric actor-critic (3-layer MLPs; the critic gets a task indicator and privileged state); actions are residual PD joint-position targets without a reference feedforward term. One policy per behavior (walk-jump, walk-climb, climb-down) is trained from a single reference motion (plus mirrored variants).

Results

On the Unitree G1 (29 DoF, 1.2 m, 35 kg) and in MuJoCo/Isaac Lab: under nominal initializations the method reaches 1.00 success on all three skills; under beyond-nominal initializations (up to ±2 m forward, ±1 m lateral, ±45° yaw) it achieves 0.62/0.76/0.98 (walk-jump/walk-climb/climb-down) vs 0.17/0.57/0.90 for ZEST mocap tracking and 0.54/0.53/0.91 for tabula rasa RL, while keeping lower root-orientation error. Ablations: removing the task curriculum or the imitation task yields 0.00 success on walk-jump/walk-climb; removing the generalization task collapses beyond-nominal success (e.g., 0.27 on walk-jump). A rule-based state machine composes the skills into long-horizon parkour sequences, evaluated sim-to-sim in MuJoCo, and hardware runs show adaptive strategies (e.g., jumping directly when initialized near the box).

Significance

Shows that dense imitation shaping and sparse goal-driven RL can be co-optimized in one policy — no adversarial objectives, distillation stages, or reference-dependent inference — resolving the naturalness-vs-generalization trade-off for athletic humanoid control. Relevant to the humanoid whole-body-control thread in Review-Humanoid-VLA.

← Back to RSS 2026 survey · RSS-2026-Papers · Home