RSS 2026 TeleGate - Heungwoo/research GitHub Wiki

TeleGate: Whole-Body Humanoid Teleoperation via Gated Expert Selection with Motion Prior

Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Humanoids · paper #25 Authors: Jie Li, Bing Tang, Feng Wu arXiv: 2602.09628 · program page

Summary compiled from the arXiv paper (v2, USTC + AnyWit Robotics); all numbers quoted from the paper. Trend context: RSS 2026 survey.

TeleGate real-time teleoperation of a Unitree G1 (Figure 1 of arXiv 2602.09628, © the authors)

Figure 1 shows an operator in an inertial motion-capture suit teleoperating the Unitree G1 in real time across four highly dynamic behaviors: (a) grasping and placing toys into a basket, (b) standing long jump, (c) prone-position stand-up, and (d) kicking a ball — the motion classes that typically break distilled single-policy controllers.

Problem

Real-time whole-body humanoid teleoperation needs one controller that handles motions with very different dynamics (walking, jumping, fall recovery), but multi-motion RL suffers catastrophic forgetting and the standard fix — distilling multiple expert policies into one generalist — loses performance on highly dynamic motions. Real-time teleoperation additionally lacks the future reference trajectory that offline trackers exploit, so anticipatory motions (jumping, standing up) are hard.

Method

TeleGate has three stages: (I) collect whole-body motion with inertial mocap and retarget it online to the robot, partitioning data into six motion categories; (II) train four expert policies (walk/run, dance/martial-arts, fall-and-recovery, jump) with PPO in MuJoCo using asymmetric actor–critic, plus a jointly trained small-Transformer VAE motion-prediction prior whose encoder reads the historical trajectory window and whose decoder reconstructs the future window, so the latent z_t injects implicit future-motion intent into the policy (loss = PPO + reconstruction + KL); a failure-rate-based curriculum reweights hard clips; (III) freeze all experts and train a lightweight gating network that scores the K experts from proprioception and reference trajectory and routes each control cycle to the top-1 expert. Training uses only ~2.5 hours of mocap (walking 40 min, running 24, dancing 24, martial arts 20, fall recovery 26, jumping 16).

Results

In 4-seed simulation comparisons TeleGate reaches 97.3% success and 17.22 mm mean per-joint position error, versus TWIST 68.9%/53.55 mm, Any2track 91.2%/18.40 mm, and GMT 92.0%/29.14 mm. Ablations show top-1 gating (96.7%) beats DAgger distillation (91.2%), a single policy (94.7%), top-2 gating (93.7%), and random expert partitioning (95.0%), and even edges the per-expert Oracle (96.3%); adding the VAE prior lifts overall SR to 97.3% and cuts Jump Empjpe by 13% (22.27→19.40 mm) and Fall-Recovery by 8% (28.23→25.94 mm). The system is deployed on a real Unitree G1 performing running, standing long jump, prone stand-up, and ball kicking.

Significance

Shows that routing among frozen experts, rather than distilling them, preserves expert-level dynamic skill in a unified real-time teleoperation controller with tiny data (2.5 h vs 42–700 h for TWIST/SONIC-scale efforts). Relevant to the humanoid control stack discussion in Review-Humanoid-VLA and the fast/slow control decomposition in Review-System-0-1-2.

← Back to RSS 2026 survey · RSS-2026-Papers · Home