RSS 2026 OmniXtreme - Heungwoo/research GitHub Wiki

OmniXtreme: Breaking the Generality Barrier in High-Dynamic Humanoid Control

Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Humanoids · paper #31 Authors: Yunshen Wang, Shaohang Zhu, Peiyuan Zhi, Yuhan Li, Jiaxin Li, Yong-Lu Li, Yuchen Xiao, Xingxing Wang, Baoxiong Jia, Siyuan Huang arXiv: 2602.23843 · program page

Summary compiled from the arXiv paper (v1); all numbers quoted from the paper. Trend context: RSS 2026 survey.

OmniXtreme extreme whole-body control (Figure 1 of arXiv 2602.23843, © the authors)

Real-world executions of the single unified policy on a Unitree G1: (a) a radar plot showing the curated extreme-motion library occupies far more challenging regimes (angular speed, linear acceleration/velocity, contact switching, airborne time) than Unitree-retargeted LAFAN1; panels (b)–(e) show extreme balance, rapid contact switching, high-speed motions up to 15 rad/s, and diverse whole-body behaviors including flips and breakdancing.

Problem

Humanoid motion-tracking policies hit a "generality barrier": as motion libraries scale in diversity and difficulty, tracking fidelity collapses, especially for high-dynamic behaviors on real hardware. The authors attribute this to two compounding factors — gradient interference in multi-motion RL optimization with limited-capacity MLP policies, and unmodeled actuator nonlinearities (torque-speed envelopes, regenerative power effects) that break sim-to-real executability.

Method

OmniXtreme is a two-stage framework. Stage 1, specialist-to-unified generative pretraining: per-motion RL experts are distilled via DAgger into a single high-capacity flow-matching policy (Beta-sampled flow timesteps), decoupling representation scaling from interference-heavy multi-motion RL. Stage 2, actuation-aware residual RL post-training: a lightweight MLP residual policy corrects the frozen base under motor torque-speed constraints, aggressive domain randomization, and power-safety regularization (overcurrent / regenerative braking penalties). Training uses full LAFAN1 plus a curated XtremeMotion set of about 60 extreme motions from LAFAN1, AMASS, MimicKit, and Reallusion; deployment runs via TensorRT at ~10 ms end-to-end latency (50 Hz) on the Unitree G1's onboard Orin NX.

Results

In simulation on LAFAN1+XtremeMotion, the full model reaches 98.54% success at 30.93 mm MPJPE versus 82.95% / 47.95 mm for from-scratch multi-motion RL and 94.91% / 33.35 mm for specialist-to-unified MLP distillation; on the XtremeMotion subset it holds 95.64% success. On the real Unitree G1, 157 trials across 24 high-dynamic motions yield 91.08% overall success (flips 96.36%, martial arts 93.33%, handspring 88.57%, breakdance 86.36%, acrobatics 80.00%). Ablations show flips need motor constraints, breakdance additionally needs aggressive domain randomization, and acrobatic landings only become stable with power-safety regularization.

Significance

Evidence that the fidelity-scalability trade-off in humanoid tracking is not inherent: a generative (flow-matching) action head scales with capacity where MLP trackers saturate, and actuation-aware refinement is what carries that scaling onto hardware. Related wiki threads: Review-Humanoid-VLA · Review-System-0-1-2.

← Back to RSS 2026 survey · RSS-2026-Papers · Home