RSS 2026 Now You See That - Heungwoo/research GitHub Wiki

Now You See That: Learning End-to-End Humanoid Locomotion from Raw Pixels

Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Humanoids · paper #27 Authors: Wandong Sun, Yongbo Su, Leoric Huang, Alex Zhang, Dwyane Wei, Mu San, Daniel Tian, Ellie Cao, Baoshi Cao, Yang Liu, Finn Yan, Ethan Xie, Zongwu Xie arXiv: 2602.06382 · program page

Summary compiled from the arXiv paper (v2); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Terrain traversal montage (Figure 1 of arXiv 2602.06382, © the authors)

Figure 1 shows the single unified depth-vision policy deployed on a full-sized humanoid across nine real terrains: high stone, debris field, high-low gap, trolley, grid holes, long staircase down and up, high platform, and a platform-slope-gap combination — all behaviors emerging from one policy trained on raw depth images.

Problem

Vision-based humanoid locomotion faces two coupled obstacles: the sim-to-real gap in depth sensing injects perception noise that ruins fine-grained tasks needing centimeter-level foot placement, and training one policy across heterogeneous terrains (extreme parkour obstacles vs precise stair traversal) suffers from conflicting learning objectives. Prior humanoid work relies on drift-prone LiDAR elevation maps or trains separate specialist policies.

Method

A two-stage end-to-end framework. Stage 1 trains a privileged teacher from 21×33 = 693-point height scans (1.6 m × 1.0 m window, 0.05 m resolution) using terrain-specific reward shaping with K = 3 critics and K = 3 AMP-style discriminators (stairs/platforms, gap crossing, rough terrain), selected by terrain label. Stage 2 distills into a deployment policy on raw depth via DAgger-style behavior cloning plus a denoising consistency loss between clean and augmented depth features and a KL feature regularizer (λ = 0.1 each). Key enabler is a realistic depth-sensor simulation: an 8-operator augmentation pipeline (stereo-fusion consistency check reproducing hole artifacts, quadratic depth-dependent Gaussian noise, 5-octave Perlin structured noise, random-convolution optical distortion, calibration/intrinsics/extrinsics randomization, pixel failures, clipping to 0.3–2.0 m, cropping 30×40→24×32, 2–4-frame delay). The final policy runs onboard at 50 Hz.

Results

On RDT-Bench — a CycleGAN-based real-noise evaluation benchmark the authors construct — the method reaches 98.9% average success with 5.8% power-degradation ratio, vs 71.0% / 30.9% for Humanoid Parkour Learning and 43.0% / 70.9% with no augmentation; single critic/discriminator drops to 82.0%, BC-only to 86.0%. Ablations show stereo-fusion simulation matters most (removing it: 90.4% SR). Real-world deployment on a full-sized humanoid with an Orbbec Gemini 336L camera scores 97.8% overall (88/90 trials), perfect on five of six scenarios including 30+-step extended staircases and 45 cm gaps, with stair descending at 86.7% the only weak spot; the pipeline also transfers to a Unitree G1 with RealSense D435i.

Significance

Demonstrates that faithful sensor-noise modeling — not just more domain randomization — is the linchpin of pixel-to-action humanoid parkour, and that multi-critic/multi-discriminator RL lets one policy master conflicting terrain objectives. A useful locomotion-side counterpart to the manipulation-focused sim-to-real threads in the RSS 2026 survey.

← Back to RSS 2026 survey · RSS-2026-Papers · Home