RSS 2026 Toward Reliable Sim to Real Predictability for - Heungwoo/research GitHub Wiki

Toward Reliable Sim-to-Real Predictability for MoE-based Robust Quadrupedal Locomotion

Venue: RSS 2026 (Sydney, Jul 13–17) · Session: RL · paper #156 Authors: Tianyang Wu, Hanwei Guo, Yuhang Wang, Junshu Yang, Xinyang Sui, Jiayi Xie, Xingyu Chen, Zeyang Liu, Xuguang Lan arXiv: 2602.00678 · program page

Summary compiled from the arXiv paper (v4); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Closed-loop MoE training + RoboGauge assessment framework (Figure 1 of arXiv 2602.00678, © the authors)

Figure 1: left, the MoE representation encoder (gating network over expert latents feeding the actor) trained on parallel multi-terrain simulation; center, the closed loop — train in IsaacGym, evaluate with the RoboGauge sim-to-sim suite (10 difficulty levels x 7 tasks, terrain and metric grids), adjust, and select the policy to deploy; right, real-world Unitree Go2 tests on curbs, sand, snow, and stone stairs.

Problem

Proprioception-only RL locomotion policies that score high training rewards in simulation often fail to transfer — reward overfitting and the sim-to-real gap make training-time terrain levels a poor predictor of hardware performance, while validating every checkpoint physically is risky and slow. The paper wants both a more robust policy representation and a trustworthy simulated predictor of real-world transferability.

Method

Two coupled components. (1) An MoE student encoder inside an asymmetric teacher–student framework: K expert subnetworks with a gating network over the observation history decompose latent terrain and command modeling, with an auxiliary load-balancing loss (ablations show applying MoE to the actor-critic instead, as in MoE-Loco/MCP, destabilizes training). (2) RoboGauge, a MuJoCo-based sim-to-sim assessment suite: 7 terrain configurations (flat, wave, slopes and stairs up/down, obstacle) x 10 difficulty levels x 9 domain randomizations, scored on 8 normalized metrics (velocity tracking, torque smoothness, orientation stability, ZMP/contact-wrench safety margins, etc.) aggregated by a weighted geometric mean and a difficulty-overlapping scoring function; a level is passable at ≥80% goal-reaching success. Training uses 8,192 parallel IsaacGym agents with a 7-terrain curriculum; a command design suite raises the peak RoboGauge score by 11%.

Results

RoboGauge's metric errors against motion-capture ground truth are markedly lower than IsaacGym training-environment estimates (average tracking error 0.0558 vs 0.0883), validating it as a transfer predictor. Under RoboGauge, the MoE policy scores 0.6713 (level 7.85) vs CTS 0.5786, HIM 0.5379, DreamWaQ 0.5054. On the real Unitree Go2: 18/20 survival under 80–100 N lateral impulses (baselines ≤11/20), 85/85 tile stairs (15.5 cm, µ=0.38), and 17/20 on 30 cm obstacles where every baseline scores 0/20; stairs at 1.31 m/s average (0.15 m/s tracking error), 30° slope at 1.53 m/s, 60 cm drop recovery, and a 4.01 m/s peak sprint within 2.16 s exhibiting an emergent narrow-base gait. Outdoor field tests (snow, sand, ice, slopes) completed with 100% success.

Significance

Elevates policy selection to a first-class problem: a calibrated sim-to-sim gauge replaces risky physical trials, and paired with the MoE encoder it closes the loop between predicted and actual transfer — a methodology transferable to humanoids and relevant to evaluation-methodology discussions in RL and Review-VLA-Evaluation.

← Back to RSS 2026 survey · RSS-2026-Papers · Home