IROS 2026 DreamMimic - Heungwoo/research GitHub Wiki

IROS 2026 — DreamMimic: Visuomotor Whole-Body Loco-Manipulation via World Model

Venue: IROS 2026 (Pittsburgh) · paper #279 · Shanghai Jiao Tong University · Tsinghua University (Yin, Lai) · code to be released. Paper: arXiv 2608.22278. Representative of: humanoid whole-body loco-manipulation × world models — a WM used as distillation supervision + representation, not as an inference-time planner. Companions: Humanoid VLA · World Models · VLA Hybrid Architectures · IROS 2026 survey.

DreamMimic — privileged specialist teachers (trained in simulation with privileged obs) distill into a visual generalist student via world-model-assisted distillation: the world model supplies latent knowledge (proprio+visual prediction) as a representation + multi-step supervision, and Performance-Conditioned Guidance balances teacher guidance vs student exploration (pipeline figure from Yin & Lai, arXiv 2608.22278, © the authors)

1. Problem

Vision-based whole-body loco-manipulation on humanoids is hard: partial observability, contact-rich dynamics, and learning long-horizon behaviors from high-dimensional visual input. The goal: a vision-based humanoid controller (no privileged state at deployment) that stays stable over long horizons.

2. Method

DreamMimic distills privileged teacher policies into vision-based student controllers via world-model-assisted distillation:

  • RSSM repurposed — instead of a Dreamer-style RSSM for planning, it learns predictive latent dynamics that serve as (a) a representation space and (b) a multi-step supervision signal, and exposes compact predictive features to the student to reduce long-term drift.
  • Interaction-Aware Prediction Heads (IAPH) — predict contact and object state to ground the latent in agent–object dynamics.
  • Reward Prediction Head (RPH) — predicts instantaneous task reward to anchor the representation to behaviorally relevant signals.
  • Performance-Conditioned Guidance (PCG) — a reward-driven adaptive distillation schedule scoring teacher and student to balance guidance vs exploration (prevents premature teacher annealing and over-interference).

3. Results

  • On OMOMO and BEHAVE datasets: improved manipulation, robust vision-based control without privileged info, and consistent cross-simulator transfer across humanoid platforms.

4. Why it matters (WAM lens)

DreamMimic is a clean IROS 2026 datapoint for the "WAM as training-time scaffold, not inference policy" trend (survey §5.1): the world model is a distillation teacher's representation + multi-step supervisor, never run to plan at deployment (the student is a plain vision policy). This is the humanoid-control cousin of AtomVLA (WM as offline critic) — both use the WM to shape training, keeping the deployed policy reactive. It also advances Humanoid VLA's whole-body thread with a concrete recipe for stabilizing visual policy distillation under contact-rich dynamics.

Limitations (reviewer): teacher–student distillation needs a privileged teacher (sim); evaluated on mocap-derived OMOMO/BEHAVE + sim transfer (not extensive real-humanoid hardware in the abstract); RSSM latent fidelity bounds the supervision quality.

5. Links

← Back to IROS 2026 survey · Home