IROS 2026 DreamMimic - Heungwoo/research GitHub Wiki
IROS 2026 — DreamMimic: Visuomotor Whole-Body Loco-Manipulation via World Model
Venue: IROS 2026 (Pittsburgh) · paper #279 · Shanghai Jiao Tong University · Tsinghua University (Yin, Lai) · code to be released. Paper: arXiv 2608.22278. Representative of: humanoid whole-body loco-manipulation × world models — a WM used as distillation supervision + representation, not as an inference-time planner. Companions: Humanoid VLA · World Models · VLA Hybrid Architectures · IROS 2026 survey.

1. Problem
Vision-based whole-body loco-manipulation on humanoids is hard: partial observability, contact-rich dynamics, and learning long-horizon behaviors from high-dimensional visual input. The goal: a vision-based humanoid controller (no privileged state at deployment) that stays stable over long horizons.
2. Method
DreamMimic distills privileged teacher policies into vision-based student controllers via world-model-assisted distillation:
- RSSM repurposed — instead of a Dreamer-style RSSM for planning, it learns predictive latent dynamics that serve as (a) a representation space and (b) a multi-step supervision signal, and exposes compact predictive features to the student to reduce long-term drift.
- Interaction-Aware Prediction Heads (IAPH) — predict contact and object state to ground the latent in agent–object dynamics.
- Reward Prediction Head (RPH) — predicts instantaneous task reward to anchor the representation to behaviorally relevant signals.
- Performance-Conditioned Guidance (PCG) — a reward-driven adaptive distillation schedule scoring teacher and student to balance guidance vs exploration (prevents premature teacher annealing and over-interference).
3. Results
- On OMOMO and BEHAVE datasets: improved manipulation, robust vision-based control without privileged info, and consistent cross-simulator transfer across humanoid platforms.
4. Why it matters (WAM lens)
DreamMimic is a clean IROS 2026 datapoint for the "WAM as training-time scaffold, not inference policy" trend (survey §5.1): the world model is a distillation teacher's representation + multi-step supervisor, never run to plan at deployment (the student is a plain vision policy). This is the humanoid-control cousin of AtomVLA (WM as offline critic) — both use the WM to shape training, keeping the deployed policy reactive. It also advances Humanoid VLA's whole-body thread with a concrete recipe for stabilizing visual policy distillation under contact-rich dynamics.
Limitations (reviewer): teacher–student distillation needs a privileged teacher (sim); evaluated on mocap-derived OMOMO/BEHAVE + sim transfer (not extensive real-humanoid hardware in the abstract); RSSM latent fidelity bounds the supervision quality.
5. Links
- Official program: IROS 2026 (paper #279) · survey: IROS 2026
- Related: Humanoid VLA · World Models · AtomVLA · 3D FlowMatch Actor
← Back to IROS 2026 survey · Home