ICML 2026 From Imagined Futures to Executable - Heungwoo/research GitHub Wiki
From Imagined Futures to Executable Actions — A mixture of latent actions bridging video imagination and policy execution
Venue: ICML 2026 (Poster) Category: World Model Affiliations: Fudan University; Imperial College London; University of Surrey (Yajie Li, Bozhou Zhang, Chun Gu, Zipei Ma, Jiahui Zhang, Jiankang Deng, Xiatian Zhu, Li Zhang) Traction (2026-06): 2 citations (arXiv)

Problem
Video generation models can imagine long-horizon future observations for robot manipulation, but turning those imagined futures into reliable actions is hard. Existing approaches either condition the policy on predicted frames or directly decode generated videos into actions. Both suffer a mismatch between visual realism and control relevance: predicted observations emphasize perceptual fidelity rather than the action-centric causes of state transitions, producing indirect and unstable control. MoLA targets this gap with a control-oriented interface that converts imagined video into executable representations.
Method
MoLA (Mixture of Latent Actions) is an imagination-based framework with three components and a three-stage training recipe.
- Video generation (imagination). Given the current RGB observation and task instruction, a Stable Video Diffusion (SVD) backbone synthesizes future frames. At inference MoLA uses a single denoising step for efficiency; the paper finds compact future hypotheses preserve the temporal structure useful for action inference without needing photorealism. Multi-view setups predict each camera stream independently.
- Mixture of inverse dynamics models (MoIDM). Because predicted frames are not inherently action-centric, MoLA passes them through a mixture of modality-aware pretrained inverse dynamics models capturing complementary semantic, depth, and flow cues, inferring a mixture of latent actions that explain the visual transitions.
- Action head. A DiT + flow-matching diffusion head decodes the latent-action mixture into executable robot control commands.
Training: Stage I fine-tunes the video generation model; Stage II pretrains the MoIDMs; Stage III performs end-to-end fine-tuning.

Results
Evaluated on CALVIN, LIBERO, LIBERO-Plus, and a real UR5e robot.
- CALVIN ABC-D (Table 1): MoLA reaches avg. length 4.55 (rollout success 98.5/95.0/91.1/88.1/82.6 over 5 chained tasks), topping DreamVLA (4.44), VPP (4.33), Seer (4.28), and π0.5* (3.97).
- LIBERO (Table 2): 97.0% average (Spatial 93.0, Object 99.5, Goal 99.5, Long 96.0), best on all four suites vs. VPP* 90.9 and CoT-VLA 83.9.
- LIBERO-Plus (Table 3): surpasses the strongest baseline OpenVLA-OFT+ by 13.2 percentage points average.
- Real-world UR5e (Table 4): 73.0% average across in- and out-of-distribution tasks, beating VPP (62.0%), OpenVLA (38.0%), and Diffusion Policy (41.5%); the company-scale π0.5 (77.5%) is included as a strong reference rather than an apples-to-apples comparator.
- Ablations. Adding modalities lifts CALVIN avg. length monotonically: baseline 4.24 → flow+depth 4.46 → all-modalities 4.55 (Table 5). DiT+flow head beats an AR Transformer head (4.55 vs 4.40). One-step video diffusion already gives the best avg. length (4.55) at lowest cost (0.13 s); more steps do not monotonically help (Table 10). A stronger Wan2.2-5B backbone nudges avg. length to 4.59.
Significance
MoLA reframes video-based imagination not as a frame-prediction problem but as a latent-action inference problem, using a structured, physically grounded interface (semantic/depth/flow inverse dynamics) between imagination and control. The finding that one-step, non-photorealistic video suffices challenges the assumption that better imagined pixels yield better control, and the modality decomposition delivers consistent gains across simulation and real hardware.
Links
- arXiv: 2605.12167
- Project: https://logosroboticsgroup.github.io/MoLA
- ICML 2026: https://icml.cc/virtual/2026/poster/66091
← Back to ICML-2026