ICML 2026 Towards Practical World Model based Reinforcement - Heungwoo/research GitHub Wiki

Towards Practical World Model-based Reinforcement Learning for Vision-Language-Action Models — VLA-MBPO: data-efficient world modeling for safe VLA RL

Venue: ICML 2026 (Poster) Category: World Model Affiliations: Zhilong Zhang, Haoxiang Ren, Yihao Sun, Yifei Sheng, Haonan Wang, Haoxin Lin, Zhichao Wu, Pierre-Luc Bacon, Yang Yu Traction (2026-06): 1 citation (arXiv)

Framework of VLA-MBPO: (A) a UMM-based world model with interleaved view decoding for multi-view observation and reward prediction; (B) a stable, scalable policy update with chunk-level branched rollout; (C) simulated and real-world task designs (Figure 1 from Zhang et al., 2026)

Problem

VLA models generalize well for robotic control, but fine-tuning them with reinforcement learning is constrained by the high cost and safety risks of real-world interaction. Training inside a learned world model sidesteps these issues, yet world-model RL for VLAs introduces three practical challenges: (i) pixel-level world modeling (VLAs consume raw images, demanding high-fidelity generation rather than low-dimensional latent rollouts); (ii) multi-view consistency across head and wrist cameras; and (iii) compounding model errors under sparse rewards, where long imagined rollouts drift and corrupt the learning signal.

Method

VLA-MBPO is a practical world model-based RL framework built from three design choices. The algorithm runs in three phases: (1) data collection with the VLA model; (2) world model fine-tuning on collected data; (3) policy optimization with RL inside the world model.

Frame-skipping scheme in the UMM-based world model, which generates high-fidelity pixel observations while keeping inference tractable (Figure 2 from Zhang et al., 2026)

  • (i) UMM-based world model. A unified multimodal model (UMM-World) is adapted for data-efficient world modeling, using a frame-skipping scheme to produce high-fidelity pixel-based observations plus reward prediction.
  • (ii) Interleaved view decoding (IVD). Enforces multi-view consistency across head and wrist views during generation.
  • (iii) Chunk-level branched rollout. Mitigates error compounding by branching short imagined rollouts of chunk size k; the GAE advantage is computed over n branched rollouts (Eq. 4) rather than full-horizon imagination.

For policy optimization the authors adopt Flow-Noise, a PPO variant for flow-matching policies, appending an MLP value head to the VLA. A theoretical analysis bounds the value gap of world-model RL and shows VLA-MBPO reduces it.

Results

World model quality (LIBERO Object, Table 1): UMM-World beats Ctrl-World on head-view LPIPS (0.094 vs 0.150), PSNR (23.29 vs 21.95) and wrist-view PSNR (18.76 vs 13.87), with lower inference time (10 vs 21) and a reward-model accuracy of 98.4% / F1 0.861 (vs Qwen3-VL-8B at 97.0% / 0.841). Ablating IVD or pretraining (PT) degrades every metric.

LIBERO policy results (Table 2): Starting from a one-trajectory-SFT π0.5 policy, VLA-MBPO reaches 85.9 average SR versus 76.8 for the SFT baseline (Δ +9.1), beating online RL (πRL, 82.6) and offline IDQL (77.5). Per suite: Spatial 87.8, Object 96.6, Goal 92.8, Long 66.8 — with the largest gain on long-horizon Long (+12.2).

Real-world (Figure 5): Five tasks across two platforms — bimanual Arx-X5 (Plug Cable, Fold Towel) and whole-body Galaxy-R1 (Pick Cup, Insert Pen, Wipe Board) — show consistent gains, including contact-rich insertion, deformable folding, and whole-body control under partial observability, on both seen (30) and unseen (20) conditions.

Ablations: A branched-rollout length of 2 chunks is best (66.8 on LIBERO-Long) versus 1 (63.9) or 4 (62.9); full-horizon rollout collapses to 52.8 due to compounding error. Success rate improves monotonically with imagined sample size.

Significance

VLA-MBPO makes world-model-based RL practical for modern flow-matching VLAs by combining a UMM world model, multi-view-consistent decoding, and branched rollouts that cap compounding error — backed by a value-gap theory. It delivers significant gains in both success rate and sample efficiency in simulation and on real bimanual and whole-body robots, pointing toward safer, cheaper VLA fine-tuning that avoids large volumes of risky real-world interaction.

Links

← Back to ICML-2026