ICML 2026 VLAW - Heungwoo/research GitHub Wiki
VLAW: Iterative Co-Improvement of VLA Policy and World Model — bootstrapping a real-world simulator from policy rollouts
Venue: ICML 2026 (Poster) Category: World Model Affiliations: Yanjiang Guo, Tony Lee, Lucy Xiaoyang Shi, Jianyu Chen, Percy Liang, Chelsea Finn (Stanford / Tsinghua) Traction (2026-06): 18 citations (arXiv)

Problem
Improving VLA models via online interaction is bottlenecked by the cost of real-world policy rollouts. A natural fix is to use a learned simulator — an action-conditioned video generation model — to manufacture extra rollout data. But existing world models lack the physical fidelity needed for policy improvement: trained predominantly on demonstration datasets that under-cover diverse physical interactions (especially failure cases), they fail to model the small but critical contact dynamics of object manipulation.
Method
VLAW is a simple iterative co-improvement loop between a VLA policy and an action-conditioned video world model. Real-world rollout data improves the world model's fidelity, and the improved world model then generates synthetic rollouts that improve the policy:
- World Model Learning with Real Roll-outs: the policy is rolled out
Ktimes in the real world to form a datasetD_real, with each trajectory assigned a sparse success/failure rewardr ∈ {0,1}at robot reset. BecauseD_realcontains both successes and failures, it counters two known pathologies — over-optimism (training data dominated by successful demos) and limited physical fidelity on contact-rich/deformable dynamics. - World model initialization: VLAW starts from a pretrained Ctrl-World diffusion-based world model (trained on the full DROID dataset) and fine-tunes it on
D_realwith the standard diffusion denoising objective. - Iterative policy improvement: the grounded world model generates large-scale synthetic rollouts used to improve the VLA policy; the paper relates this loop to regularized reinforcement learning (Appendix A).

Results
"39.2% absolute success rate improvement over the base policy". On the DROID platform (Franka Panda + Robotiq gripper, two third-person + one wrist camera) across five contact-rich task categories (stacking, opening a book, and others), VLAW improves a state-of-the-art VLA model by 39.2% absolute success rate over the base policy, with 11.6% of that coming specifically from training on the world-model-generated synthetic rollouts. World-model quality metrics (PSNR/SSIM/LPIPS/FID/FVD plus an event confusion matrix) confirm that adding expert and online rollouts sharply improves fidelity over the pretrained Ctrl-World baseline (e.g., FVD dropping from 225.13 toward ~100 after expert-rollout fine-tuning).
Significance
VLAW shows that a video world model can be turned into a useful, physically grounded data engine for VLA post-training using only a small amount of real interaction — and that including failure rollouts is key to fidelity. This offers a scalable, sample-efficient alternative to pure real-world RL for VLA improvement on contact-rich tasks.
Links
- arXiv: 2602.12063
- ICML 2026: https://icml.cc/virtual/2026/poster/66169
← Back to ICML-2026