ICLR 2026 VLA RFT - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 ยท OpenReview: Jaut99EHeu Category: RL for VLA Trend tag: Trend 3
flowchart LR
Real[Real robot data] --> WM[Train data-driven<br/>controllable world model]
Pol[Policy] --> Sim[Rollout INSIDE world model]
WM --> Sim
Sim --> Vr["Verified reward:<br/>โ(L1 + LPIPS) of rollout frames<br/>vs goal-achieving reference frames"]
Vr --> Grad[GRPO policy gradient]
Grad --> Pol
note[No real rollouts<br/>No physics simulator<br/>No sim-to-real gap]:::n
classDef n fill:#fff8e6,stroke:#e0a800
RL fine-tuning needs rollouts. Real-robot rollouts are expensive and slow. Physics simulator rollouts are cheap but have a sim-to-real gap. Alternative: roll out in a learned world model โ but this needs rewards that are actually meaningful inside a generative model.
Train a data-driven controllable world model as the RL environment: a compact ~138M-parameter LLaMA-style autoregressive transformer (12 layers, 768 hidden, VQGAN-style image tokenizer) that predicts future visual observations conditioned on actions. The base VLA policy is VLA-Adapter (a lightweight VLA with a DiT flow-matching action head). RL fine-tuning happens entirely inside the world model using GRPO (with policy-ratio clipping, auxiliary MSE loss, and entropy regularization).
The verified reward is a dense, trajectory-level reward: the negative weighted sum of per-frame reconstruction loss (L1) and perceptual similarity (LPIPS) between the world-model rollout frames and goal-achieving reference frames from expert demonstrations. The reward is "verified" by anchoring the policy's generative rollout to expert reference trajectories within the world model's space โ it is not a check validated against real-world hardware success. The ablation confirms this design (Type 3, trajectory comparison in generative space) gives +4.5 pts, versus action-level only (+1.1) and pixel-level with real images (+0.5).
On LIBERO (Spatial/Object/Goal/Long), VLA-RFT lifts average success from 86.6% (SFT base) to 91.1% (+4.5 pts) in only 400 RL steps, versus the 150K-step supervised baseline and ~40K steps for competing RL methods. Under perturbations (object/goal shifts ~2.5โ5 cm, robot-state noise), it gains +6.5 pts (minor) and +3.0 pts (major), indicating stronger failure recovery and robustness under distribution shift.
The cleanest "model-based RL ร VLA" recipe in 2026. Foreshadows a broader direction: RL environment, evaluation environment (WorldGym), and the policy itself may all share the same underlying world model, collapsing three previously separate systems.
- OpenReview: https://openreview.net/forum?id=Jaut99EHeu
- arXiv: https://arxiv.org/abs/2510.00406
- Ctrl-World (world model substrate)
- WorldGym (world-model evaluation)
- PLD (residual-RL alternative)