ICLR 2026 VLA RFT - Heungwoo/research GitHub Wiki

VLA-RFT โ€” Reinforcement Fine-Tuning with Verified Rewards in World Simulators

Venue: ICLR 2026 ยท OpenReview: Jaut99EHeu Category: RL for VLA Trend tag: Trend 3

Approach diagram

flowchart LR
  Real[Real robot data] --> WM[Train data-driven<br/>controllable world model]
  Pol[Policy] --> Sim[Rollout INSIDE world model]
  WM --> Sim
  Sim --> Vr["Verified reward:<br/>โˆ’(L1 + LPIPS) of rollout frames<br/>vs goal-achieving reference frames"]
  Vr --> Grad[GRPO policy gradient]
  Grad --> Pol
  note[No real rollouts<br/>No physics simulator<br/>No sim-to-real gap]:::n
  classDef n fill:#fff8e6,stroke:#e0a800
Loading

Problem

RL fine-tuning needs rollouts. Real-robot rollouts are expensive and slow. Physics simulator rollouts are cheap but have a sim-to-real gap. Alternative: roll out in a learned world model โ€” but this needs rewards that are actually meaningful inside a generative model.

Method

Train a data-driven controllable world model as the RL environment: a compact ~138M-parameter LLaMA-style autoregressive transformer (12 layers, 768 hidden, VQGAN-style image tokenizer) that predicts future visual observations conditioned on actions. The base VLA policy is VLA-Adapter (a lightweight VLA with a DiT flow-matching action head). RL fine-tuning happens entirely inside the world model using GRPO (with policy-ratio clipping, auxiliary MSE loss, and entropy regularization).

The verified reward is a dense, trajectory-level reward: the negative weighted sum of per-frame reconstruction loss (L1) and perceptual similarity (LPIPS) between the world-model rollout frames and goal-achieving reference frames from expert demonstrations. The reward is "verified" by anchoring the policy's generative rollout to expert reference trajectories within the world model's space โ€” it is not a check validated against real-world hardware success. The ablation confirms this design (Type 3, trajectory comparison in generative space) gives +4.5 pts, versus action-level only (+1.1) and pixel-level with real images (+0.5).

Results

On LIBERO (Spatial/Object/Goal/Long), VLA-RFT lifts average success from 86.6% (SFT base) to 91.1% (+4.5 pts) in only 400 RL steps, versus the 150K-step supervised baseline and ~40K steps for competing RL methods. Under perturbations (object/goal shifts ~2.5โ€“5 cm, robot-state noise), it gains +6.5 pts (minor) and +3.0 pts (major), indicating stronger failure recovery and robustness under distribution shift.

Significance

The cleanest "model-based RL ร— VLA" recipe in 2026. Foreshadows a broader direction: RL environment, evaluation environment (WorldGym), and the policy itself may all share the same underlying world model, collapsing three previously separate systems.

Links

Related pages

โ† Back to ICLR-2026 ยท Topic: RL

โš ๏ธ **GitHub.com Fallback** โš ๏ธ