ICLR 2026 WMPO - Heungwoo/research GitHub Wiki

WMPO — World Model-based Policy Optimization for VLA

Venue: ICLR 2026 Category: RL for VLA / World Model Trend tag: Trend 8 (world models as RL substrate) Affiliations: HKUST, ByteDance Seed

Approach diagram

flowchart LR
  S0[Initial state s0 from real env] --> Pol[VLA policy π_θold<br/>OpenVLA-OFT base, K=8 chunk]
  Pol --> WM[Pixel-space world model<br/>OpenSora + SDXL VAE<br/>+ frame-level AdaLN]
  WM --> Img[Imagined trajectory τ_i]
  Img --> RM[Reward model R_ψ<br/>VideoMAE + linear head<br/>binary success]
  RM --> Adv[Group-relative advantage Â_i<br/>= R_i - mean / std]
  Adv --> GRPO[On-policy GRPO update<br/>no KL term, dynamic sampling]
  GRPO --> Pol2[Updated π_θ]
  Pol2 -.-> Pol
Loading

Problem

VLA models trained from expert demonstrations cannot easily learn from failures or self-correct — they have never seen recovery behavior, so when out-of-distribution states arise (DAgger-style compounding errors) they fail unrecoverably. RL is the natural remedy but real-robot RL is sample-inefficient (millions of interactions, hardware/safety costs). Two prior workarounds:

  1. Human-in-the-loop RL (HIL-SERL, ConRFT) — labor-intensive, hard to scale.
  2. Sim-to-real RL (VLA-RL, SimpleVLA-RL) — limited by simulator fidelity for diverse tasks.

A third path is model-based RL with a learned world model, but classical world-models (Dreamer family) operate in a latent space that mismatches the pixel-space pretraining of VLAs. WMPO's bet: do model-based RL in pixel space so the imagined frames stay in the distribution the VLA was pretrained on, and use on-policy GRPO instead of off-policy alternatives that bias value estimation.

Detailed Method

Architecture

  • Base policy: OpenVLA-OFT fine-tuned via supervised IL on each target task. Action chunk length K = 8, action discretized into 256 bins per dimension. Robot proprioceptive state and wrist camera input are dropped for simplicity.
  • World model: Diffusion video model based on OpenSora, with:
    • 3D VAE replaced by SDXL's 2D VAE (preserves fine motion detail, avoids temporal-compression artifacts).
    • Diffusion runs in VAE latent space; output decoded back to pixels (so the VLA stays in its pretrained image distribution).
    • Noisy-frame conditioning: during training, conditioning frames are perturbed at diffusion timestep ~50 (out of 1000). This stabilizes long-horizon autoregressive generation by training the model to handle imperfect inputs.
    • Frame-level action control: extends AdaLN to inject per-frame action signals plus diffusion timestep embeddings. Each action a_i passes through an MLP that emits scale γ_1^i, shift β_1^i for LayerNorm and residual scale α_1^i for the MHA/FFN block.
  • Reward model: VideoMAE encoder + linear head, trained with binary cross-entropy on success/failure. Positives are the terminal clip c_N of successful trajectories; negatives are non-terminal clips of successes plus arbitrary clips of failures. Class balance per batch. Inference uses a sliding window of stride s=1 and clip length L=8; success if any clip exceeds threshold τ_thr (validation-set tuned). Reported F1 ≥ 0.95 on all four Mimicgen tasks.

Imagined trajectory generation

Given c=4 conditioning frames and a predicted action chunk a_{i:i+K}, the world model generates the next K=8 frames. The process iterates autoregressively to a maximum length N. Each generated trajectory is τ = {I_0:N, a_0:N} with binary reward y = R_ψ(I_0:N) ∈ {0, 1}.

Policy Behavior Alignment

The world model is first pretrained on Open X-Embodiment trajectories (broad dynamics knowledge), then fine-tuned on real rollouts collected by the current base policy itself. This is essential because OXE is success-biased — without policy-behavior alignment, the world model cannot faithfully imagine failures, and without faithful failure imagination GRPO has nothing to optimize against.

On-policy RL (GRPO)

WMPO uses GRPO with the DAPO trick of removing KL regularization (no reference model needed; reduces memory; encourages exploration):

J(θ) = E_{s_0~D, {τ_i}~π_θold} [(1/G) Σ (1/T) Σ min(r_i,t Â_i, clip(r_i,t, 1-ε_low, 1+ε_high) Â_i)]
r_i,t(θ) = π_θ(a_i,t|s_i,t) / π_θold(a_i,t|s_i,t)
Â_i = (R_i - mean({R_j})) / std({R_j})

Dynamic sampling: groups where all G trajectories have the same reward (all success or all failure) are discarded and replaced — guarantees non-vanishing gradients. Group size G = 8.

Hyperparameters

World model (Table 3): AdamW (β=0.9, 0.999), LR 1e-4, batch 128, gradient clip 0.1, EMA 0.9999, weight decay 0, 12M pretraining steps, 3M fine-tuning steps, ε-prediction.

GRPO (Table 4): AdamW, LR 5e-6, training batch 64 trajectories, group size G=8, mini-batch 128, clip ratio ε_low = 0.20, ε_high = 0.28 (asymmetric, allowing larger upside), temperature 1.6.

Hardware

SFT of OpenVLA-OFT on 8× H100; world-model training and policy optimization on 32× H100.

Comprehensive Results

Mimicgen simulation (4 fine-grained tasks, 300 expert trajectories each as base, 128 evaluation initial states)

Rollout budget P Method Coffee D0 StackThree D0 ThreePieceAssembly D0 Square D0 Mean
– Base policy (SFT) 43.8 46.9 19.5 24.2 33.6
128 GRPO (online, real rollouts) 38.3 52.3 17.2 25.0 33.2
128 DPO (offline pairs) 43.8 53.9 23.4 28.1 37.3
128 WMPO 61.7 56.3 37.5 32.8 47.1
1280 GRPO 47.7 54.7 20.3 25.8 37.1
1280 DPO 52.3 57.0 26.7 33.6 42.4
1280 WMPO 75.0 64.1 46.1 45.3 57.6

Two reads: (1) at small budget P=128, WMPO already gives +9.8 absolute mean S.R. over the strongest baseline; (2) the gap widens to +15.2 at P=1280 — WMPO scales while DPO plateaus and direct GRPO underperforms because batch sizes large enough for stable GRPO (≥64) need ≥512 real trajectories per update.

Generalization (each policy tested on its task's disruption variants)

Method Position Disruption Background Disruption Texture Disruption Mean
Base policy 14.1 46.1 10.9 23.7
GRPO 15.6 47.7 10.9 24.7
DPO 16.4 34.4 7.8 19.5
WMPO 22.3 50.0 16.4 29.6

DPO actually hurts generalization (reliance on spurious visual cues), GRPO ≈ base, only WMPO improves across all three disruption types.

Lifelong learning (StackThree)

Iteratively collect P=128 trajectories with current policy, do WMPO, repeat. Compared against re-training the base policy with 300, 428, 556 expert trajectories. WMPO improves stably across iterations whereas DPO fails to improve (training instability).

Real-world ("Insert the square into the stick", 5mm clearance, Cobot Mobile ALOHA, 200 expert demos for SFT)

  • Base policy: 53%
  • DPO: 60%
  • WMPO: 70% (over 30 trials)

The world model successfully imagines failure trajectories where the square misaligns and cannot be inserted (Fig. 9). Failure mode of the world model: subtle sticking after the square contacts the stick is sometimes missed (Fig. 10), but rare on validation.

Behavior analysis

  • Self-correction emerges: when WMPO collides with the stick, it autonomously lifts, realigns, and inserts — a behavior absent from the expert demonstrations. The base policy, never having seen collisions, just keeps pushing until timeout.
  • Faster trajectories: successful-trial trajectory lengths shorter under WMPO across all four Mimicgen tasks (Fig. 5), because WMPO discourages "stuck" timeouts.

Ablation Studies

The paper does not include a separate ablation table; ablations are folded into the comparison study:

  • Online vs offline RL with same budget: GRPO baseline (online, real rollouts) underperforms WMPO substantially even at P=1280 (37.1 vs 57.6) — establishing that the world model is the source of efficiency, not just on-policy RL.
  • Off-policy vs on-policy: DPO (offline) plateaus due to static data reuse; off-policy methods cause "value estimation errors" (cited Park et al. 2025). WMPO uses on-policy GRPO inside imagination to avoid this.
  • Reward-model F1 ≥ 0.95 across all tasks reportedly mitigates reward hacking.
  • Batch size sensitivity for GRPO baseline: authors report trying batch 8 (with proportional 8× LR drop) and batch 64; batch 64 is more stable but feasible only at P ≥ 512.

Limitations (as explicitly stated, Appendix D)

"While the WMPO framework can in principle support flow-based policies, this work focuses on discretized action representations. As future work, we plan to extend WMPO to more expressive policy classes, such as flow-matching based policies (π0), and explore policy optimization with FlowGRPO, thereby broadening its applicability across diverse action spaces."

That is the only stated limitation. Implicit limitations:

  • State assumption: s = (image observations, instruction). POMDPs (with hidden state, partial observability) deferred to future work.
  • Robot-state input dropped for simplicity — likely loses some performance on tasks requiring precise force/joint feedback.
  • World-model fine-tuning cost: P real trajectories + 32× H100 training, repeated per task domain; substantial compute even without real-robot RL.
  • Reward model failure modes: rare cases (e.g. subtle sticking) where the model misclassifies; not formally bounded.

Significance & Positioning

WMPO targets a specific point in the world-model-for-VLA design space:

  • Versus latent world models (the paper's cited contrast is the Dreamer family / RSSM, refs 16–18; it also notes large video world models such as Genie 3): pixel space keeps the VLA's pretrained vision intact. The argument is that VLAs are trained on web-scale images, so any latent space — however well-learned — is a domain shift the VLA cannot exploit.
  • Versus Ctrl-World (controllable world model infrastructure): Ctrl-World focuses on world-model design and controllability; WMPO closes the loop with on-policy RL and shows it actually improves a real VLA.
  • Versus VLA-RFT / SimpleVLA-RL: those use sim or real-robot RL; WMPO uses imagination, with explicit policy-behavior alignment as the mechanism to make imagined failures useful.
  • Versus Genie Envisioner: similar pixel-space philosophy but Envisioner is a generic world-model engine; WMPO is the dedicated RL recipe on top.

The on-policy insistence is the methodological wedge — most prior VLA-RL work used off-policy DPO/PPO; WMPO argues that off-policy methods incur estimation bias precisely when the rollout is cheap enough (in imagination) to do on-policy properly. The asymmetric clip (ε_low=0.20, ε_high=0.28) inherited from DAPO biases toward exploration, consistent with that rationale.

The emergent self-correction is the qualitative headline: a policy trained only on success demos learns recovery from imagined failures, which is the holy grail of imitation→RL bootstrapping. The mechanism is that policy-behavior-aligned rollouts surface failure states, the reward model labels them 0, GRPO advantages flip, and the policy learns to escape them.

Links

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️