ICLR 2026 WorldGym - Heungwoo/research GitHub Wiki

WorldGym — World Model as an Environment for Policy Evaluation

Venue: ICLR 2026 Category: Benchmark Trend tag: Trend 8 (evaluation)

Approach diagram

flowchart LR
  Pol[Policy under test] --> Act[Action]
  Act --> WM[Action-conditioned<br/>video world model]
  WM --> NF[Generated next frame]
  NF --> Pol
  NF --> Judge[VLM-derived reward / progress]
  Judge --> Score[Final policy score]
Loading

Problem

Physics simulators are the standard substrate for large-scale policy evaluation, but building them is engineering-heavy and they struggle with deformable, cluttered, long-horizon scenes. Alternatively, evaluate inside a learned world model — but closed-loop policy rollouts inside a generative model are non-trivial as a simulation substitute.

Method

Use an autoregressive, action-conditioned video world model as the evaluation environment. The world model is a latent Diffusion Transformer (~609M params) trained with Diffusion Forcing to enable autoregressive next-frame generation, with actions injected via AdaLN-Zero modulation; it is trained on 9 robot datasets from Open-X Embodiment (Bridge V2, RT-1, etc.) and normalizes action spaces across robots. Policy emits actions → world model generates next frame → loop repeats (Monte Carlo rollouts). Rewards/success judgments come from a VLM (GPT-4o) that inspects the generated sequence against the language instruction (reported ~81% true-positive, 97% true-negative on validation video). No physics simulator required.

Results

Policy success rates inside WorldGym closely track real-robot success rates: Pearson r = 0.78 across tasks, with mean success rates differing by only ~3.3% (e.g., RT-1-X, Octo, OpenVLA-7B). WorldGym preserves relative policy rankings across policy versions, sizes, and training checkpoints. Validation is against real-robot trials (e.g., 17 Bridge eval tasks from OpenVLA); the paper does not benchmark against physics simulators.

Significance

Fully generative evaluation pipeline: world model + VLM reward. If validity holds at scale, this is the substrate on which next-generation VLA RL fine-tuning can happen — no simulator engineering required.

Links

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️