ICLR 2026 WorldGym - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 Category: Benchmark Trend tag: Trend 8 (evaluation)
flowchart LR
Pol[Policy under test] --> Act[Action]
Act --> WM[Action-conditioned<br/>video world model]
WM --> NF[Generated next frame]
NF --> Pol
NF --> Judge[VLM-derived reward / progress]
Judge --> Score[Final policy score]
Physics simulators are the standard substrate for large-scale policy evaluation, but building them is engineering-heavy and they struggle with deformable, cluttered, long-horizon scenes. Alternatively, evaluate inside a learned world model — but closed-loop policy rollouts inside a generative model are non-trivial as a simulation substitute.
Use an autoregressive, action-conditioned video world model as the evaluation environment. The world model is a latent Diffusion Transformer (~609M params) trained with Diffusion Forcing to enable autoregressive next-frame generation, with actions injected via AdaLN-Zero modulation; it is trained on 9 robot datasets from Open-X Embodiment (Bridge V2, RT-1, etc.) and normalizes action spaces across robots. Policy emits actions → world model generates next frame → loop repeats (Monte Carlo rollouts). Rewards/success judgments come from a VLM (GPT-4o) that inspects the generated sequence against the language instruction (reported ~81% true-positive, 97% true-negative on validation video). No physics simulator required.
Policy success rates inside WorldGym closely track real-robot success rates: Pearson r = 0.78 across tasks, with mean success rates differing by only ~3.3% (e.g., RT-1-X, Octo, OpenVLA-7B). WorldGym preserves relative policy rankings across policy versions, sizes, and training checkpoints. Validation is against real-robot trials (e.g., 17 Bridge eval tasks from OpenVLA); the paper does not benchmark against physics simulators.
Fully generative evaluation pipeline: world model + VLM reward. If validity holds at scale, this is the substrate on which next-generation VLA RL fine-tuning can happen — no simulator engineering required.
- ICLR 2026 listing
- arXiv:2506.00613
- OpenReview
- Project page
← Back to ICLR-2026