ICLR 2026 Ctrl World - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 Category: World Model / Policy Evaluation & Improvement Authors: Yanjiang Guo*, Lucy Xiaoyang Shi*, Jianyu Chen, Chelsea Finn (Stanford University, Tsinghua University) arXiv: 2510.10125 Trend tag: Trend 8
flowchart LR
Img[7 history frames<br/>+ arm poses] --> WM[Ctrl-World<br/>SVD-1.5B diffusion backbone]
Act[Frame-level action chunks] --> WM
WM --> NF[Multi-view future frames<br/>3 cameras, 192x320]
NF --> Loop[Pose-conditioned memory<br/>retrieval: 20s+ consistency]
Loop --> WM
WM --> Use{Used for}
Use --> Eval[Policy evaluation<br/>rank pi0 / pi0-FAST / pi0.5<br/>without real rollouts]
Use --> Imag[Policy improvement via SFT<br/>on imagined successful rollouts]
World models for robot manipulation must satisfy partly conflicting requirements: high visual fidelity, precise action-conditioning, long-horizon consistency, physical plausibility. Existing video generative models handle the first but struggle with the rest.
A generative world model explicitly designed to be controllable, simulating multi-view long-horizon robot interactions:
- Initialized from a pretrained 1.5B Stable-Video-Diffusion (SVD) latent video-diffusion backbone (spatial-temporal transformers).
- Frame-level action conditioning: future action chunks (โ15 actions / 1 s, applied over 5 downsampled steps) are tightly coupled to generation for precise, centimeter-level control.
- Pose-conditioned memory retrieval for long-horizon consistency: k history frames (7 frames, 1โ2 s intervals) plus the corresponding robot-arm poses are injected via frame-wise cross-attention in the spatial transformer, anchoring future predictions to relevant past frames.
- Generates 3 synchronized camera views (1 wrist + 2 third-person), 192ร320 each.
- Trained on the DROID dataset (95k trajectories, 564 scenes); produces consistent rollouts for 20+ seconds under novel scenes and new camera placements.
Used for policy evaluation and policy improvement via supervised fine-tuning on imagined rollouts โ not as an RL training environment.
- Policy evaluation: instruction-following behavior in imagination closely tracks real-world rollouts; reliably ranks ฯ0, ฯ0-FAST, and ฯ0.5 across seven task families (pick-place, towel-folding, drawer, wipe-table, close-laptop, pull-tissue, stack) without physical robot rollouts. (Paper reports qualitative/side-by-side agreement; it does not state an explicit correlation coefficient.)
- Policy improvement: collecting successful synthetic rollouts (scored by human preference) and fine-tuning on them improves ฯ0.5-DROID instruction-following by +44.7% on average (38.7% โ 83.4%) on unseen objects and novel instructions.
Acknowledged gaps in low-level execution: imprecise modeling of complex contact physics (collisions, sliding, rotations), tendency to underestimate execution success, weaker performance on precise/long-horizon tasks, sensitivity to initial observation, and incomplete capture of real-world retry behavior. Reliance on human-preference scoring (no automated reward model). Horizon beyond ~20 s remains challenging.
An infrastructure piece for world-model-based policy evaluation and improvement. Together with WorldGym and VLA-RFT, it advances the "world model as universal substrate" agenda for 2026 VLA research โ though Ctrl-World itself contributes the evaluation + imagination-SFT path rather than online RL.
- arXiv: https://arxiv.org/abs/2510.10125
- OpenReview: https://openreview.net/forum?id=748bHL2BAv
- VLA-RFT (RL consumer)
- WorldGym (evaluation consumer)
- Cosmos Policy (related video-FM angle)