ICLR 2026 Ctrl World - Heungwoo/research GitHub Wiki

Ctrl-World โ€” A Controllable Generative World Model for Robot Manipulation

Venue: ICLR 2026 Category: World Model / Policy Evaluation & Improvement Authors: Yanjiang Guo*, Lucy Xiaoyang Shi*, Jianyu Chen, Chelsea Finn (Stanford University, Tsinghua University) arXiv: 2510.10125 Trend tag: Trend 8

Approach diagram

flowchart LR
  Img[7 history frames<br/>+ arm poses] --> WM[Ctrl-World<br/>SVD-1.5B diffusion backbone]
  Act[Frame-level action chunks] --> WM
  WM --> NF[Multi-view future frames<br/>3 cameras, 192x320]
  NF --> Loop[Pose-conditioned memory<br/>retrieval: 20s+ consistency]
  Loop --> WM
  WM --> Use{Used for}
  Use --> Eval[Policy evaluation<br/>rank pi0 / pi0-FAST / pi0.5<br/>without real rollouts]
  Use --> Imag[Policy improvement via SFT<br/>on imagined successful rollouts]
Loading

Problem

World models for robot manipulation must satisfy partly conflicting requirements: high visual fidelity, precise action-conditioning, long-horizon consistency, physical plausibility. Existing video generative models handle the first but struggle with the rest.

Method

A generative world model explicitly designed to be controllable, simulating multi-view long-horizon robot interactions:

  • Initialized from a pretrained 1.5B Stable-Video-Diffusion (SVD) latent video-diffusion backbone (spatial-temporal transformers).
  • Frame-level action conditioning: future action chunks (โ‰ˆ15 actions / 1 s, applied over 5 downsampled steps) are tightly coupled to generation for precise, centimeter-level control.
  • Pose-conditioned memory retrieval for long-horizon consistency: k history frames (7 frames, 1โ€“2 s intervals) plus the corresponding robot-arm poses are injected via frame-wise cross-attention in the spatial transformer, anchoring future predictions to relevant past frames.
  • Generates 3 synchronized camera views (1 wrist + 2 third-person), 192ร—320 each.
  • Trained on the DROID dataset (95k trajectories, 564 scenes); produces consistent rollouts for 20+ seconds under novel scenes and new camera placements.

Used for policy evaluation and policy improvement via supervised fine-tuning on imagined rollouts โ€” not as an RL training environment.

Results

  • Policy evaluation: instruction-following behavior in imagination closely tracks real-world rollouts; reliably ranks ฯ€0, ฯ€0-FAST, and ฯ€0.5 across seven task families (pick-place, towel-folding, drawer, wipe-table, close-laptop, pull-tissue, stack) without physical robot rollouts. (Paper reports qualitative/side-by-side agreement; it does not state an explicit correlation coefficient.)
  • Policy improvement: collecting successful synthetic rollouts (scored by human preference) and fine-tuning on them improves ฯ€0.5-DROID instruction-following by +44.7% on average (38.7% โ†’ 83.4%) on unseen objects and novel instructions.

Limitations

Acknowledged gaps in low-level execution: imprecise modeling of complex contact physics (collisions, sliding, rotations), tendency to underestimate execution success, weaker performance on precise/long-horizon tasks, sensitivity to initial observation, and incomplete capture of real-world retry behavior. Reliance on human-preference scoring (no automated reward model). Horizon beyond ~20 s remains challenging.

Significance

An infrastructure piece for world-model-based policy evaluation and improvement. Together with WorldGym and VLA-RFT, it advances the "world model as universal substrate" agenda for 2026 VLA research โ€” though Ctrl-World itself contributes the evaluation + imagination-SFT path rather than online RL.

Links

Related pages

โ† Back to ICLR-2026 ยท Topic: RL

โš ๏ธ **GitHub.com Fallback** โš ๏ธ