ICLR 2026 Cosmos Policy - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 Authors: Moo Jin Kim, Yihuai Gao, Tsung-Yi Lin, Yen-Chen Lin, Yunhao Ge, Grace Lam, Percy Liang, Shuran Song, Ming-Yu Liu, Chelsea Finn, Jinwei Gu (NVIDIA + Stanford) Base model: Cosmos-Predict2-2B (latent diffusion video foundation model) Category: VLA Architecture — Video Foundation Model Trend tag: Trend 5 (data) / Trend 1 (architecture) arXiv: 2601.16163 (Jan 2026)
flowchart LR
V[Video frames<br/>→ latent frames] --> Cosmos[Cosmos-Predict2-2B<br/>latent diffusion transformer<br/>no architectural changes]
A[Actions, future states, values<br/>encoded AS latent frames] --> Cosmos
Cosmos --> NF[Predicted future-state frames]
Cosmos --> Act[Predicted action latent frames]
Cosmos --> Val[Value latent frames<br/>expected cumulative reward → planning]
Video foundation models (NVIDIA Cosmos, etc.) are trained on orders of magnitude more video than any robot dataset and have strong world priors. But they are not directly controllable — they generate video, not actions — and their latents are optimized for generation, not decision-making.
A single stage of post-training on robot demonstrations from the target platform, with no architectural modifications. The key mechanism is latent frame injection: robot actions, future-state images, and values (expected cumulative rewards) are all encoded as latent frames within the same latent-diffusion sequence, so multiple modalities are modeled jointly by the unmodified Cosmos-Predict2 diffusion transformer. (Note: actions are encoded as latent frames, not as separate "control tokens", and the value signal is itself a latent frame, not a bolted-on regression head.) The model can run as a direct policy (only actions decoded at inference) or as a planning policy, where candidate action trajectories are scored by predicting their future states and values for test-time planning.
State-of-the-art among compared policies: LIBERO 98.5%, RoboCasa 67.1% (50 demos), and real-world ALOHA bimanual 93.6% average. On ALOHA it beats fine-tuned VLAs — π0.5 (88.6%) and OpenVLA-OFT+ (62.0%). Baselines also include Diffusion Policy, π0, UniVLA, OpenVLA-OFT, CogVLA (LIBERO) and GR00T, UVA, FLARE, Video Policy (RoboCasa). Generating predicted future-state frames additionally provides interpretability.
Validates the hypothesis that video foundation models are a useful substrate for robot control, and establishes a concrete recipe. The latent-frame-injection design is notable for unifying action generation, world-model rollout, and value estimation in one diffusion sequence with zero added modules — enabling model-based planning "for free" from the same backbone. Brings industrial-scale video models (not just LLMs) into the VLA conversation.
- Unified Diffusion VLA (also predicts future frames)
- Ctrl-World (world model for RL)
- WorldGym (world model for evaluation)
← Back to ICLR-2026