ICLR 2026 Cosmos Policy - Heungwoo/research GitHub Wiki

Cosmos Policy — Fine-Tuning Video Models for Visuomotor Control and Planning

Venue: ICLR 2026 Authors: Moo Jin Kim, Yihuai Gao, Tsung-Yi Lin, Yen-Chen Lin, Yunhao Ge, Grace Lam, Percy Liang, Shuran Song, Ming-Yu Liu, Chelsea Finn, Jinwei Gu (NVIDIA + Stanford) Base model: Cosmos-Predict2-2B (latent diffusion video foundation model) Category: VLA Architecture — Video Foundation Model Trend tag: Trend 5 (data) / Trend 1 (architecture) arXiv: 2601.16163 (Jan 2026)

Approach diagram

flowchart LR
  V[Video frames<br/>→ latent frames] --> Cosmos[Cosmos-Predict2-2B<br/>latent diffusion transformer<br/>no architectural changes]
  A[Actions, future states, values<br/>encoded AS latent frames] --> Cosmos
  Cosmos --> NF[Predicted future-state frames]
  Cosmos --> Act[Predicted action latent frames]
  Cosmos --> Val[Value latent frames<br/>expected cumulative reward → planning]
Loading

Problem

Video foundation models (NVIDIA Cosmos, etc.) are trained on orders of magnitude more video than any robot dataset and have strong world priors. But they are not directly controllable — they generate video, not actions — and their latents are optimized for generation, not decision-making.

Method

A single stage of post-training on robot demonstrations from the target platform, with no architectural modifications. The key mechanism is latent frame injection: robot actions, future-state images, and values (expected cumulative rewards) are all encoded as latent frames within the same latent-diffusion sequence, so multiple modalities are modeled jointly by the unmodified Cosmos-Predict2 diffusion transformer. (Note: actions are encoded as latent frames, not as separate "control tokens", and the value signal is itself a latent frame, not a bolted-on regression head.) The model can run as a direct policy (only actions decoded at inference) or as a planning policy, where candidate action trajectories are scored by predicting their future states and values for test-time planning.

Results

State-of-the-art among compared policies: LIBERO 98.5%, RoboCasa 67.1% (50 demos), and real-world ALOHA bimanual 93.6% average. On ALOHA it beats fine-tuned VLAs — π0.5 (88.6%) and OpenVLA-OFT+ (62.0%). Baselines also include Diffusion Policy, π0, UniVLA, OpenVLA-OFT, CogVLA (LIBERO) and GR00T, UVA, FLARE, Video Policy (RoboCasa). Generating predicted future-state frames additionally provides interpretability.

Significance

Validates the hypothesis that video foundation models are a useful substrate for robot control, and establishes a concrete recipe. The latent-frame-injection design is notable for unifying action generation, world-model rollout, and value estimation in one diffusion sequence with zero added modules — enabling model-based planning "for free" from the same backbone. Brings industrial-scale video models (not just LLMs) into the VLA conversation.

Links

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️