RSS 2026 mimic video - Heungwoo/research GitHub Wiki
mimic-video — Video-Action Models for Generalizable Robot Control Beyond VLAs
Venue: RSS 2026 (Imitation Learning session) · Authors: Jonas Pai*, Liam Achenbach*, … Oier Mees, Elvis Nava — mimic robotics × ETH Zürich × Microsoft Zürich × UC Berkeley · arXiv: 2512.15692 · project Category: Video-action model (VAM) — alternative backbone class to VLAs Trend tag: RSS 2026 thread 3 — video/world models vs the VLA backbone
Compiled from the verified RSS 2026 abstract and the paper's Fig. 1; the video backbone is NVIDIA Cosmos-Predict2 (per the project page).
Key figure

Figure 1 of the mimic-video paper. Top: the standard VLA pipeline — a VLM pretrained on static image-text pairs supplies semantics only, so large-scale robotics data must teach both dynamics and control in expensive post-training. Bottom: the Video-Action Model — a video model pretrained on video-text pairs already carries semantics plus visual dynamics, so small-scale robot data only has to teach control. Right: the payoff chart — success rate vs robot-data quantity, where the VAM at 10% of the data matches the VLA at 100% (the "10× sample efficiency" claim).
Problem
VLA backbones are pretrained on static, disconnected web data — semantically rich but "blind to physical causality." The policy must therefore infer dynamics and temporal dependencies entirely from robot trajectories, creating an unsustainable expert-data burden.
Method
- Video-Action Model (VAM): pair a pretrained Internet-scale video model (which captures semantics and visual dynamics jointly) with a flow-matching action decoder conditioned on its latent representations.
- The decoder functions as an inverse dynamics model (IDM): the video model produces latent video-space action plans; the IDM translates them into low-level robot actions.
- Division of labor: pre-training handles physics + semantics; robot data only has to teach low-level control.
Results (as reported)
- State-of-the-art on simulated and real-world manipulation benchmarks.
- 10× sample efficiency and 2× convergence speed vs traditional VLA architectures.
Significance
The cleanest statement of the VAM thesis at RSS 2026 — that the right pre-training substrate for control is video, not vision-language. It converges with LDA-1B (dynamics in latent space), Cosmos-Policy (fine-tune a video model into a policy), and the WAM line ([Review-WAM-vs-VLA-Robustness]]), and pressures the [VLM4VLA question from the other side: if the vision encoder is the VLA bottleneck, a video-dynamics encoder may be the fix. The plan-in-video-latents + IDM decomposition is also the policy-side mirror of Qwen-RobotWorld's generation-side stack.
← RSS 2026 survey · Home