ICLR 2026 Imitating Generated Videos - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 ยท Authors: Shivansh Patel, Shraddhaa Mohan, Hanlin Mai, Unnat Jain, Svetlana Lazebnik, Yunzhu Li ยท arXiv: 2507.00990 ยท Category: World model for actions ยท Trend tag: generative video as demonstration source.
flowchart LR
Cmd[Language command<br/>+ initial scene image] --> VDM[Video diffusion model<br/>Kling v1.6]
VDM --> Vids[Candidate demo videos]
Vids --> VLM[VLM filter<br/>reject off-command results]
VLM --> Pose[6D pose tracker<br/>FoundationPose]
Pose --> Traj[Object 6D trajectories]
Traj --> Retarget[Embodiment-agnostic<br/>retargeting to robot]
Retarget --> Exec[Robot execution<br/>no robot-specific training]
Imitation learning for manipulation needs costly physical demonstrations and robot-specific training. The authors ask whether AI-generated videos can replace real demonstrations entirely โ letting a robot perform tasks like pouring, lifting, placing, and sweeping with zero physical demos and no policy training.
RIGVid (Robots Imitating Generated Videos) is a training-free pipeline:
- A video diffusion model (Kling v1.6, chosen over Sora and Kling v1.5 for reliability) generates candidate demonstration videos from a language command and one initial scene image.
- A vision-language model automatically filters out videos that do not follow the command.
- A 6D pose tracker (FoundationPose, no instance-specific fine-tuning, real-time) extracts object trajectories from the chosen video.
- Trajectories are retargeted to the robot in an embodiment-agnostic fashion, so the same generated video drives execution without robot data.
- Overall success 85.0%, vs. Gen2Act 67.5%, AVDC 32.5%, 4D-DPM 35.0%, Track2Act 7.5%.
- Beats the keypoint-constraint baseline ReKep (50%) at 85% across four tasks.
- Higher video-generation quality lifts per-task success (e.g., pouring 80%โ100%, placing 50%โ90%, sweeping 20%โ70%).
- Evaluated on four real-world tasks: pouring water, lifting a lid, placing a spatula on a pan, sweeping trash.
Demonstrates that filtered generated videos can be as effective as real demonstrations, decoupling manipulation skill acquisition from physical data collection. Positions video generative models as a scalable, embodiment-agnostic demonstration source โ a complementary take on the "world/video model as substrate for actions" theme.
- arXiv: https://arxiv.org/abs/2507.00990
- Project page: https://rigvid-robot.github.io/
- OpenReview: https://openreview.net/forum?id=zjjVQDUgZr
- ICLR 2026 Survey
- VLA Architectures review (world-model category)
โ Back to ICLR-2026