ICLR 2026 Imitating Generated Videos - Heungwoo/research GitHub Wiki

RIGVid โ€” Robotic manipulation by imitating AI-generated videos, zero physical demos

Venue: ICLR 2026 ยท Authors: Shivansh Patel, Shraddhaa Mohan, Hanlin Mai, Unnat Jain, Svetlana Lazebnik, Yunzhu Li ยท arXiv: 2507.00990 ยท Category: World model for actions ยท Trend tag: generative video as demonstration source.

Approach diagram

flowchart LR
  Cmd[Language command<br/>+ initial scene image] --> VDM[Video diffusion model<br/>Kling v1.6]
  VDM --> Vids[Candidate demo videos]
  Vids --> VLM[VLM filter<br/>reject off-command results]
  VLM --> Pose[6D pose tracker<br/>FoundationPose]
  Pose --> Traj[Object 6D trajectories]
  Traj --> Retarget[Embodiment-agnostic<br/>retargeting to robot]
  Retarget --> Exec[Robot execution<br/>no robot-specific training]
Loading

Problem

Imitation learning for manipulation needs costly physical demonstrations and robot-specific training. The authors ask whether AI-generated videos can replace real demonstrations entirely โ€” letting a robot perform tasks like pouring, lifting, placing, and sweeping with zero physical demos and no policy training.

Method

RIGVid (Robots Imitating Generated Videos) is a training-free pipeline:

  • A video diffusion model (Kling v1.6, chosen over Sora and Kling v1.5 for reliability) generates candidate demonstration videos from a language command and one initial scene image.
  • A vision-language model automatically filters out videos that do not follow the command.
  • A 6D pose tracker (FoundationPose, no instance-specific fine-tuning, real-time) extracts object trajectories from the chosen video.
  • Trajectories are retargeted to the robot in an embodiment-agnostic fashion, so the same generated video drives execution without robot data.

Results

  • Overall success 85.0%, vs. Gen2Act 67.5%, AVDC 32.5%, 4D-DPM 35.0%, Track2Act 7.5%.
  • Beats the keypoint-constraint baseline ReKep (50%) at 85% across four tasks.
  • Higher video-generation quality lifts per-task success (e.g., pouring 80%โ†’100%, placing 50%โ†’90%, sweeping 20%โ†’70%).
  • Evaluated on four real-world tasks: pouring water, lifting a lid, placing a spatula on a pan, sweeping trash.

Significance

Demonstrates that filtered generated videos can be as effective as real demonstrations, decoupling manipulation skill acquisition from physical data collection. Positions video generative models as a scalable, embodiment-agnostic demonstration source โ€” a complementary take on the "world/video model as substrate for actions" theme.

Links

Related pages

โ† Back to ICLR-2026

โš ๏ธ **GitHub.com Fallback** โš ๏ธ