ICLR 2026 Self Improving Loops - Heungwoo/research GitHub Wiki

SILVR — self-improving video-planning loops for robotic tasks

Venue: ICLR 2026 · Authors: Calvin Luo, Zilai Zeng, Mingxi Jia, Yilun Du, Chen Sun · arXiv:2506.06658 · Category: training / self-improvement for visual planning · Trend tag: continual / online self-improvement.

Approach diagram

flowchart LR
  VM[In-domain video model] --> Plan[Generate visual plan for task]
  Plan --> Exec[Execute -> self-collected trajectories]
  Exec --> Filter[Keep self-produced rollouts]
  Filter --> Update[Iteratively update video model]
  Update --> VM
  Update -.improves over iterations.-> Better[Higher success on novel tasks]

Problem

Video generative models trained on expert demonstrations act as strong text-conditioned visual planners, but generalize poorly to unseen tasks. Rather than relying solely on pre-collected offline data (e.g., web-scale video), the goal is an agent that continuously improves online from its own collected behaviors.

Method

SILVR (Self-Improving Loops for Visual Robotic Planning): an in-domain video model iteratively updates itself on self-produced trajectories, steadily improving on a specified target task. The loop requires no human-provided ground-truth reward and no expert-quality demonstrations — it bootstraps from its own rollouts.

Results

  • Applied to a diverse suite of MetaWorld tasks plus two manipulation tasks on a real robot arm.
  • Performance improvements emerge continuously over multiple iterations for novel tasks unseen during initial in-domain training.
  • SILVR is robust without ground-truth rewards or expert demos, and is preferable to alternative online-experience methods in both performance and sample efficiency.

Significance

Embodies the "era of experience" view: instead of scaling offline data, the planner improves itself from self-collected behavior. A reward-free, demo-free self-improvement loop is an attractive recipe for continual adaptation to new manipulation tasks.

Links

Related pages

← Back to ICLR-2026