ICLR 2026 Self Improving Loops - Heungwoo/research GitHub Wiki
SILVR — self-improving video-planning loops for robotic tasks
Venue: ICLR 2026 · Authors: Calvin Luo, Zilai Zeng, Mingxi Jia, Yilun Du, Chen Sun · arXiv:2506.06658 · Category: training / self-improvement for visual planning · Trend tag: continual / online self-improvement.
Approach diagram
flowchart LR
VM[In-domain video model] --> Plan[Generate visual plan for task]
Plan --> Exec[Execute -> self-collected trajectories]
Exec --> Filter[Keep self-produced rollouts]
Filter --> Update[Iteratively update video model]
Update --> VM
Update -.improves over iterations.-> Better[Higher success on novel tasks]
Problem
Video generative models trained on expert demonstrations act as strong text-conditioned visual planners, but generalize poorly to unseen tasks. Rather than relying solely on pre-collected offline data (e.g., web-scale video), the goal is an agent that continuously improves online from its own collected behaviors.
Method
SILVR (Self-Improving Loops for Visual Robotic Planning): an in-domain video model iteratively updates itself on self-produced trajectories, steadily improving on a specified target task. The loop requires no human-provided ground-truth reward and no expert-quality demonstrations — it bootstraps from its own rollouts.
Results
- Applied to a diverse suite of MetaWorld tasks plus two manipulation tasks on a real robot arm.
- Performance improvements emerge continuously over multiple iterations for novel tasks unseen during initial in-domain training.
- SILVR is robust without ground-truth rewards or expert demos, and is preferable to alternative online-experience methods in both performance and sample efficiency.
Significance
Embodies the "era of experience" view: instead of scaling offline data, the planner improves itself from self-collected behavior. A reward-free, demo-free self-improvement loop is an attractive recipe for continual adaptation to new manipulation tasks.
Links
Related pages
← Back to ICLR-2026