ICLR 2026 Batch Online RL - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 ยท Authors: Perry Dong, Suvir Mirchandani, Dorsa Sadigh, Chelsea Finn (Stanford) ยท arXiv 2505.08078 ยท Category: RL for manipulation ยท Trend tag: scalable self-improvement / batch online RL
flowchart LR
P0[Initial policy] --> C[Collect large batch of<br/>autonomous rollouts]
C --> D[(Aggregated data)]
D --> AX1[Axis 1: algorithm class<br/>IL vs filtered-IL vs value-based RL]
D --> AX2[Axis 2: policy extraction<br/>explicit AWR vs implicit Q-guided]
D --> AX3[Axis 3: policy expressivity<br/>Gaussian vs diffusion]
AX1 & AX2 & AX3 --> R[Recipe: Q-functions + implicit extraction<br/>+ expressive diffusion policy]
R --> P1[Improved policy]
P1 -->|next iteration| C
Batch online RL โ learning from large batches of autonomously collected data for policy improvement โ promises truly scalable robot learning by cutting human data-collection effort while gaining from self-improvement. But it is unclear what design choices actually enable effective self-improvement in robotics. The paper is a systematic empirical study of the question.
Three axes are studied for how they affect performance and scaling with the amount of autonomously collected data:
- (i) Algorithm class โ imitation learning vs. filtered-IL vs. value-based RL.
- (ii) Policy extraction โ explicit (e.g., advantage-weighted regression) vs. implicit (choosing the best in-distribution action via the Q-function).
- (iii) Policy expressivity โ Gaussian vs. diffusion policies.
The resulting recipe combines Q-functions to guide batch online RL, implicit policy extraction (best in-distribution action), and an expressive (diffusion) policy class; the authors also note temporally-correlated noise helps exploration.
- Value-based RL with Q-functions outperforms imitation-based methods across the tested tasks.
- Implicit extraction is necessary over traditional explicit methods.
- Expressive (diffusion) policies outperform less expressive ones.
- The recipe yields up to 2x performance improvement over prior methods, and a 30% real-world success-rate gain over batch online RL iterations.
- Evaluated in simulation (Robomimic, MimicGen, Adroit) and on a real-world 7-DoF Franka vision-based task.
A "what matters" ablation study that gives a concrete, reproducible recipe for batch online RL in robotics, clarifying that the gains come from value guidance + implicit extraction + expressive policies rather than from imitation alone. Directly relevant to self-improving VLA pipelines that bootstrap from autonomous rollouts.
- arXiv 2505.08078
- OpenReview
- ICLR 2026 listing
โ Back to ICLR-2026