ICLR 2026 Batch Online RL - Heungwoo/research GitHub Wiki

Batch Online RL โ€” what matters for self-improving robot learning from autonomous data

Venue: ICLR 2026 ยท Authors: Perry Dong, Suvir Mirchandani, Dorsa Sadigh, Chelsea Finn (Stanford) ยท arXiv 2505.08078 ยท Category: RL for manipulation ยท Trend tag: scalable self-improvement / batch online RL

Approach diagram

flowchart LR
  P0[Initial policy] --> C[Collect large batch of<br/>autonomous rollouts]
  C --> D[(Aggregated data)]
  D --> AX1[Axis 1: algorithm class<br/>IL vs filtered-IL vs value-based RL]
  D --> AX2[Axis 2: policy extraction<br/>explicit AWR vs implicit Q-guided]
  D --> AX3[Axis 3: policy expressivity<br/>Gaussian vs diffusion]
  AX1 & AX2 & AX3 --> R[Recipe: Q-functions + implicit extraction<br/>+ expressive diffusion policy]
  R --> P1[Improved policy]
  P1 -->|next iteration| C
Loading

Problem

Batch online RL โ€” learning from large batches of autonomously collected data for policy improvement โ€” promises truly scalable robot learning by cutting human data-collection effort while gaining from self-improvement. But it is unclear what design choices actually enable effective self-improvement in robotics. The paper is a systematic empirical study of the question.

Method

Three axes are studied for how they affect performance and scaling with the amount of autonomously collected data:

  • (i) Algorithm class โ€” imitation learning vs. filtered-IL vs. value-based RL.
  • (ii) Policy extraction โ€” explicit (e.g., advantage-weighted regression) vs. implicit (choosing the best in-distribution action via the Q-function).
  • (iii) Policy expressivity โ€” Gaussian vs. diffusion policies.

The resulting recipe combines Q-functions to guide batch online RL, implicit policy extraction (best in-distribution action), and an expressive (diffusion) policy class; the authors also note temporally-correlated noise helps exploration.

Results

  • Value-based RL with Q-functions outperforms imitation-based methods across the tested tasks.
  • Implicit extraction is necessary over traditional explicit methods.
  • Expressive (diffusion) policies outperform less expressive ones.
  • The recipe yields up to 2x performance improvement over prior methods, and a 30% real-world success-rate gain over batch online RL iterations.
  • Evaluated in simulation (Robomimic, MimicGen, Adroit) and on a real-world 7-DoF Franka vision-based task.

Significance

A "what matters" ablation study that gives a concrete, reproducible recipe for batch online RL in robotics, clarifying that the gains come from value guidance + implicit extraction + expressive policies rather than from imitation alone. Directly relevant to self-improving VLA pipelines that bootstrap from autonomous rollouts.

Links

Related pages

โ† Back to ICLR-2026

โš ๏ธ **GitHub.com Fallback** โš ๏ธ