RSS 2026 When Life Gives You BC - Heungwoo/research GitHub Wiki

When Life Gives You BC, Make Q-functions: Extracting Q-values from Behavior Cloning for On-Robot Reinforcement Learning

Venue: RSS 2026 (Sydney, Jul 13–17) · Session: RL · paper #153 Authors: Lakshita Dodeja, Ondrej Biza, Shivam Vats, Stephen Hart, Stefanie Tellex, Robin Walters, Karl Schmeckpeper, Thomas Weng (Robotics and AI Institute; Brown University; Northeastern University) arXiv: 2605.05172 · program page

Summary compiled from the arXiv paper (v3); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Q2RL Q-Estimation and Q-Gating pipeline (Figure 1 of arXiv 2605.05172, © the authors)

Figure 1 overview. Q-Estimation extracts a Q-function (Q̂_BC) from a BC policy's value function, action distribution, and entropy. During online RL, Q-Gating uses a frozen Q̂_BC and a trainable Q_RL, executing whichever of the BC or RL action has the higher Q-value and updating the RL policy on the collected interactions. The bottom strip and right column show the real-world contact-rich tasks (Peg Insertion, Pipe Assembly, Kitting) learned in 1–2 hours.

Problem

Behavior Cloning is effective but has no mechanism for self-guided online improvement after demonstrations are collected. Existing offline-to-online methods often "unlearn" good BC actions because of the distribution mismatch between offline data and online rollouts, and training an RL policy from scratch on a robot is unsafe and sample-inefficient.

Method

Q2RL (Q-Estimation and Q-Gating from BC for RL) has two parts. Q-Estimation extracts a Q-function Q̂_BC from a trained BC policy — using only its action-selection probabilities and entropy — with a few environment interaction steps, giving a stable value reference without large offline positive/negative datasets. Q-Gating then runs online RL, initializing Q_RL from Q̂_BC and, at each step, selecting and executing the BC or RL action with the higher respective Q-value to collect samples for training the RL policy. This keeps good BC behavior while allowing targeted online improvement. Real-world BC policies are Gaussian Mixture Models with ResNet-10 encoders; on-robot RL uses a small CNN encoder and Gaussian MLP head with stochastic actions.

Results

On D4RL (Adroit Pen/Door, Kitchen-Complete) and robomimic (Lift, Can, Square), Q2RL outperforms SOTA offline-to-online and BC-to-RL baselines (WSRL, CalQL, RLPD, IBRL) on success rate and time to convergence, improving BC success from 50–60% up to 80–100%. In real, on-robot RL on a Franka FR3 (Table VI, best checkpoint over 20 trials): Peg Insertion BC 0.70 → Q2RL 1.00 (1.4×); Pipe Assembly 0.20 → 0.75 (3.75×); Kitting-Modified 0.35 → 0.70 (2×), all learned in 1–2 hours of online interaction. IBRL matched Q2RL only on Peg Insertion and scored 0.0 on the long-horizon tasks (and 0.0 without seeded demos, where Q2RL still reached 1.00). Stochastic CalQL caused 4 safety violations; IBRL caused 2 during Peg Insertion.

Significance

Q2RL turns an existing BC policy directly into a Q-function to bootstrap safe, sample-efficient on-robot RL, avoiding both large offline datasets and the safety hazards of from-scratch online RL on contact-rich, high-precision tasks. Related: RL · Review-Dexterous-Manipulation.

← Back to RSS 2026 survey · RSS-2026-Papers · Home