RSS 2026 When to Act Ask or Learn - Heungwoo/research GitHub Wiki
When to Act, Ask, or Learn: Uncertainty-Aware Policy Steering
Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Imitation learning 2 · paper #142 Authors: Jessie Yuan, Yilin Wu, Andrea Bajcsy (Carnegie Mellon University) arXiv: 2602.22474 · program page
Summary compiled from the arXiv paper (v2); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Figure 1 pipeline. A base diffusion policy proposes K action samples; a world model plus VLM narration turns each into an open-ended outcome description; a conformal-prediction-calibrated VLM verifier returns a prediction set that maps to one of three responses — execute a confident action, ask a natural-language clarification question (set size > 1), or trigger re-training via residual learning (set = "none of the above"). The three request columns illustrate capable, ambiguous, and incapable cases.
Problem
Policy steering uses a learned verifier (often a VLM) to select action samples from a pre-trained policy, but existing frameworks assume the verifier is well-calibrated. Overconfident VLM judgments degrade steering under both semantic task ambiguity and low-level policy incapability, so the robot cannot tell "confident" from "ambiguous" from "incapable."
Method
Uncertainty-Aware Policy Steering (UPS) jointly reasons about semantic task uncertainty and low-level action feasibility. A base image-conditioned diffusion policy proposes K action chunks; a Dreamer-v3 world model imagines their outcomes, which a VLM narrates (Gemini-3-flash-preview for narration, Gemini-2.0-flash for verification). A factorized, intent-conditioned "Bayesian Intent" score function feeds conformal prediction to build a calibrated prediction set — possibly including a "none of the above" option — with a 1−ε coverage guarantee (target 1−ε = 85%). Singleton sets execute; larger sets trigger natural-language clarification; a "none of the above" set triggers interactive human intervention, after which the policy is improved via residual learning (a gated residual mixed into half the samples) for continual learning.
Results
Across simulation (Robomimic square nut-on-peg) and Franka hardware (PnP Cup, Insert Block), the Bayesian-Intent + CP verifier reaches ≥85% coverage while minimizing clarification, and UPS shows a 30% improvement in ambiguous scenarios over the uncalibrated Forewarn baseline (which gains only ~15% in ambiguous cases vs 45% in straightforward). Clarification adds up to 15% success in ambiguous cases. For continual learning, UPS minimizes human effort: normalized intervention rate 0.05 on Insert Block vs 0.06 (HG-DAgger) and 0.26 (EnsembleDAgger); in simulation 0.058 vs 0.20 and 0.27. UPS w/ Clarification + Residual attains the highest post-training success across tasks.
Significance
UPS shows that calibrating the composition of a VLM verifier and a base policy — rather than trusting raw VLM confidence — lets a robot choose between acting, asking, and learning with statistical guarantees, cutting expensive human interventions in continual imitation learning. Related: Review-LBM-Cotraining · RL.
← Back to RSS 2026 survey · RSS-2026-Papers · Home