RSS 2026 ReSteer - Heungwoo/research GitHub Wiki
ReSteer: Quantifying and Refining the Steerability of Multitask Robot Policies
Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Imitation learning 2 · paper #143 Authors: Zhenyang Chen, Alan Tian, Liquan Wang, Benjamin Joffe, Yingyan Celine Lin, Yuxiao Chen, Siddharth Karamcheti, Danfei Xu arXiv: 2603.17300 · program page
Summary compiled from the arXiv paper (v1); all numbers quoted from the paper. Trend context: RSS 2026 survey.

The motivating kitchen scenario: a robot arm is executing "Place the teabag on the oven" (orange trajectory) when the user issues "Put it on the countertop" mid-execution (blue trajectory). Daily manipulation is interruptible — policies must switch behaviors from any intermediate state in response to a new prompt.
Problem
Language-conditioned multitask policies often ignore a new instruction issued mid-execution, continuing the original task as if conditioned on state alone. The authors trace this to training data collected one task at a time, so language is informative for only a small fraction of execution and policies shortcut-learn to over-index on state.
Method
ReSteer first quantifies steerability: rolling out task i, switching the prompt to task j at sampled timesteps, and averaging success gives a steerability matrix/score; conditional mutual information I(A; L|S) between actions and language given state serves as a rollout-free proxy (Pearson r = 0.74 overall, >0.9 within the π0.5 family). The improvement pipeline has three parts: (i) a CMI-based steerability estimator that flags "instruction-blind" states, (ii) SteerGen, a stage-aware generator that synthesizes short transition trajectories from low-CMI states of task A to stage-matched states of task B, and (iii) self-refining behavior cloning (SRBC) that iteratively fine-tunes on the policy's own successful steering rollouts.
Results
Evaluating π0.5, OpenVLA-OFT, and MolmoACT on LIBERO-Goal — all above 90% single-task success — steerability scores are only 0.403, 0.252, and 0.295. On LIBERO over 18k rollouts, ReSteer improves steerability by 11% over the base π0.5 (stage-aware augmentation alone +8.8%, outperforming a CAST-style baseline). On a real DROID kitchen setup with the π0.5 DROID checkpoint, ReSteer reaches 73% steering success versus 33% for standard teleoperation fine-tuning (2.2×) while keeping 85% single-task success; pregrasp-stage switches improve from 40%/20% to 100%.
Significance
Names and measures a failure mode users hit immediately in interactive deployment, and shows it is a data-distribution problem fixable with targeted synthetic switching data rather than architecture changes — with CMI as a cheap diagnostic for any language-conditioned policy. Related wiki threads: Review-LBM-Cotraining · RL.
← Back to RSS 2026 survey · RSS-2026-Papers · Home