RSS 2026 Steerable Vision Language Action Policies for Embodied - Heungwoo/research GitHub Wiki
Steerable Vision-Language-Action Policies for Embodied Reasoning and Hierarchical Control
Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Imitation learning 1 · paper #74 Authors: William Chen, Jagdeep Singh Bhatia, Catherine Glossop, Nikhil Mathihalli, Andy Tang, Ria Doshi, Danny Driess, Karl Pertsch, Sergey Levine arXiv: 2602.13193 · program page
Summary compiled from the arXiv paper (v3); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Left: a high-level VLM reasons about a WidowX tabletop scene ("put all the food on the plate") and issues detailed steering commands — pointing at the banana's pixel coordinates, a trace along waypoints, then a drop subtask. Right: the four new steering command styles beyond standard task-level language: subtasks, atomic motions, grounded points, and gripper traces.
Problem
Hierarchical robot systems usually connect a high-level VLM to a low-level VLA through natural-language task instructions, which caps how much VLM reasoning can actually steer behavior: a VLM may infer exactly where to grasp, but the policy cannot consume that inference. The authors (UC Berkeley, Stanford, Physical Intelligence) argue the bottleneck is insufficient policy steerability.
Method
Steerable Policies are VLAs trained on rich synthetic commands spanning abstraction levels — subtasks, atomic motions ("move right"), grounded pixel points, and gripper traces — generated at scale by an automatic labeling pipeline (using Molmo, DETR, SAM, and Gemini) over the Bridge WidowX dataset. Policies are instantiated on both OpenVLA (Prismatic 7B backbone) and π0.5. Two hierarchical controllers are then built on top: (1) a fine-tuned embodied reasoner VLM producing chain-of-thought followed by steering commands, and (2) an off-the-shelf VLM (Gemini 3.0) doing in-context reasoning to pick both the next behavior and the best command abstraction, re-queried every N=20 policy steps with a history of images and executed commands.
Results
A human oracle issuing unrestricted steering commands solves the Bridge task suite at nearly 100% success, and every single command style beats task-level language alone — but no single style dominates (traces/points win on semantic generalization, atomic motions on spatial). With the learned embodied reasoner, the system outperforms five baselines — standard Bridge OpenVLA and π0.5, both ECoT-Lite variants, and full ECoT — on the ECoT-Lite task suite, with the largest gains on motion and semantic generalization. The in-context-learning controller outperforms both the standard OpenVLA and a SayCan-like subtask-only baseline on multi-step long-horizon tasks (task-progression metric), and beats its own non-reasoning ablation.
Significance
Reframes the VLM-to-VLA interface question: instead of a single fixed conditioning modality, let the high level choose among abstractions per situation — a capability neither language-only nor trace-only interfaces provide. Directly relevant to the wiki's hierarchical-control threads Review-System-0-1-2 and Review-LBM-Cotraining.
← Back to RSS 2026 survey · RSS-2026-Papers · Home