ICLR 2026 Steerable Policies - Heungwoo/research GitHub Wiki
Authors: William Chen · Jagdeep Singh Bhatia · Catherine Glossop · Nikhil Mathihalli · Ria Doshi · Andy Tang · Danny Driess · Karl Pertsch · Sergey Levine Affiliations: UC Berkeley · Stanford · Physical Intelligence (Driess, Pertsch) Venue: arXiv preprint · Feb 13, 2026 · arXiv 2602.13193 Project page: https://steerable-policies.github.io/ Category: VLA Architecture (hierarchical / dual-system) · Reasoning-augmented Trend tag: Trend 1 (system architecture) · Trend 5 (reasoning)
flowchart LR
subgraph Old["Prior hierarchical VLA — NL bottleneck"]
direction LR
O1["High-level VLM"] --> O2["NL subtask:<br/>'pick up the carrot'"]
O2 --> O3["Low-level policy"]
end
subgraph New["Steerable Policies — 6-category command vocabulary"]
direction LR
N1["High-level controller<br/>(reasoner OR Gemini ICL)"] --> N2["Steering command<br/>at any of 6 categories"]
N2 --> N3["Steerable policy<br/>OpenVLA OR π0.5"]
end
classDef old fill:#ffebee,stroke:#c62828,color:#000
classDef new fill:#e8f5e9,stroke:#2e7d32,color:#000
class O1,O2,O3 old
class N1,N2,N3 new
Standard hierarchical VLAs (GR00T, π0.5, Hi Robot, RoboBrain) use natural-language subtask strings as the System-2 → System-1 interface. The authors argue that NL is fundamentally too vague to fully steer low-level control:
- "Move arm left" — by 1 cm or 10 cm? Unitless.
- "Pick the carrot" — which carrot? Which side?
- "Reach for the cup" — overhead approach or side approach?
- Paraphrases ("grab" vs. "grasp" vs. "pick up") don't reliably hit the same skill.
NL subtasks "remain too formulaic and vague to induce the full range of physical skills" (Sec. 1).
Six steering-command categories, all expressed as text tokens so the same VLM tokenizer handles them:
- Tasks — "put the carrot in the pot" (the standard VLA task label)
- Subtasks — "reach for the carrot" (intermediate semantic skill)
- Atomic motions — "move left" / "open gripper" (low-level motion verbs)
- Gripper traces — "move from [x₁, y₁] to [x₂, y₂]" (short pixel-trajectory)
- Points — "grasp at [x, y]" / "lift above ⟨pot position⟩" (pixel-grounded targets)
- Combinations — "move left from [x, y] to the carrot at ⟨carrot position⟩" (hybrid)
A single demo trajectory yields commands at all 6 categories, generated by a foundation-model auto-labeling pipeline. The policy is trained with the steering command randomly replacing the standard task label — so at inference any of the 6 levels can drive it.
Auto-labeling pipeline (Bridge V2 → 2M commands):
- Feature extraction — Molmo (object names → masks), SAM2 (temporal tracking), DETR (gripper traces).
- Subtask decomposition — Gemini extracts motion language and segments episodes into semantic subtasks.
- Command generation — Gemini restates each subtask in all 6 styles, conditioned on the grounded features.
- Rationalization (for the reasoner) — Gemini emits post-hoc explanations of why each command is appropriate.
This expands Bridge V2 from 38k task-level labels → 206k subtasks → ~2M total steering commands.
Two backbones tested (no architecture changes):
| Variant | Backbone | Action head | Steering input |
|---|---|---|---|
| Steerable-OpenVLA | Prismatic 7B (Llama-2 + SigLIP/DINO) | AR discretized tokens | Text tokens prepended to AR stream |
| Steerable-π0.5 | PaliGemma-3B | Flow-matching expert | Text tokens in the prefix |
Two ways to drive at inference:
- Variant A — fine-tuned reasoner. A separate VLM is trained to emit a grounded rationale + steering command per observation. Queries every N=5 env steps. Closer to MolmoAct's pointing-as-action than to ECoT's pure-text CoT.
- Variant B — off-the-shelf Gemini ICL. No fine-tune on the high level. Gemini gets the observation + a few in-context examples and emits a steering command. Queries every N=20 env steps. Works because the Steerable Policy accepts any of the 6 levels — Gemini picks the level it finds easiest.
Real-robot platform: Bridge V2 WidowX 250 (5 Hz control).
Eval axes: in-distribution / motion / spatial / semantic generalization (§VI-A, VI-B), plus long-horizon multi-step (§VI-C).
Headline reads (paper reports as bar charts; exact percentages not in tables):
- Steerable + reasoner beats all of {OpenVLA, π0.5, ECoT, ECoT-Lite, non-reasoning hierarchical} across all four generalization axes — for both OpenVLA and π0.5 backbones.
- Steerable + Gemini ICL beats OpenVLA and a SayCan-like baseline on long-horizon multi-step tasks — with no fine-tune on the high level.
- Non-reasoning ablation still beats baselines, but with a smaller margin — confirming the gain comes partly from the interface, partly from the reasoning.
- Human-oracle upper bound (Fig. 4): unrestricted steering reaches near 100% success — establishing that the policy can execute these tasks given the right command, so the bottleneck is the high-level controller's command quality.
Important caveat on cross-method comparison: the reasoner (§VI-B) and the Gemini-ICL (§VI-C) variants are evaluated on different task suites, so the paper does not directly compare them head-to-head.
The strongest empirical statement to date that NL is a leaky S2↔S1 interface. This paper isolates the interface as an architectural surface that can be improved independently of the backbone — the same recipe ports from AR-token (OpenVLA) to flow-matching (π0.5) without changing attention or the action head.
Backbone-agnostic. Where π0.7 makes the prompt the integration surface (subgoal images, metadata, controls) and adds CFG + dropout, this paper makes the command vocabulary the integration surface and adds 6 categories of grounded primitives. The two are orthogonal axes and could in principle stack.
Off-the-shelf VLM ICL drives a robot. The Gemini-ICL variant beating SayCan on long-horizon is the most surprising result — it implies that frozen general-purpose VLMs can steer custom-trained low-level policies when the interface is rich enough.
Foundation-model auto-labeling as data infrastructure. Together with π0.7's metadata-via-annotation and GR00T N1.7's EgoScale, this paper confirms a 2026 pattern: VLA training is partly a labeling-pipeline-engineering problem — data scale comes from foundation-model relabelers, not human annotators.
- arXiv: 2602.13193 (Feb 13, 2026)
- HTML: https://arxiv.org/html/2602.13193
- Project page: https://steerable-policies.github.io/
- In-depth review: Steerable Policies (long-form)
- π0.7 — the contemporary "steerable" paper from PI; orthogonal mechanism
- Embodied-R1 — pointing/trace primitive lineage
- Hybrid Training — train-with-CoT, infer-without-CoT recipe
- VLM↔Action Connection Review — interface taxonomy
- System 0/1/2 Review — concrete S2-on-S1 instance
- Survey: VLA & Manipulation
← Back to ICLR-2026