Review Steerable Policies - Heungwoo/research GitHub Wiki
Paper: Steerable Vision-Language-Action Policies for Embodied Reasoning and Hierarchical Control Authors: William Chen, Jagdeep Singh Bhatia, Catherine Glossop, Nikhil Mathihalli, Ria Doshi, Andy Tang, Danny Driess, Karl Pertsch, Sergey Levine Affiliations: UC Berkeley Β· Stanford Β· Physical Intelligence (Driess, Pertsch) arXiv: 2602.13193 Β· Feb 13, 2026 Project page: https://steerable-policies.github.io/
This page sits in the VLA architectures review as a Category F (hierarchical / dual-system) + Category G (reasoning-augmented) instance, and complements System 0/1/2 review as a concrete realization of explicit S2 steering on top of S1.
Standard hierarchical VLA stacks (GR00T, Ο0.5, Hi Robot, RoboBrain) use natural-language subtask strings as the System-2 β System-1 interface. The authors argue this NL bottleneck is too vague to actually steer low-level control β "move arm left" doesn't distinguish 1 cm from 10 cm, and "pick the carrot" doesn't say which carrot or which side to grasp. Their fix: train Steerable Policies that accept a 6-category command vocabulary (Tasks Β· Subtasks Β· Atomic motions Β· Points Β· Gripper traces Β· Combinations) generated automatically by a foundation-model labeling pipeline (Molmo + SAM2 + DETR + Gemini), expanding Bridge V2 from 38k task-level labels β 206k subtasks β ~2M total steering commands. They then drive the steered policy two different ways β (a) a fine-tuned embodied-reasoning VLM that queries every 5 env steps, (b) an off-the-shelf Gemini API via in-context learning that queries every 20 env steps β and beat ECoT, ECoT-Lite, OpenVLA, Ο0.5, and SayCan-like baselines on real-robot Bridge WidowX evaluations.
The headline contribution is the interface, not a new architecture β Steerable Policies are instantiated on existing backbones (OpenVLA AR-token and Ο0.5 flow-matching) without changing their attention or action head. All steering commands are emitted as plain text tokens so the policy's existing tokenizer handles them.
flowchart LR
subgraph Old[Prior hierarchical VLA β NL subtask interface]
direction LR
S2A[High-level VLM] --> NL["NL subtask:<br/>'pick up the carrot'"]
NL --> S1A[Low-level policy]
end
subgraph New[Steerable Policies β multi-level command interface]
direction LR
S2B[High-level VLM<br/>or Gemini ICL] --> CMD[Steering command at<br/>any of 6 categories]
CMD --> S1B[Steerable low-level policy<br/>OpenVLA / Ο0.5]
end
classDef old fill:#ffebee,stroke:#c62828,color:#000
classDef new fill:#e8f5e9,stroke:#2e7d32,color:#000
class S2A,NL,S1A old
class S2B,CMD,S1B new
Specific failure modes the NL interface causes:
- Too coarse β "reach the carrot" can't distinguish overhead approach vs. side approach.
- Compositionally rigid β subtask vocabulary is fixed at training time.
- Spatially ungrounded β "move left" is unitless; the policy guesses scale.
- Brittle to phrasing β paraphrases don't reliably hit the same skill.
The authors quote Sec. 1: NL subtasks "remain too formulaic and vague to induce the full range of physical skills."
The paper defines six categories of steering commands (Sec. IV-A), all emitted as text tokens so the same VLM tokenizer handles them. Earlier descriptions (including a prior version of this wiki page) under-counted to 5 by collapsing the default Task label into the trained vocabulary; the paper treats Tasks as a category in its own right.
flowchart TB
subgraph Levels["6 categories β concurrent during training, randomly substituted for the standard task label"]
direction TB
L0["1. Tasks<br/>'put the carrot in the pot'<br/>(default VLA task label)"]
L1["2. Subtasks<br/>'reach for the carrot'<br/>(intermediate semantic skill)"]
L2["3. Atomic motions<br/>'move left' / 'open gripper'<br/>(low-level motion verb)"]
L3["4. Points<br/>'grasp at [x, y]' /<br/>'lift above β¨pot positionβ©'"]
L4["5. Gripper traces<br/>'move from [xβ, yβ] to [xβ, yβ]'<br/>(short pixel-trajectory)"]
L5["6. Combinations<br/>'move left from [x, y] to<br/>the carrot at β¨carrot positionβ©'<br/>(hybrid)"]
L0 --> L1 --> L2 --> L3 --> L4 --> L5
end
classDef l fill:#fff9c4,stroke:#f57f17,color:#000
class L0,L1,L2,L3,L4,L5 l
| # | Category | Example | Strength |
|---|---|---|---|
| 1 | Tasks | "put the carrot in the pot" | Default β most semantic, least spatially grounded |
| 2 | Subtasks | "reach for the carrot" | Intermediate β like ECoT/Hi Robot subgoal strings |
| 3 | Atomic motions | "move left", "open gripper" | Direct motion verbs, but unitless |
| 4 | Points | "grasp at [x, y]" | Pixel-grounded target β disambiguates which object/where |
| 5 | Gripper traces | "move from [xβ, yβ] to [xβ, yβ]" | Proto-trajectory β shape AND spatial scale |
| 6 | Combinations | "move left from [x, y] to the carrot at β¨carrot positionβ©" | Hybrid β mixes language semantics with pixel grounding |
A single demo trajectory yields commands at all 6 categories simultaneously through the auto-labeling pipeline (Β§5). During training the policy is trained with the steering command randomly replacing the standard task label β so at inference any of the 6 categories can drive it. The high-level controller picks the right category per step: language for easy semantic steps, points or traces when fine spatial control matters.
Structural claim: the right abstraction is task- and step-dependent, so the policy must accept all of them.
The paper instantiates Steerable Policies on two existing VLA backbones to show the recipe transfers. No architectural surgery β the policy is "modified to map steering commands (rather than standard task labels) and third-person RGB images to end effector actions" via behavior cloning, with the steering command randomly substituted for the task label during training.
| Instantiation | Backbone | Action head | Attention | What's added |
|---|---|---|---|---|
| Steerable-OpenVLA | Prismatic 7B (Llama-2 + SigLIP/DINO) | AR discretized action tokens | causal LM | Steering text tokens prepended to the AR stream β same tokenizer, no learned input projector |
| Steerable-Ο0.5 | PaliGemma-3B | flow-matching expert | block-causal (PaliGemma prefix-LM + bidir on actions) | Steering text tokens go into the prefix β same tokenizer |
No new attention pattern. No new position embeddings. No new tokenizer. All capability gain is from the input-side conditioning, similar in spirit to Ο0.7's "prompt is the new integration surface" claim β except where Ο0.7 adds modalities (subgoal images, metadata), this paper adds abstraction levels of language/spatial commands within the same text channel.
Why this matters. Because the gain is interface-only and not architectural, the recipe inherits the strengths of whichever backbone it's bolted onto: Steerable-OpenVLA gets OpenVLA's open-weights story, Steerable-Ο0.5 gets Ο0.5's flow-matching action quality. A practitioner can adopt the multi-level vocabulary on top of any future VLA without retraining the backbone from scratch β only the BC fine-tune on the relabeled data.
flowchart TB
subgraph S2[System 2 β high-level controller, 2 variants]
direction TB
R1[Variant A: fine-tuned CoT reasoner<br/>VLM emits grounded rationale + steering command]
R2[Variant B: Gemini API in-context learning<br/>off-the-shelf, no fine-tune]
end
S2 --> CMD[Steering command<br/>any of 6 categories β task / subtask / motion / point / trace / hybrid]
subgraph S1[System 1 β Steerable low-level policy]
direction TB
BB[Backbone β OpenVLA OR Ο0.5<br/>untouched architecture]
BB --> ACT[Action chunk]
end
CMD --> BB
OBS[Robot observation] --> BB
ACT --> ROB[WidowX robot 5 Hz]
classDef s2 fill:#bbdefb,stroke:#1565c0,color:#000
classDef cmd fill:#fff9c4,stroke:#f57f17,color:#000
classDef s1 fill:#c8e6c9,stroke:#2e7d32,color:#000
class R1,R2 s2
class CMD cmd
class BB,ACT s1
The training data isn't human-labeled at multiple levels β that would not scale. Instead, a four-stage foundation-model pipeline re-labels Bridge V2 demonstrations. Crucially, this is not just Gemini β it's a stack of specialist vision models feeding Gemini (2.0) grounded features:
flowchart TB
D["Bridge V2 demos<br/>WidowX trajectories<br/>38k task-level labels"] --> S1
subgraph S1["Stage 1 β Feature extraction (specialist VLMs)"]
direction LR
F1["Molmo<br/>object names β masks"]
F2["SAM2<br/>temporal mask tracking"]
F3["DETR<br/>gripper traces"]
end
S1 --> S2["Stage 2 β Subtask decomposition<br/>(Gemini)<br/>extract motion language,<br/>segment episodes into semantic subtasks<br/>β 206k subtasks"]
S2 --> S3["Stage 3 β Command generation<br/>(Gemini conditioned on grounded features)<br/>restate each subtask in all 6 categories<br/>β ~2M total steering commands"]
S3 --> S4["Stage 4 β Rationalization<br/>(Gemini, only for reasoner training)<br/>post-hoc explanations of<br/>why each command is appropriate"]
S3 --> POL["Steerable Policy<br/>BC fine-tune on Bridge V2 +<br/>random command substitution"]
S4 --> RZN["Embodied Reasoner VLM<br/>fine-tuned on rationale + command pairs"]
classDef data fill:#e1bee7,stroke:#6a1b9a,color:#000
classDef stage fill:#bbdefb,stroke:#1565c0,color:#000
classDef out fill:#c8e6c9,stroke:#2e7d32,color:#000
class D data
class S1,S2,S3,S4,F1,F2,F3 stage
class POL,RZN out
Pipeline quantification:
| Stage | Tool | Input β Output |
|---|---|---|
| 1 β feature extraction | Molmo, SAM2, DETR | demo frames β object masks, tracked masks, gripper-trace pixels |
| 2 β subtask decomposition | Gemini | episode β ordered subtasks (38k tasks β 206k subtasks) |
| 3 β command generation | Gemini + grounded features | each subtask β command in all 6 styles (~2M total commands) |
| 4 β rationalization | Gemini | command β grounded rationale (only for reasoner training data) |
The pipeline produces all 6 categories of steering commands for every demo. During training the policy is conditioned on a randomly sampled category per training example β so at inference any category can drive it.
The 38k β 206k β 2M cascade is the data side of why this paper works: Bridge V2's 38k human-supplied task labels become 2M synthetic steering targets without any new robot data collection. This is the same data-scaling pattern as Ο0.7's metadata-via-annotation and GR00T N1.7's EgoScale 20kh β 2026 VLA training is partly a labeling-pipeline-engineering problem.
Practical bottleneck. The multi-level command quality is upper-bounded by the Gemini labeler's quality + the specialist-VLM quality (Molmo / SAM2 / DETR). If Molmo mis-segments the carrot, all downstream point/trace commands for that frame are wrong. The paper acknowledges this as a limitation.
A separate VLM is fine-tuned on the Stage-4 rationalization data to emit, for each robot observation:
- A grounded rationale β "the carrot is on the left side of the cutting board, the gripper needs to approach from above to avoid the knife"
- A steering command in any of the 6 categories β typically a point or trace for fine spatial steps, language for higher-level steps.
Inference cadence: the reasoner is queried every N=5 environment steps. Between queries, the steerable policy continues executing under the most recent command at its native ~5 Hz control rate.
This is in the ECoT lineage but with a key difference: rationales are followed by structured grounded primitives (points, traces) rather than only text β closer to MolmoAct's pointing-as-action than to ECoT's pure-text CoT.
No fine-tuning on the high level. The authors prompt the Gemini API with:
- The current robot observation
- A handful of in-context examples of (observation β multi-category steering command)
- A request to emit the next steering command
Inference cadence: Gemini is queried every N=20 environment steps β 4Γ less frequent than the fine-tuned reasoner, because each Gemini API call is much slower and (critically) costs money per call.
This works because the Steerable Policy itself accepts any of the 6 categories β so Gemini can match its emission to whatever category it finds easiest to reason about. Empirically Gemini gravitates to points and short subtask language.
| Axis | Fine-tuned reasoner (A) | Gemini ICL (B) |
|---|---|---|
| High-level fine-tune required | β | β |
| Inference cadence | every 5 steps | every 20 steps |
| Per-call cost | self-hosted (free) | Gemini API (paid) |
| Failure mode | overfits to Bridge | API rate limits / hallucinated points |
| Where evaluated in paper | Β§VI-B (embodied-reasoning suite) | Β§VI-C (long-horizon suite) |
Because the two variants are evaluated on different task suites, the paper does not directly compare them head-to-head. This is one of the bigger missed opportunities (see Β§10 limitations).
- Platform: Bridge V2 WidowX 250 β 6-DoF arm + parallel gripper, single third-person RGB camera.
- Control rate: 5 Hz end-effector deltas.
- Observation: third-person RGB only β no wrist camera, no depth, no proprioceptive state.
flowchart LR
subgraph SuiteA["Suite A β embodied-reasoning suite (Β§VI-B)"]
direction TB
A1["In-distribution"]
A2["Motion generalization<br/>new motion patterns on seen objects"]
A3["Spatial generalization<br/>new positions / orientations"]
A4["Semantic generalization<br/>new object categories / paraphrases"]
end
subgraph SuiteB["Suite B β long-horizon multi-step suite (Β§VI-C)"]
direction TB
B1["Multi-step compositional tasks<br/>(carrot-in-pot β wipe β bag)"]
B2["OOD object combinations"]
B3["Long-horizon with sub-failures"]
end
classDef a fill:#fff9c4,stroke:#f57f17,color:#000
classDef b fill:#bbdefb,stroke:#1565c0,color:#000
class A1,A2,A3,A4 a
class B1,B2,B3 b
The reasoner variant is evaluated on Suite A, the Gemini-ICL variant on Suite B. This is the cleanest empirical separation in the paper and also its main weakness for cross-method comparison.
| Baseline | Where used | What it tests |
|---|---|---|
| OpenVLA (vanilla) | both suites | "is the steering interface needed at all?" |
| Ο0.5 (vanilla) | both suites | "is the steering interface needed for flow-matching too?" |
| ECoT (full chain-of-thought) | Suite A | "does ECoT-style text CoT match the 6-category vocabulary?" |
| ECoT-Lite (CoT distilled to a small head) | Suite A | "does the cheap-CoT version match?" |
| Non-reasoning hierarchical ablation | Suite A | "is the gain from the reasoner, or from the interface?" |
| SayCan-like baseline (subtask-only NL) | Suite B | "does NL-only multi-step planning match the 6-category vocabulary?" |
| Non-reasoning ICL (Gemini emits commands without rationale) | Suite B | "is the gain from the rationale, or from the steering interface?" |
| Human oracle (Fig. 4) | sanity check | "given perfect commands, can the policy execute these tasks?" |
The paper presents results almost entirely as bar charts (Figs. 4β6) without exact numeric tables. The reads are:
Suite A β Steerable + reasoner vs. baselines:
- Steerable + reasoner > all baselines (OpenVLA, Ο0.5, ECoT, ECoT-Lite, non-reasoning hierarchical) across all four generalization axes.
- Holds for both backbones β Steerable-OpenVLA and Steerable-Ο0.5 each beat their respective vanilla counterparts and the ECoT variants.
- The non-reasoning hierarchical ablation is better than vanilla baselines but worse than full Steerable + reasoner β implying both the interface and the reasoning supply gain.
Suite B β Steerable + Gemini ICL vs. baselines:
- Steerable + Gemini ICL > OpenVLA + SayCan-like baseline on long-horizon tasks. No fine-tune on the high level.
- Non-reasoning ICL ablation (Gemini emits commands without rationales) still beats the SayCan baseline but with a smaller margin β confirming part of the gain is from the interface alone, part from the rationale.
Fig. 4 β Human Oracle upper bound:
- A human picks the best steering category per step (unrestricted): near 100% success.
- Each individual category alone (Tasks-only, Subtasks-only, Atomic-only, Points-only, Traces-only) succeeds on a different subset of tasks. No single category dominates.
This last finding is structurally important: it confirms the paper's central claim that the right abstraction is task- and step-dependent, so the policy needs to accept all 6.
- No numeric success-rate tables. All results are bar charts; readers cannot easily compare exact deltas across methods.
- No head-to-head reasoner vs. Gemini-ICL. Different task suites β so we don't know which high-level strategy wins on equal footing.
- No per-category ablation of the trained policy β Fig. 4 isolates each category for the human oracle, but the trained policy is never evaluated as "Steerable trained on Tasks-only", "trained on Points-only", etc. So we don't know which categories carry the weight at training time.
- No latency numbers. The paper notes the hierarchical setup "speeds up inference" vs. end-to-end embodied reasoning "even without compilation techniques," but no absolute milliseconds.
The single biggest missing experiment: a clean ablation isolating interface richness from reasoning quality β Steerable + non-reasoning, or non-Steerable + Gemini ICL. The paper provides partial versions of both but they're scattered across Β§VI-B and Β§VI-C with different task suites.
| Variant | High-level call frequency | Low-level control rate |
|---|---|---|
| Reasoner (A) | every 5 env steps | 5 Hz |
| Gemini ICL (B) | every 20 env steps | 5 Hz |
| Vanilla OpenVLA / Ο0.5 | n/a (no high level) | 5 Hz |
| Axis | Steerable Policies | Ο0.7 | GR00T N1.7 | ECoT | RT-H | MolmoAct |
|---|---|---|---|---|---|---|
| New architecture? | β (backbone-agnostic) | β (same as Ο0.6) | β (vl_self_attention) | β | β | β (Molmo + pointing) |
| Hierarchy | Explicit S2/S1 | Explicit S2/S1 | Explicit S2/S1 | Implicit (single model) | Implicit | Implicit |
| Steering interface | 6-category multi-modal commands | task + subtask + subgoal img + metadata + ctrl | NL subtask + image | NL CoT trace | NL motion vocabulary | Pointing as action |
| Reasoning style | CoT + grounded primitives | metadata-conditioned | Cosmos-Reason CoT | Pure text CoT | Motion vocab | Reasoning in 3D |
| Backbone | OpenVLA / Ο0.5 | Gemma3-4B + 860M expert | Qwen3-VL-2B + DiT | OpenVLA | RT-2 family | Molmo |
| Data | Bridge V2 + Gemini auto-labels | Ο0.5/0.6 mix + autonomous + RL | thousand-h teleop + EgoScale 20kh | Bridge | Multi-task | Molmo training |
| Headline OOD claim | beats ECoT/ECoT-Lite/Ο0.5/OpenVLA | air fryer / RL specialist parity | language-following improvement | strong on Bridge | mid-level abstraction wins | pointing as primitive |
The interesting positioning vs. Ο0.7: both papers use the word "steerable," but they mean different things. Ο0.7's steerability = multi-modal prompt with independent dropout + CFG β one foundation model that can be coaxed to different modes via prompt. This paper's steerability = explicit S2/S1 split where the S1 policy accepts a richer command vocabulary than NL alone. Orthogonal axes β could in principle stack.
- Multi-level annotation depends on Gemini quality β the labeler ceiling is the policy ceiling.
- NL subtasks remain too vague for the full range of physical skills β motivates the work but flags that NL alone isn't fixed.
- Single embodiment β Bridge WidowX, tabletop only.
- Closed API dependency β both annotation and one of two inference modes use a closed Gemini API.
- No comparison to Ο0.7 or GR00T N1.7 in the headline tables β both are concurrent (Apr 2026) so this is understandable, but the positioning vs. them is left to the reader.
- No ablation isolating "interface richness" from "reasoning quality" β Steerable + non-reasoning, or non-Steerable + Gemini ICL would help disentangle.
- Bridge V2 evaluation is the standard but limited β generalizing to mobile, bimanual, or long-horizon home tasks (where Ο0.5/0.7 operate) is the natural next test.
- The 6-category vocabulary is hand-designed. Would the recipe scale to 10 or 20 categories? Or would diminishing returns kick in?
- No reported latency β adding multi-level steering to OpenVLA's AR stream presumably increases sequence length; impact on inference time is unreported.
- NL is a leaky S2/S1 interface. This is the strongest empirical statement to date that natural language alone is insufficient for hierarchical steering β multi-level commands measurably help.
- Backbone-agnostic interface. The same recipe works on AR-token (OpenVLA) and flow-matching (Ο0.5) backbones β the contribution is the conditioning layer, not architecture.
- Off-the-shelf VLM ICL can drive a robot policy. The Gemini-ICL variant beating SayCan on long-horizon is striking β implies frozen general VLMs can steer custom-trained low-level policies if the interface is right.
- Foundation-model auto-labeling as data infrastructure. Like Ο0.7's metadata-via-annotation and GR00T N1.7's EgoScale, this paper assumes that a strong VLM labeler is an infrastructure-level resource. The trend is now clear: 2026 VLA training is partly a labeling-pipeline-engineering problem.
- Does the multi-level vocabulary scale to dexterous hands or whole-body humanoid control where atomic motions are higher-dimensional?
- Could the steering vocabulary be learned instead of hand-designed (e.g., autoencoder-discovered command primitives)?
- How does Steerable + Ο0.7's metadata + CFG compose? They're orthogonal axes β both "steerable" by different mechanisms.
- The pointing/trace primitives overlap with MolmoAct's and Embodied-R1's pointing-as-output. Where's the dividing line between "pointing as action" and "pointing as steering command"?
- Can the reasoner be self-trained by distilling Gemini's ICL outputs back into the fine-tuned reasoner β closing the loop without staying API-dependent?
- arXiv abstract: https://arxiv.org/abs/2602.13193
- arXiv HTML: https://arxiv.org/html/2602.13193
- arXiv PDF: https://arxiv.org/pdf/2602.13193
- Project page: https://steerable-policies.github.io/
- Review-VLA-Architecture β Β§5.F (hierarchical) + Β§5.G (reasoning-augmented) place this paper
- Review-System-0-1-2 β System 2 / System 1 framing this paper instantiates concretely
- Review-pi07 β the contemporary "steerable" paper from PI; orthogonal mechanism (multi-modal prompt vs. multi-level command)
- Review-VLM-Action-Connection β interface taxonomy; Steerable Policies fit "command-stream conditioning" between KV-share and unified-stream
- ICLR-2026-Embodied-R1 Β· CVPR-2025-VLA-Manipulation-Survey β pointing/trace primitive lineage (ECoT β MolmoAct β Embodied-R1 β Steerable)
β Back to Home