Review Steerable Policies - Heungwoo/research GitHub Wiki

In-Depth Review β€” Steerable VLA Policies for Embodied Reasoning and Hierarchical Control

Paper: Steerable Vision-Language-Action Policies for Embodied Reasoning and Hierarchical Control Authors: William Chen, Jagdeep Singh Bhatia, Catherine Glossop, Nikhil Mathihalli, Ria Doshi, Andy Tang, Danny Driess, Karl Pertsch, Sergey Levine Affiliations: UC Berkeley Β· Stanford Β· Physical Intelligence (Driess, Pertsch) arXiv: 2602.13193 Β· Feb 13, 2026 Project page: https://steerable-policies.github.io/

This page sits in the VLA architectures review as a Category F (hierarchical / dual-system) + Category G (reasoning-augmented) instance, and complements System 0/1/2 review as a concrete realization of explicit S2 steering on top of S1.


1. TL;DR

Standard hierarchical VLA stacks (GR00T, Ο€0.5, Hi Robot, RoboBrain) use natural-language subtask strings as the System-2 ↔ System-1 interface. The authors argue this NL bottleneck is too vague to actually steer low-level control β€” "move arm left" doesn't distinguish 1 cm from 10 cm, and "pick the carrot" doesn't say which carrot or which side to grasp. Their fix: train Steerable Policies that accept a 6-category command vocabulary (Tasks Β· Subtasks Β· Atomic motions Β· Points Β· Gripper traces Β· Combinations) generated automatically by a foundation-model labeling pipeline (Molmo + SAM2 + DETR + Gemini), expanding Bridge V2 from 38k task-level labels β†’ 206k subtasks β†’ ~2M total steering commands. They then drive the steered policy two different ways β€” (a) a fine-tuned embodied-reasoning VLM that queries every 5 env steps, (b) an off-the-shelf Gemini API via in-context learning that queries every 20 env steps β€” and beat ECoT, ECoT-Lite, OpenVLA, Ο€0.5, and SayCan-like baselines on real-robot Bridge WidowX evaluations.

The headline contribution is the interface, not a new architecture β€” Steerable Policies are instantiated on existing backbones (OpenVLA AR-token and Ο€0.5 flow-matching) without changing their attention or action head. All steering commands are emitted as plain text tokens so the policy's existing tokenizer handles them.


2. Motivation β€” what the NL bottleneck breaks

flowchart LR
  subgraph Old[Prior hierarchical VLA β€” NL subtask interface]
    direction LR
    S2A[High-level VLM] --> NL["NL subtask:<br/>'pick up the carrot'"]
    NL --> S1A[Low-level policy]
  end

  subgraph New[Steerable Policies β€” multi-level command interface]
    direction LR
    S2B[High-level VLM<br/>or Gemini ICL] --> CMD[Steering command at<br/>any of 6 categories]
    CMD --> S1B[Steerable low-level policy<br/>OpenVLA / Ο€0.5]
  end

  classDef old fill:#ffebee,stroke:#c62828,color:#000
  classDef new fill:#e8f5e9,stroke:#2e7d32,color:#000
  class S2A,NL,S1A old
  class S2B,CMD,S1B new
Loading

Specific failure modes the NL interface causes:

  • Too coarse β€” "reach the carrot" can't distinguish overhead approach vs. side approach.
  • Compositionally rigid β€” subtask vocabulary is fixed at training time.
  • Spatially ungrounded β€” "move left" is unitless; the policy guesses scale.
  • Brittle to phrasing β€” paraphrases don't reliably hit the same skill.

The authors quote Sec. 1: NL subtasks "remain too formulaic and vague to induce the full range of physical skills."


3. The 6-category steering vocabulary

The paper defines six categories of steering commands (Sec. IV-A), all emitted as text tokens so the same VLM tokenizer handles them. Earlier descriptions (including a prior version of this wiki page) under-counted to 5 by collapsing the default Task label into the trained vocabulary; the paper treats Tasks as a category in its own right.

flowchart TB
  subgraph Levels["6 categories β€” concurrent during training, randomly substituted for the standard task label"]
    direction TB
    L0["1. Tasks<br/>'put the carrot in the pot'<br/>(default VLA task label)"]
    L1["2. Subtasks<br/>'reach for the carrot'<br/>(intermediate semantic skill)"]
    L2["3. Atomic motions<br/>'move left' / 'open gripper'<br/>(low-level motion verb)"]
    L3["4. Points<br/>'grasp at [x, y]' /<br/>'lift above ⟨pot position⟩'"]
    L4["5. Gripper traces<br/>'move from [x₁, y₁] to [xβ‚‚, yβ‚‚]'<br/>(short pixel-trajectory)"]
    L5["6. Combinations<br/>'move left from [x, y] to<br/>the carrot at ⟨carrot position⟩'<br/>(hybrid)"]
    L0 --> L1 --> L2 --> L3 --> L4 --> L5
  end

  classDef l fill:#fff9c4,stroke:#f57f17,color:#000
  class L0,L1,L2,L3,L4,L5 l
Loading
# Category Example Strength
1 Tasks "put the carrot in the pot" Default β€” most semantic, least spatially grounded
2 Subtasks "reach for the carrot" Intermediate β€” like ECoT/Hi Robot subgoal strings
3 Atomic motions "move left", "open gripper" Direct motion verbs, but unitless
4 Points "grasp at [x, y]" Pixel-grounded target β€” disambiguates which object/where
5 Gripper traces "move from [x₁, y₁] to [xβ‚‚, yβ‚‚]" Proto-trajectory β€” shape AND spatial scale
6 Combinations "move left from [x, y] to the carrot at ⟨carrot position⟩" Hybrid β€” mixes language semantics with pixel grounding

A single demo trajectory yields commands at all 6 categories simultaneously through the auto-labeling pipeline (Β§5). During training the policy is trained with the steering command randomly replacing the standard task label β€” so at inference any of the 6 categories can drive it. The high-level controller picks the right category per step: language for easy semantic steps, points or traces when fine spatial control matters.

Structural claim: the right abstraction is task- and step-dependent, so the policy must accept all of them.


4. Architecture β€” backbone-agnostic

The paper instantiates Steerable Policies on two existing VLA backbones to show the recipe transfers. No architectural surgery β€” the policy is "modified to map steering commands (rather than standard task labels) and third-person RGB images to end effector actions" via behavior cloning, with the steering command randomly substituted for the task label during training.

Instantiation Backbone Action head Attention What's added
Steerable-OpenVLA Prismatic 7B (Llama-2 + SigLIP/DINO) AR discretized action tokens causal LM Steering text tokens prepended to the AR stream β€” same tokenizer, no learned input projector
Steerable-Ο€0.5 PaliGemma-3B flow-matching expert block-causal (PaliGemma prefix-LM + bidir on actions) Steering text tokens go into the prefix β€” same tokenizer

No new attention pattern. No new position embeddings. No new tokenizer. All capability gain is from the input-side conditioning, similar in spirit to Ο€0.7's "prompt is the new integration surface" claim β€” except where Ο€0.7 adds modalities (subgoal images, metadata), this paper adds abstraction levels of language/spatial commands within the same text channel.

Why this matters. Because the gain is interface-only and not architectural, the recipe inherits the strengths of whichever backbone it's bolted onto: Steerable-OpenVLA gets OpenVLA's open-weights story, Steerable-Ο€0.5 gets Ο€0.5's flow-matching action quality. A practitioner can adopt the multi-level vocabulary on top of any future VLA without retraining the backbone from scratch β€” only the BC fine-tune on the relabeled data.

flowchart TB
  subgraph S2[System 2 β€” high-level controller, 2 variants]
    direction TB
    R1[Variant A: fine-tuned CoT reasoner<br/>VLM emits grounded rationale + steering command]
    R2[Variant B: Gemini API in-context learning<br/>off-the-shelf, no fine-tune]
  end

  S2 --> CMD[Steering command<br/>any of 6 categories β€” task / subtask / motion / point / trace / hybrid]

  subgraph S1[System 1 β€” Steerable low-level policy]
    direction TB
    BB[Backbone β€” OpenVLA OR Ο€0.5<br/>untouched architecture]
    BB --> ACT[Action chunk]
  end

  CMD --> BB
  OBS[Robot observation] --> BB
  ACT --> ROB[WidowX robot 5 Hz]

  classDef s2 fill:#bbdefb,stroke:#1565c0,color:#000
  classDef cmd fill:#fff9c4,stroke:#f57f17,color:#000
  classDef s1 fill:#c8e6c9,stroke:#2e7d32,color:#000
  class R1,R2 s2
  class CMD cmd
  class BB,ACT s1
Loading

5. Method β€” automated multi-level annotation pipeline

The training data isn't human-labeled at multiple levels β€” that would not scale. Instead, a four-stage foundation-model pipeline re-labels Bridge V2 demonstrations. Crucially, this is not just Gemini β€” it's a stack of specialist vision models feeding Gemini (2.0) grounded features:

flowchart TB
  D["Bridge V2 demos<br/>WidowX trajectories<br/>38k task-level labels"] --> S1
  subgraph S1["Stage 1 β€” Feature extraction (specialist VLMs)"]
    direction LR
    F1["Molmo<br/>object names β†’ masks"]
    F2["SAM2<br/>temporal mask tracking"]
    F3["DETR<br/>gripper traces"]
  end
  S1 --> S2["Stage 2 β€” Subtask decomposition<br/>(Gemini)<br/>extract motion language,<br/>segment episodes into semantic subtasks<br/>β†’ 206k subtasks"]
  S2 --> S3["Stage 3 β€” Command generation<br/>(Gemini conditioned on grounded features)<br/>restate each subtask in all 6 categories<br/>β†’ ~2M total steering commands"]
  S3 --> S4["Stage 4 β€” Rationalization<br/>(Gemini, only for reasoner training)<br/>post-hoc explanations of<br/>why each command is appropriate"]
  S3 --> POL["Steerable Policy<br/>BC fine-tune on Bridge V2 +<br/>random command substitution"]
  S4 --> RZN["Embodied Reasoner VLM<br/>fine-tuned on rationale + command pairs"]

  classDef data fill:#e1bee7,stroke:#6a1b9a,color:#000
  classDef stage fill:#bbdefb,stroke:#1565c0,color:#000
  classDef out fill:#c8e6c9,stroke:#2e7d32,color:#000
  class D data
  class S1,S2,S3,S4,F1,F2,F3 stage
  class POL,RZN out
Loading

Pipeline quantification:

Stage Tool Input β†’ Output
1 β€” feature extraction Molmo, SAM2, DETR demo frames β†’ object masks, tracked masks, gripper-trace pixels
2 β€” subtask decomposition Gemini episode β†’ ordered subtasks (38k tasks β†’ 206k subtasks)
3 β€” command generation Gemini + grounded features each subtask β†’ command in all 6 styles (~2M total commands)
4 β€” rationalization Gemini command β†’ grounded rationale (only for reasoner training data)

The pipeline produces all 6 categories of steering commands for every demo. During training the policy is conditioned on a randomly sampled category per training example β€” so at inference any category can drive it.

The 38k β†’ 206k β†’ 2M cascade is the data side of why this paper works: Bridge V2's 38k human-supplied task labels become 2M synthetic steering targets without any new robot data collection. This is the same data-scaling pattern as Ο€0.7's metadata-via-annotation and GR00T N1.7's EgoScale 20kh β€” 2026 VLA training is partly a labeling-pipeline-engineering problem.

Practical bottleneck. The multi-level command quality is upper-bounded by the Gemini labeler's quality + the specialist-VLM quality (Molmo / SAM2 / DETR). If Molmo mis-segments the carrot, all downstream point/trace commands for that frame are wrong. The paper acknowledges this as a limitation.


6. Two ways to steer at inference

6.1 Fine-tuned embodied reasoner (Variant A Β· Β§VI-B)

A separate VLM is fine-tuned on the Stage-4 rationalization data to emit, for each robot observation:

  1. A grounded rationale β€” "the carrot is on the left side of the cutting board, the gripper needs to approach from above to avoid the knife"
  2. A steering command in any of the 6 categories β€” typically a point or trace for fine spatial steps, language for higher-level steps.

Inference cadence: the reasoner is queried every N=5 environment steps. Between queries, the steerable policy continues executing under the most recent command at its native ~5 Hz control rate.

This is in the ECoT lineage but with a key difference: rationales are followed by structured grounded primitives (points, traces) rather than only text β€” closer to MolmoAct's pointing-as-action than to ECoT's pure-text CoT.

6.2 Off-the-shelf Gemini API via in-context learning (Variant B Β· Β§VI-C)

No fine-tuning on the high level. The authors prompt the Gemini API with:

  • The current robot observation
  • A handful of in-context examples of (observation β†’ multi-category steering command)
  • A request to emit the next steering command

Inference cadence: Gemini is queried every N=20 environment steps β€” 4Γ— less frequent than the fine-tuned reasoner, because each Gemini API call is much slower and (critically) costs money per call.

This works because the Steerable Policy itself accepts any of the 6 categories β€” so Gemini can match its emission to whatever category it finds easiest to reason about. Empirically Gemini gravitates to points and short subtask language.

6.3 The Reasoner-vs-ICL trade

Axis Fine-tuned reasoner (A) Gemini ICL (B)
High-level fine-tune required βœ“ βœ—
Inference cadence every 5 steps every 20 steps
Per-call cost self-hosted (free) Gemini API (paid)
Failure mode overfits to Bridge API rate limits / hallucinated points
Where evaluated in paper Β§VI-B (embodied-reasoning suite) Β§VI-C (long-horizon suite)

Because the two variants are evaluated on different task suites, the paper does not directly compare them head-to-head. This is one of the bigger missed opportunities (see Β§10 limitations).


7. Experimental setup

7.1 Robot

  • Platform: Bridge V2 WidowX 250 β€” 6-DoF arm + parallel gripper, single third-person RGB camera.
  • Control rate: 5 Hz end-effector deltas.
  • Observation: third-person RGB only β€” no wrist camera, no depth, no proprioceptive state.

7.2 Two evaluation suites (different tasks per variant)

flowchart LR
  subgraph SuiteA["Suite A β€” embodied-reasoning suite (Β§VI-B)"]
    direction TB
    A1["In-distribution"]
    A2["Motion generalization<br/>new motion patterns on seen objects"]
    A3["Spatial generalization<br/>new positions / orientations"]
    A4["Semantic generalization<br/>new object categories / paraphrases"]
  end
  subgraph SuiteB["Suite B β€” long-horizon multi-step suite (Β§VI-C)"]
    direction TB
    B1["Multi-step compositional tasks<br/>(carrot-in-pot β†’ wipe β†’ bag)"]
    B2["OOD object combinations"]
    B3["Long-horizon with sub-failures"]
  end

  classDef a fill:#fff9c4,stroke:#f57f17,color:#000
  classDef b fill:#bbdefb,stroke:#1565c0,color:#000
  class A1,A2,A3,A4 a
  class B1,B2,B3 b
Loading

The reasoner variant is evaluated on Suite A, the Gemini-ICL variant on Suite B. This is the cleanest empirical separation in the paper and also its main weakness for cross-method comparison.

7.3 Baselines

Baseline Where used What it tests
OpenVLA (vanilla) both suites "is the steering interface needed at all?"
Ο€0.5 (vanilla) both suites "is the steering interface needed for flow-matching too?"
ECoT (full chain-of-thought) Suite A "does ECoT-style text CoT match the 6-category vocabulary?"
ECoT-Lite (CoT distilled to a small head) Suite A "does the cheap-CoT version match?"
Non-reasoning hierarchical ablation Suite A "is the gain from the reasoner, or from the interface?"
SayCan-like baseline (subtask-only NL) Suite B "does NL-only multi-step planning match the 6-category vocabulary?"
Non-reasoning ICL (Gemini emits commands without rationale) Suite B "is the gain from the rationale, or from the steering interface?"
Human oracle (Fig. 4) sanity check "given perfect commands, can the policy execute these tasks?"

8. Results β€” what the paper actually reports

8.1 Headline reads

The paper presents results almost entirely as bar charts (Figs. 4–6) without exact numeric tables. The reads are:

Suite A β€” Steerable + reasoner vs. baselines:

  • Steerable + reasoner > all baselines (OpenVLA, Ο€0.5, ECoT, ECoT-Lite, non-reasoning hierarchical) across all four generalization axes.
  • Holds for both backbones β€” Steerable-OpenVLA and Steerable-Ο€0.5 each beat their respective vanilla counterparts and the ECoT variants.
  • The non-reasoning hierarchical ablation is better than vanilla baselines but worse than full Steerable + reasoner β€” implying both the interface and the reasoning supply gain.

Suite B β€” Steerable + Gemini ICL vs. baselines:

  • Steerable + Gemini ICL > OpenVLA + SayCan-like baseline on long-horizon tasks. No fine-tune on the high level.
  • Non-reasoning ICL ablation (Gemini emits commands without rationales) still beats the SayCan baseline but with a smaller margin β€” confirming part of the gain is from the interface alone, part from the rationale.

Fig. 4 β€” Human Oracle upper bound:

  • A human picks the best steering category per step (unrestricted): near 100% success.
  • Each individual category alone (Tasks-only, Subtasks-only, Atomic-only, Points-only, Traces-only) succeeds on a different subset of tasks. No single category dominates.

This last finding is structurally important: it confirms the paper's central claim that the right abstraction is task- and step-dependent, so the policy needs to accept all 6.

8.2 What the paper does not report (and why this matters)

  • No numeric success-rate tables. All results are bar charts; readers cannot easily compare exact deltas across methods.
  • No head-to-head reasoner vs. Gemini-ICL. Different task suites β€” so we don't know which high-level strategy wins on equal footing.
  • No per-category ablation of the trained policy β€” Fig. 4 isolates each category for the human oracle, but the trained policy is never evaluated as "Steerable trained on Tasks-only", "trained on Points-only", etc. So we don't know which categories carry the weight at training time.
  • No latency numbers. The paper notes the hierarchical setup "speeds up inference" vs. end-to-end embodied reasoning "even without compilation techniques," but no absolute milliseconds.

The single biggest missing experiment: a clean ablation isolating interface richness from reasoning quality β€” Steerable + non-reasoning, or non-Steerable + Gemini ICL. The paper provides partial versions of both but they're scattered across Β§VI-B and Β§VI-C with different task suites.

8.3 Inference cadence summary

Variant High-level call frequency Low-level control rate
Reasoner (A) every 5 env steps 5 Hz
Gemini ICL (B) every 20 env steps 5 Hz
Vanilla OpenVLA / Ο€0.5 n/a (no high level) 5 Hz

9. Comparison table β€” Steerable Policies vs. peers

Axis Steerable Policies Ο€0.7 GR00T N1.7 ECoT RT-H MolmoAct
New architecture? βœ— (backbone-agnostic) βœ— (same as Ο€0.6) βœ“ (vl_self_attention) βœ— βœ— βœ“ (Molmo + pointing)
Hierarchy Explicit S2/S1 Explicit S2/S1 Explicit S2/S1 Implicit (single model) Implicit Implicit
Steering interface 6-category multi-modal commands task + subtask + subgoal img + metadata + ctrl NL subtask + image NL CoT trace NL motion vocabulary Pointing as action
Reasoning style CoT + grounded primitives metadata-conditioned Cosmos-Reason CoT Pure text CoT Motion vocab Reasoning in 3D
Backbone OpenVLA / Ο€0.5 Gemma3-4B + 860M expert Qwen3-VL-2B + DiT OpenVLA RT-2 family Molmo
Data Bridge V2 + Gemini auto-labels Ο€0.5/0.6 mix + autonomous + RL thousand-h teleop + EgoScale 20kh Bridge Multi-task Molmo training
Headline OOD claim beats ECoT/ECoT-Lite/Ο€0.5/OpenVLA air fryer / RL specialist parity language-following improvement strong on Bridge mid-level abstraction wins pointing as primitive

The interesting positioning vs. Ο€0.7: both papers use the word "steerable," but they mean different things. Ο€0.7's steerability = multi-modal prompt with independent dropout + CFG β†’ one foundation model that can be coaxed to different modes via prompt. This paper's steerability = explicit S2/S1 split where the S1 policy accepts a richer command vocabulary than NL alone. Orthogonal axes β€” could in principle stack.


10. Limitations

Authors' (per Sec. 1 + Appendix E.D)

  1. Multi-level annotation depends on Gemini quality β€” the labeler ceiling is the policy ceiling.
  2. NL subtasks remain too vague for the full range of physical skills β€” motivates the work but flags that NL alone isn't fixed.
  3. Single embodiment β€” Bridge WidowX, tabletop only.
  4. Closed API dependency β€” both annotation and one of two inference modes use a closed Gemini API.

Reviewer's concerns

  1. No comparison to Ο€0.7 or GR00T N1.7 in the headline tables β€” both are concurrent (Apr 2026) so this is understandable, but the positioning vs. them is left to the reader.
  2. No ablation isolating "interface richness" from "reasoning quality" β€” Steerable + non-reasoning, or non-Steerable + Gemini ICL would help disentangle.
  3. Bridge V2 evaluation is the standard but limited β€” generalizing to mobile, bimanual, or long-horizon home tasks (where Ο€0.5/0.7 operate) is the natural next test.
  4. The 6-category vocabulary is hand-designed. Would the recipe scale to 10 or 20 categories? Or would diminishing returns kick in?
  5. No reported latency β€” adding multi-level steering to OpenVLA's AR stream presumably increases sequence length; impact on inference time is unreported.

11. Takeaways β€” what this paper changes

  1. NL is a leaky S2/S1 interface. This is the strongest empirical statement to date that natural language alone is insufficient for hierarchical steering β€” multi-level commands measurably help.
  2. Backbone-agnostic interface. The same recipe works on AR-token (OpenVLA) and flow-matching (Ο€0.5) backbones β€” the contribution is the conditioning layer, not architecture.
  3. Off-the-shelf VLM ICL can drive a robot policy. The Gemini-ICL variant beating SayCan on long-horizon is striking β€” implies frozen general VLMs can steer custom-trained low-level policies if the interface is right.
  4. Foundation-model auto-labeling as data infrastructure. Like Ο€0.7's metadata-via-annotation and GR00T N1.7's EgoScale, this paper assumes that a strong VLM labeler is an infrastructure-level resource. The trend is now clear: 2026 VLA training is partly a labeling-pipeline-engineering problem.

12. Open questions

  1. Does the multi-level vocabulary scale to dexterous hands or whole-body humanoid control where atomic motions are higher-dimensional?
  2. Could the steering vocabulary be learned instead of hand-designed (e.g., autoencoder-discovered command primitives)?
  3. How does Steerable + Ο€0.7's metadata + CFG compose? They're orthogonal axes β€” both "steerable" by different mechanisms.
  4. The pointing/trace primitives overlap with MolmoAct's and Embodied-R1's pointing-as-output. Where's the dividing line between "pointing as action" and "pointing as steering command"?
  5. Can the reasoner be self-trained by distilling Gemini's ICL outputs back into the fine-tuned reasoner β€” closing the loop without staying API-dependent?

13. Sources

Related wiki pages

← Back to Home

⚠️ **GitHub.com Fallback** ⚠️