Review Behavior Prompting - Heungwoo/research GitHub Wiki
In-Depth Review â Behavior Prompting Policy: Demonstrations as Prompts for Manipulation
Paper: "Behavior Prompting Policy: Demonstrations as Prompts for Manipulation" â arXiv 2606.30457 (Jun 2026) ¡ Stanford ¡ UC Berkeley (Patel, Pekarek, Castro Hernandez, Shuran Song) ¡ project ¡ code. The "demo-as-prompt" datapoint â perform a new task from a single human demonstration ("behavior prompt") at inference, and it identifies task diversity as the primary driver of the prompting ability. Companions: In-Context Imitation ¡ ICRT ¡ MimicDroid.

1. Problem
In-context imitation should let a robot do a new task from a single demonstration at test time â but two things block it: (1) the architecture must translate an arbitrary demo + the current observation into actions, and (2) what training data actually confers the prompting ability was unclear.
2. Method â three contributions
(Algorithm) Behavior Prompting Policy (BPP) â an in-context visuomotor architecture:
- The behavior prompt (one demo) is chunked into
(observation, action-segment, proprioception)chunks; each chunk is attention-pooled into a prompt token. - A prompt encoder (transformer) has the current observation self-attend, then cross-attend to the prompt chunks, yielding a prompt embedding.
- An action decoder (diffusion) denoises the current action chunk conditioned on
current obs + prompt embedding.
(Data) Task diversity is the driver. The paper's key empirical finding: task diversity â not sheer quantity â is the primary driver of prompting capability. To collect diverse data cheaply, they introduce iPhUMI, a handheld manipulation interface (a UMI-style device â cf. DexUMI/YUBI).
(Evaluation) New test-time-adaptation benchmarks: DrawAnything (unseen drawing tasks) and LIBERO-Gen (unseen tabletop manipulation) â probes for genuine test-time adaptation to unseen tasks.
3. Why it matters (in-context lens)
BPP advances the In-Context Imitation family on the axis ICRT left open â where the prompting ability comes from. Its answer, task diversity > quantity, is a training-data recipe that reframes the whole family: to get in-context generalization you need many different tasks, and a cheap diverse-data interface (iPhUMI) is the enabler â connecting in-context imitation to the Dexterous-Hand Data Pyramid's L3 capture-interface thread. Architecturally it's a cross-attention (prompt-encoder) + diffusion (decoder) hybrid â sitting between cluster A (cross-attention on the demo) and the diffusion-head mainstream, distinct from ICRT's pure next-token sequence. The DrawAnything / LIBERO-Gen benchmarks give the field a cleaner test of unseen-task adaptation.
Limitations (reviewer). Single-demo prompting is powerful but sensitive to prompt-demo quality/coverage; the diversity finding is shown on their tasks/benchmarks; handheld-interface data still needs collection (cheaper than teleop, not free); diffusion decoder adds inference cost vs a single-step head.
4. Links
- Paper: arXiv 2606.30457 ¡ project ¡ code ¡ HF
- Family: In-Context Imitation ¡ ICRT ¡ MimicDroid ¡ RoboSSM
- Data interface kin: DexUMI ¡ YUBI ¡ Dexterous-Hand Data Pyramid
â Back to In-Context Imitation ¡ Reviews ¡ Home