Review In Context Imitation - Heungwoo/research GitHub Wiki
In-Depth Survey β In-Context Imitation & Demo-Following for Robot Policies
Question: how does a policy watch a demonstration and reproduce it at test time β ideally without per-task fine-tuning β by holding that demo (and its own history) in memory/context? Framing: in-context imitation is fundamentally a memory-conditioning problem. The demo is external context the policy attends to; long-horizon history is self-context. The controlled benchmark for both is RoboMME (its Imitation suite = replicate a demonstrated motion strategy, i.e. procedural memory). Companion: VLA Memory Β· RoboTTT.
1. The idea β demo as context, not as new weights
Classic imitation learning trains on demonstrations. In-context imitation instead conditions on a demo held in the model's context/memory and produces actions that copy its strategy β the way in-context learning works for LLMs. Two properties define the family:
- Test-time skill acquisition β a new task is specified by showing one (or a few) demos, no gradient step per task.
- Memory is the mechanism β the demo must be encoded and retrieved during rollout, so every method is really a choice of how to store and attend to context.
This is exactly the faculty RoboMME's Imitation suite isolates (MoveCube / InsertPeg / PatternLock / RouteStick β pick-place vs push vs hook, linear vs circular), scored as procedural memory.
2. Taxonomy β how the demo is conditioned
A. Cross-attention on the demo (video/demo-conditioned policies)
Encode the demo video and cross-attend from the current robot state to it.
- Vid2Robot (2403.12943) β prompt-video encoder + robot-state encoder joined by cross-attention transformers; imitates a human video without per-task fine-tuning, +20% over BC-Z, with cross-object motion transfer. The canonical demo-conditioned policy.
- VLBiMan β vision-language-anchored one-shot demonstration β generalizable bimanual manipulation.
- See Once, Then Act (2512.07582) β VLA that learns a task from a single video demonstration at test time.
B. Recurrent / query memory (history-aware VLAs)
Compress rollout (and demo) history into a recurrent latent or a set of learnable queries.
- HAMLET β "switch your VLA into a history-aware policy."
- MemoryVLA β perceptual-cognitive memory in a VLA.
- RememVLA (2026) β memory via dual-level recurrent queries.
- ContextVLA (2510.04246) β amortized multi-frame context.
C. Fast-weight / Test-Time-Training (compress the demo into weights)
Turn context into a parametric recurrent state updated by gradient descent at inference.
- RoboTTT (NVIDIA GEAR) β 8K-timestep context via TTT fast weights in GR00T N1.7's DiT; enables one-shot in-context imitation from a human video (Circuit 6/10 vs a baseline's 0/10) at constant latency. The strongest recent unification of long memory + demo-following.
D. Retrieval-augmented (fetch the relevant memory/demo on demand)
Store many experiences/keyframes and retrieve the ones relevant to the current step.
- Memory Retrieval in Visuomotor Policies (long-horizon control) Β· MemER (scale memory via experience retrieval) Β· MAP-VLA (memory-augmented prompting) Β· KEMO (2606.23589, event-driven keyframe memory) Β· Long-Context IL (focus on key history frames).
E. Token-sequence ICL (next-token prediction over demo + rollout)
Treat [demo tokens; rollout tokens] as one sequence and predict the next action token β LLM-style ICL.
- In-Context Robot Transformer / ICRT (ICRA 2025) β in-context imitation via next-token prediction over sensorimotor tokens (the cluster-E anchor). Keypoint Action Tokens (KAT) and Instant Policy are related graph/token-ICL variants.
- Behavior Prompting Policy (2606.30457) β demonstrations as prompts for manipulation; finds task diversity (not quantity) drives the prompting ability.
F. Play-video in-context (learn from unstructured human play)
- MimicDroid (ICRA 2026) β in-context learning for humanoid manipulation from human play videos β learns to ICL from unlabeled video, no teleop; ~2Γ real success.
3. The benchmark that ties it together β RoboMME
RoboMME turns "which memory helps which task?" into a controlled sweep: 14 memory-augmented policies on one Ο0.5 backbone (3 memory representations Γ 3 integration mechanisms). Its Imitation (procedural) suite is the in-context-imitation testbed; its temporal/spatial/object suites test self-history memory. Headline: no single memory design wins everywhere (best non-oracle = FrameSamp+Modulator, 44.5% avg; humans 90.5%) β so the Β§2 mechanisms are complementary, not competing. Design-space detail: VLA Memory.
4. Comparison
| Approach | Demo / context form | Conditioning mechanism | No per-task FT? | Memory representation |
|---|---|---|---|---|
| Vid2Robot | human prompt video | cross-attention (A) | β | encoded demo tokens |
| [VLBiMan](/Heungwoo/research/wiki/ICLR-2026-VLBiMan) | one-shot demo | vision-language anchor (A) | β | anchored demo |
| [RoboTTT](/Heungwoo/research/wiki/Review-RoboTTT) | human video (one-shot) | fast-weight TTT (C) | β | parametric (fast weights) |
| [HAMLET](/Heungwoo/research/wiki/ICLR-2026-HAMLET) / [MemoryVLA](/Heungwoo/research/wiki/ICLR-2026-MemoryVLA) | own history | recurrent / cognitive memory (B) | n/a (history) | latent / token buffer |
| [MemER](/Heungwoo/research/wiki/ICLR-2026-Memory-Experience-Retrieval) / [MAP-VLA](/Heungwoo/research/wiki/ICRA-2026-MAP-VLA) | stored experiences | retrieval (D) | β (retrieve) | external memory store |
| [ICRT](/Heungwoo/research/wiki/Review-ICRT) / [Behavior Prompting](/Heungwoo/research/wiki/Review-Behavior-Prompting) | demo token sequence | next-token / cross-attn+diffusion ICL (E) | β | in-context tokens |
| [MimicDroid](/Heungwoo/research/wiki/Review-MimicDroid) | human play video | in-context (F) | β | in-context |
(β = specifies a new task by showing a demo, no gradient step per task.)
5. Design axes β when to use what
- Clean expert video demo, want motion transfer? β cross-attention (Vid2Robot, VLBiMan).
- Very long demo/rollout, must stay real-time? β fast-weight/TTT (RoboTTT) β memory is parametric, latency constant.
- Many past experiences, need the relevant one? β retrieval (MemER, MAP-VLA, KEMO).
- Long-horizon own-history dependence (counting, re-finding)? β recurrent/keyframe memory (HAMLET, Long-Context-IL) β and consult RoboMME for which representation fits which memory type.
- Unstructured/noisy demos (human play)? β play-video ICL (MimicDroid).
Unifying view: in-context imitation and history-memory are the same operation β attend to non-current context β differing only in whose context (a demonstrator's vs the robot's own) and where it's stored (attention tokens Β· recurrent latent Β· fast weights Β· external store).
6. Open challenges
- One demo β robust policy β sensitivity to demo viewpoint, object identity, and speed; generalization beyond the shown instance is uneven.
- Context length vs latency β attention over long demos is quadratic; fast-weight/retrieval are the escape hatches but under-benchmarked at scale (RoboTTT Β§6).
- Which memory for which task is unsolved β RoboMME shows no universal winner; humans still lead 90.5% vs ~44%.
- Humanβrobot demo gap β video demos carry the same embodiment/retargeting gap as egocentric-video pretraining.
- Evaluation β few benchmarks isolate in-context skill acquisition from ordinary multi-task competence (RoboMME's Imitation suite is a start).
7. Links
- Demo-conditioned / one-shot: RoboTTT Β· VLBiMan Β· Vid2Robot (2403.12943) Β· See-Once-Then-Act (2512.07582) Β· Behavior Prompting (2606.30457) Β· MimicDroid (2026)
- IROS 2026 π: ICLR: In-Context Imitation with Visual Reasoning (EΓvisual-reasoning β image-space intent traces, USC) Β· RoboSSM (scalable in-context imitation via state-space models β linear-cost long context) β context: IROS 2026 survey Β§5.2
- Memory-augmented VLAs: HAMLET Β· MemoryVLA Β· MemER Β· MAP-VLA Β· Memory Retrieval Β· Long-Context IL Β· KEMO (2606.23589)
- Benchmark & synthesis: RoboMME Β· VLA Memory Β· VLA Architectures