Review In Context Imitation - Heungwoo/research GitHub Wiki

In-Depth Survey β€” In-Context Imitation & Demo-Following for Robot Policies

Question: how does a policy watch a demonstration and reproduce it at test time β€” ideally without per-task fine-tuning β€” by holding that demo (and its own history) in memory/context? Framing: in-context imitation is fundamentally a memory-conditioning problem. The demo is external context the policy attends to; long-horizon history is self-context. The controlled benchmark for both is RoboMME (its Imitation suite = replicate a demonstrated motion strategy, i.e. procedural memory). Companion: VLA Memory Β· RoboTTT.


1. The idea β€” demo as context, not as new weights

Classic imitation learning trains on demonstrations. In-context imitation instead conditions on a demo held in the model's context/memory and produces actions that copy its strategy β€” the way in-context learning works for LLMs. Two properties define the family:

  • Test-time skill acquisition β€” a new task is specified by showing one (or a few) demos, no gradient step per task.
  • Memory is the mechanism β€” the demo must be encoded and retrieved during rollout, so every method is really a choice of how to store and attend to context.

This is exactly the faculty RoboMME's Imitation suite isolates (MoveCube / InsertPeg / PatternLock / RouteStick β€” pick-place vs push vs hook, linear vs circular), scored as procedural memory.


2. Taxonomy β€” how the demo is conditioned

A. Cross-attention on the demo (video/demo-conditioned policies)

Encode the demo video and cross-attend from the current robot state to it.

  • Vid2Robot (2403.12943) β€” prompt-video encoder + robot-state encoder joined by cross-attention transformers; imitates a human video without per-task fine-tuning, +20% over BC-Z, with cross-object motion transfer. The canonical demo-conditioned policy.
  • VLBiMan β€” vision-language-anchored one-shot demonstration β†’ generalizable bimanual manipulation.
  • See Once, Then Act (2512.07582) β€” VLA that learns a task from a single video demonstration at test time.

B. Recurrent / query memory (history-aware VLAs)

Compress rollout (and demo) history into a recurrent latent or a set of learnable queries.

  • HAMLET β€” "switch your VLA into a history-aware policy."
  • MemoryVLA β€” perceptual-cognitive memory in a VLA.
  • RememVLA (2026) β€” memory via dual-level recurrent queries.
  • ContextVLA (2510.04246) β€” amortized multi-frame context.

C. Fast-weight / Test-Time-Training (compress the demo into weights)

Turn context into a parametric recurrent state updated by gradient descent at inference.

  • RoboTTT (NVIDIA GEAR) β€” 8K-timestep context via TTT fast weights in GR00T N1.7's DiT; enables one-shot in-context imitation from a human video (Circuit 6/10 vs a baseline's 0/10) at constant latency. The strongest recent unification of long memory + demo-following.

D. Retrieval-augmented (fetch the relevant memory/demo on demand)

Store many experiences/keyframes and retrieve the ones relevant to the current step.

E. Token-sequence ICL (next-token prediction over demo + rollout)

Treat [demo tokens; rollout tokens] as one sequence and predict the next action token β€” LLM-style ICL.

  • In-Context Robot Transformer / ICRT (ICRA 2025) β€” in-context imitation via next-token prediction over sensorimotor tokens (the cluster-E anchor). Keypoint Action Tokens (KAT) and Instant Policy are related graph/token-ICL variants.
  • Behavior Prompting Policy (2606.30457) β€” demonstrations as prompts for manipulation; finds task diversity (not quantity) drives the prompting ability.

F. Play-video in-context (learn from unstructured human play)

  • MimicDroid (ICRA 2026) β€” in-context learning for humanoid manipulation from human play videos β€” learns to ICL from unlabeled video, no teleop; ~2Γ— real success.

3. The benchmark that ties it together β€” RoboMME

RoboMME turns "which memory helps which task?" into a controlled sweep: 14 memory-augmented policies on one Ο€0.5 backbone (3 memory representations Γ— 3 integration mechanisms). Its Imitation (procedural) suite is the in-context-imitation testbed; its temporal/spatial/object suites test self-history memory. Headline: no single memory design wins everywhere (best non-oracle = FrameSamp+Modulator, 44.5% avg; humans 90.5%) β€” so the Β§2 mechanisms are complementary, not competing. Design-space detail: VLA Memory.


4. Comparison

Approach Demo / context form Conditioning mechanism No per-task FT? Memory representation
Vid2Robot human prompt video cross-attention (A) βœ… encoded demo tokens
[VLBiMan](/Heungwoo/research/wiki/ICLR-2026-VLBiMan) one-shot demo vision-language anchor (A) βœ… anchored demo
[RoboTTT](/Heungwoo/research/wiki/Review-RoboTTT) human video (one-shot) fast-weight TTT (C) βœ… parametric (fast weights)
[HAMLET](/Heungwoo/research/wiki/ICLR-2026-HAMLET) / [MemoryVLA](/Heungwoo/research/wiki/ICLR-2026-MemoryVLA) own history recurrent / cognitive memory (B) n/a (history) latent / token buffer
[MemER](/Heungwoo/research/wiki/ICLR-2026-Memory-Experience-Retrieval) / [MAP-VLA](/Heungwoo/research/wiki/ICRA-2026-MAP-VLA) stored experiences retrieval (D) βœ… (retrieve) external memory store
[ICRT](/Heungwoo/research/wiki/Review-ICRT) / [Behavior Prompting](/Heungwoo/research/wiki/Review-Behavior-Prompting) demo token sequence next-token / cross-attn+diffusion ICL (E) βœ… in-context tokens
[MimicDroid](/Heungwoo/research/wiki/Review-MimicDroid) human play video in-context (F) βœ… in-context

(βœ… = specifies a new task by showing a demo, no gradient step per task.)


5. Design axes β€” when to use what

  • Clean expert video demo, want motion transfer? β†’ cross-attention (Vid2Robot, VLBiMan).
  • Very long demo/rollout, must stay real-time? β†’ fast-weight/TTT (RoboTTT) β€” memory is parametric, latency constant.
  • Many past experiences, need the relevant one? β†’ retrieval (MemER, MAP-VLA, KEMO).
  • Long-horizon own-history dependence (counting, re-finding)? β†’ recurrent/keyframe memory (HAMLET, Long-Context-IL) β€” and consult RoboMME for which representation fits which memory type.
  • Unstructured/noisy demos (human play)? β†’ play-video ICL (MimicDroid).

Unifying view: in-context imitation and history-memory are the same operation β€” attend to non-current context β€” differing only in whose context (a demonstrator's vs the robot's own) and where it's stored (attention tokens Β· recurrent latent Β· fast weights Β· external store).


6. Open challenges

  1. One demo β‰  robust policy β€” sensitivity to demo viewpoint, object identity, and speed; generalization beyond the shown instance is uneven.
  2. Context length vs latency β€” attention over long demos is quadratic; fast-weight/retrieval are the escape hatches but under-benchmarked at scale (RoboTTT Β§6).
  3. Which memory for which task is unsolved β€” RoboMME shows no universal winner; humans still lead 90.5% vs ~44%.
  4. Human→robot demo gap — video demos carry the same embodiment/retargeting gap as egocentric-video pretraining.
  5. Evaluation β€” few benchmarks isolate in-context skill acquisition from ordinary multi-task competence (RoboMME's Imitation suite is a start).

7. Links

← Back to Reviews Β· Home