Review ICRT - Heungwoo/research GitHub Wiki
In-Depth Review ā ICRT: In-Context Imitation Learning via Next-Token Prediction
Paper: "In-Context Imitation Learning via Next-Token Prediction" ā arXiv 2408.15980 Ā· ICRA 2025 Ā· UC Berkeley (Letian Fu et al.) Ā· code. The foundational token-sequence datapoint of in-context imitation ā a causal transformer over sensorimotor trajectories that learns a new task from a prompt of demonstrations at test time, no fine-tuning, no language, no reward. Companions: In-Context Imitation (this is its cluster-E anchor) Ā· RoboSSM Ā· Behavior Prompting Ā· MimicDroid.

1. Problem
Robots should learn a new task by being shown a few demonstrations at test time ā the way LLMs do in-context learning ā without updating policy parameters and without language or reward supervision. The question ICRT answers: can next-token prediction over raw sensorimotor trajectories give a real robot this in-context ability?
2. Method
ICRT is a causal transformer trained by autoregressive prediction on sensorimotor trajectories (image observations, proprioceptive states, actions) ā no linguistic data, no reward.
- Tokenization: multi-view images (third-person + wrist) ā ViT ā attention pooling ā a state token
f_s; proprioception ā MLP; each action ā MLP āf_a. A trajectory becomes an interleaved[f_s¹, f_a¹, f_s², f_a², ā¦]sequence. - Prompt-then-rollout as one sequence: at test time the model is prompted with the new task's trajectory tokens (
šÆ_prompts), then continues the sequence ā autoregressively emitting actions for the current rollout (šÆā, šÆā, ā¦). This is training-free task specification ā the task is the prompt, not a gradient step. - Training data: teleoperated multi-task trajectories; the model learns how to imitate from context, not any single task.
3. Results
- On a Franka Emika robot, ICRT adapts to new tasks specified purely by the prompt, even in environment configurations different from both the prompt and the training data.
- In a multitask setup, ICRT significantly outperforms prior next-token-prediction robot models on generalizing to unseen tasks.
4. Why it matters (in-context lens)
ICRT is the anchor of cluster E (token-sequence ICL) in In-Context Imitation ā it established that LLM-style next-token prediction over [demo; rollout] tokens gives a real robot genuine in-context imitation, with no language, no reward, no test-time training. Everything after refines this recipe:
- RoboSSM swaps ICRT's Transformer for an SSM to fix its long-prompt collapse (in RoboSSM's own experiments ICRT degrades to ~0 as demos grow past training length ā the direct motivation).
- ICLR: Visual Reasoning adds image-space intent traces to ICRT-style prompts (action-only ā action+intent).
- Behavior Prompting and MimicDroid attack ICRT's data bottleneck (teleop ā handheld interface / human play video).
So ICRT is best read as the reference point the newer in-context papers improve on ā architecture (SSM), prompt content (visual reasoning), and data source (video/handheld).
Limitations. Transformer prompt length caps how many/long demos help ā degrades past the training length (RoboSSM's finding); needs teleoperated training trajectories (the scalability wall Behavior Prompting/MimicDroid target); no language conditioning, so tasks must be shown, not told.
5. Links
- Paper: arXiv 2408.15980 Ā· HF Ā· code Ā· IEEE Xplore
- Family: In-Context Imitation (cluster E) Ā· RoboSSM Ā· Behavior Prompting Ā· MimicDroid Ā· ICLR: Visual Reasoning
ā Back to In-Context Imitation Ā· Reviews Ā· Home