ICML 2026 FOCA - Heungwoo/research GitHub Wiki
Venue: ICML 2026 (Poster) Category: VLA Architecture Affiliations: Duc Nguyen, Nghiem Diep, Binh Nguyen Gia, Trong-Bao Ho, Doanh Le Thien, Quang Nguyen, Thien-Loc Ha, Tran Van Nhiem, Bao Thach, An Thai Le, Daniel Sonntag, Mathias Niepert, Vien Ngo, et al.
Vision–Language–Action (VLA) models enable general-purpose robotic control via large-scale multimodal pretraining, yet their effectiveness under few-shot imitation learning remains limited. When only a handful of demonstrations are available per task, adaptation quality degrades sharply — the model struggles with long-horizon reasoning and tends to overfit the short, narrow demonstration set rather than learning where the task is headed.
FOCA (Future-Oriented Conditioning for data-efficient Adaptation) augments VLA fine-tuning with a forward-looking objective that combines two complementary signals:
- Explicit prediction of task-grounded future interaction embeddings — the model is trained to anticipate compact embeddings of upcoming interactions rather than raw pixels.
- Implicit alignment to future goal observations, giving the policy long-horizon reasoning without any expensive pixel-level video prediction.
Conceptually, this amounts to learning a future-conditioned, value-like representation: the policy is shaped by where the trajectory is going, not just the immediate next action. Crucially, the framework also supports action-free co-training with synthetic videos generated by video world models, letting FOCA absorb additional future-dynamics signal without needing extra teleoperated action labels.
flowchart LR
O[Current obs + language] --> VLA[VLA backbone]
VLA --> A[Action head]
VLA --> F[Future interaction embedding predictor]
F -. explicit prediction .-> FE[Task-grounded future embedding]
VLA -. implicit alignment .-> G[Future goal observation]
WM[Video world model<br/>synthetic videos] -. action-free co-training .-> F
- 95.7% success with only 20 demonstrations on LIBERO, demonstrating strong data efficiency in the few-shot regime.
- 7–12% improvements on RoboCasa over baselines.
- Up to 26% absolute gains on real robots, indicating the future-oriented conditioning transfers beyond simulation.
These results consistently show that conditioning on future interaction/goal representations recovers much of the performance lost when demonstrations are scarce.
FOCA targets one of the most practical bottlenecks for deploying VLAs: adapting a pretrained model to a new task from only a few demonstrations. By replacing costly pixel-prediction world-modeling with lightweight future-embedding prediction and goal alignment — and by allowing action-free co-training on synthetic world-model videos — FOCA offers a data-efficient adaptation recipe that improves both simulated and real-robot few-shot imitation.
- ICML 2026: https://icml.cc/virtual/2026/poster/66754
← Back to ICML-2026