ICML 2026 FOCA - Heungwoo/research GitHub Wiki

FOCA — Future-Oriented Conditioning for data-efficient VLA adaptation

Venue: ICML 2026 (Poster) Category: VLA Architecture Affiliations: Duc Nguyen, Nghiem Diep, Binh Nguyen Gia, Trong-Bao Ho, Doanh Le Thien, Quang Nguyen, Thien-Loc Ha, Tran Van Nhiem, Bao Thach, An Thai Le, Daniel Sonntag, Mathias Niepert, Vien Ngo, et al.

Problem

Vision–Language–Action (VLA) models enable general-purpose robotic control via large-scale multimodal pretraining, yet their effectiveness under few-shot imitation learning remains limited. When only a handful of demonstrations are available per task, adaptation quality degrades sharply — the model struggles with long-horizon reasoning and tends to overfit the short, narrow demonstration set rather than learning where the task is headed.

Method

FOCA (Future-Oriented Conditioning for data-efficient Adaptation) augments VLA fine-tuning with a forward-looking objective that combines two complementary signals:

  1. Explicit prediction of task-grounded future interaction embeddings — the model is trained to anticipate compact embeddings of upcoming interactions rather than raw pixels.
  2. Implicit alignment to future goal observations, giving the policy long-horizon reasoning without any expensive pixel-level video prediction.

Conceptually, this amounts to learning a future-conditioned, value-like representation: the policy is shaped by where the trajectory is going, not just the immediate next action. Crucially, the framework also supports action-free co-training with synthetic videos generated by video world models, letting FOCA absorb additional future-dynamics signal without needing extra teleoperated action labels.

flowchart LR
    O[Current obs + language] --> VLA[VLA backbone]
    VLA --> A[Action head]
    VLA --> F[Future interaction embedding predictor]
    F -. explicit prediction .-> FE[Task-grounded future embedding]
    VLA -. implicit alignment .-> G[Future goal observation]
    WM[Video world model<br/>synthetic videos] -. action-free co-training .-> F
Loading

Results

  • 95.7% success with only 20 demonstrations on LIBERO, demonstrating strong data efficiency in the few-shot regime.
  • 7–12% improvements on RoboCasa over baselines.
  • Up to 26% absolute gains on real robots, indicating the future-oriented conditioning transfers beyond simulation.

These results consistently show that conditioning on future interaction/goal representations recovers much of the performance lost when demonstrations are scarce.

Significance

FOCA targets one of the most practical bottlenecks for deploying VLAs: adapting a pretrained model to a new task from only a few demonstrations. By replacing costly pixel-prediction world-modeling with lightweight future-embedding prediction and goal alignment — and by allowing action-free co-training on synthetic world-model videos — FOCA offers a data-efficient adaptation recipe that improves both simulated and real-robot few-shot imitation.

Links

← Back to ICML-2026

⚠️ **GitHub.com Fallback** ⚠️