ICLR 2026 VITA - Heungwoo/research GitHub Wiki

VITA — Zero-Shot Value Functions via Test-Time Adaptation of VLMs

Venue: ICLR 2026 Authors: Christos Ziakas, Alessandra Russo (Imperial College London) Source: arXiv 2506.10085 Ā· project page Ā· OpenReview Category: RL for VLA — Reward Modeling Trend tag: Trend 3

Approach diagram

flowchart LR
  Traj[Trajectory frames<br/>+ goal description] --> TTA[Test-time adaptation<br/>gradient step on meta-learned<br/>self-supervised loss, no labels]
  VLM[Frozen CLIP encoder +<br/>small adaptation MLP] --> TTA
  TTA --> AdaptedVLM[Adapted module<br/>= zero-shot goal-conditioned<br/>value function]
  AdaptedVLM --> R[Estimated task progress<br/>= reward signal]
  R --> RL[Reward shaping for<br/>offline RL policy]
Loading

Problem

RL needs reward functions, but hand-designing rewards for each new task is slow and error-prone. Learned rewards typically need lots of labeled data, defeating the purpose of zero-shot deployment.

Method

Turn a frozen contrastive VLM (OpenCLIP ViT-B/32) into a zero-shot goal-conditioned value function that estimates task progress from a goal description and current observation. The key idea is test-time adaptation (TTA): a small adaptation module (a two-layer residual MLP with GELU, projection dim d′=64) is updated at inference via a gradient step on a meta-learned self-supervised loss. The self-supervised objective is itself meta-learned (gradient-based meta-learning) so that each test-time update provably improves downstream value estimation, rather than relying on a hand-picked auxiliary task.

By applying these updates sequentially over a trajectory, VITA encodes history into the module's parameters, giving the otherwise frame-independent CLIP encoder temporal reasoning. To prevent the adaptation from latching onto spurious shortcuts, training uses a dissimilarity-based sampling strategy that selects semantically diverse trajectory segments.

Results

  • Real-world manipulation (Value-Order Correlation, VOC): generalizing from a single training environment, VITA scores 0.782 in-distribution, 0.725 under environment shift, and 0.820 under embodiment shift — far above the autoregressive-VLM baseline GVL (Gemini 1.5 Pro), which sits around 0.21–0.31 across the same settings.
  • Meta-World MT10 offline RL: using VITA's zero-shot value estimates for reward shaping yields a multi-task policy with IQM 0.815 [0.785, 0.838], exceeding both CLIP-based baselines (VLM-CL, VLM-RM, CLIP-FT, CLIP-GRU) and the simulator's hand-engineered fuzzy-logic dense rewards (0.779).

The decisive result is that a zero-shot VLM-derived reward beats a task-specific hand-designed dense reward, validating that VLMs encode enough task-progress structure to serve as reward models — but only once TTA supplies the missing temporal/generalization signal.

Significance

Removes the reward-design bottleneck that limited RL-based VLA fine-tuning. Combined with VLA-RFT (environment) and PLD (training loop), VITA provides the missing third component — the reward.

Links

Related pages

← Back to ICLR-2026 Ā· Topic: RL

āš ļø **GitHub.com Fallback** āš ļø