ICML 2026 RA VLA - Heungwoo/research GitHub Wiki
RA-VLA: Retrieval-Augmented VLA for Test-Time Adaptation — Training-free adaptation via behavior-aligned retrieval
Venue: ICML 2026 (Poster) Category: VLA Architecture
Vision-Language-Action (VLA) models provide a versatile foundation for general robotic manipulation, yet they exhibit significant brittleness when confronted with novel task distributions. In-Context Imitation Learning (ICIL) frameworks promise training-free adaptation by conditioning on expert demonstrations at inference time, but they suffer from an adaptation bottleneck that hinders the effective translation of expert context into actions. The authors trace this bottleneck to two causes: inadequate retrieval mechanisms that surface poorly matched context, and behavioral inertia — the model's tendency to ignore retrieved cues and fall back on its prior policy.
RA-VLA is a retrieval-augmented VLA framework for training-free test-time adaptation. It combines behavior-aligned context retrieval with a grounded execution pipeline. Rather than retrieving demonstrations by surface visual similarity alone, the retrieval is aligned to behavior so that the surfaced context is functionally relevant to the action the policy must take. The execution pipeline then enforces strict adherence to functional cues within a scalable architecture, counteracting behavioral inertia so that retrieved expert context is actually reflected in the emitted actions — all while preserving inference efficiency.
flowchart LR
A[Novel task observation] --> B[Behavior-aligned<br/>context retrieval]
C[(Expert<br/>demonstrations)] --> B
B --> D[Grounded execution pipeline<br/>strict adherence to functional cues]
D --> E[Adapted action]
RA-VLA is evaluated on the LIBERO benchmark and a real-world UR5e environment. Across these settings it achieves superior success rates and computational efficiency compared to prior in-context imitation learning approaches, demonstrating effective training-free robotic adaptation that scales without sacrificing inference speed.
RA-VLA reframes test-time adaptation for VLA models as a retrieval-plus-grounding problem rather than a fine-tuning problem. By diagnosing the adaptation bottleneck (weak retrieval and behavioral inertia) and addressing both with behavior-aligned retrieval and a cue-adherent execution pipeline, it offers a training-free path to deploy general VLA policies on novel tasks — validated in both simulation and on real hardware.
- ICML 2026: https://icml.cc/virtual/2026/poster/60980
← Back to ICML-2026