ICML 2026 Contrastive Representation Regularization for Vision Language Action - Heungwoo/research GitHub Wiki
Contrastive Representation Regularization for VLAs — aligning VLA features with proprioceptive state
Venue: ICML 2026 (Poster) Category: VLA Architecture Traction (2026-06): 7 citations (arXiv)

Problem
Vision-Language-Action (VLA) models inherit rich visual and semantic grounding from pre-trained Vision-Language Models (VLMs), but those representations "arguably remain suboptimal, lacking sensitivity to robotic signals such as control actions and proprioceptive states." Standard VLAs train a generative action decoder conditioned on VLM embeddings purely through an action-prediction objective, which provides no explicit pressure for the upstream representation to encode control-relevant geometry. The paper argues this gap hurts precise behaviors such as positioning during grasping and placing.
Method
The paper introduces Robot State-aware Contrastive Loss (RS-CL), a lightweight representation regularizer added alongside the usual action-prediction loss and fully compatible with standard VLA training pipelines. Key ingredients:
- Soft, state-distance supervision. RS-CL is a weighted variant of InfoNCE. Rather than hard positive/negative labels, the soft weights
w_ijare computed from the relative distance between proprioceptive statesq_i, q_jwithin a batch, so representations of frames with similar robot states are pulled together in proportion to state similarity. - Projection head. A 2-layer MLP projection head
g_ψ(hidden dim 2048, projection dim 128) maps VLM embeddings into the contrastive space, where cosine similarity and a temperatureτcontrol sharpness. - Amortized VLM embeddings. Contrastive learning is applied over the sequence of VLM embeddings, with a representation-level augmentation strategy for cheap representation learning.
- Annealed weighting. The RS-CL coefficient
λis initialized to 1.0 and decayed to 0 on a cosine schedule, emphasizing representation refinement early in training and handing off to the action objective later.

The method is applied on top of state-of-the-art VLAs including π0 and π0-FAST (reproduced from openpi pi0_fast_base, trained 60K / 30K steps, batch size 64, lr 2.5e-5 cosine-decayed, action horizon H=16).
Results
RS-CL "pushes the prior art from 30.8% to 41.5% on pick-and-place tasks in RoboCasa-Kitchen," and on the overall benchmark reaches 69.7% (+4.0%) with 30/100/300 demonstrations. The gain is concentrated on precision-critical pick-and-place, improving 30.3% to 41.5% (+11.2%). On real-robot hardware the success rate rises from 45.0% to 58.3% (+13.3%). Headline: "69.7% on RoboCasa-Kitchen; real-robot success 45.0% to 58.3%".
Significance
RS-CL is a drop-in, near-zero-cost auxiliary loss that injects proprioceptive structure into VLM-derived representations without changing the decoder or data pipeline. In a 2026 landscape where most VLA progress comes from scaling data or swapping action heads, it shows that simply regularizing representations toward robot state yields large, precision-relevant gains—especially for fine grasping and placing where current VLAs are weakest.
Links
- arXiv: 2510.01711
- ICML 2026: https://icml.cc/virtual/2026/poster/62819
← Back to ICML-2026