ICLR 2026 Policy Contrastive Decoding - Heungwoo/research GitHub Wiki

Policy Contrastive Decoding โ€” training-free decoding to kill spurious visual correlations in robot policies

Venue: ICLR 2026 ยท Authors: Shihan Wu, Xu Luo, Ji Zhang, Junlin Xie, Jingkuan Song, Heng Tao Shen, Lianli Gao ยท arXiv:2505.13255 ยท Category: reasoning / inference-time decoding for VLA ยท Trend tag: generalization & robustness.

Approach diagram

flowchart LR
  Img[Original visual input] --> Pol1[Robot policy]
  ImgM[Object-masked visual input] --> Pol2[Same policy]
  Pol1 --> D1[Action distribution p_orig]
  Pol2 --> D2[Action distribution p_masked]
  D1 --> C{Contrast: p_orig - p_masked}
  D2 --> C
  C --> Act[Object-grounded action]

Problem

Generalist robot policies (robotic foundation models) tend to learn spurious correlations from pre-training trajectories โ€” e.g., latching onto backgrounds or table layouts rather than the task-relevant object. This hurts generalization beyond the training distribution.

Method

Policy Contrastive Decoding (PCD) redirects the policy's focus toward object-relevant visual clues by contrasting two action probability distributions: one from the original image and one from an object-masked image. The difference amplifies action components that genuinely depend on the manipulated object and suppresses background-driven ones.

PCD is training-free and works as a plug-in: it requires no fine-tuning and no access to model weights, so it can wrap heterogeneous policy types โ€” both autoregressive and diffusion-based.

Results

  • Evaluated on three open-source policies: OpenVLA (autoregressive), Octo (diffusion), and ฯ€โ‚€ (diffusion/flow).
  • On ฯ€โ‚€: +8.9% in simulation and +108% in real-world manipulation (relative improvement reported by the authors).
  • Consistent gains across policy families, in both simulation and real-world settings.

Significance

Demonstrates that contrastive-decoding ideas from VLM hallucination mitigation transfer to action models, offering a cheap, model-agnostic robustness boost without retraining โ€” a practical lever for deploying existing robot foundation models more reliably.

Links

Related pages

โ† Back to ICLR-2026