ICLR 2026 Policy Contrastive Decoding - Heungwoo/research GitHub Wiki
Policy Contrastive Decoding โ training-free decoding to kill spurious visual correlations in robot policies
Venue: ICLR 2026 ยท Authors: Shihan Wu, Xu Luo, Ji Zhang, Junlin Xie, Jingkuan Song, Heng Tao Shen, Lianli Gao ยท arXiv:2505.13255 ยท Category: reasoning / inference-time decoding for VLA ยท Trend tag: generalization & robustness.
Approach diagram
flowchart LR
Img[Original visual input] --> Pol1[Robot policy]
ImgM[Object-masked visual input] --> Pol2[Same policy]
Pol1 --> D1[Action distribution p_orig]
Pol2 --> D2[Action distribution p_masked]
D1 --> C{Contrast: p_orig - p_masked}
D2 --> C
C --> Act[Object-grounded action]
Problem
Generalist robot policies (robotic foundation models) tend to learn spurious correlations from pre-training trajectories โ e.g., latching onto backgrounds or table layouts rather than the task-relevant object. This hurts generalization beyond the training distribution.
Method
Policy Contrastive Decoding (PCD) redirects the policy's focus toward object-relevant visual clues by contrasting two action probability distributions: one from the original image and one from an object-masked image. The difference amplifies action components that genuinely depend on the manipulated object and suppresses background-driven ones.
PCD is training-free and works as a plug-in: it requires no fine-tuning and no access to model weights, so it can wrap heterogeneous policy types โ both autoregressive and diffusion-based.
Results
- Evaluated on three open-source policies: OpenVLA (autoregressive), Octo (diffusion), and ฯโ (diffusion/flow).
- On ฯโ: +8.9% in simulation and +108% in real-world manipulation (relative improvement reported by the authors).
- Consistent gains across policy families, in both simulation and real-world settings.
Significance
Demonstrates that contrastive-decoding ideas from VLM hallucination mitigation transfer to action models, offering a cheap, model-agnostic robustness boost without retraining โ a practical lever for deploying existing robot foundation models more reliably.
Links
Related pages
โ Back to ICLR-2026