ICLR 2026 VP Fail - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 Authors: Jingxian Lu · Wenke Xia · Yuxuan Wu · Zhiwu Lu · Di Hu (Gaoling School of AI, Renmin University of China — GeWu-Lab) arXiv: 2602.12032 Category: Robustness / security for VLA Trend tag: Modality collaboration · gradient balancing · proprioception dominance
flowchart LR
subgraph DIAG[Diagnosis via temporally controlled experiments]
P[Proprioception<br/>concise · fast loss reduction]
V[Vision<br/>needed for target localization]
P -->|dominates optimization| SUP[Visual learning suppressed<br/>during motion-transition phases]
V --> SUP
end
SUP --> FAIL[Policy fails when vision is<br/>actually required]
subgraph GAP[GAP: Gradient Adjustment with Phase-guidance]
EST["Use proprioception to estimate<br/>P(timestep in motion-transition phase)"]
EST --> ADJ[Reduce proprioception gradient<br/>magnitude by estimated probability]
ADJ --> BAL[Dynamic vision-proprioception collaboration]
end
FAIL --> GAP
BAL --> OUT[Robust, generalizable VP policy]
Proprioception gives real-time robot state for precise servo control, and combining it with vision is expected to boost manipulation in complex tasks. Yet prior work reports inconsistent generalization for vision-proprioception policies. The paper diagnoses why through temporally controlled experiments: during motion-transition sub-phases that require target localization, the vision modality plays only a limited role. The root cause is an optimization pathology — the policy gravitates toward the concise proprioceptive signal because it yields faster loss reduction, so proprioception dominates training and suppresses learning of the visual modality exactly when vision matters.
GAP (Gradient Adjustment with Phase-guidance) adaptively modulates proprioception's optimization to enable dynamic collaboration:
- Phase estimation — use proprioception to capture robot state and estimate, for each trajectory timestep, the probability that it belongs to a motion-transition phase.
- Fine-grained gradient adjustment — during policy learning, reduce the magnitude of proprioception's gradient in proportion to the estimated motion-transition probability, forcing the policy to lean on vision where localization is needed.
GAP is a training-time intervention, not an architecture change, so it slots into existing pipelines.
GAP is validated across simulated and real-world environments, one-arm and dual-arm setups, and is compatible with both conventional policies and Vision-Language-Action (VLA) models (MLP-based, diffusion-based, and transformer-based heads). The paper reports robust, generalizable vision-proprioception policies across these settings. (Specific success-rate numbers omitted here pending the camera-ready.)
This work reframes a long-standing puzzle — why adding proprioception sometimes hurts generalization — as a modality-imbalance / gradient-dominance problem during specific task phases, rather than an architecture or data issue. The phase-guided gradient remedy is a lightweight, broadly compatible fix that complements robustness work focused on perturbations rather than modality collaboration.
- arXiv: https://arxiv.org/abs/2602.12032
- OpenReview: https://openreview.net/forum?id=Nbj1GFCKB3
- RobustVLA — perturbation robustness across modalities
- Survey: VLA & Manipulation
← Back to ICLR-2026