ICLR 2026 VP Fail - Heungwoo/research GitHub Wiki

VP-Fail — When would Vision-Proprioception Policies Fail in Robotic Manipulation?

Venue: ICLR 2026 Authors: Jingxian Lu · Wenke Xia · Yuxuan Wu · Zhiwu Lu · Di Hu (Gaoling School of AI, Renmin University of China — GeWu-Lab) arXiv: 2602.12032 Category: Robustness / security for VLA Trend tag: Modality collaboration · gradient balancing · proprioception dominance

Approach diagram

flowchart LR
  subgraph DIAG[Diagnosis via temporally controlled experiments]
    P[Proprioception<br/>concise · fast loss reduction]
    V[Vision<br/>needed for target localization]
    P -->|dominates optimization| SUP[Visual learning suppressed<br/>during motion-transition phases]
    V --> SUP
  end
  SUP --> FAIL[Policy fails when vision is<br/>actually required]
  subgraph GAP[GAP: Gradient Adjustment with Phase-guidance]
    EST["Use proprioception to estimate<br/>P(timestep in motion-transition phase)"]
    EST --> ADJ[Reduce proprioception gradient<br/>magnitude by estimated probability]
    ADJ --> BAL[Dynamic vision-proprioception collaboration]
  end
  FAIL --> GAP
  BAL --> OUT[Robust, generalizable VP policy]
Loading

Problem

Proprioception gives real-time robot state for precise servo control, and combining it with vision is expected to boost manipulation in complex tasks. Yet prior work reports inconsistent generalization for vision-proprioception policies. The paper diagnoses why through temporally controlled experiments: during motion-transition sub-phases that require target localization, the vision modality plays only a limited role. The root cause is an optimization pathology — the policy gravitates toward the concise proprioceptive signal because it yields faster loss reduction, so proprioception dominates training and suppresses learning of the visual modality exactly when vision matters.

Method

GAP (Gradient Adjustment with Phase-guidance) adaptively modulates proprioception's optimization to enable dynamic collaboration:

  1. Phase estimation — use proprioception to capture robot state and estimate, for each trajectory timestep, the probability that it belongs to a motion-transition phase.
  2. Fine-grained gradient adjustment — during policy learning, reduce the magnitude of proprioception's gradient in proportion to the estimated motion-transition probability, forcing the policy to lean on vision where localization is needed.

GAP is a training-time intervention, not an architecture change, so it slots into existing pipelines.

Results

GAP is validated across simulated and real-world environments, one-arm and dual-arm setups, and is compatible with both conventional policies and Vision-Language-Action (VLA) models (MLP-based, diffusion-based, and transformer-based heads). The paper reports robust, generalizable vision-proprioception policies across these settings. (Specific success-rate numbers omitted here pending the camera-ready.)

Significance

This work reframes a long-standing puzzle — why adding proprioception sometimes hurts generalization — as a modality-imbalance / gradient-dominance problem during specific task phases, rather than an architecture or data issue. The phase-guided gradient remedy is a lightweight, broadly compatible fix that complements robustness work focused on perturbations rather than modality collaboration.

Links

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️