RSS 2026 GuidedVLA - Heungwoo/research GitHub Wiki

GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization

Venue: RSS 2026 (Sydney, Jul 13–17) · Session: VLA Models · paper #84 Authors: Xiaosong Jia, Bowen Yang, Zuhao Ge, Xian Nie, Yuchen Zhou, Cunxin Fan, Yufeng Li, Yilin Chai, Chao Jing, Zijian Liang, Qingwen Bu, Haidong Cao, Chao Wu, Qifeng Li, Zhenjie Yang, Chenhe Zhang, Hongyang Li, Zuxuan Wu, Junchi Yan, Yu-Gang Jiang ...more> arXiv: 2605.12369 · program page

Summary compiled from the arXiv paper (v2); all numbers quoted from the paper. Trend context: RSS 2026 survey.

GuidedVLA overview (Figure 1 of arXiv 2605.12369, © the authors)

Top: benchmark gains over the π0 baseline on LIBERO-Plus, RoboTwin 2.0 (77.38→90.63) and two real-world platforms. Bottom panels contrast baseline vs. GuidedVLA on the three specialized factors: skill recognition (correct Move→Sweep→Dump sequencing), object grounding (attention concentrated on the target object instead of scattered), and geometry perception (clean depth prediction).

Problem

Existing VLAs learn task-relevant features only implicitly through end-to-end action supervision, so the action decoder often latches onto spurious correlations — visual shortcuts, background noise — that hurt out-of-domain generalization. GuidedVLA (Fudan TEAI, SJTU, OpenDriveLab/HKU) asks whether explicitly guiding the action decoder toward task-relevant factors fixes this.

Method

The core idea treats the action decoder not as a monolithic learner but as an assembly of functional components: individual cross-attention heads are supervised with auxiliary signals while the remaining heads stay free. The initial instantiation (on a π0 base policy) uses three specialized heads: an object head whose attention is constrained to Grounded-SAM-annotated target/destination regions; a skill head aligning internal features with temporal sub-skill phases; and a depth head distilling features from a depth encoder (no depth annotation needed — injected architecturally). A largely automatic annotation pipeline (Qwen3-VL + SAM2, ~11× faster than manual: ~4 min vs. ~43.5 min per 50 episodes; 95.2% auto accuracy for object masks, 87.3% for skill labels) produces the guidance data.

Results

On LIBERO-Plus, GuidedVLA reaches 75.4% average vs. 68.2% for π0, topping 12 baselines including OpenVLA-OFT (69.6%) and DreamVLA (69.9%); single-head ablations show the object head strongest on the Object suite (82.5%, +8.4 over π0), skill head best on Goal (68.9%), depth head best on Spatial. On RoboTwin 2.0 the full model lifts average success from 77.38% to 90.63% (e.g., Click Bell 35%→63% with the depth head, Beat Hammer Block 78%→96%). On real robots (ALOHA AgileX household tasks; PSI-Bot RealMan chemistry-lab tasks, 20 trials each) it beats the base policy in every setting: 75.8% vs. 55.8% in-domain, 67.5% vs. 44.2% with scene distractors, 79.2% vs. 57.5% under lighting shifts. Factor-quality analyses show success rises monotonically with head quality (e.g., depth-feature ratio: 15.0%→74.2%).

Significance

A plug-and-play middle path between fully implicit end-to-end VLAs and hand-built modular stacks: supervise a few attention heads, keep the rest free. The demonstrated correlation between factor quality and success supports interpretable-by-construction action decoding — relevant to Review-VLA-Architecture and the robustness themes of Review-VLA-Evaluation.

← Back to RSS 2026 survey · RSS-2026-Papers · Home