RSS 2026 GuidedVLA - Heungwoo/research GitHub Wiki
GuidedVLA: Specifying Task-Relevant Factors via Plug-and-Play Action Attention Specialization
Venue: RSS 2026 (Sydney, Jul 13–17) · Session: VLA Models · paper #84 Authors: Xiaosong Jia, Bowen Yang, Zuhao Ge, Xian Nie, Yuchen Zhou, Cunxin Fan, Yufeng Li, Yilin Chai, Chao Jing, Zijian Liang, Qingwen Bu, Haidong Cao, Chao Wu, Qifeng Li, Zhenjie Yang, Chenhe Zhang, Hongyang Li, Zuxuan Wu, Junchi Yan, Yu-Gang Jiang ...more> arXiv: 2605.12369 · program page
Summary compiled from the arXiv paper (v2); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Top: benchmark gains over the π0 baseline on LIBERO-Plus, RoboTwin 2.0 (77.38→90.63) and two real-world platforms. Bottom panels contrast baseline vs. GuidedVLA on the three specialized factors: skill recognition (correct Move→Sweep→Dump sequencing), object grounding (attention concentrated on the target object instead of scattered), and geometry perception (clean depth prediction).
Problem
Existing VLAs learn task-relevant features only implicitly through end-to-end action supervision, so the action decoder often latches onto spurious correlations — visual shortcuts, background noise — that hurt out-of-domain generalization. GuidedVLA (Fudan TEAI, SJTU, OpenDriveLab/HKU) asks whether explicitly guiding the action decoder toward task-relevant factors fixes this.
Method
The core idea treats the action decoder not as a monolithic learner but as an assembly of functional components: individual cross-attention heads are supervised with auxiliary signals while the remaining heads stay free. The initial instantiation (on a π0 base policy) uses three specialized heads: an object head whose attention is constrained to Grounded-SAM-annotated target/destination regions; a skill head aligning internal features with temporal sub-skill phases; and a depth head distilling features from a depth encoder (no depth annotation needed — injected architecturally). A largely automatic annotation pipeline (Qwen3-VL + SAM2, ~11× faster than manual: ~4 min vs. ~43.5 min per 50 episodes; 95.2% auto accuracy for object masks, 87.3% for skill labels) produces the guidance data.
Results
On LIBERO-Plus, GuidedVLA reaches 75.4% average vs. 68.2% for π0, topping 12 baselines including OpenVLA-OFT (69.6%) and DreamVLA (69.9%); single-head ablations show the object head strongest on the Object suite (82.5%, +8.4 over π0), skill head best on Goal (68.9%), depth head best on Spatial. On RoboTwin 2.0 the full model lifts average success from 77.38% to 90.63% (e.g., Click Bell 35%→63% with the depth head, Beat Hammer Block 78%→96%). On real robots (ALOHA AgileX household tasks; PSI-Bot RealMan chemistry-lab tasks, 20 trials each) it beats the base policy in every setting: 75.8% vs. 55.8% in-domain, 67.5% vs. 44.2% with scene distractors, 79.2% vs. 57.5% under lighting shifts. Factor-quality analyses show success rises monotonically with head quality (e.g., depth-feature ratio: 15.0%→74.2%).
Significance
A plug-and-play middle path between fully implicit end-to-end VLAs and hand-built modular stacks: supervise a few attention heads, keep the rest free. The demonstrated correlation between factor quality and success supports interpretable-by-construction action decoding — relevant to Review-VLA-Architecture and the robustness themes of Review-VLA-Evaluation.
← Back to RSS 2026 survey · RSS-2026-Papers · Home