RSS 2026 SkillVLA - Heungwoo/research GitHub Wiki

SkillVLA: Tackling Combinatorial Diversity in Dual-Arm Manipulation via Skill Reuse

Venue: RSS 2026 (Sydney, Jul 13–17) · Session: VLA Models · paper #82 Authors: Xuanran Zhai, Zekai Huang, Longyan Wu, Qianyou Zhao, Qiaojun Yu, Jieji Ren, Ce Hao, Harold Soh arXiv: 2603.03836 · program page

Summary compiled from the arXiv paper (v1); all numbers quoted from the paper. Trend context: RSS 2026 survey.

SkillVLA overview (Figure 1 of arXiv 2603.03836, © the authors)

Figure 1: from a dataset of dual-arm demonstrations SkillVLA identifies per-arm skills (stir/cake/smash objects; lift/shake per arm), learns them with a high-level VLM plus separate left/right low-level VLMs and action experts, then zero-shot recomposes them into unseen left-right combinations (e.g., "put cake in cup", "smash box") or adapts few-shot to dual-arm variants.

Problem

Bimanual behaviors are combinatorial: the number of left-right skill pairings grows quadratically with the skill set, yet mainstream bimanual VLAs predict a single concatenated action vector (or share one latent across arms), entangling the arms so the policy can only reproduce pairings seen in demonstrations. The paper names this failure "skill entanglement" — action-level entanglement from joint supervision and latent entanglement in action-expert VLAs — and formulates the bimanual skill-reuse problem.

Method

SkillVLA uses a two-level reasoning pipeline built on the PaliGemma backbone released with π0.5: a frozen high-level VLM emits per-arm natural-language sub-prompts (u_L, u_R) as skill descriptors, and two separate low-level VLM streams (LoRA-fine-tuned) each encode the scene with their arm's prompt and drive per-arm flow-matching action experts. An adaptive cross-attention channel between the experts carries inter-arm messages, gated by a cooperation estimator that predicts a scalar α ∈ [0,1] from the high-level latent — identifying whether the current behavior is a composition of single-arm skills (α≈0, arms disentangled) or a genuine dual-arm skill (α≈1). Gate training combines a VLM prior loss, a sticky temporal-smoothness loss, and a sparsity term, with optional discretization of α to {0,1} for stability.

Results

On nine unseen left-right recompositions of six learned skills (real dual-arm robot, Table I), SkillVLA averages 51% success while π0.5 and π0-FAST get 0% (idle-arm reversion, jitter) and TwinVLA 4%; on the seen skills all methods are comparable (SkillVLA 0.78 avg, Table II). On three highly cooperative tasks (Shake, Ball, Align), SkillVLA averages 0.48 vs. 0.47 for π0.5, and its no-cross-attention ablation collapses to 0.17 (Table III) — the communication channel matters. On the long-horizon Tubes and CollectItems tasks it matches baselines' progress score while cutting completion time by about 21% by parallelizing per-arm subtasks; in continual learning it reaches 20% zero-shot success on new dual-arm skills built from known single-arm parts.

Significance

Directly targets a scaling flaw in bimanual VLAs — quadratic pairing growth — with an architectural fix (per-arm disentanglement plus gated communication) rather than more data, complementing the session's data-scaling entries. Related threads: Review-VLA-Architecture · Review-VLA-Architecture-Categories.

← Back to RSS 2026 survey · RSS-2026-Papers · Home