ICLR 2026 X VLA - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 Category: VLA Architecture โ Cross-Embodiment Trend tag: Trend 4
flowchart LR
D1[Dataset 1: Franka] --> SP1[Soft prompt 1<br/>few learned tokens]
D2[Dataset 2: UR5] --> SP2[Soft prompt 2]
D3[Dataset 3: dex hand] --> SP3[Soft prompt 3]
SP1 --> B[SHARED Transformer Backbone]
SP2 --> B
SP3 --> B
B --> A[Action]
Mixing robots with different dynamics, sensors, and action spaces during cross-embodiment training pulls a shared backbone in conflicting directions. The challenge is to exploit the heterogeneous cross-embodiment features rather than letting them interfere โ converting potential conflict into positive cross-domain transfer while keeping the architecture simple and scalable.
A flow-matching VLA built exclusively on soft-prompted standard Transformer encoders (the 0.9B instantiation uses a 24-layer encoder, hidden size 1024, with a frozen Florence-Large VLM as the visionโlanguage encoder). Each data source / embodiment is assigned its own small set of learned soft-prompt embeddings, prepended to the input sequence (following the Lester et al. prompt-tuning recipe); the backbone is fully shared across embodiments, with only the soft prompt and the action-related input/output linear projections being embodiment-specific (together ~0.04% of total parameters). The prompt gives the backbone a clean "which robot am I?" signal so it can specialize internal computation without sacrificing shared representations. Actions are produced by an ODE-integrated flow-matching velocity field. Adaptation to a new robot is a two-phase recipe (Phase I generalist pretraining โ Phase II domain adaptation): a fresh prompt is first warmed up with the backbone frozen, then optionally joint-finetuned; the PEFT variant tunes only 9M params (~1%) and matches ฯโ on LIBERO/Simpler-WidowX despite ~300ร fewer trainable params.
X-VLA-0.9B reaches state-of-the-art across a wide sweep โ 6 simulation suites (LIBERO, SimplerEnv-WidowX, CALVIN, RoboTwin-2.0, VLABench, plus NAVSIM) and 3 real-world platforms (WidowX, the unseen-during-pretraining AIRBOT, and an Agilex bi-manual robot). Headline numbers include LIBERO โ98% and CALVIN ABCโD โ4.43 average sequence length (out of 5). The model also won 1st place (Champion) at the AgiBot World Challenge (IROS 2025). Joint multi-domain adaptation preserves and in some cases improves per-domain success vs. single-domain fine-tuning, evidencing positive transfer. Trained on 290K episodes from 7 data sources, the scaling trend (model size, data diversity, data volume) shows no sign of saturation.
Converts cross-embodiment scaling from an architecture problem into a data problem. Soft prompts are extremely lightweight (a few tokens per embodiment), so adding a new robot to an existing policy is cheap. Probably the cleanest scaling story in 2026 VLA research.
- ICLR 2026 listing
For how X-VLA compares to the full cross-embodiment landscape (soft-prompt vs. unified tokens vs. invariant latents vs. human-video bridging vs. morphology-aware vs. scale+prompt vs. world-model): Cross-Embodiment Training Review.
โ Back to ICLR-2026