ICLR 2026 X VLA - Heungwoo/research GitHub Wiki

X-VLA โ€” Soft-Prompted Transformer as Scalable Cross-Embodiment VLA

Venue: ICLR 2026 Category: VLA Architecture โ€” Cross-Embodiment Trend tag: Trend 4

Approach diagram

flowchart LR
  D1[Dataset 1: Franka] --> SP1[Soft prompt 1<br/>few learned tokens]
  D2[Dataset 2: UR5] --> SP2[Soft prompt 2]
  D3[Dataset 3: dex hand] --> SP3[Soft prompt 3]
  SP1 --> B[SHARED Transformer Backbone]
  SP2 --> B
  SP3 --> B
  B --> A[Action]
Loading

Problem

Mixing robots with different dynamics, sensors, and action spaces during cross-embodiment training pulls a shared backbone in conflicting directions. The challenge is to exploit the heterogeneous cross-embodiment features rather than letting them interfere โ€” converting potential conflict into positive cross-domain transfer while keeping the architecture simple and scalable.

Method

A flow-matching VLA built exclusively on soft-prompted standard Transformer encoders (the 0.9B instantiation uses a 24-layer encoder, hidden size 1024, with a frozen Florence-Large VLM as the visionโ€“language encoder). Each data source / embodiment is assigned its own small set of learned soft-prompt embeddings, prepended to the input sequence (following the Lester et al. prompt-tuning recipe); the backbone is fully shared across embodiments, with only the soft prompt and the action-related input/output linear projections being embodiment-specific (together ~0.04% of total parameters). The prompt gives the backbone a clean "which robot am I?" signal so it can specialize internal computation without sacrificing shared representations. Actions are produced by an ODE-integrated flow-matching velocity field. Adaptation to a new robot is a two-phase recipe (Phase I generalist pretraining โ†’ Phase II domain adaptation): a fresh prompt is first warmed up with the backbone frozen, then optionally joint-finetuned; the PEFT variant tunes only 9M params (~1%) and matches ฯ€โ‚€ on LIBERO/Simpler-WidowX despite ~300ร— fewer trainable params.

Results

X-VLA-0.9B reaches state-of-the-art across a wide sweep โ€” 6 simulation suites (LIBERO, SimplerEnv-WidowX, CALVIN, RoboTwin-2.0, VLABench, plus NAVSIM) and 3 real-world platforms (WidowX, the unseen-during-pretraining AIRBOT, and an Agilex bi-manual robot). Headline numbers include LIBERO โ‰ˆ98% and CALVIN ABCโ†’D โ‰ˆ4.43 average sequence length (out of 5). The model also won 1st place (Champion) at the AgiBot World Challenge (IROS 2025). Joint multi-domain adaptation preserves and in some cases improves per-domain success vs. single-domain fine-tuning, evidencing positive transfer. Trained on 290K episodes from 7 data sources, the scaling trend (model size, data diversity, data volume) shows no sign of saturation.

Significance

Converts cross-embodiment scaling from an architecture problem into a data problem. Soft prompts are extremely lightweight (a few tokens per embodiment), so adding a new robot to an existing policy is cheap. Probably the cleanest scaling story in 2026 VLA research.

Links

  • ICLR 2026 listing

๐Ÿ“– In-depth cross-paper review

For how X-VLA compares to the full cross-embodiment landscape (soft-prompt vs. unified tokens vs. invariant latents vs. human-video bridging vs. morphology-aware vs. scale+prompt vs. world-model): Cross-Embodiment Training Review.

Related pages

โ† Back to ICLR-2026

โš ๏ธ **GitHub.com Fallback** โš ๏ธ