ICLR 2026 XR 1 - Heungwoo/research GitHub Wiki

XR-1 โ€” Unified Vision-Motion Codes (UVMC) for Cross-Embodiment VLA

Venue: ICLR 2026 Category: VLA Architecture โ€” Cross-Embodiment Trend tag: Trend 4

Approach diagram

flowchart LR
  V[Visual dynamics] --> Vbr[Vision branch VQ-VAE]
  M[Motion trajectory] --> Mbr[Motion branch VQ-VAE]
  Vbr --> CB((SHARED codebook<br/>UVMC))
  Mbr --> CB
  CB --> Lat[Common embodiment-free latent]
  Lat --> Df[Decode for Franka]
  Lat --> Du[Decode for UR5]
  Lat --> Dh[Decode for hand]
Loading

Problem

Vision and motion are typically learned as separate modalities in VLAs โ€” a vision encoder for pixels, an action head for motor output. This separation wastes the strong correlation between visual dynamics and the motions that caused them โ€” a correlation that is embodiment-independent.

Method

Unified Vision-Motion Codes (UVMC): a dual-branch VQ-VAE in which a vision branch (visual dynamics) and a motion branch (robot motion) quantize into one shared discrete codebook, forcing the two modalities into a common latent space. UVMC acts as an intermediate representation between observations and actions. Crucially, for human videos (no action labels), the objective reduces to the vision-reconstruction loss only (โ„’_total^human = โ„’_vis), so Internet-scale human video (e.g. Ego4D) trains the vision branch alongside robot data.

The VLM backbone is PaliGemma (SigLIP visual encoder + Gemma transformer, ~2.6B params); a lightweight variant XR-1-Light uses Florence-2 (~230M trainable params). The "X" denotes XR-1's three "cross" capacities: cross-data (web human video + robot data), cross-modality (visionโ€“motion alignment), and cross-embodiment control.

XR-1 uses a three-stage paradigm:

  1. Self-supervised UVMC learning โ€” train the dual-branch VQ-VAE on large-scale robot + human video.
  2. UVMC-guided generalist pre-training โ€” inject UVMC knowledge into a VLM via learnable tokens on the cross-embodiment XR-D dataset (~158k trajectories, ~69.1M frames).
  3. Task-specific post-training โ€” refine for deployment.

Results

Validated with >14,000 real-world rollouts across six robot embodiments and 120+ manipulation tasks, consistently beating ฯ€0.5, ฯ€0, RDT, UniVLA, and GR00T-N1.5 with strong generalization to novel objects, backgrounds, distractors, and illumination. Reported gains include Dual-Arm UR-5e ~72% vs ฯ€0.5 ~62% / ฯ€0 ~43%, Tien Kung 2.0 72% vs ฯ€0.5 ~41%, and CALVIN 4.256 vs ฯ€0.5 3.885.

Significance

The strongest version of the "learn in a common embodiment-free latent" idea in 2026. The shared codebook is a commitment: it forces vision and motion to negotiate a joint vocabulary, and the quality of that vocabulary determines downstream cross-embodiment transfer.

Links

  • ICLR 2026 listing

Related pages

โ† Back to ICLR-2026

โš ๏ธ **GitHub.com Fallback** โš ๏ธ