ICLR 2026 XR 1 - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 Category: VLA Architecture โ Cross-Embodiment Trend tag: Trend 4
flowchart LR
V[Visual dynamics] --> Vbr[Vision branch VQ-VAE]
M[Motion trajectory] --> Mbr[Motion branch VQ-VAE]
Vbr --> CB((SHARED codebook<br/>UVMC))
Mbr --> CB
CB --> Lat[Common embodiment-free latent]
Lat --> Df[Decode for Franka]
Lat --> Du[Decode for UR5]
Lat --> Dh[Decode for hand]
Vision and motion are typically learned as separate modalities in VLAs โ a vision encoder for pixels, an action head for motor output. This separation wastes the strong correlation between visual dynamics and the motions that caused them โ a correlation that is embodiment-independent.
Unified Vision-Motion Codes (UVMC): a dual-branch VQ-VAE in which a vision branch (visual dynamics) and a motion branch (robot motion) quantize into one shared discrete codebook, forcing the two modalities into a common latent space. UVMC acts as an intermediate representation between observations and actions. Crucially, for human videos (no action labels), the objective reduces to the vision-reconstruction loss only (โ_total^human = โ_vis), so Internet-scale human video (e.g. Ego4D) trains the vision branch alongside robot data.
The VLM backbone is PaliGemma (SigLIP visual encoder + Gemma transformer, ~2.6B params); a lightweight variant XR-1-Light uses Florence-2 (~230M trainable params). The "X" denotes XR-1's three "cross" capacities: cross-data (web human video + robot data), cross-modality (visionโmotion alignment), and cross-embodiment control.
XR-1 uses a three-stage paradigm:
- Self-supervised UVMC learning โ train the dual-branch VQ-VAE on large-scale robot + human video.
- UVMC-guided generalist pre-training โ inject UVMC knowledge into a VLM via learnable tokens on the cross-embodiment XR-D dataset (~158k trajectories, ~69.1M frames).
- Task-specific post-training โ refine for deployment.
Validated with >14,000 real-world rollouts across six robot embodiments and 120+ manipulation tasks, consistently beating ฯ0.5, ฯ0, RDT, UniVLA, and GR00T-N1.5 with strong generalization to novel objects, backgrounds, distractors, and illumination. Reported gains include Dual-Arm UR-5e ~72% vs ฯ0.5 ~62% / ฯ0 ~43%, Tien Kung 2.0 72% vs ฯ0.5 ~41%, and CALVIN 4.256 vs ฯ0.5 3.885.
The strongest version of the "learn in a common embodiment-free latent" idea in 2026. The shared codebook is a commitment: it forces vision and motion to negotiate a joint vocabulary, and the quality of that vocabulary determines downstream cross-embodiment transfer.
- ICLR 2026 listing
โ Back to ICLR-2026