ICLR 2026 UniVLA - Heungwoo/research GitHub Wiki

UniVLA / Unified VLA โ€” Native Multimodal VLA at 8.5B Parameters

Venue: ICLR 2026 ยท OpenReview: PklMD8PwUy ยท arXiv: 2506.19850 Category: VLA Architecture Trend tag: Trend 4 (scale) Authors / affiliations: Yuqi Wang, Xinghang Li, Wenxuan Wang, Junbo Zhang, Yingyan Li, Yuntao Chen, Xinlong Wang (corresponding), Zhaoxiang Zhang (corresponding) โ€” BAAI / CASIA / Tsinghua / HKISI ยท Code: baaivision/UniVLA

Approach diagram

flowchart LR
  V[Vision tokens] --> S[Single interleaved token stream]
  L[Language tokens] --> S
  A[Action tokens] --> S
  S --> M[8.5B-param native multimodal model<br/>autoregressive over the stream]
  M --> Pred[Next token prediction<br/>could be vision/language/action]
  M --> WM[Post-training:<br/>world-model objective<br/>learn to predict future video]
Loading

Problem

Most VLAs are built by bolting an action head onto a pretrained VLM and jointly fine-tuning. This "retrofit" approach treats action modeling as a second-class citizen โ€” the action head's parameters are small relative to the backbone, and it is added late.

Method

Native multimodal model built on Emu3 and initialized from pretrained Emu3 weights (not from scratch). Vision, language, and action are autoregressively modeled as a single interleaved stream of discrete tokens: images are tokenized with a VQ/MOVQ encoder (spatial compression factor 8, following Emu3), and continuous actions are converted via Discrete Cosine Transform (DCT) into discrete tokens using the FAST tokenizer (vocabulary 1024). 8.5B parameters, action treated as a first-class modality. A post-training world-model stage โ€” predicting future visual content given current observation + instruction, with the loss computed solely from vision tokens so it can learn from unlabeled video โ€” captures causal dynamics before downstream policy fine-tuning.

Results

New SOTA across simulated manipulation benchmarks:

  • LIBERO average 95.5% (vs ฯ€โ‚€-FAST 85.5%); LIBERO-Long 94.0% (vs CoT-VLA 69.0%).
  • CALVIN (ABCDโ†’D) 4.63 avg task length (vs RoboVLMs 4.49).
  • SimplerEnv-WidowX/Bridge 69.8% success (vs SpatialVLA 42.7%).

Also evaluated on real-world ALOHA tasks and autonomous-driving settings, demonstrating the single-architecture multi-task scope (perception grounding, world modeling, policy learning).

Significance

The most committed "action is a first-class modality" architecture in 2026, and the clearest 8.5B-parameter VLA in the open literature. Native-multimodal philosophy contrasts directly with ฯ€0.6's "VLM + separate flow-matching expert" design โ€” sets up a scale-vs-modularity debate for 2026โ€“2027.

Links

Related pages

โ† Back to ICLR-2026

โš ๏ธ **GitHub.com Fallback** โš ๏ธ