ICML 2026 UniCoD - Heungwoo/research GitHub Wiki
UniCoD — Unified continuous-and-discrete representation learning for robot policy
Venue: ICML 2026 (Poster) Category: VLA Architecture Affiliations: Jianke Zhang, Yucheng Hu, Yanjiang Guo, Xiaoyu Chen, Yichen Liu, Wenna Chen, Chaochao Lu, Jianyu Chen Traction (2026-06): 6 citations (arXiv)
Problem
Building generalist robot policies that handle diverse tasks in open-ended environments is a central challenge in robotics. To leverage large-scale pretraining, prior Vision-Language-Action (VLA) work typically builds generalist policies either on top of vision-language models (VLMs) or on generative models. However, the authors argue that both semantic understanding (from vision-language pretraining) and visual dynamics modeling (from visual-generation pretraining) are crucial for embodied robots, and neither family alone captures both.
Method
UniCoD builds on recent unified models of generation and understanding, which have shown strong capabilities in both comprehension and generation through large-scale pretraining. The premise is that robotic policy learning can likewise benefit from the combined strengths of understanding, planning, and continuous future-representation learning. Concretely:
- UniCoD acquires the ability to dynamically model high-dimensional visual features by pretraining on over 1M internet-scale instructional manipulation videos, giving it predictive/generative representations of future visual states.
- It is then fine-tuned on data collected from the robot embodiment, learning the mapping from these predictive representations to action tokens.
- The "unified continuous and discrete" framing combines continuous future visual representation learning with discrete action-token prediction within one model.
flowchart LR
A[1M+ internet manipulation videos] --> B[Unified generation + understanding pretraining]
B --> C[Continuous predictive visual representations]
C --> D[Fine-tune on robot embodiment data]
D --> E[Map predictive reps -> discrete action tokens]
E --> F[Generalist robot policy]
Results
"9% and 12% over baselines in sim and real-world OOD tasks". The paper reports that UniCoD consistently outperforms baseline methods by 9% in simulation environments and by 12% on real-world out-of-distribution tasks. (The verbatim ICML abstract does not break these numbers down further per benchmark.)
Significance
UniCoD argues that the two dominant VLA pretraining paradigms—VLM-based semantic grounding and generative visual-dynamics modeling—are complementary rather than competing, and that a single unified generation-and-understanding backbone can inherit both. Leveraging 1M+ internet manipulation videos for predictive representation learning, then grounding to action tokens, points toward more data-efficient and OOD-robust generalist policies in the 2026 VLA landscape.
Links
- ICML 2026: https://icml.cc/virtual/2026/poster/63426
← Back to ICML-2026