CoRL 2024 OpenVLA - Heungwoo/research GitHub Wiki
Venue: CoRL 2024 Β· Authors: Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, et al. β Stanford, UC Berkeley, Toyota Research Institute, Google DeepMind, Physical Intelligence, MIT Β· arXiv: 2406.09246 Category: VLA Architecture (Baseline Β· autoregressive) Trend tag: Open-source AR VLA baseline
flowchart LR
IMG[RGB observation<br/>224px] --> VE[Fused visual encoder<br/>DINOv2 + SigLIP]
INST[Language instruction] --> LLM
VE --> PROJ[2-layer MLP projector] --> LLM[Llama-2-7B<br/>Prismatic-7B VLM]
LLM --> TOK[Discretized action tokens<br/>256 bins/dim, AR decode]
TOK --> ACT[7-DoF action<br/>EE delta pose + gripper]
Prior generalist manipulation policies were either closed and inaccessible (RT-2-X, 55B, behind Google's API) or small from-scratch models with limited language grounding (Octo, 93M). The community lacked an open VLA β open weights, open code, open training recipe β that could be downloaded, fine-tuned on consumer hardware, and used as a reproducible baseline. OpenVLA targets that gap directly.
- Backbone: built on the Prismatic-7B VLM β a Llama-2-7B language model with a fused dual-stream visual encoder (pretrained DINOv2 + SigLIP features concatenated channel-wise) and a 2-layer MLP projector into the LLM embedding space. The dual encoder pairs DINOv2's spatial features with SigLIP's semantics for grounding.
- Action representation: continuous 7-DoF actions (end-effector delta pose + gripper) are discretized into 256 bins per dimension using quantile bounds (1stβ99th percentile, robust to outliers vs. min-max). Each bin maps to a token in the LLM vocabulary, so action prediction is plain autoregressive next-token decoding.
- Training mix: fine-tuned on ~970k real robot episodes from Open X-Embodiment, a curated cross-embodiment mixture spanning many robots, tasks, and scenes. Trained on 64 A100s for ~15 days.
- Efficiency: LoRA fine-tuning reaches full-fine-tune quality while updating only ~1.4% of parameters; 4-bit quantization drops the memory footprint to ~7 GB with performance comparable to bfloat16 β making adaptation and inference feasible on consumer GPUs.
- Headline: "outperforming closed models such as RT-2-X (55B) by 16.5% in absolute task success rate across 29 tasks and multiple robot embodiments, with 7x fewer parameters."
- Outperforms Octo (93M) across both out-of-the-box multi-robot evaluation and fine-tuning, especially on multi-object language-grounding tasks.
- LoRA adaptation matches full fine-tuning at ~1.4% trainable params; 4-bit inference fits in ~7 GB.
OpenVLA is the canonical open autoregressive VLA baseline. By releasing weights, code, and the OXE recipe, it became the default reference that nearly every subsequent VLA paper fine-tunes from or benchmarks against. Its discrete-token design is the AR lineage that later tokenizer and decoding work iterates on β see FASTER (the OpenVLA-OFT / faster-decoding lineage). It anchors the AR-vs-flow comparison axis that flow-matching policies like Ο0.5 are measured against, and it is the standard "open baseline" entry in the VLA Architectures review.
- arXiv: 2406.09246
- Project page: openvla.github.io
- Code: github.com/openvla/openvla
β Back to Home