CoRL 2024 OpenVLA - Heungwoo/research GitHub Wiki

OpenVLA β€” Open-source 7B autoregressive VLA, the field's reference baseline

Venue: CoRL 2024 Β· Authors: Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, et al. β€” Stanford, UC Berkeley, Toyota Research Institute, Google DeepMind, Physical Intelligence, MIT Β· arXiv: 2406.09246 Category: VLA Architecture (Baseline Β· autoregressive) Trend tag: Open-source AR VLA baseline

Approach diagram

flowchart LR
  IMG[RGB observation<br/>224px] --> VE[Fused visual encoder<br/>DINOv2 + SigLIP]
  INST[Language instruction] --> LLM
  VE --> PROJ[2-layer MLP projector] --> LLM[Llama-2-7B<br/>Prismatic-7B VLM]
  LLM --> TOK[Discretized action tokens<br/>256 bins/dim, AR decode]
  TOK --> ACT[7-DoF action<br/>EE delta pose + gripper]
Loading

Problem

Prior generalist manipulation policies were either closed and inaccessible (RT-2-X, 55B, behind Google's API) or small from-scratch models with limited language grounding (Octo, 93M). The community lacked an open VLA β€” open weights, open code, open training recipe β€” that could be downloaded, fine-tuned on consumer hardware, and used as a reproducible baseline. OpenVLA targets that gap directly.

Method

  • Backbone: built on the Prismatic-7B VLM β€” a Llama-2-7B language model with a fused dual-stream visual encoder (pretrained DINOv2 + SigLIP features concatenated channel-wise) and a 2-layer MLP projector into the LLM embedding space. The dual encoder pairs DINOv2's spatial features with SigLIP's semantics for grounding.
  • Action representation: continuous 7-DoF actions (end-effector delta pose + gripper) are discretized into 256 bins per dimension using quantile bounds (1st–99th percentile, robust to outliers vs. min-max). Each bin maps to a token in the LLM vocabulary, so action prediction is plain autoregressive next-token decoding.
  • Training mix: fine-tuned on ~970k real robot episodes from Open X-Embodiment, a curated cross-embodiment mixture spanning many robots, tasks, and scenes. Trained on 64 A100s for ~15 days.
  • Efficiency: LoRA fine-tuning reaches full-fine-tune quality while updating only ~1.4% of parameters; 4-bit quantization drops the memory footprint to ~7 GB with performance comparable to bfloat16 β€” making adaptation and inference feasible on consumer GPUs.

Results

  • Headline: "outperforming closed models such as RT-2-X (55B) by 16.5% in absolute task success rate across 29 tasks and multiple robot embodiments, with 7x fewer parameters."
  • Outperforms Octo (93M) across both out-of-the-box multi-robot evaluation and fine-tuning, especially on multi-object language-grounding tasks.
  • LoRA adaptation matches full fine-tuning at ~1.4% trainable params; 4-bit inference fits in ~7 GB.

Significance

OpenVLA is the canonical open autoregressive VLA baseline. By releasing weights, code, and the OXE recipe, it became the default reference that nearly every subsequent VLA paper fine-tunes from or benchmarks against. Its discrete-token design is the AR lineage that later tokenizer and decoding work iterates on β€” see FASTER (the OpenVLA-OFT / faster-decoding lineage). It anchors the AR-vs-flow comparison axis that flow-matching policies like Ο€0.5 are measured against, and it is the standard "open baseline" entry in the VLA Architectures review.

Links

Related pages

← Back to Home

⚠️ **GitHub.com Fallback** ⚠️