CVPR 2026 OptimusVLA - Heungwoo/research GitHub Wiki

OptimusVLA — Global Prior Meets Local Consistency: Dual-Memory Augmented VLA

Venue: CVPR 2026 Category: VLA Architecture Trend tag: Trend 1 (VLA is the modal CVPR 2026 manipulation contribution)

Approach diagram

flowchart LR
  OBS["current obs"] --> VLM["frozen π0.5 VLM backbone"]
  TASK["task instruction"] --> VLM
  VLM --> AH["flow-matching action head"]
  GPM["Global Prior Memory<br/>retrieved prior replaces<br/>Gaussian noise init"] --> AH
  LCM["Local Consistency Memory<br/>self-attn + Mamba<br/>consistency bias"] --> AH
  AH --> ACT["action chunk<br/>(adaptive NFE)"]
Loading

Problem

Hierarchical flow-matching VLAs (e.g. π0.5) have two coupled bottlenecks:

  1. Low inference efficiency. The flow policy maps isotropic Gaussian noise to the action distribution. The large prior–target gap forces many denoising steps (high NFE — number of function evaluations).
  2. Poor robustness. Conditioning solely on the current observation gives the policy no temporal awareness, so it cannot distinguish task phases that look visually similar and produces jittery, inconsistent control.

The dual-memory design targets these directly: a global prior shrinks the prior–target gap (efficiency), and a local memory injects temporal continuity (robustness).

Method

Built on a frozen π0.5 backbone (3.6B params total), augmented with two memory modules that are trained in stages (VLA pre-training → GPM training → LCM training) without modifying the pre-trained VLA weights:

  • Global Prior Memory (GPM) — replaces the Gaussian-noise initialization of the flow policy with a task-level prior retrieved from a memory bank of past trajectories. The bank stores (task-embedding, trajectory) key-value pairs; given the current multimodal input it retrieves the k nearest trajectories by cosine similarity and builds a Gaussian prior from their weighted mean/variance, positioned near the target distribution. Similarity-adaptive noise scaling and adaptive NFE scheduling then spend fewer denoising steps on confident (high-similarity) cases and more on novel ones. This shrinks the prior–target gap and cuts NFE from ~10.0 (π0.5) to ~3.2.
  • Local Consistency Memory (LCM) — adds temporal awareness with minimal overhead via (i) a self-attention consistency layer over recent action chunks capturing inter-action dependencies and (ii) a Mamba-based dynamic awareness module modeling inter-chunk temporal dynamics. It outputs a learned consistency bias injected into the policy input to enforce smooth, temporally coherent trajectories.

The two memories are complementary: global shrinks the generative path (efficiency), local enforces continuity across chunks (robustness).

Results

Benchmark OptimusVLA Baseline
LIBERO (avg success) 98.6% 96.9% (π0.5)
CALVIN ABC→D (avg length) 4.45 +13.5% over π0
RoboTwin 2.0 Hard (avg success) 38% best rank
Real-world Generalization 85.0% +42.9% over π0
Real-world Long-Horizon 64.0% +52.4% over π0
Inference 2.9× speedup ~3.1× fewer NFE (10.0 → 3.2)

Real-world evaluation uses a Galaxea R1 Lite (14-DoF bimanual) platform across 4 generalization and 4 long-horizon tasks.

Significance

The notable twist is that "memory" here is repurposed as a flow-matching efficiency mechanism: GPM is not a context prompt but a retrieved initialization for the action-generation ODE, turning a trajectory memory bank into a near-target prior that slashes denoising steps (NFE 10.0 → 3.2, 2.9× speedup) while also improving success rates. This is a different use of memory than MemoryVLA and HAMLET, which use memory primarily for long-horizon context rather than sampler initialization. LCM adds an orthogonal temporal-consistency signal (self-attention + Mamba) without touching frozen π0.5 weights. Likely to inform VLA Memory taxonomies going forward.

Links

  • arXiv: 2602.20200
  • Code: iLearn-Lab/CVPR26-OptimusVLA
  • Authors: Zaijing Li, Bing Hu, Rui Shao, Gongwei Chen, Dongmei Jiang, Pengwei Xie, Jianye Hao, Liqiang Nie (HIT Shenzhen · PengCheng Lab · Shenzhen Loop Area Institute · Huawei Noah's Ark Lab)

Related pages

← Back to CVPR-2026

⚠️ **GitHub.com Fallback** ⚠️