CVPR 2026 OptimusVLA - Heungwoo/research GitHub Wiki
Venue: CVPR 2026 Category: VLA Architecture Trend tag: Trend 1 (VLA is the modal CVPR 2026 manipulation contribution)
flowchart LR
OBS["current obs"] --> VLM["frozen π0.5 VLM backbone"]
TASK["task instruction"] --> VLM
VLM --> AH["flow-matching action head"]
GPM["Global Prior Memory<br/>retrieved prior replaces<br/>Gaussian noise init"] --> AH
LCM["Local Consistency Memory<br/>self-attn + Mamba<br/>consistency bias"] --> AH
AH --> ACT["action chunk<br/>(adaptive NFE)"]
Hierarchical flow-matching VLAs (e.g. π0.5) have two coupled bottlenecks:
- Low inference efficiency. The flow policy maps isotropic Gaussian noise to the action distribution. The large prior–target gap forces many denoising steps (high NFE — number of function evaluations).
- Poor robustness. Conditioning solely on the current observation gives the policy no temporal awareness, so it cannot distinguish task phases that look visually similar and produces jittery, inconsistent control.
The dual-memory design targets these directly: a global prior shrinks the prior–target gap (efficiency), and a local memory injects temporal continuity (robustness).
Built on a frozen π0.5 backbone (3.6B params total), augmented with two memory modules that are trained in stages (VLA pre-training → GPM training → LCM training) without modifying the pre-trained VLA weights:
- Global Prior Memory (GPM) — replaces the Gaussian-noise initialization of the flow policy with a task-level prior retrieved from a memory bank of past trajectories. The bank stores (task-embedding, trajectory) key-value pairs; given the current multimodal input it retrieves the k nearest trajectories by cosine similarity and builds a Gaussian prior from their weighted mean/variance, positioned near the target distribution. Similarity-adaptive noise scaling and adaptive NFE scheduling then spend fewer denoising steps on confident (high-similarity) cases and more on novel ones. This shrinks the prior–target gap and cuts NFE from ~10.0 (π0.5) to ~3.2.
- Local Consistency Memory (LCM) — adds temporal awareness with minimal overhead via (i) a self-attention consistency layer over recent action chunks capturing inter-action dependencies and (ii) a Mamba-based dynamic awareness module modeling inter-chunk temporal dynamics. It outputs a learned consistency bias injected into the policy input to enforce smooth, temporally coherent trajectories.
The two memories are complementary: global shrinks the generative path (efficiency), local enforces continuity across chunks (robustness).
| Benchmark | OptimusVLA | Baseline |
|---|---|---|
| LIBERO (avg success) | 98.6% | 96.9% (π0.5) |
| CALVIN ABC→D (avg length) | 4.45 | +13.5% over π0 |
| RoboTwin 2.0 Hard (avg success) | 38% | best rank |
| Real-world Generalization | 85.0% | +42.9% over π0 |
| Real-world Long-Horizon | 64.0% | +52.4% over π0 |
| Inference | 2.9× speedup | ~3.1× fewer NFE (10.0 → 3.2) |
Real-world evaluation uses a Galaxea R1 Lite (14-DoF bimanual) platform across 4 generalization and 4 long-horizon tasks.
The notable twist is that "memory" here is repurposed as a flow-matching efficiency mechanism: GPM is not a context prompt but a retrieved initialization for the action-generation ODE, turning a trajectory memory bank into a near-target prior that slashes denoising steps (NFE 10.0 → 3.2, 2.9× speedup) while also improving success rates. This is a different use of memory than MemoryVLA and HAMLET, which use memory primarily for long-horizon context rather than sampler initialization. LCM adds an orthogonal temporal-consistency signal (self-attention + Mamba) without touching frozen π0.5 weights. Likely to inform VLA Memory taxonomies going forward.
- arXiv: 2602.20200
- Code:
iLearn-Lab/CVPR26-OptimusVLA - Authors: Zaijing Li, Bing Hu, Rui Shao, Gongwei Chen, Dongmei Jiang, Pengwei Xie, Jianye Hao, Liqiang Nie (HIT Shenzhen · PengCheng Lab · Shenzhen Loop Area Institute · Huawei Noah's Ark Lab)
- VLA Memory — where dual-memory designs sit in the broader taxonomy
- MemoryVLA · HAMLET
- CVPR 2026 survey
← Back to CVPR-2026