ICLR 2026 MetaVLA - Heungwoo/research GitHub Wiki

MetaVLA โ€” Unified Meta Co-Training for Efficient Embodied Adaptation

Venue: ICLR 2026 Authors: Chen Li, Zhantao Yang, Han Zhang, Fangyi Chen, Anudeepsekhar Bolimera, Marios Savvides (Carnegie Mellon University); Chenchen Zhu (Meta Reality Labs, USA) Category: VLA Training โ€” Post-training / adaptation Trend tag: Efficient adaptation ยท meta-learning ยท context-aware co-training

Approach diagram

flowchart LR
  Pre[OpenVLA-7B<br/>or NORA-Long 3B Qwen2.5-VL] --> ANP[Action-ANP module<br/>self-attention over context<br/>+ cross-attention to target]
  CB[Context Bank<br/>4 LIBERO suites + GR00T aux<br/>refreshed every K=200 steps<br/>b_C=32 examples/task] --> ANP
  TB[Target Bank<br/>4 LIBERO suites unified] --> ANP
  ANP --> Hide[Hidden states<br/>concat with Llama2 hidden]
  Hide --> LM[LM head]
  LM --> A[Discrete action tokens]
Loading

Problem

Mainstream VLA post-training fine-tunes one model per task (OpenVLA: 4 separate 240K-step runs across LIBERO-Goal/Spatial/Object/Long). Naively switching to multi-task SFT helps a little; naively adding diverse auxiliary tasks (different camera views, different DoF, bimanual) degrades performance due to optimization instability from heterogeneous distributions. MetaVLA's question: can meta-learning extract benefit from auxiliary tasks without destabilizing in-domain training?

Detailed Method

Architecture: Action-ANP atop OpenVLA's Llama2 decoder

Built on Attentive Neural Processes (ANP, Kim '19), which model conditional p(y_T | x_T, x_C, y_C) as a meta-learner over functions parameterized by both global and target-specific latents.

For target feature x_T and context pairs (x_C_i, y_C_i):

  • Self-attention over context produces per-context representations r_C_i and s_C_i.
  • Cross-attention with target query x_T โ†’ r_T (deterministic, target-aware).
  • Mean over s_C_i โ†’ sฬ„_C; stochastic latent z ~ q(z | sฬ„_C) (Gaussian via reparameterization).
  • ELBO: log p(y_T | x_T, x_C, y_C) โ‰ฅ E_{q(z|sฬ„_T)}[log p(y_T | x_T, r_T, z)] โˆ’ D_KL(q(z|sฬ„_T) โˆฅ q(z|sฬ„_C)).
  • r_T and z are concatenated with Llama2 hidden states before the LM head; rest of the OpenVLA pipeline is unchanged.

The KL term regularizes target distribution toward context distribution. Two variants:

  • Deterministic (default): drops the stochastic / KL term โ†’ reconstruction loss only. Better on harder tasks (esp. LIBERO-Long).
  • +Stochastic: full ELBO. Slightly better on Spatial, but underperforms on Long where domain shift > KL assumption.

Context Bank vs Target Bank

  • Context bank = in-domain (4 LIBERO suites, split as non-overlapping context vs target sets) + out-of-domain auxiliary (subset of partially open-sourced NVIDIA GR00T data).
  • Target bank = target sets of all four LIBERO suites jointly.

Auxiliary task selection deliberately includes structural mismatch:

  • Front-view 7-DoF single-arm (similar to LIBERO).
  • Side-view 7-DoF single-arm (camera mismatch).
  • Bimanual 14-DoF two-arm (action-space mismatch). This is intentionally less curated than CoT-VLA's tightly-aligned auxiliary picks โ€” argues for robustness to context diversity.

Training Protocols

  • Context refresh interval K = 200 steps.
  • Context batch size b_C = 32 examples per context task.
  • Per-task SFT cross-entropy loss on action tokens.
  • Trained for 75K steps total vs OpenVLA's 240K.
  • Hardware: 8ร— A100 80GB GPUs, ~24 hours total vs OpenVLA's ~100 hours total across the four suites โ€” a ~76% GPU-time reduction (paper's headline figure).
  • Eval: one 24GB RTX-4090 GPU.
  • Backbone: OpenVLA-7B (Llama2-7B + ViT) primary; ablation on NORA-Long 3B (Qwen2.5-VL-based) for backbone-agnostic claim.

Comprehensive Results

LIBERO success rate (Table 1)

Model Steps Goal Spatial Object Long Avg
ฯ€_{0.5} (Intelligence '25b) 30K 98.0 98.8 98.2 92.4 96.9
Diffusion Policy โ€“ 68.3 78.3 92.5 50.5 72.4
ATM โ€“ 77.8 68.5 68.0 39.3 63.4
TraceVLA โ€“ 75.1 84.6 85.2 54.1 74.8
OpenVLA (4 separate models) 240K 76.2 84.7 87.0 51.8 74.9
SFT-4LIBERO (vanilla multi-task) 75K 77.8 84.8 87.4 54.7 76.2
SFT-4LIBERO+1single+1bi 75K 59.7 68.0 65.2 30.0 55.7
SFT-4LIBERO+3single 75K 24.6 16.8 9.7 1.5 13.2
SFT-4LIBERO+5single+1bi 75K 15.2 5.6 12.0 1.6 8.6
SFT-4LIBERO+5single+1bi (extended) 187.5K 23.4 16.7 13.6 4.4 14.5
MetaVLA-Pretrained-Context-ONLY 75K 74.4 85.4 85.4 52.3 74.4
MetaVLA (no aux, deterministic) 75K 78.9 88.5 88.5 55.3 77.8
MetaVLA + Stochastic 75K 78.9 88.9 88.5 53.0 77.3
MetaVLA+1single+1bi 75K 78.5 89.0 87.4 59.0 78.5
MetaVLA+3single 75K 78.0 88.0 87.2 59.7 78.2
MetaVLA+5single+1bi (full) 75K 78.7 89.9 88.9 59.8 79.3

Headline: MetaVLA+6 aux beats OpenVLA by +4.4% avg (and +8.0% on LIBERO-Long) using 75K steps vs 240K (-68.75% steps). vs SFT-4LIBERO: +3.1% avg, +5.1% on Long.

The contrast within Table 1 is the real story: vanilla SFT collapses catastrophically as auxiliary tasks are added (76.2 โ†’ 8.6 avg with 6 aux tasks). MetaVLA gains from the same 6 aux tasks (77.8 โ†’ 79.3).

Backbone-agnostic โ€” NORA-Long (3B Qwen2.5-VL) (Table 2)

Model Goal Spatial Object Long Avg
NORA-Long 85.4 90.5 95.0 70.6 85.4
NORA-Long-SFT-4LIBERO 87.0 92.5 94.0 75.5 87.3
NORA-Long-SFT-4LIBERO+5single+1bi 73.6 79.5 75.2 37.2 66.4
MetaVLA-NORA-Long 90.8 96.2 96.5 77.8 90.3
MetaVLA-NORA-Long+5single+1bi 93.8 95.8 97.2 80.2 91.8

+6.4% over NORA-Long avg with aux tasks; +25.4% over native SFT counterpart with aux tasks.

Per-suite training-step efficiency (MetaVLA-EACH, Table 3)

Method Total Goal step / SR Spatial step / SR Object step / SR Long step / SR
OpenVLA-120K 120K 30K / 71.4 10K / 81.2 30K / 85.8 50K / 44.4
MetaVLA-EACH-120K 120K 30K / 76.4 10K / 86.1 30K / 89.0 50K / 55.4
OpenVLA-240K 240K 60K / 76.2 50K / 84.7 50K / 87.0 80K / 51.8
MetaVLA-EACH-240K 240K 60K / 77.4 50K / 85.8 50K / 88.5 80K / 55.8

Action-ANP alone (no co-training) already beats OpenVLA at every checkpoint; per-suite Long task continues to improve through 240K.

Ablation Studies

Context batch size b_C (Fig. 4): success monotonically increases with b_C โˆˆ {4, 8, 16, 32}; b_C = 32 is the sweet spot โ€” gain saturates and memory remains modest.

Auxiliary task selection (Table 1, three rows): MetaVLA+1single+1bi (78.5), MetaVLA+3single (78.2), MetaVLA+5single+1bi (79.3) โ€” all beat their SFT counterparts by 22-70 pts, demonstrating robustness to context diversity (different views, different action spaces).

Parameter-size confound (MetaVLA-Pretrained-Context-ONLY): replace context bank with bridge_orig + fractal20220817 (already in OpenVLA pretrain) โ†’ average drops to 74.4 (vs 77.8 with LIBERO context), showing gain is from out-of-distribution informative auxiliary signals, not just added params.

Stochastic vs deterministic ELBO: stochastic helps Spatial (+0.4), comparable on Goal/Object, hurts Long (-2.3 โ†’ 53.0). KL term assumption breaks under bigger domain shift.

Multi-task co-training mechanism (MetaVLA-EACH): Action-ANP alone โ€” single-suite training without context co-training โ€” already beats OpenVLA at half the steps. Co-training collapses 4 models to 1 and saves another 45K steps overall.

Inference-cost overhead: only +0.3 ms/token latency (Fig. 9 in appendix) due to lightweight ANP module.

Limitations (as stated by authors)

  1. Backbone breadth. Only OpenVLA-7B and NORA-Long-3B tested; broader VLM/VLA spectrum (e.g. PaliGemma-3B ฯ€โ‚€-style, Qwen2-VL action experts) remains future work.
  2. Real-robot deployment. All experiments are LIBERO simulation; real-world validation deferred.
  3. No formal proof of "scaling more robustly". Authors acknowledge their stability claim relative to vanilla multi-task SFT is empirical; a rigorous proof / theoretical bound is "left to future work due to computational constraints".
  4. Web-scale context bank not yet tried. They speculate that web-scale data augmentation of the context bank (analogous to ฯ€โ‚€.5's pretraining co-training) could help further but didn't run it.
  5. Auxiliary task selection. Explore-exhaust of combinations not done due to memory/compute.

Significance & Positioning

Vs OpenVLA / OpenVLA-OFT. OpenVLA fine-tunes each LIBERO suite as a separate Hugging Face model, totaling ~240K steps across the four suites (Table 1; per-suite breakdown 60K/50K/50K/80K in Table 3). MetaVLA collapses these into one unified model in 75K steps (-68.75% steps) and beats avg by 4.4 pts. OpenVLA-OFT (Kim '25) optimizes inference speed but per the paper's intro still demands ~150Kโ€“500K SFT steps (incl. diffusion + non-diffusion parts); MetaVLA is orthogonal โ€” could in principle be stacked.

Vs ฯ€โ‚€ / ฯ€โ‚€.โ‚… / ฯ€โ‚€.โ‚† / ฯ€โ‚€.โ‚‡. ฯ€โ‚€-class models invest heavily in pretraining (large robotics corpus + flow matching). MetaVLA is purely post-training: ฯ€โ‚€.โ‚… scores 96.9 avg LIBERO via massive pretraining; MetaVLA+aux 79.3 with much cheaper compute. Different points on the cost/performance frontier.

Vs CoT-VLA / OneTwoVLA / EO-1. Those add reasoning data into pretraining (CoT-VLA, OneTwoVLA) or interleaved Vision-Text-Action (EO-1, "high inference latency"). MetaVLA stays at post-training and avoids inference-time overhead.

Vs other efficient-adaptation peers in 2026. Sits with FASTER, HyperVLA, Hybrid Training, LBM Cotraining โ€” all treat training/adaptation efficiency as a first-class objective. MetaVLA is unique in leveraging meta-learning (ANP) rather than hypernetworks, LoRA-style routing, or curriculum.

Vs GR00T-N1.5. Interestingly, MetaVLA uses GR00T data as auxiliary context โ€” flips the relationship from competitor to fuel.

The deeper claim is structural: post-training is an under-explored axis vs the dominant pretraining-scaling story. Naive multi-task SFT fails worse with more diverse auxiliary tasks (aggregate -67.6 pt collapse with 6 aux tasks). The ANP context-bank lets the model selectively attend to relevant context, turning auxiliary diversity from a liability into an asset. Reframes VLA adaptation as a meta-learning problem rather than a fresh fine-tune per task.

Links

Related pages

โ† Back to ICLR-2026

โš ๏ธ **GitHub.com Fallback** โš ๏ธ