ICLR 2026 MetaVLA - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 Authors: Chen Li, Zhantao Yang, Han Zhang, Fangyi Chen, Anudeepsekhar Bolimera, Marios Savvides (Carnegie Mellon University); Chenchen Zhu (Meta Reality Labs, USA) Category: VLA Training โ Post-training / adaptation Trend tag: Efficient adaptation ยท meta-learning ยท context-aware co-training
flowchart LR
Pre[OpenVLA-7B<br/>or NORA-Long 3B Qwen2.5-VL] --> ANP[Action-ANP module<br/>self-attention over context<br/>+ cross-attention to target]
CB[Context Bank<br/>4 LIBERO suites + GR00T aux<br/>refreshed every K=200 steps<br/>b_C=32 examples/task] --> ANP
TB[Target Bank<br/>4 LIBERO suites unified] --> ANP
ANP --> Hide[Hidden states<br/>concat with Llama2 hidden]
Hide --> LM[LM head]
LM --> A[Discrete action tokens]
Mainstream VLA post-training fine-tunes one model per task (OpenVLA: 4 separate 240K-step runs across LIBERO-Goal/Spatial/Object/Long). Naively switching to multi-task SFT helps a little; naively adding diverse auxiliary tasks (different camera views, different DoF, bimanual) degrades performance due to optimization instability from heterogeneous distributions. MetaVLA's question: can meta-learning extract benefit from auxiliary tasks without destabilizing in-domain training?
Built on Attentive Neural Processes (ANP, Kim '19), which model conditional p(y_T | x_T, x_C, y_C) as a meta-learner over functions parameterized by both global and target-specific latents.
For target feature x_T and context pairs (x_C_i, y_C_i):
- Self-attention over context produces per-context representations r_C_i and s_C_i.
- Cross-attention with target query x_T โ r_T (deterministic, target-aware).
- Mean over s_C_i โ sฬ_C; stochastic latent z ~ q(z | sฬ_C) (Gaussian via reparameterization).
- ELBO: log p(y_T | x_T, x_C, y_C) โฅ E_{q(z|sฬ_T)}[log p(y_T | x_T, r_T, z)] โ D_KL(q(z|sฬ_T) โฅ q(z|sฬ_C)).
- r_T and z are concatenated with Llama2 hidden states before the LM head; rest of the OpenVLA pipeline is unchanged.
The KL term regularizes target distribution toward context distribution. Two variants:
- Deterministic (default): drops the stochastic / KL term โ reconstruction loss only. Better on harder tasks (esp. LIBERO-Long).
- +Stochastic: full ELBO. Slightly better on Spatial, but underperforms on Long where domain shift > KL assumption.
- Context bank = in-domain (4 LIBERO suites, split as non-overlapping context vs target sets) + out-of-domain auxiliary (subset of partially open-sourced NVIDIA GR00T data).
- Target bank = target sets of all four LIBERO suites jointly.
Auxiliary task selection deliberately includes structural mismatch:
- Front-view 7-DoF single-arm (similar to LIBERO).
- Side-view 7-DoF single-arm (camera mismatch).
- Bimanual 14-DoF two-arm (action-space mismatch). This is intentionally less curated than CoT-VLA's tightly-aligned auxiliary picks โ argues for robustness to context diversity.
- Context refresh interval K = 200 steps.
- Context batch size b_C = 32 examples per context task.
- Per-task SFT cross-entropy loss on action tokens.
- Trained for 75K steps total vs OpenVLA's 240K.
- Hardware: 8ร A100 80GB GPUs, ~24 hours total vs OpenVLA's ~100 hours total across the four suites โ a ~76% GPU-time reduction (paper's headline figure).
- Eval: one 24GB RTX-4090 GPU.
- Backbone: OpenVLA-7B (Llama2-7B + ViT) primary; ablation on NORA-Long 3B (Qwen2.5-VL-based) for backbone-agnostic claim.
| Model | Steps | Goal | Spatial | Object | Long | Avg |
|---|---|---|---|---|---|---|
| ฯ_{0.5} (Intelligence '25b) | 30K | 98.0 | 98.8 | 98.2 | 92.4 | 96.9 |
| Diffusion Policy | โ | 68.3 | 78.3 | 92.5 | 50.5 | 72.4 |
| ATM | โ | 77.8 | 68.5 | 68.0 | 39.3 | 63.4 |
| TraceVLA | โ | 75.1 | 84.6 | 85.2 | 54.1 | 74.8 |
| OpenVLA (4 separate models) | 240K | 76.2 | 84.7 | 87.0 | 51.8 | 74.9 |
| SFT-4LIBERO (vanilla multi-task) | 75K | 77.8 | 84.8 | 87.4 | 54.7 | 76.2 |
| SFT-4LIBERO+1single+1bi | 75K | 59.7 | 68.0 | 65.2 | 30.0 | 55.7 |
| SFT-4LIBERO+3single | 75K | 24.6 | 16.8 | 9.7 | 1.5 | 13.2 |
| SFT-4LIBERO+5single+1bi | 75K | 15.2 | 5.6 | 12.0 | 1.6 | 8.6 |
| SFT-4LIBERO+5single+1bi (extended) | 187.5K | 23.4 | 16.7 | 13.6 | 4.4 | 14.5 |
| MetaVLA-Pretrained-Context-ONLY | 75K | 74.4 | 85.4 | 85.4 | 52.3 | 74.4 |
| MetaVLA (no aux, deterministic) | 75K | 78.9 | 88.5 | 88.5 | 55.3 | 77.8 |
| MetaVLA + Stochastic | 75K | 78.9 | 88.9 | 88.5 | 53.0 | 77.3 |
| MetaVLA+1single+1bi | 75K | 78.5 | 89.0 | 87.4 | 59.0 | 78.5 |
| MetaVLA+3single | 75K | 78.0 | 88.0 | 87.2 | 59.7 | 78.2 |
| MetaVLA+5single+1bi (full) | 75K | 78.7 | 89.9 | 88.9 | 59.8 | 79.3 |
Headline: MetaVLA+6 aux beats OpenVLA by +4.4% avg (and +8.0% on LIBERO-Long) using 75K steps vs 240K (-68.75% steps). vs SFT-4LIBERO: +3.1% avg, +5.1% on Long.
The contrast within Table 1 is the real story: vanilla SFT collapses catastrophically as auxiliary tasks are added (76.2 โ 8.6 avg with 6 aux tasks). MetaVLA gains from the same 6 aux tasks (77.8 โ 79.3).
| Model | Goal | Spatial | Object | Long | Avg |
|---|---|---|---|---|---|
| NORA-Long | 85.4 | 90.5 | 95.0 | 70.6 | 85.4 |
| NORA-Long-SFT-4LIBERO | 87.0 | 92.5 | 94.0 | 75.5 | 87.3 |
| NORA-Long-SFT-4LIBERO+5single+1bi | 73.6 | 79.5 | 75.2 | 37.2 | 66.4 |
| MetaVLA-NORA-Long | 90.8 | 96.2 | 96.5 | 77.8 | 90.3 |
| MetaVLA-NORA-Long+5single+1bi | 93.8 | 95.8 | 97.2 | 80.2 | 91.8 |
+6.4% over NORA-Long avg with aux tasks; +25.4% over native SFT counterpart with aux tasks.
| Method | Total | Goal step / SR | Spatial step / SR | Object step / SR | Long step / SR |
|---|---|---|---|---|---|
| OpenVLA-120K | 120K | 30K / 71.4 | 10K / 81.2 | 30K / 85.8 | 50K / 44.4 |
| MetaVLA-EACH-120K | 120K | 30K / 76.4 | 10K / 86.1 | 30K / 89.0 | 50K / 55.4 |
| OpenVLA-240K | 240K | 60K / 76.2 | 50K / 84.7 | 50K / 87.0 | 80K / 51.8 |
| MetaVLA-EACH-240K | 240K | 60K / 77.4 | 50K / 85.8 | 50K / 88.5 | 80K / 55.8 |
Action-ANP alone (no co-training) already beats OpenVLA at every checkpoint; per-suite Long task continues to improve through 240K.
Context batch size b_C (Fig. 4): success monotonically increases with b_C โ {4, 8, 16, 32}; b_C = 32 is the sweet spot โ gain saturates and memory remains modest.
Auxiliary task selection (Table 1, three rows): MetaVLA+1single+1bi (78.5), MetaVLA+3single (78.2), MetaVLA+5single+1bi (79.3) โ all beat their SFT counterparts by 22-70 pts, demonstrating robustness to context diversity (different views, different action spaces).
Parameter-size confound (MetaVLA-Pretrained-Context-ONLY): replace context bank with bridge_orig + fractal20220817 (already in OpenVLA pretrain) โ average drops to 74.4 (vs 77.8 with LIBERO context), showing gain is from out-of-distribution informative auxiliary signals, not just added params.
Stochastic vs deterministic ELBO: stochastic helps Spatial (+0.4), comparable on Goal/Object, hurts Long (-2.3 โ 53.0). KL term assumption breaks under bigger domain shift.
Multi-task co-training mechanism (MetaVLA-EACH): Action-ANP alone โ single-suite training without context co-training โ already beats OpenVLA at half the steps. Co-training collapses 4 models to 1 and saves another 45K steps overall.
Inference-cost overhead: only +0.3 ms/token latency (Fig. 9 in appendix) due to lightweight ANP module.
- Backbone breadth. Only OpenVLA-7B and NORA-Long-3B tested; broader VLM/VLA spectrum (e.g. PaliGemma-3B ฯโ-style, Qwen2-VL action experts) remains future work.
- Real-robot deployment. All experiments are LIBERO simulation; real-world validation deferred.
- No formal proof of "scaling more robustly". Authors acknowledge their stability claim relative to vanilla multi-task SFT is empirical; a rigorous proof / theoretical bound is "left to future work due to computational constraints".
- Web-scale context bank not yet tried. They speculate that web-scale data augmentation of the context bank (analogous to ฯโ.5's pretraining co-training) could help further but didn't run it.
- Auxiliary task selection. Explore-exhaust of combinations not done due to memory/compute.
Vs OpenVLA / OpenVLA-OFT. OpenVLA fine-tunes each LIBERO suite as a separate Hugging Face model, totaling ~240K steps across the four suites (Table 1; per-suite breakdown 60K/50K/50K/80K in Table 3). MetaVLA collapses these into one unified model in 75K steps (-68.75% steps) and beats avg by 4.4 pts. OpenVLA-OFT (Kim '25) optimizes inference speed but per the paper's intro still demands ~150Kโ500K SFT steps (incl. diffusion + non-diffusion parts); MetaVLA is orthogonal โ could in principle be stacked.
Vs ฯโ / ฯโ.โ / ฯโ.โ / ฯโ.โ. ฯโ-class models invest heavily in pretraining (large robotics corpus + flow matching). MetaVLA is purely post-training: ฯโ.โ scores 96.9 avg LIBERO via massive pretraining; MetaVLA+aux 79.3 with much cheaper compute. Different points on the cost/performance frontier.
Vs CoT-VLA / OneTwoVLA / EO-1. Those add reasoning data into pretraining (CoT-VLA, OneTwoVLA) or interleaved Vision-Text-Action (EO-1, "high inference latency"). MetaVLA stays at post-training and avoids inference-time overhead.
Vs other efficient-adaptation peers in 2026. Sits with FASTER, HyperVLA, Hybrid Training, LBM Cotraining โ all treat training/adaptation efficiency as a first-class objective. MetaVLA is unique in leveraging meta-learning (ANP) rather than hypernetworks, LoRA-style routing, or curriculum.
Vs GR00T-N1.5. Interestingly, MetaVLA uses GR00T data as auxiliary context โ flips the relationship from competitor to fuel.
The deeper claim is structural: post-training is an under-explored axis vs the dominant pretraining-scaling story. Naive multi-task SFT fails worse with more diverse auxiliary tasks (aggregate -67.6 pt collapse with 6 aux tasks). The ANP context-bank lets the model selectively attend to relevant context, turning auxiliary diversity from a liability into an asset. Reframes VLA adaptation as a meta-learning problem rather than a fresh fine-tune per task.
- OpenReview: https://openreview.net/forum?id=E1K2Ph3LtS
- HyperVLA
- FASTER
- Hybrid Training
- LBM Cotraining
- OpenVLA
โ Back to ICLR-2026