Review Stellar VLA - Heungwoo/research GitHub Wiki

In-Depth Review — Stellar VLA: Continually Evolving Skill Knowledge in Vision-Language-Action Models

Paper: Continually Evolving Skill Knowledge in Vision Language Action Model Authors: Yuxuan Wu, Guangming Wang, Zhiheng Yang, Tianchen Deng, Maoqing Yao, Brian Sheil, Hesheng Wang Affiliations: Shanghai Jiao Tong University · Shanghai Innovation Institute · University of Cambridge · Beihang · NTU · MIT SMART · AgiBot arXiv: 2511.18085 (v4 May 8, 2026; v1 Nov 2025, cs.RO) · Project: stellarvla.github.io Status: preprint — indexed here under Latest Papers

Companion reviews: RL for VLA · LBM Co-training · VLA Evaluation · VLA Architectures · VLA Memory. Related insight: ICML 2026's Pretrained VLAs Resist Forgetting (Oral) — the finding this paper builds against.


1. TL;DR

  1. Continual imitation learning (CIL) for VLAs without growing the network. Stellar VLA lets a fixed-size ~1B VLA learn a stream of tasks while retaining prior skills, using only 1% data replay — versus conventional CIL methods that bolt on adapters/task-modules (parameter growth, storage cost) and VLA-replay work that needs ~20% replay.
  2. A self-evolving knowledge space is the core idea. Task-relevant knowledge is organized as a Dirichlet-Process-based cluster space with an unbounded number of components, so new task/skill clusters emerge automatically as tasks arrive — no predefined component count, no manual annotation of task identity.
  3. Two variants, flat vs hierarchical. T-Stellar uses a Dirichlet Process Mixture Model (DPMM) for flat task-centric clusters; TS-Stellar extends to a Hierarchical Dirichlet Process (HDP) that models a task→skill structure where reusable subskills are shared across tasks (task distribution = aggregation over shared skill atoms).
  4. Knowledge-guided expert routing = specialization at no extra params. A diffusion-based MoE action head (MoDE-style) is routed by the knowledge space — via a knowledge-relation embedding (distance of the task latent to cluster centers, weighted by posterior membership) and Top-K semantic embeddings — instead of MoDE's noise-level routing, giving task-specific parameter sharing/differentiation without increasing model size.
  5. State-of-the-art CIL on LIBERO + real dual-arm. Against both VLA baselines (MoDE, UniVLA, π0, π0.5) and CIL baselines (ER, SeqLoRA, LoTUS, IsCiL), Stellar variants take the best AUC and Final SR across LIBERO-goal / -long / -30*, with >50% average AUC/Final-SR improvement in the from-scratch setting and ~20% Final-SR improvement over CIL baselines; on a real dual-arm 7-task suite TS-Stellar reaches 90.0% Final SR (7.4 NBT) vs π0.5's 72.9% (25.8 NBT).

2. Why this paper matters

  • It reframes VLA continual learning as knowledge modeling, not parameter management. The lifelong-robot problem is usually attacked with replay buffers or parameter isolation. Stellar argues the leverage is capturing task/skill relationships so related tasks share parameters and unrelated ones don't interfere — pushing the RL/continual discussion from "how much to replay / how many adapters" toward "what structure to learn."
  • It's the constructive answer to the ICML 2026 "VLAs resist forgetting" Oral. That insight paper showed large pretrained VLAs barely forget with simple replay — implicitly challenging complex CIL designs. Stellar takes the challenge seriously (fixed size, minimal replay) but shows a structured knowledge space still buys real gains, especially from scratch and on long-horizon/hierarchical tasks where plain replay is weakest.
  • Non-parametric Bayesian structure meets VLA MoE. Bringing DPMM/HDP (unbounded, auto-discovering clusters; memoized variational Bayes for incremental updates) into VLA routing is a genuinely different mechanism from the fixed-expert MoEs elsewhere in Review-VLA-Architecture — the number of task/skill clusters grows with experience rather than being a hyperparameter.
  • Hierarchical skill sharing is validated where it should help. TS-Stellar's edge concentrates on long-horizon and compositional manipulation (LIBERO-long, real bimanual handovers), the regime where reusing subskills across tasks is the natural inductive bias.

3. Method

Stellar VLA architecture (Figure 2 of arXiv 2511.18085, © the authors)

Figure 2 of the paper. A CLIP/FiLM-conditioned encoder turns language + image into embeddings; a VAE latent encoder/decoder learns a task-centric latent z (reconstruction + DP-aware KL). The Knowledge Space (center) co-evolves with z: new-task latents update Dirichlet-Process clusters while a L_KL term pulls z toward current clusters ("Co-Evolution"). The learned KS-prior routing then guides an MoE Diffusion Transformer action head (Attention + Expert blocks ×K) that outputs the action chunk. Bottom strip: past-task latents aggregate into clusters; a new task spawns/updates knowledge components.

3.1 Setting

Standard CIL: tasks {T_j} arrive sequentially, each with expert demos (language, obs, actions); the agent must learn new tasks with limited access to past data while retaining old skills. Stellar uses Experience Replay with a tiny buffer (1% of past demos; 5% for the real-world experiments) and adds structure so that small replay doesn't drift.

3.2 Dirichlet-Process knowledge space

  • DP prior G ∼ DP(α, G₀) supports clustering with an unbounded number of components — clusters are shared across tasks but new ones emerge as tasks evolve.
  • T-Stellar (DPMM): task latent z_j ∼ F_task(θ_j), θ_j ∼ G — dynamic clustering of task representations.
  • TS-Stellar (HDP): each task is a distribution over skills (z_ji ∼ F_skill(θ_ji), θ_ji ∼ G_j, G_j ∼ DP(γ,G), G ∼ DP(α,G₀)); the task-level Gaussian parameters are aggregated from shared skill atoms with task-specific mixture weights — so subskills transfer across tasks.

3.3 Self-evolution (co-training z and the knowledge space)

A hierarchical variational-inference loop (Algorithm 1): a VAE infers task-centric latents from vision-language input (reconstruction loss L_recon + DP-aware L_KL toward current clusters); periodically the knowledge distribution Θ is updated from sampled latents via memoized variational Bayes (memoVB) for efficient incremental global-statistic sharing. TS-Stellar decodes task latents → language goals and skill latents → visual observations separately, with an HDP-structured KL. The result is a self-reinforcing cycle that retains old and discovers new task/skill knowledge.

3.4 Knowledge-guided expert routing

A diffusion MoE action head (à la MoDE) where routing is conditioned on the knowledge space instead of denoising level:

  • Knowledge-relation embedding f_R = Σ_k p_k · |z − μ_k| (posterior membership p_k × distance to cluster centers) — a fixed-dim summary despite variable cluster counts.
  • Top-K semantic embeddings for the most relevant clusters. These route experts to give task-specific specialization + related-task sharing without adding parameters — the whole model stays ~1B.

4. Results

4.1 LIBERO CIL (Final SR, higher better)

Metrics: FWT (forward transfer), NBT (negative backward transfer = forgetting, lower better), AUC (success-rate-curve stability), Final SR (after all tasks). 100 trials × 50 init states, 3 seeds on -goal/-long.

Benchmark Best VLA baseline (Final SR) Best CIL baseline T-Stellar TS-Stellar
LIBERO-goal (scratch) π0 35.7 — 67.9 64.2
LIBERO-long (scratch) π0 12.0 — 34.2 35.0
LIBERO-30* (scratch) MoDE 28.5 — 42.9 42.6
LIBERO-goal (CIL, pretrained on -90) ER 55.3 ER 55.3 62.1 57.3
LIBERO-long (CIL, pretrained on -90) ER 16.1 LoTUS 31.8 36.3 40.9

Reported summary: Stellar variants take best AUC and Final SR across all scenarios; >50% average AUC/Final-SR improvement over all baselines in the scratch setting; ~20% Final-SR improvement over CIL baselines with only 1% replay and no parameter growth (vs LoTUS/IsCiL which add parameters). TS-Stellar leads specifically on long-horizon tasks. Some baselines post low NBT only because their FWT is also very low (the forward/backward transfer trade-off).

4.2 Real-world dual-arm (7 tasks, 5% replay, 10 trials each)

Metric ER UniVLA π0 π0.5 T-Stellar TS-Stellar
FWT ↑ 97.1 70.0 98.6 95.7 98.6 98.6
NBT ↓ 21.9 37.1 34.6 25.8 12.4 7.4
AUC ↑ 79.9 43.6 72.6 75.6 89.6 93.4
Final SR ↑ 70.0 21.4 57.1 72.9 84.3 90.0

TS-Stellar shows the lowest forgetting (NBT 7.4) and highest retention (Final SR 90.0) on a new embodiment with compositional/bimanual tasks (e.g., "Handover Toy" after training on "Pull Stick from Bag"), confirming the hierarchical-skill hypothesis transfers to real hardware.


5. Significance & positioning

  • A structured middle path in the CIL debate. Between "just replay, VLAs barely forget" (ICML Oral) and "add adapters/modules per task" (LoTUS, IsCiL), Stellar keeps the model fixed-size with 1% replay yet recovers large gains — strongest exactly where plain replay is weakest (from-scratch, long-horizon, hierarchical). It refines, rather than overturns, the forgetting-resistance finding.
  • Non-parametric knowledge as the routing signal. Auto-discovering task/skill clusters and using them to route a diffusion MoE is a distinct mechanism from fixed-expert or noise-routed MoEs in Review-VLA-Architecture; it grows structure with experience without growing parameters — relevant to the "lifelong VLA" thread the field is opening.
  • Ties to the data-efficiency agenda. 1% replay / ~1B params is a storage-and-compute argument aligned with the co-training and evaluation-cost concerns in Review-LBM-Cotraining and Review-VLA-Evaluation.

6. Limitations

6.1 Visible in the paper

  • LIBERO-30* is single-run (cost), so those numbers carry less statistical weight than the 3-seed -goal/-long results.
  • Cross-embodiment pretraining is harder for the method — the paper notes pretraining on LIBERO-90 (aligned embodiment) transfers more cleanly than pretraining on 1k+ cross-embodiment tasks, which needs more parameter updates; the clean-transfer story is strongest within-embodiment.
  • Task/skill count and horizon are modest (LIBERO suites + 7 real tasks); very long task streams (hundreds of tasks) aren't tested.

6.2 Reviewer's notes

  • The DP/HDP + memoVB machinery adds conceptual and implementation complexity; the paper's own framing (against "increasingly complex CIL designs") invites the question of whether the knowledge-space gains justify it versus tuned replay at larger buffers — an ablation of replay-rate × structure would sharpen this.
  • Backbone is a ~1B CLIP/FiLM+ResNet+diffusion-MoE stack, not a frontier VLM-initialized VLA; whether the knowledge-space benefit persists on top of a strong pretrained VLM (π/GR00T-class) is untested.
  • Multiple arXiv versions (v1 Nov 2025 → v4 May 2026); numbers here are from the latest revision — treat as an evolving preprint.

7. Links & related pages

← Back to Latest Papers · Home · Reviews