ICML 2026 EnsembleVLA - Heungwoo/research GitHub Wiki
Venue: ICML 2026 (Poster) Category: VLA Architecture Affiliations: Mingchen Song, Xiang Deng, Jie Wei, Dongmei Jiang, Liqiang Nie, Weili Guan
Many diverse vision-language-action (VLA) models now exist and perform well at robotic manipulation, but how to effectively ensemble them to push performance further is largely unexplored. Conventional ensemble techniques are designed for discriminative tasks and cannot be applied directly to generative action policies, which produce high-dimensional, multimodal action distributions. EnsembleVLA targets this gap: combining several pre-trained VLAs in a principled way.
EnsembleVLA is an energy-based framework for principled ensembling of diverse VLA models. Its central theoretical claim is a unified formulation showing that both diffusion-based and flow-based VLA models can be written as energy-based models (EBMs), where additive combination of their energies naturally induces policy composition at the distribution level. Under this view, multiple pre-trained policies can be aggregated into a single stronger ensemble policy without retraining the base models. On top of this compositional foundation, EnsembleVLA adds:
- Learnable composition weights for dynamic policy balancing across the constituent VLAs.
- A confidence-aware gating mechanism that adaptively modulates bounded residual corrections, which the authors credit for stable and robust task execution.
flowchart LR
O[Observation + language] --> V1[Pre-trained VLA 1<br/>diffusion/flow]
O --> V2[Pre-trained VLA 2]
O --> Vn[Pre-trained VLA N]
V1 -->|energy E1| C[Additive energy<br/>composition]
V2 -->|energy E2| C
Vn -->|energy En| C
W[Learnable composition weights] --> C
C --> G[Confidence-aware gating<br/>bounded residual correction]
G --> A[Ensembled action]
The abstract reports that "EnsembleVLA achieves competitive performance across various tasks in both simulated and real-world environments." No specific numeric results are given in the source available (ICML abstract only); the paper has no arXiv preprint listed, so concrete benchmark figures, ablations, and the set of base VLAs ensembled could not be verified here.
EnsembleVLA reframes VLA ensembling as energy composition, giving a single theory that covers both of the dominant generative action-policy families (diffusion and flow). If the empirical claims hold, it offers a training-free way to combine heterogeneous off-the-shelf VLAs into a stronger policy — an attractive direction as the 2026 landscape accumulates many specialized open-weight manipulation models.
- ICML 2026: https://icml.cc/virtual/2026/poster/61698
← Back to ICML-2026