ICML 2026 StableVLA - Heungwoo/research GitHub Wiki
StableVLA — Information-bottleneck adapters for robust VLAs without extra data
Venue: ICML 2026 (Poster) Category: VLA Architecture Affiliations: Peking University, Tsinghua University, Astribot, Nanjing University, Nankai University (Yiyang Fu, Chubin Zhang, Shukai Gong, Yufan Deng, Kaiwei Sun, Qiyang Min, Qibin Hou, Yansong Tang, Jianan Wang, Daquan Zhou) Traction (2026-06): 1 citation (arXiv)

Problem
A training set cannot enumerate every real-world visual disturbance. The authors systematically study state-of-the-art Vision-Language-Action (VLA) models under unseen visual corruptions and find a significant performance drop when disturbances absent from training (sensor noise, blur, physical occlusions) appear at test time. The goal is intrinsic robustness — without extra data, augmentation, or larger backbones.
Method
StableVLA reframes the VLA modality-alignment projector through an Information Bottleneck (IB) lens. A standard VLA uses a visual encoder, an MLP projector, and an LLM policy. The authors argue MLP projectors act as all-pass filters, indiscriminately maximizing mutual information I(X_v; Z) and propagating noise downstream. They instead minimize an IB objective I(X_v; Z) − β·I(Z; S), compressing nuisances while retaining task-relevant semantics S.
- IB-Adapter. Proposition 3.1 shows that, under Gaussian/latent-structure assumptions, the optimal IB update corresponds to a channel-wise attention operation Z = V·σ(β·QᵀK), with σ being Softmax (categorical latent) or Sigmoid (independent Bernoulli latent). Crucially, grouping is performed across the channel dimension (each channel is an information unit), since semantics and noise are heterogeneously distributed across channels. A multi-head design captures correlations across H semantic subspaces; an identity-key (K = X′) preserves high-frequency spatial cues, and Sigmoid gating suppresses noisy channels.
- Fused IB-Adapter (hybrid). Because pure IB filtering can attenuate high-frequency detail needed for precise, long-horizon manipulation, StableVLA uses a dual pathway: Z = MLP(X) + tanh(λ)·IB-Adapter(X) — an MLP high-fidelity path plus a denoising IB path. Stochastic Pathway Dropout (SPD) calibrates the balance per task.
The adapter adds fewer than 10M parameters.

Results
- Headline. IB-Adapter consistently improves over the baseline by an average of 30% with <10M added parameters; with a 14× smaller backbone (0.5B params) and no Open X-Embodiment pre-training, StableVLA matches the robustness of 7B-scale SOTA VLAs.
- LIBERO + CALVIN (Table 1). StableVLA (0.5B) is best or second-best across all suites, rivaling OpenVLA-OFT (7B) and OpenPi-0.5 (3B). Versus the same-architecture VLA-Adapter (0.5B), it improves 40.2% to 139.6% across the four LIBERO suites at the most severe corruption (severity 5) — e.g. LIBERO-Long rises from 58.5→82.0 (clean-corruption avg) and Object S5 from 29.3→70.2.
- Real-world (Table 2). On a physical robot under Noise/Blur/Oil/Shelter corruptions, StableVLA shows the smallest average performance drop (e.g. Pick-and-place avg Δ −17.5 vs. −30.1 for π₀.₅ and −49.2 for VLA-Adapter).
- Mechanism. K-means (K=2) on output features shows MLP projectors produce diffused features that conflate targets with background under noise, while the Fused IB-Adapter maintains coherent, object-centric clusters — the covariance-based Sigmoid gating suppresses low-correlation stochastic corruptions.
Significance
StableVLA turns robustness into an architectural property of the projector rather than a data problem, deriving a channel-wise attention adapter directly from the IB principle. With negligible overhead and no extra data, a 0.5B model reaches robustness competitive with 7B foundation VLAs, making it an attractive drop-in for deployment under imperfect visual conditions.
Links
- arXiv: 2605.18287
- ICML 2026: https://icml.cc/virtual/2026/poster/63066
← Back to ICML-2026