ICML 2026 LangForce - Heungwoo/research GitHub Wiki
LangForce: Bayesian Decomposition of VLA Models via Latent Action Queries — penalizing the vision shortcut to force instruction following
Venue: ICML 2026 (Poster) Category: VLA Architecture Affiliations: HUST, ZGCA, ZGCI, HIT, HKUST(GZ), ZZU, BUAA, ECNU, DeepCybo Traction (2026-06): 10 citations (arXiv)

Problem
Vision-Language-Action (VLA) models generalize poorly to new instructions and multi-task settings. The paper diagnoses a concrete pathology: goal-driven data collection produces datasets where the language instruction is highly predictable from the visual observation alone. In such data the conditional mutual information between instructions and actions vanishes — a phenomenon the authors call Information Collapse — and models degenerate into vision-only policies that ignore the language command and fail out-of-distribution (OOD). The authors demonstrate this with three motivating experiments (the vision shortcut in in-distribution testing, failure in ambiguous scenes, and catastrophic OOD failure).
Method
LangForce (named BayesianVLA in the paper) enforces instruction following through a Bayesian decomposition rather than new data. It introduces learnable Latent Action Queries Q and a dual-branch architecture built on shared VLM weights:
- Priori Branch (vision-only): input sequence
[v, Q, ℓ]. Because of the decoder's causal mask,Qcan attend to the visual observationvbut not the later languageℓ, so its hidden states encode purely visual information and are trained with a flow-matching lossL_priorto capture the dataset's inherent action bias — i.e., the vision-only priorp(a|v). - Posteriori Branch (vision + language): input
[v, ℓ, Q], whereQnow attends to both vision and language, estimating the true policyπ(a|v,ℓ)via a main flow-matching lossL_main. - Maximizing the Likelihood Ratio (LLR): the policy is optimized to maximize the conditional Pointwise Mutual Information (PMI) between actions and instructions, using the VLM's language-modeling loss as a proxy for
log p(ℓ|…). This penalizes the vision shortcut and rewards actions that explicitly "explain" the language command.

Results
"11.3% improvement on the OOD SimplerEnv benchmark". Experiments span SimplerEnv (trained on BridgeDataV2 + Fractal from OXE, evaluated over 480 trials/Avg@480 on four manipulation tasks) and the RoboCasa GR1 Tabletop benchmark (24 tasks, Avg@50). Ablations on a Qwen3-VL-4B backbone show the full BayesianVLA reaching 63.5% versus 57.5% for the "+ Action Query" architectural ablation (a +6.0% gain), confirming the core benefit comes from the dual-branch PMI objective rather than the added queries alone. The paper also shows preserved general multimodal reasoning, avoiding the catastrophic forgetting seen in the QwenGR00T baseline.
Significance
By framing instruction-following failure as an information-theoretic collapse and fixing it with a training objective rather than more data, LangForce offers a cheap, architecture-light remedy to one of the most persistent VLA failure modes — language being ignored in favor of visual shortcuts — while preserving the backbone VLM's reasoning abilities.
Links
- arXiv: 2601.15197
- ICML 2026: https://icml.cc/virtual/2026/poster/65457
← Back to ICML-2026