ICML 2026 LangForce - Heungwoo/research GitHub Wiki

LangForce: Bayesian Decomposition of VLA Models via Latent Action Queries — penalizing the vision shortcut to force instruction following

Venue: ICML 2026 (Poster) Category: VLA Architecture Affiliations: HUST, ZGCA, ZGCI, HIT, HKUST(GZ), ZZU, BUAA, ECNU, DeepCybo Traction (2026-06): 10 citations (arXiv)

Examples of the vision shortcut in RoboCasa: training data has visual diversity but limited task diversity (Figure 1 from Lian et al., 2026)

Problem

Vision-Language-Action (VLA) models generalize poorly to new instructions and multi-task settings. The paper diagnoses a concrete pathology: goal-driven data collection produces datasets where the language instruction is highly predictable from the visual observation alone. In such data the conditional mutual information between instructions and actions vanishes — a phenomenon the authors call Information Collapse — and models degenerate into vision-only policies that ignore the language command and fail out-of-distribution (OOD). The authors demonstrate this with three motivating experiments (the vision shortcut in in-distribution testing, failure in ambiguous scenes, and catastrophic OOD failure).

Method

LangForce (named BayesianVLA in the paper) enforces instruction following through a Bayesian decomposition rather than new data. It introduces learnable Latent Action Queries Q and a dual-branch architecture built on shared VLM weights:

  • Priori Branch (vision-only): input sequence [v, Q, ℓ]. Because of the decoder's causal mask, Q can attend to the visual observation v but not the later language ℓ, so its hidden states encode purely visual information and are trained with a flow-matching loss L_prior to capture the dataset's inherent action bias — i.e., the vision-only prior p(a|v).
  • Posteriori Branch (vision + language): input [v, ℓ, Q], where Q now attends to both vision and language, estimating the true policy π(a|v,ℓ) via a main flow-matching loss L_main.
  • Maximizing the Likelihood Ratio (LLR): the policy is optimized to maximize the conditional Pointwise Mutual Information (PMI) between actions and instructions, using the VLM's language-modeling loss as a proxy for log p(ℓ|…). This penalizes the vision shortcut and rewards actions that explicitly "explain" the language command.

The BayesianVLA dual-branch framework with shared VLM weights and causal masking (Figure 3 from Lian et al., 2026)

Results

"11.3% improvement on the OOD SimplerEnv benchmark". Experiments span SimplerEnv (trained on BridgeDataV2 + Fractal from OXE, evaluated over 480 trials/Avg@480 on four manipulation tasks) and the RoboCasa GR1 Tabletop benchmark (24 tasks, Avg@50). Ablations on a Qwen3-VL-4B backbone show the full BayesianVLA reaching 63.5% versus 57.5% for the "+ Action Query" architectural ablation (a +6.0% gain), confirming the core benefit comes from the dual-branch PMI objective rather than the added queries alone. The paper also shows preserved general multimodal reasoning, avoiding the catastrophic forgetting seen in the QwenGR00T baseline.

Significance

By framing instruction-following failure as an information-theoretic collapse and fixing it with a training objective rather than more data, LangForce offers a cheap, architecture-light remedy to one of the most persistent VLA failure modes — language being ignored in favor of visual shortcuts — while preserving the backbone VLM's reasoning abilities.

Links

← Back to ICML-2026