ICML 2026 Latent Reasoning VLA - Heungwoo/research GitHub Wiki
Latent Reasoning VLA (LaRA-VLA) โ internalizing multi-modal chain-of-thought into latent space
Venue: ICML 2026 (Poster) Category: Reasoning Affiliations: Shuanghao Bai, Jing Lyu, Wanqi Zhou, Cheng Chi, Badong Chen, Shanghang Zhang et al. (2026) Traction (2026-06): 7 citations (arXiv)

Problem
Chain-of-thought (CoT) reasoning improves VLA performance, but existing CoT-based VLAs face two issues. First, inference overhead: textual CoT verbalizes long reasoning traces, inflating token length, KV-cache usage, and latency โ such models can run below 5 Hz, or even ~1 Hz, which is unacceptable for real-time control. Second, representational mismatch: textual CoT is constrained to discrete language tokens and visual CoT to discrete VQ-based visual tokens, whereas embodied perception and action evolve in continuous spaces. The authors argue CoT works because it exposes structured intermediate reasoning, not because it is expressed in natural language โ so reasoning can be moved into continuous latent space.
Method
LaRA-VLA is a unified framework that internalizes both textual and visual CoT into continuous latent representations, eliminating explicit CoT generation at inference. Training is a three-stage curriculum:

- Stage I โ Explicit multi-modal CoT: jointly optimizes CoT supervision, a visual-alignment loss, and a discrete action loss (๐_cot + 0.1ยท๐_vis + ๐_act-dis). Future visual prediction latents are aligned with their annotations and supervised by an EMA target encoder to prevent representation collapse.
- Stage II โ Latent transition: progressively anneals the explicit textual CoT loss to zero, replacing textual CoT with a compact set of text latents while visual prediction objectives provide implicit supervision (0.2ยท๐_vis + ๐_act-dis).
- Stage III โ Action generation via flow matching: discrete action supervision is replaced by continuous action regression. An action expert predicts a velocity field conditioned on a multi-modal latent context h_t (current observation + instruction + text-reasoning latent + predicted future-visual latent), trained with a flow-matching loss. No separate action latent is introduced.
A tailored LaRA attention mechanism regulates cross-token flow across text, current-image, future-image, and action tokens. Two structured CoT datasets, LIBERO-LaRA and Bridge-LaRA, supply multi-modal reasoning annotations.
Results
On LIBERO, LaRA-VLA reaches a 97.9% average (96.4 Spatial, 98.6 Goal, 99.8 Object, 96.6 Long), topping the latent-CoT, visual-CoT, and textual-CoT groups including OpenVLA-OFT (97.1), ฯโ.โ (96.8), DeepThinkVLA (97.0), and Fast-ThinkAct (89.7). On SimplerEnv-WidowX it averages 68.8% (95.8 Put Spoon, 62.5 Put Carrot, 25.0 Stack Block, 91.7 Put Eggplant), ahead of UD-VLA (62.5), F1 (59.4), and CogACT (51.3). An ablation on SimplerEnv isolates the contributions: no CoT 55.21% โ explicit text-CoT 58.33% โ latent text-CoT 64.58% โ adding latent visual-CoT 68.75%, showing latent reasoning gives far larger gains than explicit textual CoT. A t-SNE analysis finds no latent collapse โ reasoning-component latents form well-separated, semantically coherent clusters distinct from instruction-token latents. On inference efficiency (A100), LaRA-VLA runs at 135 ms per rollout, an up to 90% reduction versus explicit-CoT baselines. Real-world tests on an Agilex Cobot Magic across four long-horizon tasks beat ACT and GR00T N1.5, with gains most pronounced on multi-stage tasks.
Significance
LaRA-VLA shows that CoT's benefit can be retained while discarding its inference cost: replacing verbose textual/visual reasoning with compact continuous latents yields both higher success rates and an order-of-magnitude latency reduction, aligning reasoning with the continuous nature of perception and control. The curriculum + EMA-stabilized latents address the collapse problem that plagues latent-reasoning approaches.
Links
- arXiv: 2602.01166
- ICML 2026: https://icml.cc/virtual/2026/poster/64290
โ Back to ICML-2026