ICML 2026 LaST0 - Heungwoo/research GitHub Wiki
LaST₀: Latent Spatio-Temporal Chain-of-Thought for Robotic Vision-Language-Action Models — reasoning before acting in a token-efficient latent space
Venue: ICML 2026 (Poster) Category: Reasoning
Chain-of-thought (CoT) reasoning can help Vision-Language-Action (VLA) models "think before they act," but expressing that reasoning as natural-language tokens is slow and ill-suited to the physical, spatio-temporal nature of manipulation. A robot's relevant intermediate state is not text — it is future visual dynamics, 3D structure, and proprioception. LaST₀ targets a reasoning representation that captures these physical factors while remaining cheap enough for closed-loop control.
LaST₀ is a VLA framework that reasons before acting in a token-efficient latent spatio-temporal chain-of-thought space. Instead of emitting language CoT, it reasons over a compact latent that encodes:
- future visual dynamics (how the scene will evolve),
- 3D structure of the scene, and
- proprioception (the robot's own state).
To reconcile the differing timescales of deliberation and control, LaST₀ adopts a Mixture-of-Transformers dual-system design that separates a low-frequency reasoning expert from a high-frequency action expert. The reasoning system operates over the latent spatio-temporal CoT at a slower cadence, while the action system runs at a higher rate to produce control — analogous to "System 2" deliberation feeding a fast "System 1" actuator.
flowchart LR
Obs[Visual + proprio observation] --> R[Low-frequency reasoning expert]
R --> CoT[Latent spatio-temporal CoT:<br/>future dynamics · 3D · proprioception]
CoT --> A[High-frequency action expert]
A --> ACT[Action]
subgraph MoT[Mixture-of-Transformers]
R
A
end
LaST₀ reports improvements in mean success rate of 13%, 14% and 14% over prior state-of-the-art VLA methods across its evaluated settings. (Detailed per-benchmark tables are not available, as no public arXiv/HTML source was accessible for this entry at the time of writing; the figures above are from the submission's reported headline.)
LaST₀ argues that effective robotic reasoning should be latent and physical, not textual: by compressing the chain-of-thought into a spatio-temporal latent over future dynamics, 3D structure, and proprioception, it keeps deliberation token-efficient enough for control. The Mixture-of-Transformers dual-system split — slow reasoning, fast action — is a clean architectural answer to the timescale mismatch between thinking and acting, and the consistent double-digit success-rate gains suggest the latent-CoT formulation is a meaningful alternative to language-based CoT for VLAs.
- ICML 2026: poster (link pending)
← Back to ICML-2026