ICML 2026 LaST0 - Heungwoo/research GitHub Wiki

LaST₀: Latent Spatio-Temporal Chain-of-Thought for Robotic Vision-Language-Action Models — reasoning before acting in a token-efficient latent space

Venue: ICML 2026 (Poster) Category: Reasoning

Problem

Chain-of-thought (CoT) reasoning can help Vision-Language-Action (VLA) models "think before they act," but expressing that reasoning as natural-language tokens is slow and ill-suited to the physical, spatio-temporal nature of manipulation. A robot's relevant intermediate state is not text — it is future visual dynamics, 3D structure, and proprioception. LaST₀ targets a reasoning representation that captures these physical factors while remaining cheap enough for closed-loop control.

Method

LaST₀ is a VLA framework that reasons before acting in a token-efficient latent spatio-temporal chain-of-thought space. Instead of emitting language CoT, it reasons over a compact latent that encodes:

  • future visual dynamics (how the scene will evolve),
  • 3D structure of the scene, and
  • proprioception (the robot's own state).

To reconcile the differing timescales of deliberation and control, LaST₀ adopts a Mixture-of-Transformers dual-system design that separates a low-frequency reasoning expert from a high-frequency action expert. The reasoning system operates over the latent spatio-temporal CoT at a slower cadence, while the action system runs at a higher rate to produce control — analogous to "System 2" deliberation feeding a fast "System 1" actuator.

flowchart LR
    Obs[Visual + proprio observation] --> R[Low-frequency reasoning expert]
    R --> CoT[Latent spatio-temporal CoT:<br/>future dynamics · 3D · proprioception]
    CoT --> A[High-frequency action expert]
    A --> ACT[Action]
    subgraph MoT[Mixture-of-Transformers]
      R
      A
    end
Loading

Results

LaST₀ reports improvements in mean success rate of 13%, 14% and 14% over prior state-of-the-art VLA methods across its evaluated settings. (Detailed per-benchmark tables are not available, as no public arXiv/HTML source was accessible for this entry at the time of writing; the figures above are from the submission's reported headline.)

Significance

LaST₀ argues that effective robotic reasoning should be latent and physical, not textual: by compressing the chain-of-thought into a spatio-temporal latent over future dynamics, 3D structure, and proprioception, it keeps deliberation token-efficient enough for control. The Mixture-of-Transformers dual-system split — slow reasoning, fast action — is a clean architectural answer to the timescale mismatch between thinking and acting, and the consistent double-digit success-rate gains suggest the latent-CoT formulation is a meaningful alternative to language-based CoT for VLAs.

Links

  • ICML 2026: poster (link pending)

← Back to ICML-2026

⚠️ **GitHub.com Fallback** ⚠️