ICML 2026 From Pixels to Tokens - Heungwoo/research GitHub Wiki
From Pixels to Tokens โ A systematic study of how to supervise VLAs with latent actions
Venue: ICML 2026 (Oral) Category: Analysis-Insight / VLA Architecture Traction (2026-06): 0 citations (arXiv)
![]()
Problem
Latent actions are an intermediate representation that lets VLAs train consistently across heterogeneous data โ different robot platforms, action spaces, and even human video โ by abstracting motion away from inconsistent raw action semantics. But the field's approaches to supervising a VLA with latent actions are fragmented: some use discrete token supervision, others auxiliary alignment objectives, all targeting the VLM backbone differently, with no apples-to-apples comparison. It has been unclear which formulation helps which kind of task, and whether discrete or continuous targets are better.
Method
The paper structures latent-action supervision along two perspectives and instantiates four strategies under a single unified VLA baseline (so differences reflect the integration strategy, not the latent model). Perspective 1 โ image-based latent actions (derived from visual state transitions) regularize the trajectory: Strategy 1 (LA-Align) implicit representation alignment of internal VLM features with latent embeddings; Strategy 2 (LA-Direct) explicit direct decoding, where the VLM is supervised to predict latent actions as an intermediate plan; Strategy 3 (LA-Cond) explicit conditional decoding, jointly predicting latent plans and action representations with the latter conditioned on the former. Perspective 2 โ action-based latent actions unify the target space by discretizing continuous actions into tokens: Strategy 4 (LA-Tok) action-to-token mapping, supervising the VLM to predict those tokens. The training objective combines the continuous action-head loss with a weighted latent supervision loss L = L_action + ฮปยทL_latent.
![]()
Results
A formulation-task correspondence emerges. Image-based latent actions help long-horizon reasoning and scene generalization: on LIBERO-Long image strategies gain +8.4% to +10.8% over baseline (vs. +6.8% for action-based LA-Tok), and on the real "Stack 4 Bowls" task LA-Direct scores 79 vs. 48 for baseline. Action-based latent actions excel at complex motor coordination: on RoboTwin 2.0 LA-Tok improves average success by +17.5%. On integration strategy: explicit supervision beats implicit alignment (LA-Direct 96.6% vs. LA-Align 94.8% on LIBERO-Long), and explicit direct decoding generally beats conditional decoding (LA-Direct 96.6% vs. LA-Cond 94.2%). The headline finding: directly supervising the VLM to predict discrete latent action tokens is most effective โ discrete supervision beats continuous regression by +2.7% / +2.2% average. Latent actions also reduce negative transfer in multi-task joint training: LA-Cond reaches 74.5% avg on 10 RoboTwin tasks (+20.9% over baseline, no per-task degradation). On real-world OOD scene generalization the baseline collapses to a 17 OOD average while LA-Cond reaches 70 and LA-Direct gets the best total average (75). Ablations confirm gains come from latent integration, not longer action placeholders.
Significance
This is an oral that brings order to a fragmented design space: rather than proposing a new model, it offers a unified taxonomy and controlled comparison that tells practitioners which latent-action supervision to use for which task regime, and provides strong evidence that discrete latent-action token supervision of the VLM is the most consistent, effective signal โ including for mixed-data training. Code is released.
Links
- arXiv: 2605.04678
- ICML 2026: https://icml.cc/virtual/2026/poster/63621
โ Back to ICML-2026