ICLR 2026 FASTER - Heungwoo/research GitHub Wiki

FASTer — Powerful and Efficient AR VLAs with Learnable Action Tokenizer and Block-wise Decoding

Venue: ICLR 2026 Category: Action Tokenizer Trend tag: Trend 4 (cross-embodiment / efficiency) arXiv: 2512.04952 ("FASTer: Toward ... with Learnable Action Tokenizer and Block-wise Decoding") Authors / affiliations: Yicheng Liu, Shiduo Zhang, Zibin Dong, et al. — Tsinghua University, Fudan University, Shanghai Innovation Institute, Galaxea AI, UCSD, et al.

Approach diagram

flowchart LR
  A[Raw action chunk<br/>non-uniform semantic patches:<br/>EE pos / orientation / gripper] --> E[Hybrid transformer encoder]
  E --> R[RVQ<br/>residual vector quantization]
  R --> T[Discrete action tokens<br/>high compression ratio]
  R --> Loss{Reconstruction loss}
  Loss --> L1[Time-domain L1<br/>local dynamics]
  Loss --> DCT[Frequency-domain DCT L1<br/>global trends]
Loading

Problem

The FAST action tokenizer (2025) compressed action chunks into discrete tokens for AR or diffusion-based VLAs, but its compression ratio was modest and it treated all action dimensions equally. Higher compression directly reduces the number of tokens generated per control step.

FASTer has two components: FASTerVQ (the learnable tokenizer) and FASTerVLA (the autoregressive policy built on it).

  • Non-uniform semantic patchification: action dimensions are grouped non-uniformly by physical meaning (end-effector position, orientation, gripper) and partitioned temporally, rather than treating all dimensions uniformly like FAST.
  • Hybrid transformer encoder + Residual Vector Quantization (RVQ): the encoder extracts latents that are quantized through N_c codebook levels (layered quantization → higher capacity at the same vocabulary size). Action chunks are encoded as single-channel "images" to capture global spatio-temporal dependencies.
  • Dual-domain reconstruction loss: an L1 loss on the temporal (time-domain) action signal for local dynamics, plus an L1 loss on the Discrete Cosine Transform (DCT / frequency-domain) to capture global trends.
  • Block-wise decoding (FASTerVLA): discrete codes are split into J contiguous blocks of size B; within a block tokens attend to preceding and intra-block tokens, reducing forward passes from N to ≈N/B. Decoding proceeds codebook-first, then across the horizon.
  • Lightweight action expert: a parameter-efficient expert architecturally aligned with the VLM backbone bridges language and continuous control.

Results

  • LIBERO: 97.9% success vs. π0-FAST 94.2%.
  • Simpler-Bridge: 87.9% success (≈12.9 points over the next-best baseline).
  • Inference latency: ~112 ms (LIBERO single-arm) and ~237 ms (R1Lite whole-body) for FASTer vs. ~197–556 ms / ~1,100–3,000 ms for π0-FAST — up to ≈3× speedup from block-wise decoding (typically 3 blocks).
  • Higher compression ratio than FAST while maintaining strong reconstruction (>99% codebook utilization), enabling larger backbones at the same inference budget.

Significance

Action tokenizer design is becoming a competitive sub-area. FASTer establishes that domain-specific priors — dual time/frequency (DCT) reconstruction and non-uniform semantic action grouping — meaningfully improve tokenizer quality, and that co-designing the tokenizer with the decoder (high compression + block-wise decoding) yields both higher task success and faster inference (up to ~3×) over the FAST-based π0-FAST baseline.

Links

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️