RSS 2026 OAT - Heungwoo/research GitHub Wiki

OAT โ€” Ordered Action Tokenization

Venue: RSS 2026 (Imitation Learning session) ยท Authors: Chaoqi Liu, Xiaoshen Han, Jiawei Gao, Yue Zhao, Haonan Chen, Yilun Du โ€” Harvard ร— Stanford ยท arXiv: 2602.04215 ยท project ยท code Category: Action tokenization for autoregressive policies Trend tag: RSS 2026 thread 7 โ€” action representation & inference mechanics

Compiled from the verified RSS 2026 abstract and the paper's Fig. 1.

Key figure

OAT overview (Figure 1 of arXiv 2602.04215, ยฉ the authors)

Figure 1 of the paper. Left: the three-desiderata Venn diagram โ€” Bin satisfies decodability, FAST compression, QueST sits between, and only OAT occupies the intersection of high compression + total decodability + causal ordering. Middle: the inference-behavior comparison that motivates the design: diffusion refines whole trajectories over iterations, AR-Bin decodes but too slowly, AR-QueST/FAST pass through undecodable intermediate states (grey), while AR-OAT is decodable at every prefix โ€” early tokens give a coarse but valid action chunk that later tokens refine ("progressive JPEG for motion"). Right: aggregate over 20+ tasks โ€” OAT with 8 tokens hits 52.3%, above QueST (45.9), Diffusion Policy (41.0), Bin (20.5) and FAST (20.3); even 1-, 2-, and 4-token prefixes score 24.1 / 39.2 / 45.0, quantifying the anytime compute-fidelity dial.

Problem

Autoregressive policies need discrete action tokens, but current schemes fail one way or the other: analytical discretizers (ร  la FAST) produce prohibitively long sequences; learned latent tokenizers lack the structure next-token prediction wants.

Method

Three stated desiderata โ€” high compression, total decodability, left-to-right causal order โ€” met by a learned tokenizer combining:

  • Transformer with registers + finite scalar quantization (FSQ) over action chunks;
  • Ordering-inducing training mechanisms that make the token sequence causally ordered โ€” so earlier tokens carry coarse action content and later tokens refine it;
  • Prefix-based detokenization: decode any prefix into a valid (coarser) action โ€” an anytime trade-off between inference cost and action fidelity.

Results (as reported)

  • Across 20+ tasks over four simulation benchmarks and real-world settings, AR policies with OAT consistently beat prior tokenization schemes and diffusion-based baselines, with greater inference-time flexibility.

Significance

A pointed counter-timing: the same conference's LBM study reconfirms that discrete action tokens don't help as a co-training signal โ€” while OAT argues they can win as the primary action representation if the token space is properly ordered. The anytime prefix-decoding property is genuinely new for action tokenizers (a compute-quality dial diffusion heads lack), and the ordered-FSQ recipe is a direct upgrade candidate for every Cat-A autoregressive VLA in Review-VLA-Architecture. Whether ordered tokens revive the AR-policy line against flow matching is now a live question.

โ† RSS 2026 survey ยท Home