ICLR 2026 OmniSAT - Heungwoo/research GitHub Wiki

OmniSAT — Compact Action Token, Faster Auto Regression

Venue: ICLR 2026 (arXiv:2510.09667, Oct 2025) Category: Action Tokenizer Trend tag: Trend 4 Authors: Huaihai Lyu, Chaofan Chen, Senwei Xie, Pengwei Wang, Xiansheng Chen, Shanghang Zhang, Changsheng Xu

Approach diagram

flowchart LR
  A[Variable-length action chunk] --> CE[Consistency Encoding<br/>B-Spline fit to fixed-length control points]
  CE --> SP[Split into subspaces:<br/>position / rotation / gripper]
  SP --> RQ[Multi-stage residual VQ<br/>per-subspace codebooks<br/>pos 256 / rot 256 / grip 64]
  RQ --> T[Compact discrete tokens<br/>6.8x sequence compression]
  T --> AR[Faster auto-regressive VLA training]
Loading

Problem

Auto-regressive (AR) VLA policies are appealing for large-scale pretraining but are bottlenecked by long action-token sequences, which slow convergence and raise target entropy. Existing tokenizers trade off compression against reconstruction fidelity (e.g., FAST achieves near-lossless reconstruction but only ~3.7x compression). OmniSAT targets higher-rate compression with preserved fidelity to accelerate AR training, while also providing a unified action-pattern space for cross-embodiment learning from robot and human demonstrations.

Method

Two-stage tokenization:

  1. Consistency Encoding — variable-length, variable-rate trajectories are fit with a B-Spline into temporally aligned, fixed-length control-point representations (normalizes value range and temporal horizon).
  2. Multi-stage residual quantization — control-point features are partitioned into three separate subspaces (position, rotation, gripper), each quantized independently with its own residual VQ codebook (position 256, rotation 256, gripper 64 entries), yielding coarse-to-fine discrete tokens.

Note: the codebooks are subspace-specific, not a single shared global vocabulary; cross-embodiment sharing comes from the unified normalized action-pattern space, not from one codebook.

Results

  • Compression (DROID, Table 1): OmniSAT-8 reaches 6.8x sequence-length reduction vs FAST 3.7x and BEAST 4.6x, while keeping reconstruction MAE at 9.4e-4 (mm-level).
  • LIBERO: 93.4% average success (Object 98.7 / Goal 94.6 / Spatial 94.1 / Long 86.0), rank 1.
  • SimplerEnv-WidowX: 55.2% overall vs BEAST 37.5%.
  • Real robot (dual-arm AgileX Cobot Magic): PlaceObj 73% / ZipSeal 63% / TubeRack 48% (vs BEAST 63/45/23); adding human egocentric video data lifts these to 80/66/58.
  • Compression translates to faster AR convergence (e.g., ~2.5k vs ~4k steps to converge on LIBERO).

Significance

Complements X-VLA (soft prompts) and XR-1 (UVMC) — all three address cross-embodiment scaling at the tokenizer / interface level rather than the backbone level. OmniSAT's distinctive lever is amplifying AR efficiency via high-rate B-Spline + residual-VQ compression while retaining mm-level reconstruction, jointly leveraging robot and human (egocentric video) demonstrations in a unified action-pattern space.

Links

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️