ICLR 2026 OmniSAT - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 (arXiv:2510.09667, Oct 2025) Category: Action Tokenizer Trend tag: Trend 4 Authors: Huaihai Lyu, Chaofan Chen, Senwei Xie, Pengwei Wang, Xiansheng Chen, Shanghang Zhang, Changsheng Xu
flowchart LR
A[Variable-length action chunk] --> CE[Consistency Encoding<br/>B-Spline fit to fixed-length control points]
CE --> SP[Split into subspaces:<br/>position / rotation / gripper]
SP --> RQ[Multi-stage residual VQ<br/>per-subspace codebooks<br/>pos 256 / rot 256 / grip 64]
RQ --> T[Compact discrete tokens<br/>6.8x sequence compression]
T --> AR[Faster auto-regressive VLA training]
Auto-regressive (AR) VLA policies are appealing for large-scale pretraining but are bottlenecked by long action-token sequences, which slow convergence and raise target entropy. Existing tokenizers trade off compression against reconstruction fidelity (e.g., FAST achieves near-lossless reconstruction but only ~3.7x compression). OmniSAT targets higher-rate compression with preserved fidelity to accelerate AR training, while also providing a unified action-pattern space for cross-embodiment learning from robot and human demonstrations.
Two-stage tokenization:
- Consistency Encoding — variable-length, variable-rate trajectories are fit with a B-Spline into temporally aligned, fixed-length control-point representations (normalizes value range and temporal horizon).
- Multi-stage residual quantization — control-point features are partitioned into three separate subspaces (position, rotation, gripper), each quantized independently with its own residual VQ codebook (position 256, rotation 256, gripper 64 entries), yielding coarse-to-fine discrete tokens.
Note: the codebooks are subspace-specific, not a single shared global vocabulary; cross-embodiment sharing comes from the unified normalized action-pattern space, not from one codebook.
- Compression (DROID, Table 1): OmniSAT-8 reaches 6.8x sequence-length reduction vs FAST 3.7x and BEAST 4.6x, while keeping reconstruction MAE at 9.4e-4 (mm-level).
- LIBERO: 93.4% average success (Object 98.7 / Goal 94.6 / Spatial 94.1 / Long 86.0), rank 1.
- SimplerEnv-WidowX: 55.2% overall vs BEAST 37.5%.
- Real robot (dual-arm AgileX Cobot Magic): PlaceObj 73% / ZipSeal 63% / TubeRack 48% (vs BEAST 63/45/23); adding human egocentric video data lifts these to 80/66/58.
- Compression translates to faster AR convergence (e.g., ~2.5k vs ~4k steps to converge on LIBERO).
Complements X-VLA (soft prompts) and XR-1 (UVMC) — all three address cross-embodiment scaling at the tokenizer / interface level rather than the backbone level. OmniSAT's distinctive lever is amplifying AR efficiency via high-rate B-Spline + residual-VQ compression while retaining mm-level reconstruction, jointly leveraging robot and human (egocentric video) demonstrations in a unified action-pattern space.
- ICLR 2026 listing
- arXiv:2510.09667
- OpenReview
- Project page
← Back to ICLR-2026