RSS 2026 Mimic Intent Not Just Trajectories - Heungwoo/research GitHub Wiki
Mimic Intent, Not Just Trajectories
Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Imitation learning 3 · paper #206 Authors: Renming Huang, Chendong Zeng, Wenjing Tang, Jintian Cai, Cewu Lu, Panpan Cai arXiv: 2602.08602 · program page
Summary compiled from the arXiv paper (v3); all numbers quoted from the paper. Trend context: RSS 2026 survey.

The Spectrally Disentangled Action Tokenizer encodes an action chunk (via DCT into the frequency domain) into multi-scale tokens: the coarsest S1 "Intent token" is forced to reconstruct dominant low-frequency structure, while finer S2..Sk "Execution tokens" capture high-frequency residuals. The t-SNE (right) shows S1 tokens forming coherent clusters for semantically consistent behaviors like "Pick up", "Move forward", and "Clockwise rotation".
Problem
VLA/imitation policies mimic raw trajectories without modeling why actions are executed, so they overfit to surface correlations and generalize poorly to environmental changes and skill transfer. Existing action tokenizers act mostly as compressors and don't align the token space with interpretable intent.
Method
MINT (SJTU/Shanghai Innovation Institute) disentangles intent from execution via multi-scale frequency-space tokenization. Action chunks are transformed with a Discrete Cosine Transform; a multi-scale VQ-VAE is trained with a progressive scale-wise spectral reconstruction objective so early scales explain low-frequency global structure (the "Intent token" S1) and later scales encode high-frequency residuals ("Execution tokens"). The policy performs next-scale autoregression (coarse-to-fine intent-to-execution reasoning, tokens within a scale generated in parallel) and uses an intent-based action ensemble weighting overlapping predictions by S1 cosine similarity. Two variants: a lightweight MINT-30M transformer trained from scratch and a MINT-4B built on a PaliGemma-2.6B VLM. One-shot transfer is done by injecting the S1 intent token extracted from a single demonstration (MINT-Zero).
Results
On LIBERO, MINT-4B averages 98.3 (SPATIAL/OBJECT/GOAL/LONG 97.4/99.6/98.2/97.8) vs π0.5 96.9 and OpenVLA-OFT 95.4; MINT-30M without pretraining hits 97.1 avg. On CALVIN (ABCD→D) MINT-4B reaches average length 4.57. On LIBERO-Plus disturbances, MINT-4B averages 80.1 vs OpenVLA-OFT 71.4 and π0.5 65.0 (the abstract's ~15% gain over the strongest baseline). One-shot transfer via intent injection (MINT-Zero-30M) scores 0.90/0.68/0.72 on new-specification/new-layout/extended-horizon (avg 0.77) vs 0.17 for language fine-tuning — the ~60% relative transfer gain. Real-world on a 6-DOF Piper-X arm (BridgeDataV2 pretrain, ~20 demos/task) MINT-4B beats π0.5* by 29% and is statistically distinguishable from ACT/π0 baselines including on an unseen Stack Cups task. Ablations confirm scale-wise spectral loss (93.4 LIBERO-Long) beats time-domain variants, and intent-based ensembling (93.2) beats temporal/action ensembling.
Significance
Reframes action tokenization from compression to semantic intent abstraction, giving a clean handle for one-shot transfer that sidesteps ambiguous language conditioning — a notable entry for Review-Human-Video-Transfer and the coarse-to-fine/streaming-generation direction in Review-Realtime-Execution.
← Back to RSS 2026 survey · RSS-2026-Papers · Home