ICLR 2026 Masked Action Chunking - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 · Authors: Haoxuan Wang, Gengyu Zhang, Yan Yan, Yuzhang Shang, Ramana Rao Kompella, Gaowen Liu (UIC · UCF · Cisco Research) · arXiv: 2601.20130 Category: Efficient VLA / inference acceleration Trend tag: Training the policy to be robust to asynchronous-inference mismatch, not just smoothing chunk boundaries
flowchart LR
OBS[Current observation] --> POLICY[Pretrained chunking policy<br/>+ REMAC corrective adjustments]
PREV[Partially executed chunk<br/>intended vs actual] --> POLICY
POLICY --> MASK[Masked action chunking<br/>learn corrections under<br/>intra-chunk mismatch]
MASK --> PREFIX[Prefix-preserved sampling<br/>reinforce inter-chunk continuity]
PREFIX --> CHUNK[Next action chunk]
CHUNK --> EXEC[Execute current chunk]
EXEC -.async: predict next while executing.-> POLICY
Real-time execution is essential for robots in dynamic environments, where small delays undermine responsiveness. Asynchronous inference has become a system-level paradigm for real-time manipulation: the next action chunk is predicted while the current chunk is still executing. But naive asynchronous integration often causes execution failure. Prior work blamed inter-chunk discontinuity (jumps at chunk boundaries) and proposed test-time smoothing algorithms. This paper identifies a second, overlooked failure mode: intra-chunk inconsistency, where the robot's executed action chunk partially misaligns with its current perception because the chunk was predicted from a stale observation.
REMAC learns corrective adjustments on top of a pretrained policy via masked action chunking. By masking parts of the action chunk during training, the policy learns to remain resilient when the actions it already committed to no longer match the actual execution state during asynchronous inference — i.e., it is trained to tolerate the intended-vs-actual mismatch rather than relying on a post-hoc test-time fix. To additionally address the previously studied boundary problem, REMAC adds a prefix-preserved sampling procedure that reinforces inter-chunk continuity by conditioning the new chunk on the committed prefix of the previous one. The result is a more reliable policy that incurs no additional latency at inference.
Experiments in both simulation and real-world settings show that REMAC enables faster task execution, maintains robustness across varying delays, and consistently achieves higher completion rates than baselines, while adding no extra inference latency.
REMAC reframes the real-time-chunking problem: instead of treating asynchronous-inference failures purely as a boundary-smoothing issue handled at test time, it bakes robustness to mid-chunk perception mismatch into the policy through training-time masking. This complements latency-reduction methods (caching, pruning, parallel decoding) by attacking the consistency side of real-time execution, making asynchronous action-chunk pipelines reliable rather than just fast.
- arXiv: 2601.20130
- OpenReview: forum?id=r0RGJ1j9on
← Back to ICLR-2026