ICLR 2026 Chunking Augmentation - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 · Authors: Thomas T. Zhang, Daniel Pfrommer, Chaoyi Pan, Nikolai Matni, Max Simchowitz · Paper: arXiv 2507.09061 — Action Chunking and Exploratory Data Collection Yield Exponential Improvements in Behavior Cloning for Continuous Control (Jul 2025) · Category: Theory / empirical study · Trend tag: Control-theoretic foundations of imitation learning
Note: the ICLR index lists this under the working title "Action Chunking and Data Augmentation Yield Exponential Improvements..."; the confirmed arXiv title uses "Exploratory Data Collection" and "Continuous Control."
flowchart LR
BC[Behavior cloning<br/>continuous control] --> Comp[Error compounds<br/>EXPONENTIALLY with horizon H]
Comp --> I1[Intervention 1:<br/>Action chunking<br/>open-loop action sequences]
Comp --> I2[Intervention 2:<br/>Exploratory data collection<br/>augment expert demos]
I1 --> Stab{Control-theoretic<br/>stability}
I2 --> Stab
Stab --> Poly[Compounding error<br/>circumvented in different regimes]
Poly --> Bounds[Tighter statistical<br/>guarantees on IL error]
In continuous control, behavior cloning can suffer error that compounds exponentially with the task horizon — a worst-case that information-theoretic analyses capture but do not explain mechanistically. Two interventions are widely used in practice (action chunking; augmenting/exploring around expert demonstrations) and empirically help, but lacked a rigorous account of why and when.
A theoretical analysis with a control-theoretic lens:
- Action chunking = predicting sequences of actions executed open-loop. The paper shows this circumvents exponential compounding error in certain regimes.
- Exploratory data collection = augmenting expert demonstrations with exploratory data. This similarly avoids exponential blow-up, but in a different operating regime.
- The unifying mechanism identified is control-theoretic stability of the underlying dynamics; stability is what prevents the error cascade. This yields fine-grained insight into how compounding error arises and tighter statistical guarantees on imitation-learning error than information-theoretic bounds alone.
Theoretical predictions are validated on popular robot-learning benchmarks, matching the regimes in which each intervention is predicted to help. (The work is primarily analytical; precise bound constants and benchmark scores are omitted here pending the full tables.)
Provides a principled, stability-based explanation for two of the most impactful empirical tricks in modern robot imitation learning (chunking as used in ACT/π0-style policies; data augmentation/DAgger-like exploration). Reframes the compounding-error story from worst-case pessimism to a regime-dependent picture governed by control-theoretic stability — guidance for when chunking vs. exploratory data is the right lever.
- arXiv: https://arxiv.org/abs/2507.09061
- OpenReview: https://openreview.net/forum?id=jiWXDvw1Lf
← Back to ICLR-2026