ICLR 2026 Chunking Augmentation - Heungwoo/research GitHub Wiki

Action Chunking & Exploratory Data — exponential improvements in BC

Venue: ICLR 2026 · Authors: Thomas T. Zhang, Daniel Pfrommer, Chaoyi Pan, Nikolai Matni, Max Simchowitz · Paper: arXiv 2507.09061 — Action Chunking and Exploratory Data Collection Yield Exponential Improvements in Behavior Cloning for Continuous Control (Jul 2025) · Category: Theory / empirical study · Trend tag: Control-theoretic foundations of imitation learning

Note: the ICLR index lists this under the working title "Action Chunking and Data Augmentation Yield Exponential Improvements..."; the confirmed arXiv title uses "Exploratory Data Collection" and "Continuous Control."

Approach diagram

flowchart LR
  BC[Behavior cloning<br/>continuous control] --> Comp[Error compounds<br/>EXPONENTIALLY with horizon H]
  Comp --> I1[Intervention 1:<br/>Action chunking<br/>open-loop action sequences]
  Comp --> I2[Intervention 2:<br/>Exploratory data collection<br/>augment expert demos]
  I1 --> Stab{Control-theoretic<br/>stability}
  I2 --> Stab
  Stab --> Poly[Compounding error<br/>circumvented in different regimes]
  Poly --> Bounds[Tighter statistical<br/>guarantees on IL error]
Loading

Problem

In continuous control, behavior cloning can suffer error that compounds exponentially with the task horizon — a worst-case that information-theoretic analyses capture but do not explain mechanistically. Two interventions are widely used in practice (action chunking; augmenting/exploring around expert demonstrations) and empirically help, but lacked a rigorous account of why and when.

Method

A theoretical analysis with a control-theoretic lens:

  • Action chunking = predicting sequences of actions executed open-loop. The paper shows this circumvents exponential compounding error in certain regimes.
  • Exploratory data collection = augmenting expert demonstrations with exploratory data. This similarly avoids exponential blow-up, but in a different operating regime.
  • The unifying mechanism identified is control-theoretic stability of the underlying dynamics; stability is what prevents the error cascade. This yields fine-grained insight into how compounding error arises and tighter statistical guarantees on imitation-learning error than information-theoretic bounds alone.

Results

Theoretical predictions are validated on popular robot-learning benchmarks, matching the regimes in which each intervention is predicted to help. (The work is primarily analytical; precise bound constants and benchmark scores are omitted here pending the full tables.)

Significance

Provides a principled, stability-based explanation for two of the most impactful empirical tricks in modern robot imitation learning (chunking as used in ACT/π0-style policies; data augmentation/DAgger-like exploration). Reframes the compounding-error story from worst-case pessimism to a regime-dependent picture governed by control-theoretic stability — guidance for when chunking vs. exploratory data is the right lever.

Links

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️