ICML 2026 Drift is a Sampling Error - Heungwoo/research GitHub Wiki
Drift is a Sampling Error: CAPS — SNR-gated power-distribution search for long-horizon VLA planning
Venue: ICML 2026 (Poster) Category: Efficiency Affiliations: Kewei Chen, Yayu Long, Mingsheng Shang Traction (2026-06): 0 citations (arXiv)

Problem
Vision-Language-Action (VLA) models excel at short-horizon control but suffer instruction drift on long-horizon tasks: as a task progresses, irrelevant observations dilute attention to the original instruction, and the robot executes actions that are locally reasonable but globally wrong. The paper reconceptualizes drift as a systematic sampling error — local greedy sampling collapses into "Negative Pivotal Windows": irreversible local optima with high local probability that sever global success pathways. Prior fixes (CoT prompt engineering; generate-verify rerankers like TACO) either stay open-loop or rely on parallel independent sampling and reranking, lacking iterative refinement of a single trajectory.
Method
The authors propose Context-Aware Power Sampling (CAPS), a training-free, inference-time framework requiring no parameter updates.
- Power distributions for global sharpening. CAPS raises the model's conditional trajectory distribution to a power (sharpening coefficient
α) to sharpen global trajectory probabilities, enabling lookahead search over the generative trajectory distribution. The paper argues (and proves) this is not equivalent to local temperature sampling. - Metacognitive SNR control. A KL-divergence-based Signal-to-Noise Ratio signal estimates drift risk. When contextual SNR/entropy is high-certainty (
H ≤ γ), the system stays in fast "System 1" greedy decoding; when SNR drops below a critical thresholdγ_SNR, it triggers slow "System 2" search. - Block-based autoregressive MCMC. Once triggered, an intra-inference MCMC loop performs stochastic resampling of future trajectory chunks (Proposal) followed by an Acceptance step, iteratively refining a single trajectory rather than reranking independent samples. This is framed as adaptive computation / optimal control and provides an effective horizon extension.

Results
- RoboTwin 1.0 (π₀): CAPS lifts π₀ to 47.4% (+15.2%), vs TACO's +6.1%; on the drift-prone "Dual Bottles Pick Hard," success rises from 48.0% to 61.0%.
- Simpler-WindowX (OOD): CAPS reaches 60.5% average with π₀ (+12.5% over π₀'s 48.0%), beating SpatialVLA (42.7%); "Carrot on Plate" +19.0%.
- RoboTwin 2.0 (large-scale, 100 seeds/100 scenes, π₀.₅): average 66.2% vs π₀.₅ 59.3% and π₀.₅+TACO 64.0%.
- Libero-long (π₀.₅): average 97.6% vs π₀.₅ 94.8%, TACO 96.6%, OpenVLA 49.8%, Robomonkey 56.5%.
- Ablation (RoboTwin 1.0): Full adaptive CAPS = 47.4% at 2.15× latency, vs Always-On CAPS 48.1% at 8.50×, Rejection Sampling 40.2%, w/o Sharpening 38.5%, base π₀ 32.2% (1.00×) — the SNR gate recovers nearly all of the always-on accuracy at a fraction of the compute.
Significance
CAPS reframes long-horizon robustness as an inference-time sampling/search problem rather than a training problem, and its SNR gate makes "slow thinking" cost-adaptive — spending extra compute only when drift is detected. As a model-agnostic, training-free add-on that improves π₀, π₀.₅, and OpenVLA-class policies, it is a clean instance of the 2026 trend toward adaptive inference-time scaling for embodied control.
Links
- arXiv: 2605.09537
- ICML 2026: https://icml.cc/virtual/2026/poster/64055
← Back to ICML-2026