ICML 2026 TapSampling - Heungwoo/research GitHub Wiki

TapSampling — Inference-time sampling with a task-progress verifier for robot policies

Venue: ICML 2026 (Poster) Category: Efficiency / Inference-Time Scaling Affiliations: (from arXiv) Sizhe Zhao, Shengping Zhang, Shuo Yang, Weiyu Zhao, Shuigen Wang, Xiangyang Ji Traction (2026-06): 0 citations (arXiv)

TapSampling extends single-shot inference with multi-sample generation and task-progress-guided verification (Figure 1 from Zhao et al., 2026)

Problem

Generalist robot policies built on diffusion or autoregressive (next-token) backbones are non-deterministic: under identical conditions they succeed on some trials and fail on others. Because deployment typically commits to a single action per decision step (single-shot inference), there is no mechanism to correct this stochastic variance, fundamentally capping performance. A motivating study on CALVIN with VPP shows the policy reaches 4.39 average success length without retries but climbs to ~4.72 when allowed up to 4 retries — evidence of a large untapped upper bound. The paper asks whether inference-time computation, rather than more training data or parameters, can close this gap. Doing so requires two hard components: (1) a candidate sampling strategy that produces plausible alternatives cheaply, and (2) a verifier that can judge low-level actions — for which no off-the-shelf evaluator exists.

Method

TapSampling is a plug-and-play, policy-agnostic framework with two parts.

TapSampling overview: Action-VAE samples candidates from a learned posterior; a task-progress verifier scores them (Figure 2 from Zhao et al., 2026)

Action sampling with a learned posterior. An Action-VAE (transformer encoder/decoder) compresses each policy-generated action chunk into a low-dimensional Gaussian posterior, trained with reconstruction + KL loss. At inference, a small set of N=4 policy actions is encoded into a mixed posterior q_mix = (1/N) Σ q_E(z|a_i), from which an arbitrary number of latents are drawn and decoded into candidates. Because the posterior captures intra-chunk dimensional correlations, candidates stay close to the true policy distribution — and sampling is ~5× faster than re-querying the policy.

Action selection via task-progress verification. Under a linear-progress assumption, each timestep i in a trajectory of length t is labeled p_i = i/t. Forward action sub-sequences yield positive labels (+k/t); reversed sub-sequences yield negatives (−k/t), giving automatic positive/negative pairs with no manual annotation. A verifier built on VLA-Adapter (Qwen2.5-0.5B backbone) regresses the task-progress change Δp of a query action via L1 loss. At inference, candidates below a threshold are discarded and a score-weighted average of survivors is executed. The VLM backbone runs once; candidates are scored in a batch through the lightweight head.

Results

  • CALVIN ABC→D (avg. success length): Diffusion Policy 2.41→2.58, OpenVLA 3.30→3.51, VPP 4.39→4.46 — consistent plug-and-play gains across diffusion, autoregressive, and video-prediction policies.
  • LIBERO-Long: π0.5 improves 96.8%→98.0%.
  • Real-world (Franka FR3, 3 task families): π0 average 78.3%→83.3%, with the largest gains on Stack-Unseen (+10.0) and Knock-Down-Unseen (+6.6).
  • Efficiency: Policy Sampling gives the best quality (avg. len 4.50) but ~20× latency; Gaussian and Learned-Posterior Sampling add <0.01 s. Learned-Posterior best balances the two (4.46) and yields lower MMD to the policy distribution than Gaussian (e.g., 0.064 vs 0.098 at γ=2.0). Verification is ~12× faster than RoboMonkey at k=16.

Significance

TapSampling reframes test-time scaling — well established for LLMs and image diffusion — for robotic control, supplying the two missing ingredients (a distribution-faithful sampler and an interpretable, semantically grounded action verifier) without fine-tuning the base policy. The progress-change verifier is notable for producing scores with explicit meaning (negative = impedes task, larger positive = faster completion), enabling interpretable selection and reusable across heterogeneous policies.

Links

← Back to ICML-2026