ICML 2026 STEP - Heungwoo/research GitHub Wiki

STEP: Warm-Started Visuomotor Policies with Spatiotemporal Consistency Prediction — Two-step diffusion-policy inference via predicted warm-start actions

Venue: ICML 2026 (Poster) Category: Diffusion-Flow Policy Traction (2026-06): 2 citations (arXiv)

Comparison of inference pipelines for diffusion-based visuomotor policies; STEP uses spatiotemporally consistent action prediction to reduce sampling to only two steps (Figure 1 from Li et al., 2026)

Problem

Diffusion policies have become a powerful paradigm for visuomotor control because they model the full distribution of action sequences and capture multimodality. But their iterative denoising incurs substantial inference latency, which caps control frequency in real-time closed-loop systems. Existing accelerators either reduce sampling steps, bypass diffusion through direct prediction, or reuse past actions — yet they typically struggle to jointly preserve action quality and achieve consistently low latency, often collapsing in the low-step regime.

Method

STEP keeps the original diffusion policy intact but gives it a high-quality starting point, so very few denoising steps suffice.

Spatiotemporal Consistency Prediction. For planning horizon H, at timestep t the policy emits an action sequence At = (a_t, …, a{t+H-1}). STEP trains a predictor f_θ: O × A^H → A^H that maps the current observation ot and the previous timestep's action sequence A{t-H} to a predicted warm-start Â_t. The predictor is a multi-layer Transformer with cross-attention: historical actions and current observations are projected to a shared 128-dimensional embedding space, then fused via cross-attention to incorporate temporal context. The warm-start action is distributionally close to the target and temporally consistent, without compromising the base policy's generative capability.

Model architecture of the predictor in STEP (Figure 2 from Li et al., 2026)

Velocity-aware Perturbation Injection. In real-world execution the robot can enter a deadlock where consecutive actions barely change, leaving too little actuation to overcome static friction and control dead zones. STEP measures action variation ΔA_t between consecutive cached steps and flags stagnation when ‖ΔA_t‖ < ε_a, then switches between normal and perturbed execution by adjusting an action scaling factor and perturbation magnitude — restoring excitation only when needed.

Theory. The authors prove the proposed prediction induces a locally contractive mapping, guaranteeing convergence of action errors during diffusion refinement.

Results

Evaluated on nine simulated benchmarks (Push-T, RoboMimic — Lift/Transport/Can/Square/ToolHang —, ManiSkill2) and two real-world tasks:

  • On the image-based RoboMimic/Push-T suite, STEP with 2 steps matches or beats DDIM, BRIDGER, DPM-Solver++, and Falcon, e.g., 0.86 on Push-T and 0.76 on ToolHang at 2 steps, where DDIM collapses (0.5 on ToolHang) and DPM-Solver++/Falcon fall to ≈0 at low steps.
  • STEP with 2 steps achieves on average 21.6% and 27.5% higher success rate than BRIDGER and DDIM on the RoboMimic benchmark and real-world tasks, respectively.
  • Real-world: STEP sustains high success with only 2 denoising steps at ~20 ms end-to-end latency, yielding 105.7× speedup over vanilla DDPM and 4.8× over DDIM at matched success rate, while 4-step DDIM collapses.

STEP consistently advances the Pareto frontier of inference latency vs. success rate.

Significance

STEP shows that a small, separately trained predictor exploiting spatiotemporal consistency can supply warm-start actions good enough to make diffusion policies usable at 2 steps — preserving the multimodal generative quality of the base policy while cutting latency by orders of magnitude. The velocity-aware perturbation mechanism addresses a practical real-world failure (execution stall), and the contractivity analysis gives a convergence guarantee, making STEP a drop-in accelerator for real-time visuomotor control.

Links

← Back to ICML-2026