ICML 2026 Sparse ActionGen - Heungwoo/research GitHub Wiki

Sparse ActionGen — Rollout-Adaptive Pruning for Real-Time Diffusion Policy

Venue: ICML 2026 (Poster) Category: Efficiency Affiliations: Tsinghua University Traction (2026-06): 2 citations (arXiv)

Framework of Sparse ActionGen: a prune-then-reuse pipeline coupled with rollout iterations, with an observation-conditioned real-time diffusion pruner (Figure 2 from Ji et al., 2026)

Problem

Diffusion Policy is the dominant action-generation head for visuomotor and VLA control because it models multi-modal action distributions, but its multi-step denoising is too slow for real-time control. The paper notes that on an RTX 4090, 50 denoising steps at ~1 ms/step take 50 ms, capping execution at 20 Hz — far below the 50–1000 Hz a Franka arm needs. Existing caching-based accelerators (e.g. EfficientVLA's uniform schedule, BAC's task-specific block-wise schedule) rely on static caching schedules fixed offline that do not adapt to the dynamics of robot–environment interaction. A leave-one-out study on the Square task (Figure 1) shows that a fixed schedule performs inconsistently across rollout iterations and that the optimal schedule differs per iteration, so any static schedule constrains the performance/efficiency tradeoff.

Method

Sparse ActionGen (SAG) is a rollout-adaptive prune-then-reuse mechanism operating over three nested levels — rollout (robot–environment interaction), denoising (diffusion inference), and block (DiT forward). In each rollout iteration it globally identifies prunable computations and substitutes them on the fly with cached activations.

  • Real-time diffusion pruner. SAG parameterizes a pruner G_ψ that, instead of profiling post-forward activation similarities (which would negate any speedup), learns to predict the sparsity pattern a priori. Because the computational pattern of action generation is strongly correlated with the visual input, the pruner is conditioned on the current observation o_t, giving environment-aware adaptation. It is built with a parameter- and inference-efficient design that serves all blocks with a single network in a single forward pass for the whole denoising trajectory.
  • Global sparsity loss. An end-to-end objective guides the pruner to non-uniformly allocate compute across both timesteps and blocks under a strict budget, challenging the standard block-wise caching paradigm and its overlooked inter-block redundancy.
  • One-for-all reusing strategy. Motivated by strong cross-block activation similarity (Figure 4a), SAG reuses cached activations across both blocks and timesteps in a zig-zag manner, minimizing global redundancy. The model parameters are never updated.

Results

On Diffusion Policy (transformer variant, DP-T) over robomimic and Kitchen tasks (Lift, Can, Square, Transport, Tool hang, Kitchen), SAG prunes over 90% of computations and achieves a 3.6–4× speedup without sacrificing performance:

  • Proficient Human (PH) data (Table 1): SAG reaches Lift 100, Can 98, Square 89, Transport 85, Tool 50 — all at ~3.4–3.7× speedup — with an average performance gain of 13% over the full-precision baseline, matching or beating BAC and far exceeding EfficientVLA/L2C.
  • Mixed Human (MH) data (Table 2): average performance gain of 29%, with >3.7× speedup across tasks (e.g. Square 79, Transport 50).
  • Multi-stage Kitchen (Table 3): a lossless 4.03× speedup, retaining p1–p4 success of 100/100/100/99, where competitors like CP collapse at high speedups and EfficientVLA fails entirely.

Ablations confirm each component contributes — the real-time pruner, the one-for-all reusing strategy, and the global sparsity loss — and a real-world pick-and-release task (Figure 5) validates improved inference frequency at maintained success.

Significance

SAG reframes diffusion-policy acceleration from offline static caching to an online, observation-conditioned pruning problem aligned with closed-loop control. By predicting sparsity before inference and reusing activations across both timesteps and blocks, it delivers near-lossless ~4× speedups without retraining, making diffusion policies far more practical as high-frequency action heads for robotics and VLA systems.

Links

← Back to ICML-2026