ICLR 2026 Block Caching - Heungwoo/research GitHub Wiki

BAC โ€” Block-wise Adaptive Caching for accelerating Diffusion Policy

Venue: ICLR 2026 ยท Authors: Kangye Ji, Yuan Meng, Hanyun Cui, Ye Li, Jianbo Zhou, Shengjia Hua, Lei Chen, Zhi Wang (Tsinghua University) ยท arXiv: 2506.13456 Category: Efficient VLA / inference acceleration Trend tag: Training-free feature caching across denoising steps for diffusion-based control

Approach diagram

flowchart LR
  X[Noisy action tokens<br/>+ obs conditioning] --> SCHED[Adaptive Caching Scheduler<br/>pick update timesteps by<br/>maximizing cached vs skipped<br/>feature similarity]
  SCHED -->|update step| COMPUTE[Recompute transformer blocks<br/>attn + FFN]
  SCHED -->|skip step| CACHE[Reuse cached<br/>block features]
  COMPUTE --> BUA[Bubbling Union Algorithm<br/>refresh upstream blocks with<br/>high caching error before<br/>downstream FFNs]
  CACHE --> BUA
  BUA --> DENOISE[Denoised action features]
  DENOISE -->|next denoising step| X
  DENOISE --> ACT[Action chunk -> robot]
Loading

Problem

Diffusion Policy has strong visuomotor modeling capability, but its iterative denoising makes per-action inference compute-heavy, which is impractical for real-time robotic control. There is large redundancy across the repetitive denoising steps, yet general-purpose diffusion acceleration techniques (designed for image/video diffusion) fail to transfer to Diffusion Policy because of architectural and data divergences specific to action generation.

Method

Block-wise Adaptive Caching (BAC) accelerates Diffusion Policy by caching and reusing intermediate action features at the level of individual transformer blocks, rather than caching whole-network outputs. The key observation is that feature similarity across denoising steps is non-uniform in time and shows distinct block-specific patterns, so a single global skip schedule is suboptimal. BAC introduces two components: (1) an Adaptive Caching Scheduler that selects the optimal timesteps to refresh each block by maximizing the global feature similarity between cached and skipped features; and (2) a Bubbling Union Algorithm that controls error growth โ€” naively caching every block causes error surges from inter-block propagation, especially within Feed-Forward Network (FFN) blocks, so BUA truncates this by updating upstream blocks with significant caching error before their downstream FFNs. BAC is a training-free plug-in that drops into existing transformer-based Diffusion Policy and vision-language-action models.

Results

As a training-free plugin requiring no retraining, BAC reports up to 3x inference speedup with lossless action generation across multiple robotic benchmarks. The authors frame the acceleration as "for free," meaning it preserves task performance relative to the uncached base policy while cutting denoising compute.

Significance

BAC adapts the now-standard idea of diffusion feature caching to the control setting, where the failure mode is not visual artifacts but degraded action accuracy and error accumulation across blocks. By making caching block-adaptive and adding an error-truncation pass over FFNs, it offers a portable, retraining-free speedup for the diffusion- and flow-based action heads that dominate current VLA stacks โ€” complementary to token-pruning and scheduling approaches.

Links

Related pages

โ† Back to ICLR-2026

โš ๏ธ **GitHub.com Fallback** โš ๏ธ