ICML 2026 Speedup Patch - Heungwoo/research GitHub Wiki
Speedup Patch (SuP) — Plug-and-play offline-RL acceleration for frozen manipulation policies
Venue: ICML 2026 (Poster) Category: Efficiency Affiliations: Zhichao Wu, Junyin Ye, Zhilong Zhang, Yihao Sun, Haoxin Lin, Jiaheng Luo, Haoxiang Ren, Lei Yuan, Yang Yu (Nanjing University) Traction (2026-06): 2 citations (arXiv)

Problem
Modern embodied policies (ACT, diffusion policies, VLAs such as π0.5) inherit the "tardy pacing" of the human teleoperation data they imitate, so even when they succeed they execute slowly. Existing acceleration methods either retrain the policy on entropy-resampled demonstrations (e.g., DemoSpeedup) or require costly online interaction — both of which do not scale to large frozen foundation models whose weights are downloaded and never touched. The paper asks: can we accelerate an arbitrary, frozen embodied policy using only offline data and no policy retraining?
Method
SuP (SpeedUp Patch) wraps a frozen base policy π_base with a lightweight external scheduler that adaptively decides, per action chunk, a downsampling rate k — keeping every k-th action to produce a shorter chunk. The key is choosing k aggressively where motion is redundant but conservatively near contact-rich or precise phases.
The authors formalize scheduler learning as a Constrained Markov Decision Process (S, K, P, r, c, h, γ): maximize an efficiency reward while keeping a safety cost below threshold. Because true task success cannot be evaluated offline, SuP introduces a world-model-based state deviation surrogate: a learned (recurrent) world model predicts the counterfactual trajectory the un-downsampled policy would have produced, and the deviation between downsampled and counterfactual states acts as the constraint cost. Training is purely offline in three phases — (1) recurrent world-model learning, (2) data synthesis, (3) scheduler optimization via IQL (offline RL).

Results
Evaluated across diverse frozen architectures (ACT and Diffusion Policy on BiGym humanoid tasks; pre-trained π0.5 and VLA-Adapter on LIBERO), plus three real-world dual-arm (Aloha-like) tasks. Cells report success rate and average steps-to-completion.
- Overall ~1.8× execution speedup while preserving original success rates.
- LIBERO (π0.5): base 0.969 avg SR at 1.00×; SuP reaches 1.94× on Long suite and maintains comparable success, vs. fixed downsample (-ds2) which drops to 0.928 SR at 1.72×.
- BiGym (ACT): SuP improves both success and speed (e.g., Sandwich 0.45→0.64 SR), outperforming Vanilla Downsample and DemoSpeedup, which suffer >5% success drops.
- Real-world (π0.5, dual-arm): average success 0.589→0.611 at 2.17× speedup, beating fixed -ds2/-ds3 (which collapse to 0.356 SR at 2.19×) and matching DemoSpeedup on speed without retraining.
A case study links violation count (world-model deviation events) to conditional success-rate drops, validating the surrogate cost.
Significance
SuP decouples acceleration from the policy itself: a small scheduler trained offline can be patched onto any frozen manipulation policy — including downloaded VLA foundation models — to roughly halve execution time without sacrificing success or touching the base weights. The CMDP-plus-world-model formulation gives a principled, retraining-free path to deploy-time efficiency.
Links
- arXiv: 2603.20658
- ICML 2026: https://icml.cc/virtual/2026/poster/62253
← Back to ICML-2026