NeurIPS 2025 Real Time Chunking - Heungwoo/research GitHub Wiki

Real-Time Chunking (RTC) โ€” Async Inpainting for Action-Chunk Flow Policies

Paper title: Real-Time Execution of Action Chunking Flow Policies Venue: NeurIPS 2025 ยท Authors: Kevin Black, Manuel Y. Galliker, Sergey Levine (Physical Intelligence + UC Berkeley) ยท arXiv: 2506.07339 Category: Diffusion / Flow Efficiency

Approach diagram

flowchart LR
  E1[Chunk 1 executing...] --> E2[Chunk 2 executing]
  C[Computing Chunk 2<br/>in parallel]
  C -- freeze committed actions --> IP[Inpaint remaining]
  IP -- hand off at boundary --> E2
  note[No pause at chunk boundaries โ€”<br/>next chunk ready just-in-time]
Loading

Problem

Flow-matching and diffusion VLAs emit action chunks (e.g., 50 steps at a time). At chunk boundaries, the robot has to pause while the next chunk is computed โ€” or the chunk is computed during the previous one's execution, but committed actions from the previous chunk are visible mid-flight, causing jitter.

Method

Async inpainting during chunk execution:

  1. While chunk $k$ is executing (e.g., steps 0โ€“15 of a 50-step chunk), start computing chunk $k+1$.
  2. Freeze the actions that will still be executed from chunk $k$ when chunk $k+1$ starts (steps 15โ€“50).
  3. Inpaint the beginning of chunk $k+1$ to continue smoothly from the frozen tail โ€” using the diffusion/flow model's denoising process conditioned on those frozen steps (implemented via training-free pseudoinverse / guided-inpainting, the known strength of iterative denoising).
  4. Hand off at the boundary โ€” no pause, no discontinuity.

Hard vs. soft masking. The first $d$ actions (those guaranteed to execute during inference delay $d$) are hard-frozen (weight 1); the final $s$ actions are freely generated (weight 0); the overlap region in between uses a soft mask with exponentially decaying guidance weights. The paper finds soft masking is critical for full cross-chunk continuity and outperforms hard masking, particularly at small $d$ (Section 5.1).

Works on any diffusion or flow-matching VLA without retraining โ€” it is a purely inference-time algorithm with no training-recipe changes.

Results

  • Evaluated on a new benchmark of 12 highly dynamic tasks in the Kinetix simulator plus 6 challenging real-world bimanual manipulation tasks (e.g. lighting a match, plugging in an Ethernet cable).
  • ~20% faster robot motion than synchronous inference for the same task, and smoother than all competing baselines (including temporal ensembling).
  • Uniquely robust to inference delay: maintains high success on precise tasks even with delays exceeding 300 ms (>30% of the model's prediction horizon).
  • Plug-and-play on ฯ€0, ฯ€0-FAST, ฯ€0.5, and other flow-matching VLAs; real-world experiments used ฯ€0.5 (H=50, 50 Hz control, 5 denoising steps).
  • A later training-time variant (Training-Time Action Conditioning, arXiv:2512.05964) folds the delay simulation into training, removing RTC's inference-time overhead, and is used in ฯ€0.6.

Significance

The training-time analog of this idea is the 240 ms delay simulation used in ฯ€0.7. RTC itself is inference-time and model-agnostic โ€” a template for how ICLR 2026's diffusion VLAs (Unified Diffusion VLA, DIVA & Fast-dVLA) handle real-time execution.

Extends the CoRL 2025 Streaming Flow Policy thread โ€” both papers tackle the chunk-boundary problem with different mechanisms.

Links

Related pages

โ† Back to NeurIPS-2025

โš ๏ธ **GitHub.com Fallback** โš ๏ธ