Review Realtime Execution - Heungwoo/research GitHub Wiki

Cross-Paper Review โ€” Real-Time Execution for VLA Policies

Topic survey (updated Aug 2026) ยท how big policies act smoothly at control rate: chunking, continuation, anytime decoding, async systems. Structure: ๐Ÿ“ˆ trend ยท โš–๏ธ approaches ยท โš ๏ธ limitations. Companion pages: RTC ยท Legato ยท OAT ยท Fast-in-Slow ยท AsyncVLA ยท VLA Attention ยท ฮจโ‚€.

1. The problem

A 2โ€“7B VLA takes 100โ€“300 ms per forward pass; robots need 10โ€“50 Hz control. Action chunking (predict H steps, execute while computing the next chunk) is the universal answer โ€” and its failure mode is the chunk boundary: inference delay plus flow-policy multimodality make consecutive chunks disagree, producing visible hesitation, jitter, and collisions.

2. Trend arc โ€” from test-time patch to trained-in property

Stage Method Mechanism Status
2025 [RTC](/Heungwoo/research/wiki/NeurIPS-2025-Real-Time-Chunking) (test-time) Inference-time inpainting/guidance constrains the new chunk to overlap the executing one Widely adopted; external to the policy โ€” spurious multimodal switching, trajectories never intrinsically smooth
Late 2025 Training-time RTC (PI) Simulate inference delay during training: expose first d ~ U(0, d_max) actions un-noised, mask from loss Adopted by [ฮจโ‚€](/Heungwoo/research/wiki/Review-Psi0) after test-time guidance proved unstable on their model โ€” a documented negative result for test-time steering
RSS 2026 [Legato](/Heungwoo/research/wiki/RSS-2026-Legato) (native continuation) Denoising initialized from a schedule-shaped mixture of committed actions + noise; flow dynamics reshaped for train/inference consistency; randomized schedules โ†’ controllable smoothness, variable delays Beats RTC by โ‰ˆ10% on both smoothness (NSPARC) and completion time across five real tasks
RSS 2026 [OAT](/Heungwoo/research/wiki/RSS-2026-OAT) (anytime decoding) Ordered token space โ†’ every prefix decodes to a valid coarse action; more tokens refine it A compute-quality dial rather than a boundary fix; AR-policies only

ICML 2026 โ€” the efficiency cluster (9+ papers), strongest of any venue. Streaming: Reflex achieves 50 Hz stable streaming (2.58ร— speedup, โˆ’54% reaction latency) via timestep-invariant attention partitioning with O(1) cache updates. Fewer/faster denoise steps: OMP one-step MeanFlow, STEP warm-started actions (+21.6% over baselines at 2 steps), Sparse ActionGen (4ร—). Token/channel pruning: GridS (โˆ’76% FLOPs, no drop), SpecPrune-VLA (1.57โ€“1.70ร—), EcoVLA. Reasoning latency: Latent Reasoning VLA internalizes CoT (โˆ’90% inference latency), AVA-VLA adds confidence-gated early exit. Policy-agnostic: Speedup Patch (1.8ร— via safe chunk downsampling). Chunk coherence gets its own treatments (FocalPolicy frequency-optimized chunking).

Parallel line โ€” architectural asynchrony: dual-rate systems (Fast-in-Slow: slow VLM cadence + fast action decoding), AsyncVLA, persistent-context AR experts (AR-VLA, RSS #85: the expert keeps its own history, VL prefixes refresh out-of-band), and action-to-action flow (RSS #209: warm-start denoising from the previous action instead of Gaussian noise).

3. Approach comparison

Approach Pros Cons Choose when
Test-time RTC Drop-in, no retraining Multimodal switching; can be unstable (ฮจโ‚€'s experience) Frozen checkpoint you can't retrain
Training-time continuation (Legato-class) Intrinsically smooth; delay-robust by randomization; faster completion Requires retraining; delay range fixed at training You own the training loop (the 2026 default)
Anytime tokens (OAT) Graceful degradation under compute pressure AR-family only Variable compute budget / early-commit control
Async dual-rate Decouples semantic and control rates System complexity; two-model coherence Reasoning-heavy tasks with fast reflex needs
Single-step action head (IROS 2026 ๐Ÿ†•) [IMLE-VLA](/Heungwoo/research/wiki/IROS-2026-IMLE-VLA) replaces the iterative diffusion/flow head with a 1-step cIMLE generator โ€” 55 Hz (3.67ร— ฯ€0.5), LIBERO 98.0%, keeps multimodality + robustness Head-swap needs retraining; single-step on high-precision contact untested You want to kill the multi-step sampling latency at the source
Remote/cloud + RTC Big models off-robot Network variance (Qwen-RobotNav: server 196 ms avg but spiky vs Jetson-Thor 204 ms stable) Fleet ops with reliable links; edge for latency-critical tasks

4. Interactions that bite

  • History conditioning raises the denoising bill: RobotManip's in-context variant needs 10 steps where the base needs 4 โ€” robustness features and latency budgets trade off directly.
  • Chunk length H couples everything: smoothness machinery, RL credit assignment (RECAP's N-step advantages), and reactivity limits (RobotManip lists fixed chunk length as a stated limitation for sub-second reactive control).
  • Wiring choice sets the floor: prefix-KV reuse (ฯ€) vs full re-attention per denoise step (concatenation) โ€” argued but never measured head-to-head (Review-VLM-Action-Connection).

5. โš ๏ธ Limitations

  • Latency is the least-reported number in the field. Among 2026 flagships, only scattered figures exist (ฯ€0.6 63 ms/chunk-class reports; ฮจโ‚€ ~160 ms/pass; Qwen-RobotNav's deployment table); the Qwen manipulation suite reports none. No standard metric (ms/chunk? effective Hz? jerk?) exists.
  • Edge deployment measurement is finally starting: the XPU characterization study profiles VLAs across GPUs/NPUs (compute-bound VLM phase, memory-bound expert phase; 2.9โ€“3.3ร— via phase-aware scheduling) โ€” but manipulation-side FP8/TensorRT deployment reports remain rare.
  • Smoothness metrics (NSPARC etc.) are young; no benchmark scores hesitation/recovery under forced delay spikes.
  • All continuation methods assume the policy is the bottleneck; perception-latency (multi-view encoding) is unaddressed.

โ† Back to Home ยท Reviews