Review Realtime Execution - Heungwoo/research GitHub Wiki
Cross-Paper Review โ Real-Time Execution for VLA Policies
Topic survey (updated Aug 2026) ยท how big policies act smoothly at control rate: chunking, continuation, anytime decoding, async systems. Structure: ๐ trend ยท โ๏ธ approaches ยท โ ๏ธ limitations. Companion pages: RTC ยท Legato ยท OAT ยท Fast-in-Slow ยท AsyncVLA ยท VLA Attention ยท ฮจโ.
1. The problem
A 2โ7B VLA takes 100โ300 ms per forward pass; robots need 10โ50 Hz control. Action chunking (predict H steps, execute while computing the next chunk) is the universal answer โ and its failure mode is the chunk boundary: inference delay plus flow-policy multimodality make consecutive chunks disagree, producing visible hesitation, jitter, and collisions.
2. Trend arc โ from test-time patch to trained-in property
| Stage | Method | Mechanism | Status |
|---|---|---|---|
| 2025 | [RTC](/Heungwoo/research/wiki/NeurIPS-2025-Real-Time-Chunking) (test-time) | Inference-time inpainting/guidance constrains the new chunk to overlap the executing one | Widely adopted; external to the policy โ spurious multimodal switching, trajectories never intrinsically smooth |
| Late 2025 | Training-time RTC (PI) | Simulate inference delay during training: expose first d ~ U(0, d_max) actions un-noised, mask from loss | Adopted by [ฮจโ](/Heungwoo/research/wiki/Review-Psi0) after test-time guidance proved unstable on their model โ a documented negative result for test-time steering |
| RSS 2026 | [Legato](/Heungwoo/research/wiki/RSS-2026-Legato) (native continuation) | Denoising initialized from a schedule-shaped mixture of committed actions + noise; flow dynamics reshaped for train/inference consistency; randomized schedules โ controllable smoothness, variable delays | Beats RTC by โ10% on both smoothness (NSPARC) and completion time across five real tasks |
| RSS 2026 | [OAT](/Heungwoo/research/wiki/RSS-2026-OAT) (anytime decoding) | Ordered token space โ every prefix decodes to a valid coarse action; more tokens refine it | A compute-quality dial rather than a boundary fix; AR-policies only |
ICML 2026 โ the efficiency cluster (9+ papers), strongest of any venue. Streaming: Reflex achieves 50 Hz stable streaming (2.58ร speedup, โ54% reaction latency) via timestep-invariant attention partitioning with O(1) cache updates. Fewer/faster denoise steps: OMP one-step MeanFlow, STEP warm-started actions (+21.6% over baselines at 2 steps), Sparse ActionGen (4ร). Token/channel pruning: GridS (โ76% FLOPs, no drop), SpecPrune-VLA (1.57โ1.70ร), EcoVLA. Reasoning latency: Latent Reasoning VLA internalizes CoT (โ90% inference latency), AVA-VLA adds confidence-gated early exit. Policy-agnostic: Speedup Patch (1.8ร via safe chunk downsampling). Chunk coherence gets its own treatments (FocalPolicy frequency-optimized chunking).
Parallel line โ architectural asynchrony: dual-rate systems (Fast-in-Slow: slow VLM cadence + fast action decoding), AsyncVLA, persistent-context AR experts (AR-VLA, RSS #85: the expert keeps its own history, VL prefixes refresh out-of-band), and action-to-action flow (RSS #209: warm-start denoising from the previous action instead of Gaussian noise).
3. Approach comparison
| Approach | Pros | Cons | Choose when |
|---|---|---|---|
| Test-time RTC | Drop-in, no retraining | Multimodal switching; can be unstable (ฮจโ's experience) | Frozen checkpoint you can't retrain |
| Training-time continuation (Legato-class) | Intrinsically smooth; delay-robust by randomization; faster completion | Requires retraining; delay range fixed at training | You own the training loop (the 2026 default) |
| Anytime tokens (OAT) | Graceful degradation under compute pressure | AR-family only | Variable compute budget / early-commit control |
| Async dual-rate | Decouples semantic and control rates | System complexity; two-model coherence | Reasoning-heavy tasks with fast reflex needs |
| Single-step action head (IROS 2026 ๐) | [IMLE-VLA](/Heungwoo/research/wiki/IROS-2026-IMLE-VLA) replaces the iterative diffusion/flow head with a 1-step cIMLE generator โ 55 Hz (3.67ร ฯ0.5), LIBERO 98.0%, keeps multimodality + robustness | Head-swap needs retraining; single-step on high-precision contact untested | You want to kill the multi-step sampling latency at the source |
| Remote/cloud + RTC | Big models off-robot | Network variance (Qwen-RobotNav: server 196 ms avg but spiky vs Jetson-Thor 204 ms stable) | Fleet ops with reliable links; edge for latency-critical tasks |
4. Interactions that bite
- History conditioning raises the denoising bill: RobotManip's in-context variant needs 10 steps where the base needs 4 โ robustness features and latency budgets trade off directly.
- Chunk length H couples everything: smoothness machinery, RL credit assignment (RECAP's N-step advantages), and reactivity limits (RobotManip lists fixed chunk length as a stated limitation for sub-second reactive control).
- Wiring choice sets the floor: prefix-KV reuse (ฯ) vs full re-attention per denoise step (concatenation) โ argued but never measured head-to-head (Review-VLM-Action-Connection).
5. โ ๏ธ Limitations
- Latency is the least-reported number in the field. Among 2026 flagships, only scattered figures exist (ฯ0.6 63 ms/chunk-class reports; ฮจโ ~160 ms/pass; Qwen-RobotNav's deployment table); the Qwen manipulation suite reports none. No standard metric (ms/chunk? effective Hz? jerk?) exists.
- Edge deployment measurement is finally starting: the XPU characterization study profiles VLAs across GPUs/NPUs (compute-bound VLM phase, memory-bound expert phase; 2.9โ3.3ร via phase-aware scheduling) โ but manipulation-side FP8/TensorRT deployment reports remain rare.
- Smoothness metrics (NSPARC etc.) are young; no benchmark scores hesitation/recovery under forced delay spikes.
- All continuation methods assume the policy is the bottleneck; perception-latency (multi-view encoding) is unaddressed.