ICML 2026 Reflex - Heungwoo/research GitHub Wiki
Reflex: Real-Time Vision-Language-Action Control through Streaming Inference — timestep-invariant streaming for flow-matching VLAs
Venue: ICML 2026 (Poster) Category: Efficiency
Problem
Flow-matching VLAs generate action chunks by iterating a denoising/flow process, which makes per-step inference expensive and bursty. For real-time closed-loop control, this latency translates into delayed reactions and unstable control loops, preventing flow-matching policies from running at the high, steady rates needed for responsive manipulation.
Method
Reflex enables real-time streaming inference for flow-matching VLAs by exploiting timestep-invariance: it partitions the attention computation into static, sliding, and dynamic regions so that the key/value cache can be updated in O(1) per step rather than recomputed. This is paired with:
- AdaRMSNorm, an adaptive normalization that supports the streaming formulation.
- An asynchronous pipeline that overlaps computation to sustain a stable streaming rate.
flowchart LR
A[Observation stream] --> B[Partition attention]
B --> S[Static region]
B --> W[Sliding region]
B --> D[Dynamic region]
S & W & D --> C["O(1) KV-cache update"]
C --> N[AdaRMSNorm]
N --> P[Async pipeline]
P --> O[50Hz streaming actions]
Results
Reflex achieves a reported 2.58x inference speedup and stable 50 Hz streaming, while reducing reaction latency by up to 54%. Together these allow a flow-matching VLA to run as a responsive, continuously-streaming controller rather than in slow, chunked bursts.
Significance
Reaction latency is one of the main practical barriers to deploying flow-matching VLAs in reactive, contact-rich manipulation. By reframing the policy's attention as a streaming computation with constant-time cache updates, Reflex turns a heavy generative head into something that can keep pace with a real control loop — without retraining a fundamentally different architecture. The static/sliding/dynamic decomposition and AdaRMSNorm are reusable ingredients for other flow- and diffusion-based policies seeking high-frequency, low-latency control.
(This page is grounded in the ICML 2026 program summary; no arXiv preprint or public abstract was available at the time of writing, so numbers above are from the program listing.)
Links
- ICML 2026: program listing (Reflex: Real-Time Vision-Language-Action Control through Streaming Inference)
← Back to ICML-2026