ICML 2026 Reflex - Heungwoo/research GitHub Wiki

Reflex: Real-Time Vision-Language-Action Control through Streaming Inference — timestep-invariant streaming for flow-matching VLAs

Venue: ICML 2026 (Poster) Category: Efficiency

Problem

Flow-matching VLAs generate action chunks by iterating a denoising/flow process, which makes per-step inference expensive and bursty. For real-time closed-loop control, this latency translates into delayed reactions and unstable control loops, preventing flow-matching policies from running at the high, steady rates needed for responsive manipulation.

Method

Reflex enables real-time streaming inference for flow-matching VLAs by exploiting timestep-invariance: it partitions the attention computation into static, sliding, and dynamic regions so that the key/value cache can be updated in O(1) per step rather than recomputed. This is paired with:

  • AdaRMSNorm, an adaptive normalization that supports the streaming formulation.
  • An asynchronous pipeline that overlaps computation to sustain a stable streaming rate.
flowchart LR
    A[Observation stream] --> B[Partition attention]
    B --> S[Static region]
    B --> W[Sliding region]
    B --> D[Dynamic region]
    S & W & D --> C["O(1) KV-cache update"]
    C --> N[AdaRMSNorm]
    N --> P[Async pipeline]
    P --> O[50Hz streaming actions]

Results

Reflex achieves a reported 2.58x inference speedup and stable 50 Hz streaming, while reducing reaction latency by up to 54%. Together these allow a flow-matching VLA to run as a responsive, continuously-streaming controller rather than in slow, chunked bursts.

Significance

Reaction latency is one of the main practical barriers to deploying flow-matching VLAs in reactive, contact-rich manipulation. By reframing the policy's attention as a streaming computation with constant-time cache updates, Reflex turns a heavy generative head into something that can keep pace with a real control loop — without retraining a fundamentally different architecture. The static/sliding/dynamic decomposition and AdaRMSNorm are reusable ingredients for other flow- and diffusion-based policies seeking high-frequency, low-latency control.

(This page is grounded in the ICML 2026 program summary; no arXiv preprint or public abstract was available at the time of writing, so numbers above are from the program listing.)

Links

  • ICML 2026: program listing (Reflex: Real-Time Vision-Language-Action Control through Streaming Inference)

← Back to ICML-2026