RSS 2026 Action to Action Flow Matching - Heungwoo/research GitHub Wiki
Action-to-Action Flow Matching
Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Imitation learning 3 · paper #209 Authors: Jindou Jia, Gen Li, Xiangyu Chen, Tuo An, Yuxuan Hu, Jingliang Li, Xinying Guo, Jianfei Yang arXiv: 2602.07322 · program page
Summary compiled from the arXiv paper (v2, 7 May 2026); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Fig. 1: Comparison of policy paradigms — (a) a regression policy maps inputs deterministically to actions, (b) a diffusion policy denoises from Gaussian noise N(0,I), and (c) the proposed A2A flow policy transports the previous action history a_≤t to future actions a_>t. The right insets contrast the long "noise-to-action" transport against the much shorter "action-to-action" path, which enables single-step generation.
Problem
Diffusion-based policies sample from random Gaussian noise and typically need dozens of iterative denoising steps to produce a clean action, incurring high inference latency — a major bottleneck for real-time robotic control where low cycle time and rapid feedback are essential.
Method
Action-to-Action flow matching (A2A) replaces uninformed noise sampling with initialization informed by the previous proprioceptive action. Rather than treating proprioceptive feedback as a static condition, A2A embeds historical proprioceptive action sequences into a high-dimensional latent space and learns a flow that transports the historical action distribution to future actions. The shorter action-to-action transport path bypasses iterative denoising — feasible even with a lightweight MLP — and an inference-consistency loss supports single-step generation. The authors also extend A2A to video generation.
Results
Across five Roboverse simulation tasks (Close Box, Pick Cube, Stack Cube, Open Drawer, Pick-Place Bowl; 100 demos, 30 epochs) A2A leads all nine compared methods, scoring 92% / 92% / 86% / 92% / 90% and outperforming eight state-of-the-art baselines (VITA, FM-UNet, FM-DiT, DDPM-UNet/DiT, DDIM-UNet, Score-UNet, ACT). It converges up to 20× and 5× faster than vanilla diffusion and flow-matching methods respectively, and sustains high-quality single-step generation with a latency of 0.56 ms (kept below 1 ms), while showing improved robustness to visual perturbations and generalization to unseen configurations.
Significance
Offers a fast, single-step generative policy that keeps diffusion-style expressivity at real-time latency by grounding generation in proprioceptive history. Connects to Review-Realtime-Execution and Review-Human-Video-Transfer.
← Back to RSS 2026 survey · RSS-2026-Papers · Home