IROS 2026 3D FlowMatch Actor - Heungwoo/research GitHub Wiki
IROS 2026 — 3D FlowMatch Actor: Unified 3D Policy for Single & Dual-Arm Manipulation
Venue: IROS 2026 (Pittsburgh) · paper #4030 · Carnegie Mellon University · NVIDIA · National Taiwan University (Gkanatsios, Xu, Bronars, Mousavian, Ke, Fragkiadaki). Paper: arXiv 2508.11002 · project · code. The bimanual-SOTA datapoint of IROS 2026 — one 3D flow-matching policy for both single- and dual-arm manipulation, ~30× faster than 3D-diffusion policies, +41.4% over the prior best on bimanual PerAct2. Companions: IROS 2026 survey · VLA Architectures · Real-Time Execution · World Models.

1. Idea
3DFA combines flow matching for trajectory prediction with 3D pretrained visual scene representations for learning from demonstration, using 3D relative attention between action and visual tokens during denoising (building on 3D diffusion single-arm policies). The headline is a unified architecture that handles single and dual-arm without separate designs: a frozen image encoder lifts image+depth into 3D scene tokens, and left/right noised trajectory tokens are denoised jointly by a Transformer that attends over scene, proprioception, and language tokens.
2. What's new
- ~30× faster training and inference than prior 3D-diffusion policies, via flow matching + system-level and architectural optimizations — without sacrificing performance. Concretely on PerAct2: training 21 days → 16 hours, inference 0.5 Hz → 18.2 Hz (the latency answer for 3D policies, cf. Real-Time Execution).
- 3D relative attention between action tokens and 3D visual tokens — the inductive bias that grounds actions in scene geometry.
- Dense end-effector trajectory prediction in the unimanual case — eliminates motion planning.
3. Results
- Bimanual PerAct2: new state of the art, beating the next-best by an absolute +41.4%.
- Real-world: surpasses baselines with up to 1000× more parameters and far more pretraining.
- Unimanual: new SOTA on 74 RLBench tasks by directly predicting dense EE trajectories.
- Ablations confirm the design choices drive both effectiveness and efficiency.
4. Why it matters (bimanual lens)
3DFA is IROS 2026's strongest evidence for the survey §5.3 insight that the bimanual frontier is coordination architecture, not just adding a second arm: a single 3D policy unifies single/dual-arm, and the win comes from 3D-geometry grounding + flow-matching efficiency, not scale (it beats models 1000× larger). It also lands squarely in the geometry-grounded manipulation thread (3D-foundation-aligned VLA) and the flow-matching efficiency thread — a small, fast, geometry-aware policy beating big ones is the counter-narrative to pure scaling.
Limitations (reviewer): relies on a 3D visual representation / calibrated 3D input (not raw RGB VLA); PerAct2/RLBench are the eval substrate; no language-instruction generalization tested (it's a demonstration policy, not a VLA with a VLM backbone).
5. Links
- Official program: IROS 2026 (paper #4030) · survey: IROS 2026
- Related: EquiBim (the other bimanual page) · VLA Architectures · Real-Time Execution · Humanoid VLA
← Back to IROS 2026 survey · Home