ICML 2026 RoboFlow4D - Heungwoo/research GitHub Wiki
RoboFlow4D — A lightweight end-to-end flow world model for real-time, flow-guided manipulation
Venue: ICML 2026 (Poster) Category: World Model Affiliations: Sixu Lin, Junliang Chen, Huaiyuan Xu, Zhuohao Li, Guangming Wang, Yixiong Jing, Sheng Xu, Runyi Zhao, Brian Sheil, Lap-Pui Chau, Guiliang Liu Traction (2026-06): 1 citation (arXiv)

Problem
Robust manipulation benefits from inserting an explicit planning signal between observation and action — predicting future manipulation trajectories ("flows") and then tracking them. Prior flow planners fall into two camps, both flawed. 2D image-space flow planners (e.g. predicting pixel trajectories) lack depth and geometry, so a "seemingly reasonable" 2D path can imply collisions or physically infeasible motions. Genuine-3D approaches first generate a task-conditioned video and then stack heavyweight expert modules (depth estimation, grounding, point tracking, 3D reconstruction) to lift it into 3D flow. These modular pipelines incur minute-level latency — the paper cites Dream2Flow at 3–11 minutes and NovaFlow at ~2 minutes on an H100 — making real-time replanning impractical. RoboFlow4D targets a lightweight, end-to-end alternative.
Method
RoboFlow4D unifies perception and planning in a single network that directly predicts multi-frame 3D flows across time (i.e. 4D spacetime) from RGB images and a text instruction, dropping the modular expert stack.
- Token extraction. A Vision Encoder (DINOv2 local patch tokens + SigLIP global tokens), a Text Encoder (SigLIP text encoder), and an optional Point Encoder (for 2D gripper query points) produce multimodal tokens. Local tokens are aggregated into context tokens via multi-head attention over zero-initialized learnable queries.
- 3D Perceiver. A Resampler with learnable 3D queries distills 3D geometry from an off-the-shelf 3D foundation model (VGGT) via an alignment loss, injecting 3D awareness into 2D features without RGB-D input.
- FlowDiT. A diffusion-based DiT denoises future multi-frame 3D flows conditioned on the extracted multimodal tokens.
- Slow–fast closed loop. RoboFlow4D acts as a low-frequency slow planner; a flow-conditioned action policy (DP or DiT) is the high-frequency fast executor, rolling out multiple action chunks per flow plan. The planning horizon adaptively extends from the current state to the atomic-task goal.

Results
- LIBERO (Table 1). Adding RoboFlow4D lifts a Diffusion Policy baseline from 78.9% to 85.1% average (+6.2; +8.2 Spatial, +8.0 Long) and a DiT policy from 83.7% to 87.7% (+4.0), competitive with much larger VLAs.
- ManiSkill3 (Table 2), single third-view, 100 trials/task. DP improves from 12.3% to 22.0% average (+9.7); DiT from 12.7% to 23.7% (+11.0), with ≥10-point gains on Push/Pick/Stack.
- Real-world (Table 5), ~20 trials/task. DP + RoboFlow4D reaches 40.0% average success (from 27.5%), surpassing π₀-Fast (32.5%) and matching π₀ (41.3%) while also reducing completion time.
- Efficiency. ~120× speedup over modular pipelines and >24% smaller model scale than other flow models. On a single RTX 6000, one flow plan takes 0.68 s; the action policy is 0.20 s per forward pass (H=20 chunk).
- Ablations. Removing the 3D alignment, query points, or context token each raises the 3D ℓ₂ flow error (best 0.0142); the slow–fast loop is stable across update ratios r ∈ {4,2,1}.
Significance
RoboFlow4D shows that a single end-to-end network can supply 4D motion priors that previously required minute-scale, multi-module video-generation pipelines, collapsing latency by two orders of magnitude while consistently boosting lightweight controllers. The slow–fast design makes explicit flow planning compatible with real-time, consumer-GPU robot deployment.
Links
- arXiv: 2605.17522
- ICML 2026: https://icml.cc/virtual/2026/poster/62543
← Back to ICML-2026