ICML 2026 RoboFlow4D - Heungwoo/research GitHub Wiki

RoboFlow4D — A lightweight end-to-end flow world model for real-time, flow-guided manipulation

Venue: ICML 2026 (Poster) Category: World Model Affiliations: Sixu Lin, Junliang Chen, Huaiyuan Xu, Zhuohao Li, Guangming Wang, Yixiong Jing, Sheng Xu, Runyi Zhao, Brian Sheil, Lap-Pui Chau, Guiliang Liu Traction (2026-06): 1 citation (arXiv)

System-level comparison of flow-based planning and RoboFlow4D's adaptive 4D flow horizon (Figure 1 from Lin et al., 2026)

Problem

Robust manipulation benefits from inserting an explicit planning signal between observation and action — predicting future manipulation trajectories ("flows") and then tracking them. Prior flow planners fall into two camps, both flawed. 2D image-space flow planners (e.g. predicting pixel trajectories) lack depth and geometry, so a "seemingly reasonable" 2D path can imply collisions or physically infeasible motions. Genuine-3D approaches first generate a task-conditioned video and then stack heavyweight expert modules (depth estimation, grounding, point tracking, 3D reconstruction) to lift it into 3D flow. These modular pipelines incur minute-level latency — the paper cites Dream2Flow at 3–11 minutes and NovaFlow at ~2 minutes on an H100 — making real-time replanning impractical. RoboFlow4D targets a lightweight, end-to-end alternative.

Method

RoboFlow4D unifies perception and planning in a single network that directly predicts multi-frame 3D flows across time (i.e. 4D spacetime) from RGB images and a text instruction, dropping the modular expert stack.

  • Token extraction. A Vision Encoder (DINOv2 local patch tokens + SigLIP global tokens), a Text Encoder (SigLIP text encoder), and an optional Point Encoder (for 2D gripper query points) produce multimodal tokens. Local tokens are aggregated into context tokens via multi-head attention over zero-initialized learnable queries.
  • 3D Perceiver. A Resampler with learnable 3D queries distills 3D geometry from an off-the-shelf 3D foundation model (VGGT) via an alignment loss, injecting 3D awareness into 2D features without RGB-D input.
  • FlowDiT. A diffusion-based DiT denoises future multi-frame 3D flows conditioned on the extracted multimodal tokens.
  • Slow–fast closed loop. RoboFlow4D acts as a low-frequency slow planner; a flow-conditioned action policy (DP or DiT) is the high-frequency fast executor, rolling out multiple action chunks per flow plan. The planning horizon adaptively extends from the current state to the atomic-task goal.

Overview: encoders, 3D Perceiver, FlowDiT, and the slow planner / fast executor closed loop (Figure 2 from Lin et al., 2026)

Results

  • LIBERO (Table 1). Adding RoboFlow4D lifts a Diffusion Policy baseline from 78.9% to 85.1% average (+6.2; +8.2 Spatial, +8.0 Long) and a DiT policy from 83.7% to 87.7% (+4.0), competitive with much larger VLAs.
  • ManiSkill3 (Table 2), single third-view, 100 trials/task. DP improves from 12.3% to 22.0% average (+9.7); DiT from 12.7% to 23.7% (+11.0), with ≥10-point gains on Push/Pick/Stack.
  • Real-world (Table 5), ~20 trials/task. DP + RoboFlow4D reaches 40.0% average success (from 27.5%), surpassing π₀-Fast (32.5%) and matching π₀ (41.3%) while also reducing completion time.
  • Efficiency. ~120× speedup over modular pipelines and >24% smaller model scale than other flow models. On a single RTX 6000, one flow plan takes 0.68 s; the action policy is 0.20 s per forward pass (H=20 chunk).
  • Ablations. Removing the 3D alignment, query points, or context token each raises the 3D ℓ₂ flow error (best 0.0142); the slow–fast loop is stable across update ratios r ∈ {4,2,1}.

Significance

RoboFlow4D shows that a single end-to-end network can supply 4D motion priors that previously required minute-scale, multi-module video-generation pipelines, collapsing latency by two orders of magnitude while consistently boosting lightweight controllers. The slow–fast design makes explicit flow planning compatible with real-time, consumer-GPU robot deployment.

Links

← Back to ICML-2026