ICLR 2026 Flow To Policy - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 Β· OpenReview: dQ6d5bgXtM Affiliation: Tsinghua IIIS Β· UCSD Β· WUSTL Β· Shanghai Qi Zhi Β· SJTU Category: Imitation Learning β Hierarchical / pixel-flow planner / online self-imitation Trend tag: Cross-embodiment video pretraining Β· hindsight relabeling Β· point-flow planners
flowchart LR
subgraph High[High-level Planner: 2D point flow]
Vid[Action-free videos D_h<br/>~300 demos/task] --> CT[CoTracker / video tracker Ξ¦]
CT --> Lbl["Trace labels p_{t+1:t+H}"]
Lbl --> TT[Track-Transformer F_flow<br/>ATM-style multi-modal transformer]
Img[Current frame o_t] --> TT
Sample[Task-centric sampler:<br/>random points on end-effector + objects<br/>16 / 8 / 8 / 16 per task] --> TT
end
subgraph Low[Low-level Policy: flow-conditioned imitation]
TT --> Subgoal[Subgoal G_t = future H point trajectories]
Subgoal --> Pol[Transformer policy<br/>spatial CLS + proprio + action CLS<br/>chunk size = 5, frame stack = 2]
Img --> Pol
Pol --> Act["Action chunk a_{t:t+5}"]
end
Act --> Env[Environment]
Env --> Rollout["Online rollout Ο = {o, a}_{1:T}<br/>Gaussian exploration Ο=0.1"]
Rollout --> Hind["Hindsight relabel:<br/>compute ACHIEVED flow via Ξ¦<br/>store ot, at, achieved p_{t:t+H}"]
Hind --> Buf[Replay buffer D_r]
Buf --> Pol
Cross-embodiment video pretraining is attractive: planners trained on action-free human and robot videos generalise well. Point flow (Wen et al. 2023 ATM, Bharadhwaj et al. 2024) is a particularly clean high-level representation because it filters out appearance and lighting variation while encoding motion dynamics. But translating these flow plans into reliable low-level robot actions is hard:
- Analytic / optimisation translators assume rigid-body dynamics and break on visual occlusions or non-rigid contacts.
- Data-driven low-level policies need many in-domain demonstrations, which kills the cross-embodiment scaling story.
- Online RL with video-prediction rewards (Escontrela et al. 2023) suffers from inefficient exploration on long-horizon tasks.
The paper proposes a third path: treat every online rollout β including failed ones β as a successful demonstration of whatever it actually accomplished, by relabeling the high-level goal to the achieved flow.
Following the Track Transformer from ATM (Wen et al. 2023), the flow predictor F_flow(o_t, p_t; ΞΎ) outputs a sequence of subgoals G_t = {pΜ_i}_{i=t}^{t+H} over a planning horizon H = 8 (vs. ATM's 16). The training loss is:
L_flow = E[ βF_flow(o_t, p_t) β p_{t+1:t+H}β ]
Task-centric point sampling. Crucially, instead of grid points (32 fixed points on a grid), HinFlow samples points on the end-effector and key objects, identified via simulation segmentation masks or Grounded-SAM-2 in the real world (prompted with "robot" and "white mouse" in the real-world experiment). Different point budgets per task β e.g., 16 on butter + 16 on gripper; 8 on chocolate + 8 on drawer + 16 on gripper. The wrist camera gets a fixed 32-point grid because task-relevant objects may leave its view.
Training settings (flow predictor): 1000 epochs, batch 64, length 16 / number 32, patch size 4, frame stack 1, random image mask 0.5 + ColorJitter + random flow shift.
Pseudocode (Algorithm 1):
1. Train flow predictor F_flow on D_h
2. Pretrain flow-conditioned policy Ο on action-labeled D_a
3. Initialise replay buffer D_r
4. for episode = 1, 2, β¦:
Roll out Ο = {o_1, a_1, ..., o_T} ~ F_flow β Ο
Use Ξ¦ to compute achieved flows in Ο
Add tuples (o_t, a_t, {p_i}_{t:t+H}) to D_r
Sample batch from D_r and update Ο via Eq. (3)
Training objective on the buffer:
min_ΞΈ E_{(o_t, a_t, {p_i}) ~ D_r} L( Ο(o_t, {p_i}_{t:t+H}; ΞΈ), a_t )
Key design choices:
- Short-horizon flow (H = 8). Long horizon makes the planner brittle and the relabel problem ill-posed; H=8 yields stable performance, H=4 fails (Fig. 8).
- Self-imitation framing. Failed rollouts become successful demonstrations of whatever they did. Avoids the credit-assignment problem of RL with visual rewards.
- Exploration noise Ο = 0.1 Gaussian on each action dimension, with custom gripper exploration that holds binary open/close states for several frames (avoiding rapid open-close oscillation).
Transformer-based, following Kim et al. 2021 / Wen et al. 2023:
- Encode multi-view images into spatial tokens; concatenate with a learned spatial CLS token; self-attention β spatial CLS as the per-timestep representation.
- Project proprioception into the shared embedding space.
- Interleave (spatial CLS, proprio, action CLS) across timesteps; causally-masked self-attention.
- Per-timestep action CLS + reconstructed flow β MLP β action.
Action chunking (Zhao et al. 2023) with chunk size 5 + exponential temporal ensemble w_i = exp(βm Β· i).
| Stage | Value |
|---|---|
| Flow predictor epochs | 1000 |
| Flow predictor batch | 64 |
| Flow length / number | 16 / 32 |
| Patch size | 4 |
| Policy input flow length / number | 8 / 32 |
| Policy frame stack | 2 |
| Chunk size | 5 |
| Pretraining iterations | 10,000 |
| Online interaction steps | 80,000 (= updates) |
| Batch size (online) | 64 |
| Exploration noise Ο | 0.1 |
| Augmentations | ColorJitter, random flow shift |
| Hardware | 1Γ RTX 3090 (24 GB) |
| Pretrain time | ~30 min |
| Online stage time | ~11 hours |
Demonstrations: 1 per LIBERO task and 5 per ManiSkill task.
Average success rate across all 7 tasks at 80k environment steps:
| Method | Average SR |
|---|---|
| BC | (collapses without flow) |
| ATM (grid) | varies |
| ATM (seg) | ~58% (strongest baseline implied) |
| Online VPT | weak; idiom errors dominate |
| HinFlow (Ours) | 84.0% |
The paper reports 1.45Γ improvement over the strongest baseline and >2Γ improvement over the base policy. Concrete tasks where HinFlow lifts near-zero baseline performance to ~75%: Hide Chocolate (LIBERO long-horizon) and Pull Cube Tool (ManiSkill complex object interaction). Specific per-task curves are in Fig. 4 but no numeric table is provided.
Place Butter, Place Book, Hide Chocolate, Close Microwave. 1 action-labeled demo per task. HinFlow rises from low/zero initial success to ~80-100% on each. ATM-seg eventually catches up on Place Butter, but lags on Hide Chocolate and Close Microwave. BC remains near zero.
Place Sphere, Pull Cube Tool, Poke Cube. 5 action-labeled demos per task. HinFlow significantly outperforms ATM variants on Pull Cube Tool (near-zero baseline β ~75%).
Pick-and-place a mouse onto a pad, 15 cm Γ 15 cm randomisation region, 10 Hz control with ZED2 wrist + third-person cameras at 128Γ128, 2 action-labeled demos + 100 action-free videos + 86 online episodes (~1 h):
| Method | Success |
|---|---|
| BC | 4 / 20 = 20% |
| ATM (seg) | 8 / 20 = 40% |
| HinFlow | 8 / 20 β 19 / 20 (95%) after 10k online steps |
HinFlow starts at 40% (same as ATM-seg) but improves to 95% within ~1 hour of interleaved data collection and updates.
Planner trained on a large action-free Franka dataset; only 5 labeled demos on the target arm:
| Task | Target arm | Cross-embodiment data | Success Rate |
|---|---|---|---|
| Place Book | Kinova Gen3 | with | 48.1% |
| Place Book | Kinova Gen3 | without | 0.6% |
| Poke Cube | xArm6 | with | 61.3% |
| Poke Cube | xArm6 | without | 24.4% |
Gains exceed 40 percentage points β cross-embodiment video pretraining gives the planner the needed motion prior.
Zero-shot evaluation on perturbations not seen during low-level training:
| Setting | BC | HinFlow |
|---|---|---|
| Original | 67.5% | 100.0% |
| Extra distractors | 0.0% | 92.8% |
| Unseen target (chocolate pudding) | 6.5% | 96.2% |
Flow's invariance to appearance is what gives HinFlow this near-zero-shot generalisation when the planner has seen the variations.
- Number of action-labeled demos (Fig. 7). With 0 demos, low-level exploration fails completely. With β₯1 demo on LIBERO or β₯2 on ManiSkill, final performance converges regardless of initial demo count β the online imitation phase washes out initial demo budget. HinFlow is not sensitive to initial policy quality as long as bootstrapping is non-trivial.
- Flow length (Fig. 8). H β {4, 8, 12, 16}. H=8/12/16 are stable; H=4 collapses on long-horizon tasks (insufficient guidance). The paper uses H=8 as default.
- Point sampling (ATM-grid vs. ATM-seg). Both baselines use 32 points; ATM-seg with task-centric points beats ATM-grid by margins consistent with the importance of what to track.
- Online VPT comparison. Online VPT alternates IDM updates and BC retraining every 10k steps. The IDM produces unreliable pseudo-action labels on action-free video β HinFlow's point-flow goals avoid this issue entirely.
- Compute (Appendix A.1). One 3090 GPU, 30 min pretrain + 11 h online β comparable to or cheaper than offline-only baselines that need bigger datasets.
- Bootstrap demonstrations are required. Without any in-domain action-labeled data, the low-level policy fails to produce meaningful exploratory rollouts. The authors point to open-world VLAs (RT-2, OpenVLA, Οβ) as possible bootstraps that could eliminate this requirement.
- 2D point flow is ambiguous for complex 3D motions. Out-of-plane rotations, occlusions, and full 6-DoF motions are not well-captured. Extension to 3D motion fields (Yin et al. 2025) is identified as future work.
- Rotation restriction. For several tasks the authors manually disable rotational degrees of freedom (e.g., only z-axis for Place Book; all rotations disabled for the other tasks) to make the flow-prediction problem tractable. This is a non-trivial assumption that limits applicability.
- Real-world evaluation is single-task. Only mouse-pickup on a Franka β broader real-world tasks not explored.
HinFlow bridges the video-pretrained planner trend and deployable low-level policies, with three notable departures:
- Short-horizon flow (8 frames) β unlike ATM, Track2Act, or human-video-pretraining methods that predict long-horizon trajectories. Short horizons make hindsight relabel self-supervised (you can always describe what the robot just did).
- Self-imitation, not RL. Online refinement is supervised learning on relabeled goals, not reward maximisation. This sidesteps the exploration bottlenecks of Flow Matching Policy Gradients and the long-horizon credit-assignment problems of visual-reward RL.
- Task-centric point sampling with SAM-based segmentation β important enough that the ATM-seg variant of ATM with HinFlow's sampler is a strong baseline.
Versus related work:
- ATM (Wen et al. 2023): offline flow-to-policy. HinFlow adds the hindsight online loop.
- Track2Act (Bharadhwaj et al. 2024b): long-horizon flow + analytic translator. HinFlow uses learned policy + hindsight, avoiding rigidity assumptions.
- Online VPT (this paper's baseline): pseudo-action labeling via IDM. Less reliable than goal-flow relabeling.
- Human Video Pretraining / EgoDex: complementary β those provide the source video; HinFlow provides the policy-grounding mechanism.
- Guided Flow Policy / Flow Matching PG: alternative formulations of flow-driven RL/imitation.
- Align-Then-Steer: different angle on closing the planner-policy gap via latent alignment.
The cross-embodiment transfer experiment (Franka β Kinova / xArm with 5 demos) is the strongest empirical claim β converting cross-embodiment video into a 48-61% success-rate policy on the target arm is a real shift in what video pretraining can buy you.
- OpenReview: https://openreview.net/forum?id=dQ6d5bgXtM
- Project: https://dwjshift.github.io/HinFlow
- Flow Matching Policy Gradients (on-policy flow RL)
- Guided Flow Policy
- Human Video Pretraining
- EgoDex
- Align-Then-Steer