ICLR 2026 Flow To Policy - Heungwoo/research GitHub Wiki

HinFlow β€” Translating Flow to Policy via Hindsight Online Imitation

Venue: ICLR 2026 Β· OpenReview: dQ6d5bgXtM Affiliation: Tsinghua IIIS Β· UCSD Β· WUSTL Β· Shanghai Qi Zhi Β· SJTU Category: Imitation Learning β€” Hierarchical / pixel-flow planner / online self-imitation Trend tag: Cross-embodiment video pretraining Β· hindsight relabeling Β· point-flow planners

Approach diagram

flowchart LR
  subgraph High[High-level Planner: 2D point flow]
    Vid[Action-free videos D_h<br/>~300 demos/task] --> CT[CoTracker / video tracker Ξ¦]
    CT --> Lbl["Trace labels p_{t+1:t+H}"]
    Lbl --> TT[Track-Transformer F_flow<br/>ATM-style multi-modal transformer]
    Img[Current frame o_t] --> TT
    Sample[Task-centric sampler:<br/>random points on end-effector + objects<br/>16 / 8 / 8 / 16 per task] --> TT
  end
  subgraph Low[Low-level Policy: flow-conditioned imitation]
    TT --> Subgoal[Subgoal G_t = future H point trajectories]
    Subgoal --> Pol[Transformer policy<br/>spatial CLS + proprio + action CLS<br/>chunk size = 5, frame stack = 2]
    Img --> Pol
    Pol --> Act["Action chunk a_{t:t+5}"]
  end
  Act --> Env[Environment]
  Env --> Rollout["Online rollout Ο„ = {o, a}_{1:T}<br/>Gaussian exploration Οƒ=0.1"]
  Rollout --> Hind["Hindsight relabel:<br/>compute ACHIEVED flow via Ξ¦<br/>store ot, at, achieved p_{t:t+H}"]
  Hind --> Buf[Replay buffer D_r]
  Buf --> Pol
Loading

Problem

Cross-embodiment video pretraining is attractive: planners trained on action-free human and robot videos generalise well. Point flow (Wen et al. 2023 ATM, Bharadhwaj et al. 2024) is a particularly clean high-level representation because it filters out appearance and lighting variation while encoding motion dynamics. But translating these flow plans into reliable low-level robot actions is hard:

  • Analytic / optimisation translators assume rigid-body dynamics and break on visual occlusions or non-rigid contacts.
  • Data-driven low-level policies need many in-domain demonstrations, which kills the cross-embodiment scaling story.
  • Online RL with video-prediction rewards (Escontrela et al. 2023) suffers from inefficient exploration on long-horizon tasks.

The paper proposes a third path: treat every online rollout β€” including failed ones β€” as a successful demonstration of whatever it actually accomplished, by relabeling the high-level goal to the achieved flow.

Method (detailed)

4.1 High-level planner: short-horizon point flow

Following the Track Transformer from ATM (Wen et al. 2023), the flow predictor F_flow(o_t, p_t; ΞΎ) outputs a sequence of subgoals G_t = {pΜ‚_i}_{i=t}^{t+H} over a planning horizon H = 8 (vs. ATM's 16). The training loss is:

L_flow = E[ β€–F_flow(o_t, p_t) βˆ’ p_{t+1:t+H}β€– ]

Task-centric point sampling. Crucially, instead of grid points (32 fixed points on a grid), HinFlow samples points on the end-effector and key objects, identified via simulation segmentation masks or Grounded-SAM-2 in the real world (prompted with "robot" and "white mouse" in the real-world experiment). Different point budgets per task β€” e.g., 16 on butter + 16 on gripper; 8 on chocolate + 8 on drawer + 16 on gripper. The wrist camera gets a fixed 32-point grid because task-relevant objects may leave its view.

Training settings (flow predictor): 1000 epochs, batch 64, length 16 / number 32, patch size 4, frame stack 1, random image mask 0.5 + ColorJitter + random flow shift.

4.2 Low-level: Hindsight Flow-conditioned Online Imitation

Pseudocode (Algorithm 1):

1. Train flow predictor F_flow on D_h
2. Pretrain flow-conditioned policy Ο€ on action-labeled D_a
3. Initialise replay buffer D_r
4. for episode = 1, 2, …:
       Roll out Ο„ = {o_1, a_1, ..., o_T} ~ F_flow ∘ Ο€
       Use Ξ¦ to compute achieved flows in Ο„
       Add tuples (o_t, a_t, {p_i}_{t:t+H}) to D_r
       Sample batch from D_r and update Ο€ via Eq. (3)

Training objective on the buffer:

min_ΞΈ E_{(o_t, a_t, {p_i}) ~ D_r} L( Ο€(o_t, {p_i}_{t:t+H}; ΞΈ), a_t )

Key design choices:

  • Short-horizon flow (H = 8). Long horizon makes the planner brittle and the relabel problem ill-posed; H=8 yields stable performance, H=4 fails (Fig. 8).
  • Self-imitation framing. Failed rollouts become successful demonstrations of whatever they did. Avoids the credit-assignment problem of RL with visual rewards.
  • Exploration noise Οƒ = 0.1 Gaussian on each action dimension, with custom gripper exploration that holds binary open/close states for several frames (avoiding rapid open-close oscillation).

Policy architecture

Transformer-based, following Kim et al. 2021 / Wen et al. 2023:

  1. Encode multi-view images into spatial tokens; concatenate with a learned spatial CLS token; self-attention β†’ spatial CLS as the per-timestep representation.
  2. Project proprioception into the shared embedding space.
  3. Interleave (spatial CLS, proprio, action CLS) across timesteps; causally-masked self-attention.
  4. Per-timestep action CLS + reconstructed flow β†’ MLP β†’ action.

Action chunking (Zhao et al. 2023) with chunk size 5 + exponential temporal ensemble w_i = exp(βˆ’m Β· i).

Hyperparameters

Stage Value
Flow predictor epochs 1000
Flow predictor batch 64
Flow length / number 16 / 32
Patch size 4
Policy input flow length / number 8 / 32
Policy frame stack 2
Chunk size 5
Pretraining iterations 10,000
Online interaction steps 80,000 (= updates)
Batch size (online) 64
Exploration noise Οƒ 0.1
Augmentations ColorJitter, random flow shift
Hardware 1Γ— RTX 3090 (24 GB)
Pretrain time ~30 min
Online stage time ~11 hours

Demonstrations: 1 per LIBERO task and 5 per ManiSkill task.

Comprehensive Results

Main benchmark (Fig. 4, average of 5 seeds)

Average success rate across all 7 tasks at 80k environment steps:

Method Average SR
BC (collapses without flow)
ATM (grid) varies
ATM (seg) ~58% (strongest baseline implied)
Online VPT weak; idiom errors dominate
HinFlow (Ours) 84.0%

The paper reports 1.45Γ— improvement over the strongest baseline and >2Γ— improvement over the base policy. Concrete tasks where HinFlow lifts near-zero baseline performance to ~75%: Hide Chocolate (LIBERO long-horizon) and Pull Cube Tool (ManiSkill complex object interaction). Specific per-task curves are in Fig. 4 but no numeric table is provided.

LIBERO tasks (Fig. 4a, 80k env steps, 5 seeds)

Place Butter, Place Book, Hide Chocolate, Close Microwave. 1 action-labeled demo per task. HinFlow rises from low/zero initial success to ~80-100% on each. ATM-seg eventually catches up on Place Butter, but lags on Hide Chocolate and Close Microwave. BC remains near zero.

ManiSkill3 tasks (Fig. 4b, 80k env steps, 5 seeds)

Place Sphere, Pull Cube Tool, Poke Cube. 5 action-labeled demos per task. HinFlow significantly outperforms ATM variants on Pull Cube Tool (near-zero baseline β†’ ~75%).

Real-world Franka Panda (Table 1, 20 trials)

Pick-and-place a mouse onto a pad, 15 cm Γ— 15 cm randomisation region, 10 Hz control with ZED2 wrist + third-person cameras at 128Γ—128, 2 action-labeled demos + 100 action-free videos + 86 online episodes (~1 h):

Method Success
BC 4 / 20 = 20%
ATM (seg) 8 / 20 = 40%
HinFlow 8 / 20 β†’ 19 / 20 (95%) after 10k online steps

HinFlow starts at 40% (same as ATM-seg) but improves to 95% within ~1 hour of interleaved data collection and updates.

Cross-embodiment transfer (Fig. 6)

Planner trained on a large action-free Franka dataset; only 5 labeled demos on the target arm:

Task Target arm Cross-embodiment data Success Rate
Place Book Kinova Gen3 with 48.1%
Place Book Kinova Gen3 without 0.6%
Poke Cube xArm6 with 61.3%
Poke Cube xArm6 without 24.4%

Gains exceed 40 percentage points β€” cross-embodiment video pretraining gives the planner the needed motion prior.

Policy generalisation (Table 2; Place Butter, 5 seeds)

Zero-shot evaluation on perturbations not seen during low-level training:

Setting BC HinFlow
Original 67.5% 100.0%
Extra distractors 0.0% 92.8%
Unseen target (chocolate pudding) 6.5% 96.2%

Flow's invariance to appearance is what gives HinFlow this near-zero-shot generalisation when the planner has seen the variations.

Ablation Studies

  1. Number of action-labeled demos (Fig. 7). With 0 demos, low-level exploration fails completely. With β‰₯1 demo on LIBERO or β‰₯2 on ManiSkill, final performance converges regardless of initial demo count β€” the online imitation phase washes out initial demo budget. HinFlow is not sensitive to initial policy quality as long as bootstrapping is non-trivial.
  2. Flow length (Fig. 8). H ∈ {4, 8, 12, 16}. H=8/12/16 are stable; H=4 collapses on long-horizon tasks (insufficient guidance). The paper uses H=8 as default.
  3. Point sampling (ATM-grid vs. ATM-seg). Both baselines use 32 points; ATM-seg with task-centric points beats ATM-grid by margins consistent with the importance of what to track.
  4. Online VPT comparison. Online VPT alternates IDM updates and BC retraining every 10k steps. The IDM produces unreliable pseudo-action labels on action-free video β€” HinFlow's point-flow goals avoid this issue entirely.
  5. Compute (Appendix A.1). One 3090 GPU, 30 min pretrain + 11 h online β€” comparable to or cheaper than offline-only baselines that need bigger datasets.

Limitations (stated by authors, Β§6)

  1. Bootstrap demonstrations are required. Without any in-domain action-labeled data, the low-level policy fails to produce meaningful exploratory rollouts. The authors point to open-world VLAs (RT-2, OpenVLA, Ο€β‚€) as possible bootstraps that could eliminate this requirement.
  2. 2D point flow is ambiguous for complex 3D motions. Out-of-plane rotations, occlusions, and full 6-DoF motions are not well-captured. Extension to 3D motion fields (Yin et al. 2025) is identified as future work.
  3. Rotation restriction. For several tasks the authors manually disable rotational degrees of freedom (e.g., only z-axis for Place Book; all rotations disabled for the other tasks) to make the flow-prediction problem tractable. This is a non-trivial assumption that limits applicability.
  4. Real-world evaluation is single-task. Only mouse-pickup on a Franka β€” broader real-world tasks not explored.

Significance & Positioning

HinFlow bridges the video-pretrained planner trend and deployable low-level policies, with three notable departures:

  • Short-horizon flow (8 frames) β€” unlike ATM, Track2Act, or human-video-pretraining methods that predict long-horizon trajectories. Short horizons make hindsight relabel self-supervised (you can always describe what the robot just did).
  • Self-imitation, not RL. Online refinement is supervised learning on relabeled goals, not reward maximisation. This sidesteps the exploration bottlenecks of Flow Matching Policy Gradients and the long-horizon credit-assignment problems of visual-reward RL.
  • Task-centric point sampling with SAM-based segmentation β€” important enough that the ATM-seg variant of ATM with HinFlow's sampler is a strong baseline.

Versus related work:

  • ATM (Wen et al. 2023): offline flow-to-policy. HinFlow adds the hindsight online loop.
  • Track2Act (Bharadhwaj et al. 2024b): long-horizon flow + analytic translator. HinFlow uses learned policy + hindsight, avoiding rigidity assumptions.
  • Online VPT (this paper's baseline): pseudo-action labeling via IDM. Less reliable than goal-flow relabeling.
  • Human Video Pretraining / EgoDex: complementary β€” those provide the source video; HinFlow provides the policy-grounding mechanism.
  • Guided Flow Policy / Flow Matching PG: alternative formulations of flow-driven RL/imitation.
  • Align-Then-Steer: different angle on closing the planner-policy gap via latent alignment.

The cross-embodiment transfer experiment (Franka β†’ Kinova / xArm with 5 demos) is the strongest empirical claim β€” converting cross-embodiment video into a 48-61% success-rate policy on the target arm is a real shift in what video pretraining can buy you.

Links

Related pages

← Back to ICLR-2026 Β· Topic: RL

⚠️ **GitHub.com Fallback** ⚠️