ICLR 2026 RFS - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 Authors: Entong Su, Tyler Westenbroek, Anusha Nagabandi, Abhishek Gupta (University of Washington, Paul G. Allen School; Amazon Frontier AI & Robotics) arXiv: 2602.01789 · Project: weirdlabuw.github.io/rfs/ Category: Dexterous Manipulation / RL Trend tag: Trends 3 + 6
flowchart LR
Obs[Observation] --> Base[Base flow-matching policy v_θ<br/>FROZEN]
Obs --> Noise[Latent-noise policy π_H<br/>RL: picks initial noise a₀<br/>GLOBAL exploration]
Obs --> Res[Residual policy π_r<br/>RL: action correction a_r<br/>LOCAL refinement]
Noise --> Base
Base --> Sum((+))
Res --> Sum
Sum --> Act["Final action<br/>a = Des(s, a₀, v_θ) + a_r"]
Contact-rich dexterous manipulation (deformable objects, in-hand) is hard to learn by pure imitation — small contact-dynamics deviations compound quickly. Full policy RL is unstable because it destabilizes pretrained weights.
RFS adapts a frozen pretrained flow-matching policy v_θ via RL using two complementary steering mechanisms that are jointly optimized (the base policy parameters are never updated):
-
Latent-noise steering (input modulation) — a learned policy
π_Hselects the initial latent noisea₀fed into the denoising process, driving global behavioral exploration. This is the DSRL-style mechanism (DSRL). -
Residual action (output modulation) — a learned policy
π_radds a small correctiona_rto the denoised base action for local, contact-level refinement (the residual-RL pattern, same family as PLD).
The executed action is a = Des(s, a₀, v_θ) + a_r. RL is PPO for online fine-tuning in simulation and TD3+BC (offline, with imitation regularization) for real-world fine-tuning. The paper's thesis is that residual-only and noise-only methods each cover only half the exploration space; combining them is the contribution.
- Simulation (6 dexterous tasks, Isaac Sim/ORBIT, Franka + LEAP hand, ~400 VR-teleop demos/task): average success ~0.86 (e.g., Stacking 0.95, Pick&Place 0.94, Grasping 0.89, Pouring 0.87, Packing 0.78, Push-to-Grasp 0.72), vs. the strongest single-component baseline DSRL ~0.48.
- Real-world (Franka + LEAP, 50 human corrective demos, offline TD3+BC): seen-object Grasping 80%, Pick-and-Place 90%, vs. zero-shot base policy 43.3% / 50%; generalizes to novel deformable objects.
- Ablation: residual-only (Policy Decorator) and noise-only (DSRL) each win on different tasks (DSRL strong on grasping 0.74 but weak on stacking 0.15); RFS dominates across all, confirming complementary local/global exploration.
- Baselines: DPPO, ReinFlow, IQL, AWAC, Flow Q-Learning, RLPD, IBRL, Policy Decorator, ResiP, DSRL.
Unifies the two previously separate RL-adaptation strategies for frozen generative policies — residual actions (PLD, Policy Decorator) and latent-noise steering (DSRL) — showing they are complementary (local vs. global exploration) and applying the combination to contact-rich dexterous manipulation. Combined with EgoDex (data) and UniHM (generalist backbone), RFS is the fine-tuning pillar of the "dexterity enters data + RL era" story.
- arXiv:2602.01789 · Project page: weirdlabuw.github.io/rfs/ · OpenReview: Kt9tJeOwjy
- ICLR 2026 listing