ICML 2026 Focus Then Contact - Heungwoo/research GitHub Wiki
Focus-Then-Contact (FTC) — Affordance-guided residual RL that speeds up real-world contact-rich learning
Venue: ICML 2026 (Poster) Category: RL for VLA Affiliations: Guanren Qiao, Ruixiang Ouyang, Sheng Xu, Ruixing Jin, Yueci Deng, Yunxin Tai, Kui Jia, Guiliang Liu
Real-world reinforcement learning has shown significant potential for robotic manipulation, but contact-rich tasks remain costly to learn directly on hardware. In particular, many methods still require substantial human-in-the-loop involvement to complete contact-rich tasks, and this burden grows when the environment introduces disruptions such as changed visual backgrounds or shifted object positions. The result is slow convergence and heavy human effort, limiting the practicality of real-world RL for tasks that depend on sustained physical contact.
The authors propose Focus Then Contact (FTC), a lightweight and low-cost method designed to accelerate the convergence of human-in-the-loop real-world RL for contact-rich tasks. FTC combines three components:
- Residual RL base actions. FTC leverages residual RL to provide base actions that help the system quickly reach the target regions, improving sample efficiency by reducing the amount of exploration the RL agent must do from scratch.
- Affordance-guided reward. FTC integrates an affordance-guided reward that drives the real-world RL system to quickly focus on key regions of interest. This lets the robotic arm continuously engage with the goal areas through force-control feedback, supporting the "contact" phase after the "focus" phase has localized where to act.
- Optimized human-in-the-loop control. FTC optimizes the human-in-the-loop implementation to prevent conflicts with the RL agent over control of the robotic arm, so human interventions and policy actions do not fight for control authority.
The naming captures the two-stage intuition: focus on the affordance-relevant region (guided by the reward) and then contact it under force control (driven by the residual policy).
flowchart LR
A[Observation] --> B[Affordance-guided reward<br/>focus on key region]
A --> C[Residual RL base action<br/>reach target region]
B --> D[Force-control engagement]
C --> D
H[Human-in-the-loop<br/>conflict-free control] --> D
D --> E[Contact-rich task success]
FTC is demonstrated on 6 contact-rich tasks in a real-world RL setting. Across these tasks it outperforms baseline methods in achieving high success rates and speeds up robotic contact-rich task learning. The combination of residual base actions (for sample efficiency) and the affordance-guided reward (for fast focusing) is credited with the faster convergence, while the optimized human-in-the-loop scheme reduces conflict between operator and policy. (The ICML abstract reports these qualitative outcomes; per-task numeric success rates are not stated in the available abstract.)
FTC targets a practical bottleneck in real-world robot learning: the high human cost and slow convergence of human-in-the-loop RL on contact-rich tasks. By layering an affordance-guided reward and residual base actions onto an interaction-safe human-in-the-loop loop, it offers a low-cost path to faster, more sample-efficient learning under real-world disruptions such as background and position changes — without requiring simulation or heavy reward engineering.
- ICML 2026: https://icml.cc/virtual/2026/poster/66289
← Back to ICML-2026