RSS 2026 Tune to Learn - Heungwoo/research GitHub Wiki

Tune to Learn: How Controller Gains Shape Robot Policy Learning

Venue: RSS 2026 (Sydney, Jul 13–17) · Session: Imitation learning 2 · paper #139 Authors: Antonia Bronars, Younghyo Park, Pulkit Agrawal arXiv: 2604.02523 · program page

Summary compiled from the arXiv paper (v1, 2 Apr 2026); all numbers quoted from the paper. Trend context: RSS 2026 survey.

Learning paradigms prefer different controller-gain regimes (Figure 1 of arXiv 2604.02523, © the authors)

Fig. 1: Position control law τ = Kp(q_d − q) − Kd·q̇ + τ_ff with proportional (Kp, stiffness) and derivative (Kd, damping) gains. The three heatmaps sweep the Kp–Kd grid and shade the regimes where each paradigm succeeds: (a) behavior cloning favors the compliant, overdamped corner (upper-left), (b) RL adapts to nearly any setting, and (c) sim-to-real is degraded by stiff and overdamped gains.

Problem

Position controllers are the dominant interface for executing learned manipulation policies, yet how to choose controller gains for policy learning is understudied. The conventional wisdom — pick gains for desired task compliance or stiffness — breaks down for state-conditioned policies, where effective stiffness emerges from the interplay of learned reactions and control dynamics rather than from gains alone.

Method

The paper reframes gain selection around learnability — how amenable a gain setting is to the learning algorithm — rather than desired task behavior. It systematically studies how proportional (Kp) and derivative (Kd) gains affect three core pipelines: offline imitation (behavior cloning), reinforcement learning from scratch, and sim-to-real transfer, running extensive experiments across multiple tasks and robot embodiments and reporting closed-loop success across Kp–Kd grids.

Results

Three findings emerge: (1) behavior cloning benefits from compliant and overdamped gain regimes; (2) RL can succeed across all gain regimes given compatible hyperparameter tuning; and (3) sim-to-real transfer is harmed by stiff and overdamped gain regimes. The conclusion is that optimal gain selection depends on the learning paradigm employed, not on the desired task behavior.

Significance

Provides practical guidance for a widely made but poorly understood design choice in robot learning pipelines, cutting across imitation, RL, and sim-to-real. Connects to Review-LBM-Cotraining and RL.

← Back to RSS 2026 survey · RSS-2026-Papers · Home