ICLR 2026 Constrained Demonstrators - Heungwoo/research GitHub Wiki

Learning from Constrained Demonstrators — when a robot can beat its (constrained) teacher

Venue: ICLR 2026 · Authors: Xinhu Li, Ayush Jain, Zhaojing Yang, Yigit Korkmaz, Erdem Bıyık (USC) · arXiv 2510.09096 · Category: RL for manipulation · Trend tag: imitation beyond the expert / suboptimal demonstrations

Approach diagram

flowchart LR
  E[Constrained expert<br/>e.g. joystick = 2D plane] --> D[Suboptimal demonstrations]
  D --> RW[Infer state-only reward<br/>measuring task progress]
  RW --> SL[Self-label reward for unknown states<br/>via temporal interpolation]
  SL --> EX[Agent explores shorter,<br/>more efficient trajectories]
  EX --> P[Policy better than the demonstrated one]
Loading

Problem

Demonstration interfaces — kinesthetic teaching, joystick control, sim-to-real — often constrain the expert from showing optimal behavior because of indirect control, setup restrictions, and hardware-safety limits. A joystick may move an arm only in a 2D plane even though the robot acts in a higher-dimensional space, so collected demonstrations are suboptimal. Key question: can a robot learn a better policy than the one its constrained expert demonstrated?

Method

Rather than directly imitating expert actions, the agent is allowed to go beyond imitation and explore shorter, more efficient trajectories:

  • Use demonstrations to infer a state-only reward signal that measures task progress.
  • Self-label the reward for unknown / unvisited states using temporal interpolation.

Because the reward is state-only (not action-imitating), the policy is free to find more efficient paths than the constrained demonstrator could execute.

Results

  • Outperforms common imitation learning in both sample efficiency and task completion time.
  • On a real WidowX robotic arm, completes the task in 12 seconds — 10x faster than behavioral cloning.

Significance

Reframes learning-from-demonstration for the regime where the robot is more capable than the demonstration interface: instead of capping performance at the (constrained) expert, it uses demonstrations only to define task progress and lets RL-style exploration exceed them. Practically relevant wherever cheap but limited teleoperation interfaces are the data source.

Links

Related pages

← Back to ICLR-2026

⚠️ **GitHub.com Fallback** ⚠️