ICLR 2026 Constrained Demonstrators - Heungwoo/research GitHub Wiki
Venue: ICLR 2026 · Authors: Xinhu Li, Ayush Jain, Zhaojing Yang, Yigit Korkmaz, Erdem Bıyık (USC) · arXiv 2510.09096 · Category: RL for manipulation · Trend tag: imitation beyond the expert / suboptimal demonstrations
flowchart LR
E[Constrained expert<br/>e.g. joystick = 2D plane] --> D[Suboptimal demonstrations]
D --> RW[Infer state-only reward<br/>measuring task progress]
RW --> SL[Self-label reward for unknown states<br/>via temporal interpolation]
SL --> EX[Agent explores shorter,<br/>more efficient trajectories]
EX --> P[Policy better than the demonstrated one]
Demonstration interfaces — kinesthetic teaching, joystick control, sim-to-real — often constrain the expert from showing optimal behavior because of indirect control, setup restrictions, and hardware-safety limits. A joystick may move an arm only in a 2D plane even though the robot acts in a higher-dimensional space, so collected demonstrations are suboptimal. Key question: can a robot learn a better policy than the one its constrained expert demonstrated?
Rather than directly imitating expert actions, the agent is allowed to go beyond imitation and explore shorter, more efficient trajectories:
- Use demonstrations to infer a state-only reward signal that measures task progress.
- Self-label the reward for unknown / unvisited states using temporal interpolation.
Because the reward is state-only (not action-imitating), the policy is free to find more efficient paths than the constrained demonstrator could execute.
- Outperforms common imitation learning in both sample efficiency and task completion time.
- On a real WidowX robotic arm, completes the task in 12 seconds — 10x faster than behavioral cloning.
Reframes learning-from-demonstration for the regime where the robot is more capable than the demonstration interface: instead of capping performance at the (constrained) expert, it uses demonstrations only to define task progress and lets RL-style exploration exceed them. Practically relevant wherever cheap but limited teleoperation interfaces are the data source.
- arXiv 2510.09096
- ICLR 2026 listing
← Back to ICLR-2026