ICML 2026 Uncertainty Guided Exploration and Stable Planning - Heungwoo/research GitHub Wiki
QUEST: Uncertainty-Guided Exploration and Stable Planning for Sparse-Reward Manipulation from Limited Demonstrations — adaptive exploration/exploitation switching driven by model uncertainty
Venue: ICML 2026 (Poster) Category: RL for VLA
Problem
Reinforcement learning from demonstrations (RLfD) is a promising route to robotic manipulation under sparse rewards, but it breaks down when demonstrations are limited. With few demonstrations, agents repeatedly drift into out-of-distribution states where learned world models produce poor predictions. The situation worsens in multi-stage tasks: jointly optimizing a learned reward function alongside the policy creates a moving target problem, and the resulting non-stationarity amplifies the effect of model uncertainty on policy learning. The agent must therefore decide, on the fly, whether to trust its model and exploit, or to treat its predictions as unreliable and explore.
Method
The authors propose QUEST, a model-based RL framework that adaptively switches between exploration and exploitation guided by uncertainty to achieve stable and efficient learning. It combines three components:
- Intrinsic rewards that capture environmental stochasticity, supplying signal where the task reward is sparse.
- Ensemble dynamics models that provide an uncertainty estimate used for uncertainty-guided planning — when the ensemble disagrees, the agent treats the region as uncertain and shifts its behaviour accordingly.
- A hybrid sampling strategy that prioritizes rare successful stage transitions, ensuring that the scarce but critical moments of progress through a multi-stage task are not drowned out during learning.
Together these let QUEST modulate between exploring uncertain, out-of-distribution regions and exploiting confident predictions, directly countering the non-stationarity introduced by jointly learning reward and policy.
flowchart TD
A[Limited demonstrations] --> B[Ensemble dynamics model]
B --> C{Uncertainty estimate}
C -->|High| D[Explore + intrinsic reward]
C -->|Low| E[Exploit via planning]
D --> F[Hybrid sampling: rare stage transitions]
E --> F
F --> G[Stable policy update]
G --> B
Results
QUEST is evaluated on challenging sparse-reward manipulation tasks with limited expert demonstrations. It outperforms state-of-the-art methods by 17% on average, with gains increasing to 60% on difficult tasks, indicating that the uncertainty-guided switching is especially valuable where world-model error and sparse reward are most punishing. The authors further demonstrate successful zero-shot sim-to-real transfer on three real-world tasks, showing the learned policies hold up outside simulation.
Significance
QUEST tackles a fundamental tension in RLfD — exploration versus exploitation under unreliable, non-stationary world models — by making model uncertainty the explicit signal that arbitrates between the two. The combination of intrinsic rewards, ensemble-based uncertainty, and transition-aware sampling offers a recipe for stable learning from few demonstrations, and the zero-shot sim-to-real results suggest the approach yields policies robust enough for physical deployment. Its largest gains on the hardest tasks point to uncertainty-guided planning as a particularly useful tool for long-horizon, multi-stage manipulation.
Links
- ICML 2026: https://icml.cc/virtual/2026/poster/64857
← Back to ICML-2026