Uncertainty-Guided Exploration and Stable Planning for Sparse-Reward Manipulation from Limited Demonstrations
Haowen Sun, Liqi Huang, Mingyang Li, Sihua Ren, Xinzhe Chen, Chengzhong Ma, Zeyang Liu, Xingyu Chen, Xuguang Lan
Abstract
Reinforcement learning from demonstrations (RLfD) offers a promising method for robotic manipulation with sparse rewards. However, limited demonstrations often cause agents to encounter out-of-distribution states where world models produce poor predictions. In multi-stage tasks, jointly optimizing a learned reward function and policy introduces a moving target problem, and the resulting non-stationarity intensifies the impact of uncertainty on policy learning. In this work, we propose QUEST, a model-based RL framework that adaptively switches between exploration and exploitation guided by uncertainty to achieve stable and efficient learning. Specifically, our approach employs intrinsic rewards to encourage exploration, leverages ensemble dynamics for uncertainty-guided planning, and introduces a hybrid sampling strategy to prioritize rare successful stage transitions. We evaluate QUEST on challenging sparse-reward manipulation tasks with limited expert demonstrations. Results show that QUEST outperforms state-of-the-art methods by 17% on average, with gains increasing to 60% on difficult tasks. We further demonstrate successful zero-shot sim-to-real transfer on five real-world tasks. Project website: https://quest-official.github.io/QUEST/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5d7d49c1-51e2-491a-af08-7597eb039a4dBuilds on16
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville et al.NeurIPS 2021 · 1,067 citations
- MOPO: Model-based Offline Policy OptimizationTianhe Yu, Garrett Thomas, Lantao Yu, Stefano Ermon et al.NeurIPS 2020 · 989 citations
- MOReL: Model-Based Offline Reinforcement LearningRahul Kidambi, Aravind Rajeswaran, Praneeth Netrapalli, Thorsten JoachimsNeurIPS 2020 · 870 citations
- Planning to Explore via Self-Supervised World ModelsRamanan Sekar, Oleh Rybkin, Kostas Daniilidis, Pieter Abbeel et al.ICML 2020 · 489 citations
Related papers
- Uncertainty-Based Smooth Policy Regularisation for Reinforcement Learning with Few DemonstrationsYujie Zhu, Charles A. Hepburn, Matthew Thorpe, Giovanni MontanaNeurIPS 2025
- Hybrid Policy Optimization from Imperfect DemonstrationsHanlin Yang, Chao Yu, Peng Sun, Siji ChenNeurIPS 2023 · 14 citations
- Reinforcement Learning from Imperfect Demonstrations under Soft Expert GuidanceMingxuan Jing, Xiaojian Ma, Wenbing Huang, Fuchun Sun et al.AAAI 2020 · 70 citations
- Multi-Stage Manipulation with Demonstration-Augmented Reward, Policy, and World Model LearningAdrià López Escoriza, Nicklas Hansen, Stone Tao, Tongzhou Mu et al.ICML 2025
- DUO: Diverse, Uncertain, On-Policy Query Generation and Selection for Reinforcement Learning from Human FeedbackXuening Feng, Zhaohui Jiang, Timo Kaufmann, Puchen Xu et al.AAAI 2025 · 7 citations
