Making Efficient Use of Demonstrations to Solve Hard Exploration Problems
Çaglar Gülçehre, Tom Le Paine, Bobak Shahriari, Misha Denil, Matt Hoffman, Hubert Soyer, Richard Tanburn, Steven Kapturowski, Neil C. Rabinowitz, Duncan Williams, Gabriel Barth-Maron, Ziyu Wang
Abstract
This paper introduces R2D3, an agent that makes efficient use of demonstrations to solve hard exploration problems in partially observable environments with highly variable initial conditions. We also introduce a suite of eight tasks that combine these three properties, and show that R2D3 can solve several of the tasks where other state of the art methods (both with and without demonstrations) fail to see even a single successful trajectory after tens of billions of steps of exploration.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b603c3ec-a5ad-41c1-807f-0d719d1ce542Cited by top-tier papers17
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 1,292 citations
- What Matters for Adversarial Imitation Learning?Manu Orsini, Anton Raichuk, Léonard Hussenot, Damien Vincent et al.NeurIPS 2021 · 106 citations
- Hit and Lead Discovery with Explorative RL and Fragment-based Molecule GenerationSoojung Yang, Doyeong Hwang, Seul Lee, Seongok Ryu et al.NeurIPS 2021 · 106 citations
- BYOL-Explore: Exploration by Bootstrapped PredictionZhaohan Guo, Shantanu Thakoor, Miruna Pislar, Bernardo Ávila Pires et al.NeurIPS 2022 · 104 citations
- Reincarnating Reinforcement Learning: Reusing Prior Computation to Accelerate ProgressRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville et al.NeurIPS 2022 · 95 citations
Related papers
- Guided Exploration with Proximal Policy Optimization using a Single DemonstrationGabriele Libardi, Gianni De Fabritiis, Sebastian DittertICML 2021 · 32 citations
- Variational Recurrent Models for Solving Partially Observable Control TasksDongqi Han, Kenji Doya, Jun TaniICLR 2020 · 75 citations
- MAMBA: an Effective World Model Approach for Meta-Reinforcement LearningZohar Rimon, Tom Jurgenson, Orr Krupnik, Gilad Adler et al.ICLR 2024 · 15 citations
- Multi-Stage Manipulation with Demonstration-Augmented Reward, Policy, and World Model LearningAdrià López Escoriza, Nicklas Hansen, Stone Tao, Tongzhou Mu et al.ICML 2025
- DISCOVER: Automated Curricula for Sparse-Reward Reinforcement LearningLeander Diaz-Bone, Marco Bagatella, Jonas Hübotter, Andreas KrauseNeurIPS 2025 · 14 citations
