Making Efficient Use of Demonstrations to Solve Hard Exploration Problems
Çaglar Gülçehre, Tom Le Paine, Bobak Shahriari, Misha Denil, Matt Hoffman, Hubert Soyer, Richard Tanburn, Steven Kapturowski, Neil C. Rabinowitz, Duncan Williams, Gabriel Barth-Maron, Ziyu Wang
2020年份
97被引次数
17顶会引用
摘要
This paper introduces R2D3, an agent that makes efficient use of demonstrations to solve hard exploration problems in partially observable environments with highly variable initial conditions. We also introduce a suite of eight tasks that combine these three properties, and show that R2D3 can solve several of the tasks where other state of the art methods (both with and without demonstrations) fail to see even a single successful trajectory after tens of billions of steps of exploration.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- A Minimalist Approach to Offline Reinforcement LearningScott Fujimoto, Shixiang Shane GuNeurIPS 2021 · 被引用 1,292 次
- What Matters for Adversarial Imitation Learning?Manu Orsini, Anton Raichuk, Léonard Hussenot, Damien Vincent 等NeurIPS 2021 · 被引用 106 次
- Hit and Lead Discovery with Explorative RL and Fragment-based Molecule GenerationSoojung Yang, Doyeong Hwang, Seul Lee, Seongok Ryu 等NeurIPS 2021 · 被引用 106 次
- BYOL-Explore: Exploration by Bootstrapped PredictionZhaohan Guo, Shantanu Thakoor, Miruna Pislar, Bernardo Ávila Pires 等NeurIPS 2022 · 被引用 104 次
- Reincarnating Reinforcement Learning: Reusing Prior Computation to Accelerate ProgressRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville 等NeurIPS 2022 · 被引用 95 次
相关 Paper
- Guided Exploration with Proximal Policy Optimization using a Single DemonstrationGabriele Libardi, Gianni De Fabritiis, Sebastian DittertICML 2021 · 被引用 32 次
- Variational Recurrent Models for Solving Partially Observable Control TasksDongqi Han, Kenji Doya, Jun TaniICLR 2020 · 被引用 75 次
- MAMBA: an Effective World Model Approach for Meta-Reinforcement LearningZohar Rimon, Tom Jurgenson, Orr Krupnik, Gilad Adler 等ICLR 2024 · 被引用 15 次
- Multi-Stage Manipulation with Demonstration-Augmented Reward, Policy, and World Model LearningAdrià López Escoriza, Nicklas Hansen, Stone Tao, Tongzhou Mu 等ICML 2025
- DISCOVER: Automated Curricula for Sparse-Reward Reinforcement LearningLeander Diaz-Bone, Marco Bagatella, Jonas Hübotter, Andreas KrauseNeurIPS 2025 · 被引用 14 次
