Guided Exploration with Proximal Policy Optimization using a Single Demonstration
Gabriele Libardi, Gianni De Fabritiis, Sebastian Dittert
Abstract
Solving sparse reward tasks through exploration is one of the major challenges in deep reinforcement learning, especially in three-dimensional, partially-observable environments. Critically, the algorithm proposed in this article is capable of using a single human demonstration to solve hardexploration problems. We train an agent on a combination of demonstrations and own experience to solve problems with variable initial conditions and we integrate it with proximal policy optimization (PPO). The agent is also able to increase its performance and to tackle harder problems by replaying its own past trajectories prioritizing them based on the obtained reward and the maximum value of the trajectory. We finally compare variations of this algorithm to different imitation learning algorithms on a set of hard-exploration tasks in the Animal-AI Olympics environment. To the best of our knowledge, learning a task in a three-dimensional environment with comparable difficulty has never been considered before using only one human demonstration.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Reinforcement Learning with Sparse Rewards using Guidance from Offline DemonstrationDesik Rengarajan, Gargi Vaidya, Akshay Sarvesh, Dileep M. Kalathil et al.ICLR 2022 · 86 citations
- Training a Scientific Reasoning Model for ChemistrySiddharth Narayanan, James D. Braza, Ryan-Rhys Griffiths, Albert Bou et al.NeurIPS 2025 · 62 citations
Builds on3
- Emergent Tool Use From Multi-Agent AutocurriculaBowen Baker, Ingmar Kanitscheider, Todor M. Markov, Yi Wu et al.ICLR 2020 · 751 citations
- Never Give Up: Learning Directed Exploration StrategiesAdrià Puigdomènech Badia, Pablo Sprechmann, Alex Vitvitskyi, Zhaohan Daniel Guo et al.ICLR 2020 · 349 citations
- Making Efficient Use of Demonstrations to Solve Hard Exploration ProblemsÇaglar Gülçehre, Tom Le Paine, Bobak Shahriari, Misha Denil et al.ICLR 2020 · 97 citations
Related papers
- Proximal Policy Gradient Arborescence for Quality Diversity Reinforcement LearningSumeet Batra, Bryon Tjanaka, Matthew Christopher Fontaine, Aleksei Petrenko et al.ICLR 2024 · 26 citations
- Discovering Hierarchical Achievements in Reinforcement Learning via Contrastive LearningSeungyong Moon, Junyoung Yeom, Bumsoo Park, Hyun Oh SongNeurIPS 2023 · 12 citations
- Accelerating Exploration with Unlabeled Prior DataQiyang Li, Jason Zhang, Dibya Ghosh, Amy Zhang et al.NeurIPS 2023 · 21 citations
- Expert Proximity as Surrogate Rewards for Single Demonstration Imitation LearningChia-Cheng Chiang, Li-Cheng Lan, Wei-Fang Sun, Chien Feng et al.ICML 2024
- Learning to Reach Goals via Iterated Supervised LearningDibya Ghosh, Abhishek Gupta, Ashwin Reddy, Justin Fu et al.ICLR 2021 · 222 citations
