Guided Exploration with Proximal Policy Optimization using a Single Demonstration
Gabriele Libardi, Gianni De Fabritiis, Sebastian Dittert
摘要
Solving sparse reward tasks through exploration is one of the major challenges in deep reinforcement learning, especially in three-dimensional, partially-observable environments. Critically, the algorithm proposed in this article is capable of using a single human demonstration to solve hardexploration problems. We train an agent on a combination of demonstrations and own experience to solve problems with variable initial conditions and we integrate it with proximal policy optimization (PPO). The agent is also able to increase its performance and to tackle harder problems by replaying its own past trajectories prioritizing them based on the obtained reward and the maximum value of the trajectory. We finally compare variations of this algorithm to different imitation learning algorithms on a set of hard-exploration tasks in the Animal-AI Olympics environment. To the best of our knowledge, learning a task in a three-dimensional environment with comparable difficulty has never been considered before using only one human demonstration.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Reinforcement Learning with Sparse Rewards using Guidance from Offline DemonstrationDesik Rengarajan, Gargi Vaidya, Akshay Sarvesh, Dileep M. Kalathil 等ICLR 2022 · 被引用 86 次
- Training a Scientific Reasoning Model for ChemistrySiddharth Narayanan, James D. Braza, Ryan-Rhys Griffiths, Albert Bou 等NeurIPS 2025 · 被引用 62 次
它引用的顶会 Paper3
- Emergent Tool Use From Multi-Agent AutocurriculaBowen Baker, Ingmar Kanitscheider, Todor M. Markov, Yi Wu 等ICLR 2020 · 被引用 751 次
- Never Give Up: Learning Directed Exploration StrategiesAdrià Puigdomènech Badia, Pablo Sprechmann, Alex Vitvitskyi, Zhaohan Daniel Guo 等ICLR 2020 · 被引用 349 次
- Making Efficient Use of Demonstrations to Solve Hard Exploration ProblemsÇaglar Gülçehre, Tom Le Paine, Bobak Shahriari, Misha Denil 等ICLR 2020 · 被引用 97 次
相关 Paper
- Proximal Policy Gradient Arborescence for Quality Diversity Reinforcement LearningSumeet Batra, Bryon Tjanaka, Matthew Christopher Fontaine, Aleksei Petrenko 等ICLR 2024 · 被引用 26 次
- Discovering Hierarchical Achievements in Reinforcement Learning via Contrastive LearningSeungyong Moon, Junyoung Yeom, Bumsoo Park, Hyun Oh SongNeurIPS 2023 · 被引用 12 次
- Accelerating Exploration with Unlabeled Prior DataQiyang Li, Jason Zhang, Dibya Ghosh, Amy Zhang 等NeurIPS 2023 · 被引用 21 次
- Expert Proximity as Surrogate Rewards for Single Demonstration Imitation LearningChia-Cheng Chiang, Li-Cheng Lan, Wei-Fang Sun, Chien Feng 等ICML 2024
- Learning to Reach Goals via Iterated Supervised LearningDibya Ghosh, Abhishek Gupta, Ashwin Reddy, Justin Fu 等ICLR 2021 · 被引用 222 次
