Improving Intrinsic Exploration by Creating Stationary Objectives
Roger Creus Castanyer, Joshua Romoff, Glen Berseth
摘要
Exploration bonuses in reinforcement learning guide long-horizon exploration by defining custom intrinsic objectives. Several exploration objectives like countbased bonuses, pseudo-counts, and state-entropy maximization are non-stationary and hence are difficult to optimize for the agent. While this issue is generally known, it is usually omitted and solutions remain under-explored. The key contribution of our work lies in transforming the original non-stationary rewards into stationary rewards through an augmented state representation. For this purpose, we introduce the Stationary Objectives For Exploration (SOFE) framework. SOFE requires identifying sufficient statistics for different exploration bonuses and finding an efficient encoding of these statistics to use as input to a deep network. SOFE is based on proposing state augmentations that expand the state space but hold the promise of simplifying the optimization of the agent's objective. We show that SOFE improves the performance of several exploration objectives, including count-based bonuses, pseudo-counts, and state-entropy maximization. Moreover, SOFE outperforms prior methods that attempt to stabilize the optimization of intrinsic objectives. We demonstrate the efficacy of SOFE in hard-exploration problems, including sparse-reward tasks, pixel-based observations, 3D navigation, and procedurally generated environments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper19
- Habitat: A Platform for Embodied AI ResearchManolis Savva, Jitendra Malik, Devi Parikh, Dhruv Batra 等ICCV 2019 · 被引用 1,863 次
- Leveraging Procedural Generation to Benchmark Reinforcement LearningKarl Cobbe, Christopher Hesse, Jacob Hilton, John SchulmanICML 2020 · 被引用 685 次
- Reinforcement Learning with Prototypical RepresentationsDenis Yarats, Rob Fergus, Alessandro Lazaric, Lerrel PintoICML 2021 · 被引用 262 次
- Count-Based Exploration with the Successor RepresentationMarlos C. Machado, Marc G. Bellemare, Michael BowlingAAAI 2020 · 被引用 206 次
- Behaviour Suite for Reinforcement LearningIan Osband, Yotam Doron, Matteo Hessel, John Aslanides 等ICLR 2020 · 被引用 204 次
相关 Paper
- RIDE: Rewarding Impact-Driven Exploration for Procedurally-Generated EnvironmentsRoberta Raileanu, Tim RocktäschelICLR 2020 · 被引用 198 次
- Successor-Predecessor Intrinsic ExplorationChangmin Yu, Neil Burgess, Maneesh Sahani, Samuel J. GershmanNeurIPS 2023 · 被引用 12 次
- MADE: Exploration via Maximizing Deviation from Explored RegionsTianjun Zhang, Paria Rashidinejad, Jiantao Jiao, Yuandong Tian 等NeurIPS 2021 · 被引用 51 次
- Optimistic Exploration even with a Pessimistic InitialisationTabish Rashid, Bei Peng, Wendelin Boehmer, Shimon WhitesonICLR 2020 · 被引用 50 次
- Flipping Coins to Estimate Pseudocounts for Exploration in Reinforcement LearningSam Lobel, Akhil Bagaria, George KonidarisICML 2023 · 被引用 29 次
