When should agents explore?
Miruna Pislar, David Szepesvari, Georg Ostrovski, Diana L. Borsa, Tom Schaul
摘要
Exploration remains a central challenge for reinforcement learning (RL). Virtually all existing methods share the feature of a monolithic behaviour policy that changes only gradually (at best). In contrast, the exploratory behaviours of animals and humans exhibit a rich diversity, namely including forms of switching between modes. This paper presents an initial study of mode-switching, non-monolithic exploration for RL. We investigate different modes to switch between, at what timescales it makes sense to switch, and what signals make for good switching triggers. We also propose practical algorithmic components that make the switching mechanism adaptive and robust, which enables flexibility without an accompanying hyper-parameter-tuning burden. Finally, we report a promising and detailed analysis on Atari, using two-mode exploration and switching at sub-episodic time-scales.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Human-Timescale Adaptation in an Open-Ended Task SpaceJakob Bauer, Kate Baumli, Feryal M. P. Behbahani, Avishkar Bhoopchand 等ICML 2023 · 被引用 155 次
- Planning Goals for ExplorationEdward S. Hu, Richard Chang, Oleh Rybkin, Dinesh JayaramanICLR 2023 · 被引用 152 次
- The Phenomenon of Policy ChurnTom Schaul, André Barreto, John Quan, Georg OstrovskiNeurIPS 2022 · 被引用 38 次
- DEP-RL: Embodied Exploration for Reinforcement Learning in Overactuated and Musculoskeletal SystemsPierre Schumacher, Daniel F. B. Haeufle, Dieter Büchler, Syn Schmitt 等ICLR 2023 · 被引用 6 次
- Meta-Explore: Exploratory Hierarchical Vision-and-Language Navigation Using Scene Object Spectrum GroundingMinyoung Hwang, Jaeyeon Jeong, Minsoo Kim, Yoonseon Oh 等CVPR 2023
它引用的顶会 Paper3
- Exploration in Reinforcement Learning with Deep Covering OptionsYuu Jinnai, Jee Won Park, Marlos C. Machado, George Dimitri KonidarisICLR 2020 · 被引用 64 次
- Rank the Episodes: A Simple Approach for Exploration in Procedurally-Generated EnvironmentsDaochen Zha, Wenye Ma, Lei Yuan, Xia Hu 等ICLR 2021 · 被引用 47 次
- Temporally-Extended ε-Greedy ExplorationWill Dabney, Georg Ostrovski, André BarretoICLR 2021 · 被引用 2 次
相关 Paper
- LESSON: Learning to Integrate Exploration Strategies for Reinforcement Learning via an Option FrameworkWoojun Kim, Jeonghye Kim, Youngchul SungICML 2023 · 被引用 5 次
- Continuously Discovering Novel Strategies via Reward-Switching Policy OptimizationZihan Zhou, Wei Fu, Bingliang Zhang, Yi WuICLR 2022 · 被引用 34 次
- Discovering Policies with DOMiNO: Diversity Optimization Maintaining Near OptimalityTom Zahavy, Yannick Schroecker, Feryal M. P. Behbahani, Kate Baumli 等ICLR 2023 · 被引用 2 次
- Near-Optimal Adversarial Reinforcement Learning with Switching CostsMing Shi, Yingbin Liang, Ness B. ShroffICLR 2023
- Reinforcement Learning in Reward-Mixing MDPsJeongyeol Kwon, Yonathan Efroni, Constantine Caramanis, Shie MannorNeurIPS 2021 · 被引用 23 次
