Temporally-Extended ε-Greedy Exploration
Will Dabney, Georg Ostrovski, André Barreto
摘要
Recent work on exploration in reinforcement learning (RL) has led to a series of increasingly complex solutions to the problem. This increase in complexity often comes at the expense of generality. Recent empirical studies suggest that, when applied to a broader set of domains, some sophisticated exploration methods are outperformed by simpler counterparts, such as ε-greedy. In this paper we propose an exploration algorithm that retains the simplicity of ε-greedy while reducing dithering. We build on a simple hypothesis: the main limitation of ε-greedy exploration is its lack of temporal persistence, which limits its ability to escape local optima. We propose a temporally extended form of ε-greedy that simply repeats the sampled action for a random duration. It turns out that, for many duration distributions, this suffices to improve exploration on a large set of domains. Interestingly, a class of distributions inspired by ecological models of animal foraging behaviour yields particularly strong performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- Motif: Intrinsic Motivation from Artificial Intelligence FeedbackMartin Klissarov, Pierluca D'Oro, Shagun Sodhani, Roberta Raileanu 等ICLR 2024 · 被引用 97 次
- On the Importance of Exploration for Generalization in Reinforcement LearningYiding Jiang, J. Zico Kolter, Roberta RaileanuNeurIPS 2023 · 被引用 48 次
- TempoRL: Learning When to ActAndré Biedenkapp, Raghu Rajan, Frank Hutter, Marius LindauerICML 2021 · 被引用 38 次
- The Phenomenon of Policy ChurnTom Schaul, André Barreto, John Quan, Georg OstrovskiNeurIPS 2022 · 被引用 38 次
- Deep Reinforcement Learning Policies Learn Shared Adversarial Features across MDPsEzgi KorkmazAAAI 2022 · 被引用 33 次
它引用的顶会 Paper4
- Never Give Up: Learning Directed Exploration StrategiesAdrià Puigdomènech Badia, Pablo Sprechmann, Alex Vitvitskyi, Zhaohan Daniel Guo 等ICLR 2020 · 被引用 349 次
- Fast Task Inference with Variational Intrinsic Successor FeaturesSteven Hansen, Will Dabney, André Barreto, David Warde-Farley 等ICLR 2020 · 被引用 176 次
- On Bonus Based Exploration Methods In The Arcade Learning EnvironmentAdrien Ali Taïga, William Fedus, Marlos C. Machado, Aaron C. Courville 等ICLR 2020 · 被引用 72 次
- Exploration in Reinforcement Learning with Deep Covering OptionsYuu Jinnai, Jee Won Park, Marlos C. Machado, George Dimitri KonidarisICLR 2020 · 被引用 64 次
相关 Paper
- Learning Uncertainty-Aware Temporally-Extended ActionsJoongkyu Lee, Seung Joon Park, Yunhao Tang, Min-hwan OhAAAI 2024 · 被引用 3 次
- Locally Persistent Exploration in Continuous Control Tasks with Sparse RewardsSusan Amin, Maziar Gomrokchi, Hossein Aboutalebi, Harsh Satija 等ICML 2021 · 被引用 17 次
- Brain Bandit: A Biologically Grounded Neural Network for Efficient Control of ExplorationChen Jiang, Jiahui An, Yating Liu, Ni JiICLR 2025
- Guarantees for Epsilon-Greedy Reinforcement Learning with Function ApproximationChristoph Dann, Yishay Mansour, Mehryar Mohri, Ayush Sekhari 等ICML 2022 · 被引用 76 次
- Control Frequency Adaptation via Action Persistence in Batch Reinforcement LearningAlberto Maria Metelli, Flavio Mazzolini, Lorenzo Bisi, Luca Sabbioni 等ICML 2020 · 被引用 43 次
