Flipping Coins to Estimate Pseudocounts for Exploration in Reinforcement Learning
Sam Lobel, Akhil Bagaria, George Konidaris
Abstract
We propose a new method for count-based exploration in high-dimensional state spaces. Unlike previous work which relies on density models, we show that counts can be derived by averaging samples from the Rademacher distribution (or coin flips). This insight is used to set up a simple supervised learning objective which, when optimized, yields a state's visitation count. We show that our method is significantly more effective at deducing ground-truth visitation counts than previous work; when used as an exploration bonus for a model-free reinforcement learning algorithm, it outperforms existing approaches on most of 9 challenging exploration tasks, including the Atari game Montezuma's Revenge.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2ac7b2ac-0696-4829-bfa4-e38c6242e15aCited by top-tier papers13
- Exploration and Anti-Exploration with Distributional Random Network DistillationKai Yang, Jian Tao, Jiafei Lyu, Xiu LiICML 2024 · 37 citations
- Count Counts: Motivating Exploration in LLM Reasoning with Count-based Intrinsic RewardsXuan Zhang, Ruixiao Li, Zhijian Zhou, Long Li et al.ICLR 2026 · 11 citations
- Beyond Optimism: Exploration With Partially Observable RewardsSimone Parisi, Alireza Kazemipour, Michael BowlingNeurIPS 2024 · 7 citations
- Emergence of Exploration in Policy Gradient Reinforcement Learning via RetryingSoichiro Nishimori, Paavo Parmas, Sotetsu Koyamada, Tadashi Kozuno et al.ICML 2026 · 6 citations
- Improving Intrinsic Exploration by Creating Stationary ObjectivesRoger Creus Castanyer, Joshua Romoff, Glen BersethICLR 2024 · 4 citations
Builds on15
- Agent57: Outperforming the Atari Human BenchmarkAdrià Puigdomènech Badia, Bilal Piot, Steven Kapturowski, Pablo Sprechmann et al.ICML 2020 · 584 citations
- Never Give Up: Learning Directed Exploration StrategiesAdrià Puigdomènech Badia, Pablo Sprechmann, Alex Vitvitskyi, Zhaohan Daniel Guo et al.ICLR 2020 · 349 citations
- Skew-Fit: State-Covering Self-Supervised Reinforcement LearningVitchyr Pong, Murtaza Dalal, Steven Lin, Ashvin Nair et al.ICML 2020 · 303 citations
- Count-Based Exploration with the Successor RepresentationMarlos C. Machado, Marc G. Bellemare, Michael BowlingAAAI 2020 · 206 citations
- RIDE: Rewarding Impact-Driven Exploration for Procedurally-Generated EnvironmentsRoberta Raileanu, Tim RocktäschelICLR 2020 · 198 citations
Related papers
- On Bonus Based Exploration Methods In The Arcade Learning EnvironmentAdrien Ali Taïga, William Fedus, Marlos C. Machado, Aaron C. Courville et al.ICLR 2020 · 72 citations
- Random Latent Exploration for Deep Reinforcement LearningSrinath Mahankali, Zhang-Wei Hong, Ayush Sekhari, Alexander Rakhlin et al.ICML 2024 · 8 citations
- Cell-Free Latent Go-ExploreQuentin Gallouédec, Emmanuel DellandréaICML 2023 · 4 citations
- Redeeming intrinsic rewards via constrained optimizationEric Chen, Zhang-Wei Hong, Joni Pajarinen, Pulkit AgrawalNeurIPS 2022 · 48 citations
- Just Cluster It: An Approach for Exploration in High-Dimensions using Clustering and Pre-Trained RepresentationsStefan Sylvius Wagner, Stefan HarmelingICML 2024 · 2 citations
