Emergence of Exploration in Policy Gradient Reinforcement Learning via Retrying
Soichiro Nishimori, Paavo Parmas, Sotetsu Koyamada, Tadashi Kozuno, Toshinori Kitamura, Shin Ishii, Yutaka Matsuo
摘要
In reinforcement learning (RL), agents benefit from exploration only because they repeatedly encounter similar states: trying different actions can improve performance or reduce uncertainty; without such retries, a greedy policy is optimal. We formalize this intuition with ReMax, an objective that evaluates a policy by the expected maximum return over samples (), while accounting for return uncertainty. Optimizing this objective induces stochastic exploration as an emergent property, without explicit bonus terms. For efficient policy optimization, we derive a new policy-gradient formulation for ReMax and introduce ReMax PPO (RePPO), a PPO variant that optimizes ReMax while generalizing the discrete retry count to a continuous parameter , enabling fine-grained control of exploration. Empirically, RePPO promotes exploration—without any explicit exploration bonuses—on the MinAtar and Craftax benchmarks. The official code is available at https://github.com/nissymori/remax-rl.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper26
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 被引用 1,852 次
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville 等NeurIPS 2021 · 被引用 1,067 次
- Uncertainty-Based Offline Reinforcement Learning with Diversified Q-EnsembleGaon An, Seungyong Moon, Jang-Hyun Kim, Hyun Oh SongNeurIPS 2021 · 被引用 430 次
- Never Give Up: Learning Directed Exploration StrategiesAdrià Puigdomènech Badia, Pablo Sprechmann, Alex Vitvitskyi, Zhaohan Daniel Guo 等ICLR 2020 · 被引用 349 次
- The NetHack Learning EnvironmentHeinrich Küttler, Nantas Nardelli, Alexander H. Miller, Roberta Raileanu 等NeurIPS 2020 · 被引用 251 次
相关 Paper
- Ready Policy One: World Building Through Active LearningPhilip J. Ball, Jack Parker-Holder, Aldo Pacchiano, Krzysztof Choromanski 等ICML 2020 · 被引用 52 次
- OVD-Explorer: Optimism Should Not Be the Sole Pursuit of Exploration in Noisy EnvironmentsJinyi Liu, Zhi Wang, Yan Zheng, Jianye Hao 等AAAI 2024 · 被引用 14 次
- SHAPO: Sharpness-Aware Policy Optimization for Safe ExplorationKaustubh Mani, Yann Pequignot, Vincent Mai, Liam PaullICLR 2026 · 被引用 3 次
- Redeeming intrinsic rewards via constrained optimizationEric Chen, Zhang-Wei Hong, Joni Pajarinen, Pulkit AgrawalNeurIPS 2022 · 被引用 48 次
- Task-Agnostic Exploration via Policy Gradient of a Non-Parametric State Entropy EstimateMirco Mutti, Lorenzo Pratissoli, Marcello RestelliAAAI 2021 · 被引用 62 次
