Efficient Exploration via Epistemic-Risk-Seeking Policy Optimization
Brendan O'Donoghue
Abstract
Exploration remains a key challenge in deep reinforcement learning (RL). Optimism in the face of uncertainty is a well-known heuristic with theoretical guarantees in the tabular setting, but how best to translate the principle to deep reinforcement learning, which involves online stochastic gradients and deep network function approximators, is not fully understood. In this paper we propose a new, differentiable optimistic objective that when optimized yields a policy that provably explores efficiently, with guarantees even under function approximation. Our new objective is a zero-sum two-player game derived from endowing the agent with an epistemic-risk-seeking utility function, which converts uncertainty into value and encourages the agent to explore uncertain states. We show that the solution to this game minimizes an upper bound on the regret, with the 'players' each attempting to minimize one component of a particular regret decomposition. We derive a new model-free algorithm which we call 'epistemic-risk-seeking actor-critic' (ERSAC), which is simply an application of simultaneous stochastic gradient ascent-descent to the game. Finally, we discuss a recipe for incorporating off-policy data and show that combining the risk-seeking objective with replay data yields a double benefit in terms of statistical efficiency. We conclude with some results showing good performance of a deep RL agent using the technique on the challenging 'DeepSea' environment, showing significant performance improvements even over other efficient exploration techniques, as well as improved performance on the Atari benchmark.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f79e484b-cdea-42b0-b4da-09f58f5d1962Cited by top-tier papers5
- Confronting Reward Model Overoptimization with Constrained RLHFTed Moskovitz, Aaditya K. Singh, DJ Strouse, Tuomas Sandholm et al.ICLR 2024 · 89 citations
- Probabilistic Inference in Reinforcement Learning Done RightJean Tarbouriech, Tor Lattimore, Brendan O'DonoghueNeurIPS 2023 · 15 citations
- Policy Mirror Descent with LookaheadKimon Protopapas, Anas BarakatNeurIPS 2024 · 7 citations
- IL-SOAR : Imitation Learning with Soft Optimistic Actor cRiticStefano Viel, Luca Viano, Volkan CevherICML 2025
- Risk-Sensitive Variational Actor-Critic: A Model-Based ApproachAlonso Granados Baca, Reza Ebrahimi, Jason PachecoICLR 2025
Builds on7
- On the Global Convergence Rates of Softmax Policy Gradient MethodsJincheng Mei, Chenjun Xiao, Csaba Szepesvári, Dale SchuurmansICML 2020 · 349 citations
- Behaviour Suite for Reinforcement LearningIan Osband, Yotam Doron, Matteo Hessel, John Aslanides et al.ICLR 2020 · 204 citations
- Sample Efficient Reinforcement Learning with REINFORCEJunzi Zhang, Jongho Kim, Brendan O'Donoghue, Stephen P. BoydAAAI 2021 · 162 citations
- Variational Bayesian Reinforcement Learning with Regret BoundsBrendan O'DonoghueNeurIPS 2021 · 48 citations
- The Neural Testbed: Evaluating Joint PredictionsIan Osband, Zheng Wen, Seyed Mohammad Asghari, Vikranth Dwaracherla et al.NeurIPS 2022 · 28 citations
Related papers
- Principled Exploration via Optimistic Bootstrapping and Backward InductionChenjia Bai, Lingxiao Wang, Lei Han, Jianye Hao et al.ICML 2021 · 46 citations
- Decoupled Policy Actor-Critic: Bridging Pessimism and Risk Awareness in Reinforcement LearningMichal Nauman, Marek CyganAAAI 2025 · 5 citations
- Risk-Sensitive Exponential Actor CriticAlonso Granados Baca, Jason PachecoAAAI 2026
- Adversarially Trained Actor Critic for Offline Reinforcement LearningChing-An Cheng, Tengyang Xie, Nan Jiang, Alekh AgarwalICML 2022 · 156 citations
- Risk-Sensitive Reinforcement Learning: Near-Optimal Risk-Sample Tradeoff in RegretYingjie Fei, Zhuoran Yang, Yudong Chen, Zhaoran Wang et al.NeurIPS 2020 · 87 citations
