Variational Bayesian Reinforcement Learning with Regret Bounds
Brendan O'Donoghue
Abstract
We consider the exploration-exploitation trade-off in reinforcement learning and we show that an agent imbued with an epistemic-risk-seeking utility function is able to explore efficiently, as measured by regret. The parameter that controls how risk-seeking the agent is can be optimized to minimize regret, or annealed according to a schedule. We call the resulting algorithm K-learning and we show that the K-values that the agent maintains are optimistic for the expected optimal Q-values at each state-action pair. The utility function approach induces a natural Boltzmann exploration policy for which the 'temperature' parameter is equal to the risk-seeking parameter. This policy achieves a Bayesian regret bound of , where L is the time horizon, S is the number of states, A is the number of actions, and T is the total number of elapsed time-steps. K-learning can be interpreted as mirror descent in the policy space, and it is similar to other well-known methods in the literature, including Q-learning, soft-Q-learning, and maximum entropy policy gradient. K-learning is simple to implement, as it only requires adding a bonus to the reward at each state-action and then solving a Bellman equation. We conclude with a numerical example demonstrating that K-learning is competitive with other state-of-the-art algorithms in practice.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers19
- Conservative Q-Learning for Offline Reinforcement LearningAviral Kumar, Aurick Zhou, George Tucker, Sergey LevineNeurIPS 2020 · 2,881 citations
- Reward is enough for convex MDPsTom Zahavy, Brendan O'Donoghue, Guillaume Desjardins, Satinder SinghNeurIPS 2021 · 96 citations
- Making Sense of Reinforcement Learning and Probabilistic InferenceBrendan O'Donoghue, Ian Osband, Catalin IonescuICLR 2020 · 54 citations
- MURAL: Meta-Learning Uncertainty-Aware Rewards for Outcome-Driven Reinforcement LearningKevin Li, Abhishek Gupta, Ashwin Reddy, Vitchyr H. Pong et al.ICML 2021 · 36 citations
- ReLOAD: Reinforcement Learning with Optimistic Ascent-Descent for Last-Iterate Convergence in Constrained MDPsTed Moskovitz, Brendan O'Donoghue, Vivek Veeriah, Sebastian Flennerhag et al.ICML 2023 · 24 citations
Builds on3
- Almost Optimal Model-Free Reinforcement Learningvia Reference-Advantage DecompositionZihan Zhang, Yuan Zhou, Xiangyang JiNeurIPS 2020 · 183 citations
- Sample Complexity of Asynchronous Q-Learning: Sharper Analysis and Variance ReductionGen Li, Yuting Wei, Yuejie Chi, Yuantao Gu et al.NeurIPS 2020 · 149 citations
- Making Sense of Reinforcement Learning and Probabilistic InferenceBrendan O'Donoghue, Ian Osband, Catalin IonescuICLR 2020 · 54 citations
Related papers
- Risk-Sensitive Reinforcement Learning: Near-Optimal Risk-Sample Tradeoff in RegretYingjie Fei, Zhuoran Yang, Yudong Chen, Zhaoran Wang et al.NeurIPS 2020 · 87 citations
- EUBRL: Epistemic Uncertainty Directed Bayesian Reinforcement LearningJianfei Ma, Wee Sun LeeICLR 2026 · 2 citations
- Efficient Exploration via Epistemic-Risk-Seeking Policy OptimizationBrendan O'DonoghueICML 2023 · 11 citations
- MaxInfoRL: Boosting exploration in reinforcement learning through information gain maximizationBhavya Sukhija, Stelian Coros, Andreas Krause, Pieter Abbeel et al.ICLR 2025
- Minimax Optimal Reinforcement Learning with Quasi-OptimismHarin Lee, Min-hwan OhICLR 2025
