Risk-Averse Bayes-Adaptive Reinforcement Learning
Marc Rigter, Bruno Lacerda, Nick Hawes
Abstract
In this work, we address risk-averse Bayes-adaptive reinforcement learning. We pose the problem of optimising the conditional value at risk (CVaR) of the total return in Bayes-adaptive Markov decision processes (MDPs). We show that a policy optimising CVaR in this setting is risk-averse to both the parametric uncertainty due to the prior distribution over MDPs, and the internal uncertainty due to the inherent stochasticity of MDPs. We reformulate the problem as a two-player stochastic game and propose an approximate algorithm based on Monte Carlo tree search and Bayesian optimisation. Our experiments demonstrate that our approach significantly outperforms baseline approaches for this problem.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 147e607c-bf45-451c-aa1b-4ee1c98e8fb9Cited by top-tier papers14
- Model-Based Offline Reinforcement Learning with Pessimism-Modulated Dynamics BeliefKaiyang Guo, Yunfeng Shao, Yanhui GengNeurIPS 2022 · 39 citations
- One Risk to Rule Them All: A Risk-Sensitive Perspective on Model-Based Offline Reinforcement LearningMarc Rigter, Bruno Lacerda, Nick HawesNeurIPS 2023 · 26 citations
- On Dynamic Programming Decompositions of Static Risk Measures in Markov Decision ProcessesJia Lin Hau, Erick Delage, Mohammad Ghavamzadeh, Marek PetrikNeurIPS 2023 · 22 citations
- Bayesian Risk Markov Decision ProcessesYifan Lin, Yuxuan Ren, Enlu ZhouNeurIPS 2022 · 18 citations
- Two steps to risk sensitivityChris Gagne, Peter DayanNeurIPS 2021 · 17 citations
Builds on2
Related papers
- Optimizing Conditional Value-At-Risk of Black-Box FunctionsQuoc Phong Nguyen, Zhongxiang Dai, Bryan Kian Hsiang Low, Patrick JailletNeurIPS 2021 · 25 citations
- Being Optimistic to Be Conservative: Quickly Learning a CVaR PolicyRamtin Keramati, Christoph Dann, Alex Tamkin, Emma BrunskillAAAI 2020 · 86 citations
- Risk-Aware Stochastic Shortest PathTobias MeggendorferAAAI 2022 · 13 citations
- Risk-Averse No-Regret Learning in Online Convex GamesZifan Wang, Yi Shen, Michael M. ZavlanosICML 2022 · 10 citations
- Boosting CVaR Policy Optimization with Quantile GradientsYudong Luo, Erick DelageICML 2026
