Risk-Averse Bayes-Adaptive Reinforcement Learning
Marc Rigter, Bruno Lacerda, Nick Hawes
摘要
In this work, we address risk-averse Bayes-adaptive reinforcement learning. We pose the problem of optimising the conditional value at risk (CVaR) of the total return in Bayes-adaptive Markov decision processes (MDPs). We show that a policy optimising CVaR in this setting is risk-averse to both the parametric uncertainty due to the prior distribution over MDPs, and the internal uncertainty due to the inherent stochasticity of MDPs. We reformulate the problem as a two-player stochastic game and propose an approximate algorithm based on Monte Carlo tree search and Bayesian optimisation. Our experiments demonstrate that our approach significantly outperforms baseline approaches for this problem.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Model-Based Offline Reinforcement Learning with Pessimism-Modulated Dynamics BeliefKaiyang Guo, Yunfeng Shao, Yanhui GengNeurIPS 2022 · 被引用 39 次
- One Risk to Rule Them All: A Risk-Sensitive Perspective on Model-Based Offline Reinforcement LearningMarc Rigter, Bruno Lacerda, Nick HawesNeurIPS 2023 · 被引用 26 次
- On Dynamic Programming Decompositions of Static Risk Measures in Markov Decision ProcessesJia Lin Hau, Erick Delage, Mohammad Ghavamzadeh, Marek PetrikNeurIPS 2023 · 被引用 22 次
- Bayesian Risk Markov Decision ProcessesYifan Lin, Yuxuan Ren, Enlu ZhouNeurIPS 2022 · 被引用 18 次
- Two steps to risk sensitivityChris Gagne, Peter DayanNeurIPS 2021 · 被引用 17 次
它引用的顶会 Paper2
相关 Paper
- Optimizing Conditional Value-At-Risk of Black-Box FunctionsQuoc Phong Nguyen, Zhongxiang Dai, Bryan Kian Hsiang Low, Patrick JailletNeurIPS 2021 · 被引用 25 次
- Being Optimistic to Be Conservative: Quickly Learning a CVaR PolicyRamtin Keramati, Christoph Dann, Alex Tamkin, Emma BrunskillAAAI 2020 · 被引用 86 次
- Risk-Aware Stochastic Shortest PathTobias MeggendorferAAAI 2022 · 被引用 13 次
- Risk-Averse No-Regret Learning in Online Convex GamesZifan Wang, Yi Shen, Michael M. ZavlanosICML 2022 · 被引用 10 次
- Boosting CVaR Policy Optimization with Quantile GradientsYudong Luo, Erick DelageICML 2026
