Bayesian Risk-Averse Q-Learning with Streaming Observations
Yuhao Wang, Enlu Zhou
Abstract
We consider a robust reinforcement learning problem, where a learning agent learns from a simulated training environment. To account for the model mis-specification between this training environment and the real environment due to lack of data, we adopt a formulation of Bayesian risk MDP (BRMDP) with infinite horizon, which uses Bayesian posterior to estimate the transition model and impose a risk functional to account for the model uncertainty. Observations from the real environment that is out of the agent's control arrive periodically and are utilized by the agent to update the Bayesian posterior to reduce model uncertainty. We theoretically demonstrate that BRMDP balances the trade-off between robustness and conservativeness, and we further develop a multi-stage Bayesian risk-averse Q-learning algorithm to solve BRMDP with streaming observations from real environment. The proposed algorithm learns a risk-averse yet optimal policy that depends on the availability of real-world observations. We provide a theoretical guarantee of strong convergence for the proposed algorithm. Keywords Q-learning • risk-averse reinforcement learning • off-policy learning • Bayesian risk Markov decision process • distributionally robust Markov decision process
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 501da7fb-80ae-4d01-9576-0bf72c954cf2Cited by top-tier papers2
- Uncertainty-based Offline Variational Bayesian Reinforcement Learning for Robustness under Diverse Data CorruptionsRui Yang, Jie Wang, Guoping Wu, Bin LiNeurIPS 2024 · 11 citations
- Resolving the Stability-Plasticity Dilemma in Reinforcement Learning via Complementary Continual CriticsBo Sun, Peixi Peng, Guang Tan, Haoran Xu et al.CVPR 2026
Builds on4
- Distributionally Robust Q-LearningZijian Liu, Qinxun Bai, Jose H. Blanchet, Perry Dong et al.ICML 2022 · 72 citations
- Distributionally Robust Policy Evaluation and Learning in Offline Contextual BanditsNian Si, Fan Zhang, Zhengyuan Zhou, Jose H. BlanchetICML 2020 · 59 citations
- Risk-Averse Bayes-Adaptive Reinforcement LearningMarc Rigter, Bruno Lacerda, Nick HawesNeurIPS 2021 · 50 citations
- Bayesian Risk Markov Decision ProcessesYifan Lin, Yuxuan Ren, Enlu ZhouNeurIPS 2022 · 18 citations
Related papers
- Approximate Bilevel Difference Convex Programming for Bayesian Risk Markov Decision ProcessesYifan Lin, Enlu ZhouAAAI 2025 · 1 citation
- Online Robust Reinforcement Learning with Model UncertaintyYue Wang, Shaofeng ZouNeurIPS 2021 · 157 citations
- Robust Anytime Learning of Markov Decision ProcessesMarnix Suilen, Thiago D. Simão, David Parker, Nils JansenNeurIPS 2022 · 31 citations
- Robust Reinforcement Learning using Offline DataKishan Panaganti, Zaiyan Xu, Dileep Kalathil, Mohammad GhavamzadehNeurIPS 2022 · 130 citations
- Risk-Averse Offline Reinforcement LearningNúria Armengol Urpí, Sebastian Curi, Andreas KrauseICLR 2021 · 81 citations
