Approximate Bilevel Difference Convex Programming for Bayesian Risk Markov Decision Processes
Yifan Lin, Enlu Zhou
Abstract
We consider infinite-horizon Markov Decision Processes where parameters, such as transition probabilities, are unknown and estimated from data. The popular distributionally robust approach to addressing the parameter uncertainty can sometimes be overly conservative. In this paper, we utilize the recently proposed formulation, Bayesian risk Markov Decision Process (BR-MDP), to address parameter (or epistemic) uncertainty in MDPs. To solve the infinite-horizon BR-MDP with a class of convex risk measures, we propose a computationally efficient approach called approximate bilevel difference convex programming (ABDCP). The optimization is performed offline and produces the optimal policy that is represented as a finite state controller with desirable performance guarantees. We also demonstrate the empirical performance of the BR-MDP formulation and the proposed algorithm.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 009a4bab-a6b1-42e9-97b5-ad04d291261bBuilds on5
- Policy Gradient Method For Robust Reinforcement LearningYue Wang, Shaofeng ZouICML 2022 · 104 citations
- Risk-Averse Bayes-Adaptive Reinforcement LearningMarc Rigter, Bruno Lacerda, Nick HawesNeurIPS 2021 · 50 citations
- Constrained Risk-Averse Markov Decision ProcessesMohamadreza Ahmadi, Ugo Rosolia, Michel D. Ingham, Richard M. Murray et al.AAAI 2021 · 31 citations
- Bayesian Risk Markov Decision ProcessesYifan Lin, Yuxuan Ren, Enlu ZhouNeurIPS 2022 · 18 citations
- Percentile Criterion Optimization in Offline Reinforcement LearningCyrus Cousins, Elita A. Lobo, Marek Petrik, Yair ZickNeurIPS 2023 · 5 citations
Related papers
- Bayesian Risk-Averse Q-Learning with Streaming ObservationsYuhao Wang, Enlu ZhouNeurIPS 2023 · 8 citations
- Efficient Solution and Learning of Robust Factored MDPsYannik Schnitzer, Alessandro Abate, David ParkerAAAI 2026 · 1 citation
- Dynamic Programming for Epistemic Uncertainty in Markov Decision ProcessesAxel Benyamine, Julien Grand-Clément, Marek Petrik, Michael Jordan et al.ICML 2026 · 1 citation
- Solving Robust Markov Decision Processes: Generic, Reliable, EfficientTobias Meggendorfer, Maximilian Weininger, Patrick WienhöftAAAI 2025
- Robust Anytime Learning of Markov Decision ProcessesMarnix Suilen, Thiago D. Simão, David Parker, Nils JansenNeurIPS 2022 · 31 citations
