Regret Bounds for Markov Decision Processes with Recursive Optimized Certainty Equivalents
Wenhao Xu, Xuefeng Gao, Xuedong He
Abstract
The optimized certainty equivalent (OCE) is a family of risk measures that cover important examples such as entropic risk, conditional value-at-risk and mean-variance models. In this paper, we propose a new episodic risk-sensitive reinforcement learning formulation based on tabular Markov decision processes with recursive OCEs. We design an efficient learning algorithm for this problem based on value iteration and upper confidence bound. We derive an upper bound on the regret of the proposed algorithm, and also establish a minimax lower bound. Our bounds show that the regret rate achieved by our proposed algorithm has optimal dependence on the number of episodes and the number of actions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f3c22e1b-44cd-4cd2-bfee-382fba1c3bd6Cited by top-tier papers7
- Conformal Risk Training: End-to-End Optimization of Conformal Risk ControlChristopher Yeh, Nicolas Christianson, Adam Wierman, Yisong YueNeurIPS 2025 · 14 citations
- Provably Efficient Iterated CVaR Reinforcement Learning with Function Approximation and Human FeedbackYu Chen, Yihan Du, Pihe Hu, Siwei Wang et al.ICLR 2024 · 12 citations
- Efficient and Sharp Off-Policy Evaluation in Robust Markov Decision ProcessesAndrew Bennett, Nathan Kallus, Miruna Oprescu, Wen Sun et al.NeurIPS 2024 · 7 citations
- Provable Risk-Sensitive Distributional Reinforcement Learning with General Function ApproximationYu Chen, Xiangcheng Zhang, Siwei Wang, Longbo HuangICML 2024 · 3 citations
- Risk-Averse Constrained Reinforcement Learning with Optimized Certainty EquivalentsJane H. Lee, Baturay Saglam, Spyridon Pougkakiotis, Amin Karbasi et al.NeurIPS 2025 · 2 citations
Builds on6
- Risk-Sensitive Reinforcement Learning: Near-Optimal Risk-Sample Tradeoff in RegretYingjie Fei, Zhuoran Yang, Yudong Chen, Zhaoran Wang et al.NeurIPS 2020 · 87 citations
- Exponential Bellman Equation and Improved Regret Bounds for Risk-Sensitive Reinforcement LearningYingjie Fei, Zhuoran Yang, Yudong Chen, Zhaoran WangNeurIPS 2021 · 70 citations
- Risk-Sensitive Reinforcement Learning with Function Approximation: A Debiasing ApproachYingjie Fei, Zhuoran Yang, Zhaoran WangICML 2021 · 53 citations
- Learning Bounds for Risk-sensitive LearningJaeho Lee, Sejun Park, Jinwoo ShinNeurIPS 2020 · 52 citations
- Regret Bounds for Risk-Sensitive Reinforcement LearningOsbert Bastani, Yecheng Jason Ma, Estelle Shen, Wanqiao XuNeurIPS 2022 · 29 citations
Related papers
- A Reductions Approach to Risk-Sensitive Reinforcement Learning with Optimized Certainty EquivalentsKaiwen Wang, Dawen Liang, Nathan Kallus, Wen SunICML 2025
- Near-Minimax-Optimal Risk-Sensitive Reinforcement Learning with CVaRKaiwen Wang, Nathan Kallus, Wen SunICML 2023 · 36 citations
- Online Learning in Risk Sensitive constrained MDPArnob Ghosh, Mehrdad MoharramiICML 2025
- Cascaded Gaps: Towards Logarithmic Regret for Risk-Sensitive Reinforcement LearningYingjie Fei, Ruitu XuICML 2022 · 13 citations
- A Distribution Optimization Framework for Confidence Bounds of Risk MeasuresHao Liang, Zhi-Quan LuoICML 2023 · 4 citations
