Regret Bounds for Markov Decision Processes with Recursive Optimized Certainty Equivalents
Wenhao Xu, Xuefeng Gao, Xuedong He
摘要
The optimized certainty equivalent (OCE) is a family of risk measures that cover important examples such as entropic risk, conditional value-at-risk and mean-variance models. In this paper, we propose a new episodic risk-sensitive reinforcement learning formulation based on tabular Markov decision processes with recursive OCEs. We design an efficient learning algorithm for this problem based on value iteration and upper confidence bound. We derive an upper bound on the regret of the proposed algorithm, and also establish a minimax lower bound. Our bounds show that the regret rate achieved by our proposed algorithm has optimal dependence on the number of episodes and the number of actions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Conformal Risk Training: End-to-End Optimization of Conformal Risk ControlChristopher Yeh, Nicolas Christianson, Adam Wierman, Yisong YueNeurIPS 2025 · 被引用 14 次
- Provably Efficient Iterated CVaR Reinforcement Learning with Function Approximation and Human FeedbackYu Chen, Yihan Du, Pihe Hu, Siwei Wang 等ICLR 2024 · 被引用 12 次
- Efficient and Sharp Off-Policy Evaluation in Robust Markov Decision ProcessesAndrew Bennett, Nathan Kallus, Miruna Oprescu, Wen Sun 等NeurIPS 2024 · 被引用 7 次
- Provable Risk-Sensitive Distributional Reinforcement Learning with General Function ApproximationYu Chen, Xiangcheng Zhang, Siwei Wang, Longbo HuangICML 2024 · 被引用 3 次
- Risk-Averse Constrained Reinforcement Learning with Optimized Certainty EquivalentsJane H. Lee, Baturay Saglam, Spyridon Pougkakiotis, Amin Karbasi 等NeurIPS 2025 · 被引用 2 次
它引用的顶会 Paper6
- Risk-Sensitive Reinforcement Learning: Near-Optimal Risk-Sample Tradeoff in RegretYingjie Fei, Zhuoran Yang, Yudong Chen, Zhaoran Wang 等NeurIPS 2020 · 被引用 87 次
- Exponential Bellman Equation and Improved Regret Bounds for Risk-Sensitive Reinforcement LearningYingjie Fei, Zhuoran Yang, Yudong Chen, Zhaoran WangNeurIPS 2021 · 被引用 70 次
- Risk-Sensitive Reinforcement Learning with Function Approximation: A Debiasing ApproachYingjie Fei, Zhuoran Yang, Zhaoran WangICML 2021 · 被引用 53 次
- Learning Bounds for Risk-sensitive LearningJaeho Lee, Sejun Park, Jinwoo ShinNeurIPS 2020 · 被引用 52 次
- Regret Bounds for Risk-Sensitive Reinforcement LearningOsbert Bastani, Yecheng Jason Ma, Estelle Shen, Wanqiao XuNeurIPS 2022 · 被引用 29 次
相关 Paper
- A Reductions Approach to Risk-Sensitive Reinforcement Learning with Optimized Certainty EquivalentsKaiwen Wang, Dawen Liang, Nathan Kallus, Wen SunICML 2025
- Near-Minimax-Optimal Risk-Sensitive Reinforcement Learning with CVaRKaiwen Wang, Nathan Kallus, Wen SunICML 2023 · 被引用 36 次
- Online Learning in Risk Sensitive constrained MDPArnob Ghosh, Mehrdad MoharramiICML 2025
- Cascaded Gaps: Towards Logarithmic Regret for Risk-Sensitive Reinforcement LearningYingjie Fei, Ruitu XuICML 2022 · 被引用 13 次
- A Distribution Optimization Framework for Confidence Bounds of Risk MeasuresHao Liang, Zhi-Quan LuoICML 2023 · 被引用 4 次
