Provably Efficient Risk-Sensitive Reinforcement Learning: Iterated CVaR and Worst Path
Yihan Du, Siwei Wang, Longbo Huang
摘要
In this paper, we study a novel episodic risk-sensitive Reinforcement Learning (RL) problem, named Iterated CVaR RL, which aims to maximize the tail of the reward-to-go at each step, and focuses on tightly controlling the risk of getting into catastrophic situations at each stage. This formulation is applicable to real-world tasks that demand strong risk avoidance throughout the decision process, such as autonomous driving, clinical treatment planning and robotics. We investigate two performance metrics under Iterated CVaR RL, i.e., Regret Minimization and Best Policy Identification. For both metrics, we design efficient algorithms ICVaR-RM and ICVaR-BPI, respectively, and provide nearly matching upper and lower bounds with respect to the number of episodes . We also investigate an interesting limiting case of Iterated CVaR RL, called Worst Path RL, where the objective becomes to maximize the minimum possible cumulative reward. For Worst Path RL, we propose an efficient algorithm with constant upper and lower bounds. Finally, our techniques for bounding the change of CVaR due to the value function shift and decomposing the regret via a distorted visitation distribution are novel, and can find applications in other risk-sensitive RL problems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Provably Efficient Iterated CVaR Reinforcement Learning with Function Approximation and Human FeedbackYu Chen, Yihan Du, Pihe Hu, Siwei Wang 等ICLR 2024 · 被引用 12 次
- Near-Minimax-Optimal Distributional Reinforcement Learning with a Generative ModelMark Rowland, Kevin Kevin Li, Rémi Munos, Clare Lyle 等NeurIPS 2024 · 被引用 9 次
- Risk-Sensitive Reward-Free Reinforcement Learning with CVaRXinyi Ni, Guanlin Liu, Lifeng LaiICML 2024 · 被引用 8 次
- Efficient and Sharp Off-Policy Evaluation in Robust Markov Decision ProcessesAndrew Bennett, Nathan Kallus, Miruna Oprescu, Wen Sun 等NeurIPS 2024 · 被引用 7 次
- Provably Efficient CVaR RL in Low-rank MDPsYulai Zhao, Wenhao Zhan, Xiaoyan Hu, Ho-fung Leung 等ICLR 2024 · 被引用 6 次
它引用的顶会 Paper6
- Fast active learning for pure exploration in reinforcement learningPierre Ménard, Omar Darwiche Domingues, Anders Jonsson, Emilie Kaufmann 等ICML 2021 · 被引用 110 次
- Risk-Sensitive Reinforcement Learning: Near-Optimal Risk-Sample Tradeoff in RegretYingjie Fei, Zhuoran Yang, Yudong Chen, Zhaoran Wang 等NeurIPS 2020 · 被引用 87 次
- Being Optimistic to Be Conservative: Quickly Learning a CVaR PolicyRamtin Keramati, Christoph Dann, Alex Tamkin, Emma BrunskillAAAI 2020 · 被引用 86 次
- Exponential Bellman Equation and Improved Regret Bounds for Risk-Sensitive Reinforcement LearningYingjie Fei, Zhuoran Yang, Yudong Chen, Zhaoran WangNeurIPS 2021 · 被引用 70 次
- Risk-Sensitive Reinforcement Learning with Function Approximation: A Debiasing ApproachYingjie Fei, Zhuoran Yang, Zhaoran WangICML 2021 · 被引用 53 次
相关 Paper
- Regret Bounds for Risk-Sensitive Reinforcement LearningOsbert Bastani, Yecheng Jason Ma, Estelle Shen, Wanqiao XuNeurIPS 2022 · 被引用 29 次
- Beyond CVaR: Leveraging Static Spectral Risk Measures for Enhanced Decision-Making in Distributional Reinforcement LearningMehrdad Moghimi, Hyejin KuICML 2025
- Near-Minimax-Optimal Risk-Sensitive Reinforcement Learning with CVaRKaiwen Wang, Nathan Kallus, Wen SunICML 2023 · 被引用 36 次
- Regret Bounds for Markov Decision Processes with Recursive Optimized Certainty EquivalentsWenhao Xu, Xuefeng Gao, Xuedong HeICML 2023 · 被引用 14 次
- Predictive CVaR Q-learningJu-Hyun Kim, Seungki MinICLR 2026
