Risk-Sensitive Policy Optimization via Predictive CVaR Policy Gradient
Ju-Hyun Kim, Seungki Min
2024年份
3被引次数
4顶会引用
摘要
This paper addresses a policy optimization task with the conditional value-at-risk (CVaR) objective. We introduce the predictive CVaR policy gradient, a novel approach that seamlessly integrates risk-neutral policy gradient algorithms with minimal modifications. Our method incorporates a reweighting strategy in gradient calculation – individual cost terms are reweighted in proportion to their predicted contribution to the objective. These weights can be easily estimated through a separate learning procedure. We provide theoretical and empirical analyses, demonstrating the validity and effectiveness of our proposed method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Reward Redistribution for CVaR MDPs using a Bellman Operator on L-infinityAneri Muni, Vincent Taboga, Esther Derman, Pierre-Luc Bacon 等ICML 2026 · 被引用 1 次
- Boosting CVaR Policy Optimization with Quantile GradientsYudong Luo, Erick DelageICML 2026
- Return Capping: Sample Efficient CVaR Policy Gradient OptimisationHarry Mead, Clarissa Costen, Bruno Lacerda, Nick HawesICML 2025
- Predictive CVaR Q-learningJu-Hyun Kim, Seungki MinICLR 2026
它引用的顶会 Paper8
- Sample Efficient Policy Gradient Methods with Recursive Variance ReductionPan Xu, Felicia Gao, Quanquan GuICLR 2020 · 被引用 99 次
- Risk-Sensitive Reinforcement Learning: Near-Optimal Risk-Sample Tradeoff in RegretYingjie Fei, Zhuoran Yang, Yudong Chen, Zhaoran Wang 等NeurIPS 2020 · 被引用 87 次
- Exponential Bellman Equation and Improved Regret Bounds for Risk-Sensitive Reinforcement LearningYingjie Fei, Zhuoran Yang, Yudong Chen, Zhaoran WangNeurIPS 2021 · 被引用 70 次
- Efficient Risk-Averse Reinforcement LearningIdo Greenberg, Yinlam Chow, Mohammad Ghavamzadeh, Shie MannorNeurIPS 2022 · 被引用 61 次
- Near-Minimax-Optimal Risk-Sensitive Reinforcement Learning with CVaRKaiwen Wang, Nathan Kallus, Wen SunICML 2023 · 被引用 36 次
相关 Paper
- Being Optimistic to Be Conservative: Quickly Learning a CVaR PolicyRamtin Keramati, Christoph Dann, Alex Tamkin, Emma BrunskillAAAI 2020 · 被引用 86 次
- Risk-Averse No-Regret Learning in Online Convex GamesZifan Wang, Yi Shen, Michael M. ZavlanosICML 2022 · 被引用 10 次
- Risk-Aware Stochastic Shortest PathTobias MeggendorferAAAI 2022 · 被引用 13 次
- Provably Efficient Iterated CVaR Reinforcement Learning with Function Approximation and Human FeedbackYu Chen, Yihan Du, Pihe Hu, Siwei Wang 等ICLR 2024 · 被引用 12 次
- On Dynamic Programming Decompositions of Static Risk Measures in Markov Decision ProcessesJia Lin Hau, Erick Delage, Mohammad Ghavamzadeh, Marek PetrikNeurIPS 2023 · 被引用 22 次
