Predictive CVaR Q-learning
Ju-Hyun Kim, Seungki Min
摘要
We propose a sample-efficient Q-learning algorithm for reinforcement learning with the Conditional Value-at-Risk (CVaR) objective. Our method introduces two key innovations. First, we propose the predictive tail value function, a novel formulation of risk-sensitive action value, admits a recursive structure as in the conventional risk-neutral Bellman equation. This novel formulation addresses the problem of noisy policy evaluation originating from the non-decomposable objective. Second, we introduce a two-way exploration strategy that explores the agent's risk-sensitivity level in addition to its actions. This technique mitigates the "blindness to success" phenomenon by preventing premature convergence to overly conservative policies. We establish a rigorous theoretical foundation for this framework, including a new Bellman optimality equation and a policy improvement theorem. Empirical results demonstrate that our algorithm significantly improves both CVaR performance and learning stability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper7
- Risk-Averse Offline Reinforcement LearningNúria Armengol Urpí, Sebastian Curi, Andreas KrauseICLR 2021 · 被引用 81 次
- Efficient Risk-Averse Reinforcement LearningIdo Greenberg, Yinlam Chow, Mohammad Ghavamzadeh, Shie MannorNeurIPS 2022 · 被引用 61 次
- Distributional Reinforcement Learning for Risk-Sensitive PoliciesShiau Hong Lim, Ilyas MalikNeurIPS 2022 · 被引用 54 次
- Near-Minimax-Optimal Risk-Sensitive Reinforcement Learning with CVaRKaiwen Wang, Nathan Kallus, Wen SunICML 2023 · 被引用 36 次
- Uncertainty-Aware Policy Optimization: A Robust, Adaptive Trust Region ApproachJames Queeney, Ioannis Ch. Paschalidis, Christos G. CassandrasAAAI 2021 · 被引用 11 次
相关 Paper
- Being Optimistic to Be Conservative: Quickly Learning a CVaR PolicyRamtin Keramati, Christoph Dann, Alex Tamkin, Emma BrunskillAAAI 2020 · 被引用 86 次
- Boosting CVaR Policy Optimization with Quantile GradientsYudong Luo, Erick DelageICML 2026
- Provably Efficient Iterated CVaR Reinforcement Learning with Function Approximation and Human FeedbackYu Chen, Yihan Du, Pihe Hu, Siwei Wang 等ICLR 2024 · 被引用 12 次
- Regret Bounds for Risk-Sensitive Reinforcement LearningOsbert Bastani, Yecheng Jason Ma, Estelle Shen, Wanqiao XuNeurIPS 2022 · 被引用 29 次
- Risk-Sensitive Policy Optimization via Predictive CVaR Policy GradientJu-Hyun Kim, Seungki MinICML 2024 · 被引用 3 次
