Boosting CVaR Policy Optimization with Quantile Gradients
Yudong Luo, Erick Delage
Abstract
Optimizing Conditional Value-at-risk (CVaR) using policy gradient (a.k.a CVaR-PG) faces significant challenges of sample inefficiency. This inefficiency stems from the fact that it focuses on tail-end performance and overlooks many sampled trajectories. We address this problem by augmenting CVaR with an expected quantile term. Quantile optimization admits a dynamic programming formulation that leverages all sampled data, thus improves sample efficiency. This does not alter the CVaR objective since CVaR corresponds to the expectation of quantile over the tail. Empirical results in domains with verifiable risk-averse behavior show that our algorithm within the Markovian policy class substantially improves upon CVaR-PG and consistently outperforms other existing methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 034ea88d-e2f3-4195-974e-1146602cf7a5Builds on5
- Offline Reinforcement Learning with Implicit Q-LearningIlya Kostrikov, Ashvin Nair, Sergey LevineICLR 2022 · 1,402 citations
- Efficient Risk-Averse Reinforcement LearningIdo Greenberg, Yinlam Chow, Mohammad Ghavamzadeh, Shie MannorNeurIPS 2022 · 61 citations
- Distributional Reinforcement Learning for Risk-Sensitive PoliciesShiau Hong Lim, Ilyas MalikNeurIPS 2022 · 54 citations
- On Dynamic Programming Decompositions of Static Risk Measures in Markov Decision ProcessesJia Lin Hau, Erick Delage, Mohammad Ghavamzadeh, Marek PetrikNeurIPS 2023 · 22 citations
- Risk-Sensitive Policy Optimization via Predictive CVaR Policy GradientJu-Hyun Kim, Seungki MinICML 2024 · 3 citations
Related papers
- Return Capping: Sample Efficient CVaR Policy Gradient OptimisationHarry Mead, Clarissa Costen, Bruno Lacerda, Nick HawesICML 2025
- Predictive CVaR Q-learningJu-Hyun Kim, Seungki MinICLR 2026
- Being Optimistic to Be Conservative: Quickly Learning a CVaR PolicyRamtin Keramati, Christoph Dann, Alex Tamkin, Emma BrunskillAAAI 2020 · 86 citations
- Reward Redistribution for CVaR MDPs using a Bellman Operator on L-infinityAneri Muni, Vincent Taboga, Esther Derman, Pierre-Luc Bacon et al.ICML 2026 · 1 citation
- Risk-Averse Bayes-Adaptive Reinforcement LearningMarc Rigter, Bruno Lacerda, Nick HawesNeurIPS 2021 · 50 citations
