Efficient Risk-Averse Reinforcement Learning
Ido Greenberg, Yinlam Chow, Mohammad Ghavamzadeh, Shie Mannor
摘要
In risk-averse reinforcement learning (RL), the goal is to optimize some risk measure of the returns. A risk measure often focuses on the worst returns out of the agent's experience. As a result, standard methods for risk-averse RL often ignore high-return strategies. We prove that under certain conditions this inevitably leads to a local-optimum barrier, and propose a mechanism we call soft risk to bypass it. We also devise a novel cross entropy module for sampling, which (1) preserves risk aversion despite the soft risk; (2) independently improves sample efficiency. By separating the risk aversion of the sampler and the optimizer, we can sample episodes with poor conditions, yet optimize with respect to successful strategies. We combine these two concepts in CeSoR -Cross-entropy Soft-Risk optimization algorithm -which can be applied on top of any risk-averse policy gradient (PG) method. We demonstrate improved risk aversion in maze navigation, autonomous driving, and resource allocation benchmarks, including in scenarios where standard risk-averse PG completely fails. Our results and CeSoR implementation are available on Github. The stand-alone cross entropy module is available on PyPI.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper27
- An Alternative to Variance: Gini Deviation for Risk-averse Policy GradientYudong Luo, Guiliang Liu, Pascal Poupart, Yangchen PanNeurIPS 2023 · 被引用 15 次
- Train Hard, Fight Easy: Robust Meta Reinforcement LearningIdo Greenberg, Shie Mannor, Gal Chechik, Eli A. MeiromNeurIPS 2023 · 被引用 15 次
- Reinforcing Long-Term Performance in Recommender Systems with User-Oriented Exploration PolicyChangshuo Zhang, Sirui Chen, Xiao Zhang, Sunhao Dai 等SIGIR 2024 · 被引用 14 次
- OVD-Explorer: Optimism Should Not Be the Sole Pursuit of Exploration in Noisy EnvironmentsJinyi Liu, Zhi Wang, Yan Zheng, Jianye Hao 等AAAI 2024 · 被引用 14 次
- Risk-Averse Fine-tuning of Large Language ModelsSapana Chaudhary, Ujwal Dinesha, Dileep Kalathil, Srinivas ShakkottaiNeurIPS 2024 · 被引用 12 次
它引用的顶会 Paper4
- Being Optimistic to Be Conservative: Quickly Learning a CVaR PolicyRamtin Keramati, Christoph Dann, Alex Tamkin, Emma BrunskillAAAI 2020 · 被引用 86 次
- RMIX: Learning Risk-Sensitive Policies for Cooperative Reinforcement Learning AgentsWei Qiu, Xinrun Wang, Runsheng Yu, Rundong Wang 等NeurIPS 2021 · 被引用 71 次
- Exponential Bellman Equation and Improved Regret Bounds for Risk-Sensitive Reinforcement LearningYingjie Fei, Zhuoran Yang, Yudong Chen, Zhaoran WangNeurIPS 2021 · 被引用 70 次
- Adaptive Sampling for Stochastic Risk-Averse LearningSebastian Curi, Kfir Y. Levy, Stefanie Jegelka, Andreas KrauseNeurIPS 2020 · 被引用 65 次
相关 Paper
- On the Global Convergence of Risk-Averse Policy Gradient Methods with Expected Conditional Risk MeasuresXian Yu, Lei YingICML 2023 · 被引用 8 次
- Off-Policy Safe Reinforcement Learning with Cost-Constrained Optimistic ExplorationGuopeng Li, Matthijs T. J. Spaan, Julian F. P. KooijICLR 2026
- Risk-Conditioned Reinforcement Learning: A Generalized Approach for Adapting to Varying Risk MeasuresGwangpyo Yoo, Jinwoo Park, Honguk WooAAAI 2024 · 被引用 3 次
- One Risk to Rule Them All: A Risk-Sensitive Perspective on Model-Based Offline Reinforcement LearningMarc Rigter, Bruno Lacerda, Nick HawesNeurIPS 2023 · 被引用 26 次
- Provably Efficient Risk-Sensitive Reinforcement Learning: Iterated CVaR and Worst PathYihan Du, Siwei Wang, Longbo HuangICLR 2023 · 被引用 2 次
