Distributional Reinforcement Learning for Risk-Sensitive Policies
Shiau Hong Lim, Ilyas Malik
摘要
We address the problem of learning a risk-sensitive policy based on the CVaR risk measure using distributional reinforcement learning. In particular, we show that the standard action-selection strategy when applying the distributional Bellman optimality operator can result in convergence to neither the dynamic, Markovian CVaR nor the static, non-Markovian CVaR. We propose modifications to the existing algorithms that include a new distributional Bellman operator and show that the proposed strategy greatly expands the utility of distributional RL in learning and representing CVaR-optimized policies. Our proposed approach is a simple extension of standard distributional RL algorithms and can therefore take advantage of many of the recent advances in deep RL. On both synthetic and real data, we empirically show that our proposed algorithm is able to learn better CVaRoptimized policies. Recently, the distributional approach to RL (Bellemare et al., 2017; Morimura et al., 2010) has received increased attention due to its ability to learn better policies than the standard approaches in many challenging tasks (Dabney et al., 2018a,b; Yang et al., 2019) . Instead of learning a value function that provides the expected return of each state-action pair, the distributional approach learns the entire return distribution of each state-action pair. The approach itself is a simple extension to standard RL and is therefore easy to implement and able to leverage many of the advances in deep RL. Since the entire distribution is available, one naturally considers exploiting this information to optimize for an objective other than the expectation. Dabney et al. (2018a) presented a simple way to 36th Conference on Neural Information Processing Systems (NeurIPS 2022).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper23
- Distributional Pareto-Optimal Multi-Objective Reinforcement LearningXin-Qiang Cai, Pushi Zhang, Li Zhao, Jiang Bian 等NeurIPS 2023 · 被引用 46 次
- RiskQ: Risk-sensitive Multi-Agent Reinforcement Learning Value FactorizationSiqi Shen, Chennan Ma, Chao Li, Weiquan Liu 等NeurIPS 2023 · 被引用 34 次
- Train Hard, Fight Easy: Robust Meta Reinforcement LearningIdo Greenberg, Shie Mannor, Gal Chechik, Eli A. MeiromNeurIPS 2023 · 被引用 15 次
- Distributional Model Equivalence for Risk-Sensitive Reinforcement LearningTyler Kastner, Murat A. Erdogdu, Amir-massoud FarahmandNeurIPS 2023 · 被引用 9 次
- Action Gaps and Advantages in Continuous-Time Distributional Reinforcement LearningHarley Wiltzer, Marc G. Bellemare, David Meger, Patrick Shafto 等NeurIPS 2024 · 被引用 9 次
它引用的顶会 Paper1
相关 Paper
- Beyond CVaR: Leveraging Static Spectral Risk Measures for Enhanced Decision-Making in Distributional Reinforcement LearningMehrdad Moghimi, Hyejin KuICML 2025
- Two steps to risk sensitivityChris Gagne, Peter DayanNeurIPS 2021 · 被引用 17 次
- Near-Minimax-Optimal Risk-Sensitive Reinforcement Learning with CVaRKaiwen Wang, Nathan Kallus, Wen SunICML 2023 · 被引用 36 次
- Pitfall of Optimism: Distributional Reinforcement Learning by Randomizing Risk CriterionTaehyun Cho, Seungyub Han, Heesoo Lee, Kyungjae Lee 等NeurIPS 2023 · 被引用 10 次
- Conservative Offline Distributional Reinforcement LearningYecheng Jason Ma, Dinesh Jayaraman, Osbert BastaniNeurIPS 2021 · 被引用 118 次
