Risk-Conditioned Reinforcement Learning: A Generalized Approach for Adapting to Varying Risk Measures
Gwangpyo Yoo, Jinwoo Park, Honguk Woo
Abstract
In application domains requiring mission-critical decision making, such as finance and robotics, the optimal policy derived by reinforcement learning (RL) often hinges on a preference for risk management. Yet, the dynamic nature of risk measures poses considerable challenges to achieving generalization and adaptation of risk-sensitive policies in the context of RL. In this paper, we propose a risk-conditioned RL model that enables rapid policy adaptation to varying risk measures via a unified risk representation, the Weighted Value-at-Risk (WV@R). To sample risk measures that avoid undue optimism, we construct a risk proposal network employing a conditional adversarial auto-encoder and a normalizing flow. This network establishes coherent representations for risk measures, preserving the continuity in terms of the Wasserstein distance on the risk measures. The normalizing flow is used to support non-crossing quantile regression that obtains valid samples for risk measures, and it is also applied to the agent’s critic to ascertain the preservation of monotonicity in quantile estimations. Through experiments with locomotion, finance, and self-driving scenarios, we show that our model is capable of adapting to a range of risk measures, achieving comparable performance to the baseline models individually trained for each measure. Our model often outperforms the baselines, especially in the cases when exploration is required during training but risk-aversion is favored during evaluation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cc7bee22-587f-4960-8c0a-ada12ace6a7fCited by top-tier papers4
- Distributional Inverse Reinforcement LearningFeiyang Wu, Ye Zhao, Anqi WuICML 2026 · 1 citation
- Risk-Averse Total-Reward Reinforcement LearningXihong Su, Jia Lin Hau, Gersi Doko, Kishan Panaganti et al.NeurIPS 2025
- Model Risk-sensitive Offline Reinforcement LearningGwangpyo Yoo, Honguk WooICLR 2025
- Knothe-Rosenblatt Quantile Regression for Risk-sensitive Multi-objective Reinforcement LearningGwangpyo Yoo, Woo Kyung Kim, Honguk WooICML 2026
Builds on2
- Controlling Overestimation Bias with Truncated Mixture of Continuous Distributional Quantile CriticsArsenii Kuznetsov, Pavel Shvechikov, Alexander Grishin, Dmitry P. VetrovICML 2020 · 266 citations
- Non-Crossing Quantile Regression for Distributional Reinforcement LearningFan Zhou, Jianing Wang, Xingdong FengNeurIPS 2020 · 63 citations
Related papers
- Learning Anisotropic Value Geometry with Finsler Reinforcement LearningJumman Hossain, Nirmalya RoyICML 2026
- Adversarial Diffusion for Robust Reinforcement LearningDaniele Foffano, Alessio Russo, Alexandre ProutièreNeurIPS 2025 · 5 citations
- Two steps to risk sensitivityChris Gagne, Peter DayanNeurIPS 2021 · 17 citations
- Beyond CVaR: Leveraging Static Spectral Risk Measures for Enhanced Decision-Making in Distributional Reinforcement LearningMehrdad Moghimi, Hyejin KuICML 2025
- PolicyFlow: Policy Optimization with Continuous Normalizing Flow in Reinforcement LearningShunpeng Yang, Ben Liu, Hua ChenICLR 2026 · 6 citations
