Provably Efficient Iterated CVaR Reinforcement Learning with Function Approximation and Human Feedback
Yu Chen, Yihan Du, Pihe Hu, Siwei Wang, Desheng Wu, Longbo Huang
Abstract
Risk-sensitive reinforcement learning (RL) aims to optimize policies that balance the expected reward and risk. In this paper, we present a novel risk-sensitive RL framework that employs an Iterated Conditional Value-at-Risk (CVaR) objective under both linear and general function approximations, enriched by human feedback. These new formulations provide a principled way to guarantee safety in each decision making step throughout the control process. Moreover, integrating human feedback into risk-sensitive RL framework bridges the gap between algorithmic decision-making and human participation, allowing us to also guarantee safety for human-in-the-loop systems. We propose provably sample-efficient algorithms for this Iterated CVaR RL and provide rigorous theoretical analysis. Furthermore, we establish a matching lower bound to corroborate the optimality of our algorithms in a linear context.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d6137ae0-eca7-491f-bfa4-727a9d80c29fCited by top-tier papers5
- Risk-Sensitive Policy Optimization via Predictive CVaR Policy GradientJu-Hyun Kim, Seungki MinICML 2024 · 3 citations
- Risk-aware Direct Preference Optimization under Nested Risk MeasureLijun Zhang, Lin Li, Yajie Qi, Huizhong Song et al.NeurIPS 2025 · 3 citations
- Adversarial Policy Optimization for Offline Preference-based Reinforcement LearningHyungkyu Kang, Min-hwan OhICLR 2025
- Offline Preference-Based Value OptimizationHyungkyu Kang, Min-hwan OhICLR 2026
- Reducing the Probability of Undesirable Outputs in Language Models Using Probabilistic InferenceStephen Zhao, Aidan Li, Rob Brekelmans, Roger Baker GrosseNeurIPS 2025
Builds on17
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Model-Based Reinforcement Learning with Value-Targeted RegressionAlex Ayoub, Zeyu Jia, Csaba Szepesvári, Mengdi Wang et al.ICML 2020 · 324 citations
- Reinforcement Learning in Feature Space: Matrix Bandit, Kernels, and Regret BoundLin Yang, Mengdi WangICML 2020 · 308 citations
- Principled Reinforcement Learning with Human Feedback from Pairwise or K-wise ComparisonsBanghua Zhu, Michael I. Jordan, Jiantao JiaoICML 2023 · 273 citations
- Reinforcement Learning with General Value Function Approximation: Provably Efficient Approach via Bounded Eluder DimensionRuosong Wang, Ruslan Salakhutdinov, Lin F. YangNeurIPS 2020 · 168 citations
Related papers
- Near-Minimax-Optimal Risk-Sensitive Reinforcement Learning with CVaRKaiwen Wang, Nathan Kallus, Wen SunICML 2023 · 36 citations
- Provably Efficient CVaR RL in Low-rank MDPsYulai Zhao, Wenhao Zhan, Xiaoyan Hu, Ho-fung Leung et al.ICLR 2024 · 6 citations
- Risk-Sensitive Reward-Free Reinforcement Learning with CVaRXinyi Ni, Guanlin Liu, Lifeng LaiICML 2024 · 8 citations
- Provably Efficient Risk-Sensitive Reinforcement Learning: Iterated CVaR and Worst PathYihan Du, Siwei Wang, Longbo HuangICLR 2023 · 2 citations
- Predictive CVaR Q-learningJu-Hyun Kim, Seungki MinICLR 2026
