Risk-Sensitive Reward-Free Reinforcement Learning with CVaR
Xinyi Ni, Guanlin Liu, Lifeng Lai
摘要
Exploration is a crucial phase in reinforcement learning (RL). The reward-free RL paradigm, as proposed by (Jin et al., 2020) , offers an efficient method to design exploration algorithms for riskneutral RL across various reward functions with a single exploration phase. However, as RL applications in safety critical settings grow, there's an increasing need for risk-sensitive RL, which takes potential risks into consideration for decisionmaking. Yet, efficient exploration strategies for risk-sensitive RL remain underdeveloped. This study presents a novel risk-sensitive reward-free framework based on Conditional Value-at-Risk (CVaR), designed to effectively address CVaR RL for any given reward function through a single exploration phase. We introduce an efficient algorithm named CVaR-RF-UCRL, which is shown to be (ϵ, p)-PAC, with a sample complexity upper bounded by Õ S 2 AH 4 ϵ 2 τ 2 with τ being the risk tolerance parameter. We also prove a Ω S 2 AH 2 ϵ 2 τ lower bound for any CVaR-RF exploration algorithm, demonstrating the near-optimality of our algorithm. Additionally, we propose the planning algorithms: CVaR-VI and its more practical variant, CVaR-VI-DISC. The effectiveness and practicality of our CVaR reward-free approach are further validated through numerical experiments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper13
- Reward-Free Exploration for Reinforcement LearningChi Jin, Akshay Krishnamurthy, Max Simchowitz, Tiancheng YuICML 2020 · 被引用 226 次
- On Reward-Free Reinforcement Learning with Linear Function ApproximationRuosong Wang, Simon S. Du, Lin F. Yang, Ruslan SalakhutdinovNeurIPS 2020 · 被引用 121 次
- Fast active learning for pure exploration in reinforcement learningPierre Ménard, Omar Darwiche Domingues, Anders Jonsson, Emilie Kaufmann 等ICML 2021 · 被引用 110 次
- Risk-Sensitive Reinforcement Learning: Near-Optimal Risk-Sample Tradeoff in RegretYingjie Fei, Zhuoran Yang, Yudong Chen, Zhaoran Wang 等NeurIPS 2020 · 被引用 87 次
- Being Optimistic to Be Conservative: Quickly Learning a CVaR PolicyRamtin Keramati, Christoph Dann, Alex Tamkin, Emma BrunskillAAAI 2020 · 被引用 86 次
相关 Paper
- Near-Minimax-Optimal Risk-Sensitive Reinforcement Learning with CVaRKaiwen Wang, Nathan Kallus, Wen SunICML 2023 · 被引用 36 次
- Provably Efficient Iterated CVaR Reinforcement Learning with Function Approximation and Human FeedbackYu Chen, Yihan Du, Pihe Hu, Siwei Wang 等ICLR 2024 · 被引用 12 次
- Provably Efficient CVaR RL in Low-rank MDPsYulai Zhao, Wenhao Zhan, Xiaoyan Hu, Ho-fung Leung 等ICLR 2024 · 被引用 6 次
- Safe Exploration Incurs Nearly No Additional Sample Complexity for Reward-Free RLRuiquan Huang, Jing Yang, Yingbin LiangICLR 2023
- Predictive CVaR Q-learningJu-Hyun Kim, Seungki MinICLR 2026
