Risk-Sensitive Reinforcement Learning with Function Approximation: A Debiasing Approach
Yingjie Fei, Zhuoran Yang, Zhaoran Wang
摘要
We study function approximation for episodic reinforcement learning with entropic risk measure. We first propose an algorithm with linear function approximation. Compared to existing algorithms, which suffer from improper regularization and regression biases, this algorithm features debiasing transformations in backward induction and regression procedures. We further propose an algorithm with general function approximation, which is shown to perform implicit debiasing transformations. We prove that both algorithms achieve a sublinear regret and demonstrate a tradeoff between generality and efficiency. Our analysis provides a unified framework for function approximation in risk-sensitive reinforcement learning, which leads to the first sub-linear regret bounds in the setting.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper27
- Exponential Bellman Equation and Improved Regret Bounds for Risk-Sensitive Reinforcement LearningYingjie Fei, Zhuoran Yang, Yudong Chen, Zhaoran WangNeurIPS 2021 · 被引用 70 次
- RiskQ: Risk-sensitive Multi-Agent Reinforcement Learning Value FactorizationSiqi Shen, Chennan Ma, Chao Li, Weiquan Liu 等NeurIPS 2023 · 被引用 34 次
- Learn to Match with No Regret: Reinforcement Learning in Markov Matching MarketsYifei Min, Tianhao Wang, Ruitu Xu, Zhaoran Wang 等NeurIPS 2022 · 被引用 31 次
- Regret Bounds for Risk-Sensitive Reinforcement LearningOsbert Bastani, Yecheng Jason Ma, Estelle Shen, Wanqiao XuNeurIPS 2022 · 被引用 29 次
- Understanding the Eluder DimensionGene Li, Pritish Kamath, Dylan J. Foster, Nati SrebroNeurIPS 2022 · 被引用 22 次
它引用的顶会 Paper5
- Model-Based Reinforcement Learning with Value-Targeted RegressionAlex Ayoub, Zeyu Jia, Csaba Szepesvári, Mengdi Wang 等ICML 2020 · 被引用 324 次
- Provably Efficient Exploration in Policy OptimizationQi Cai, Zhuoran Yang, Chi Jin, Zhaoran WangICML 2020 · 被引用 304 次
- Provably Efficient Reinforcement Learning for Discounted MDPs with Feature MappingDongruo Zhou, Jiafan He, Quanquan GuICML 2021 · 被引用 143 次
- Risk-Sensitive Reinforcement Learning: Near-Optimal Risk-Sample Tradeoff in RegretYingjie Fei, Zhuoran Yang, Yudong Chen, Zhaoran Wang 等NeurIPS 2020 · 被引用 87 次
- Optimism in Reinforcement Learning with Generalized Linear Function ApproximationYining Wang, Ruosong Wang, Simon Shaolei Du, Akshay KrishnamurthyICLR 2021 · 被引用 54 次
相关 Paper
- Risk-Aware Reinforcement Learning with Coherent Risk Measures and Non-linear Function ApproximationThanh Lam, Arun Verma, Bryan Kian Hsiang Low, Patrick JailletICLR 2023
- Provably Efficient Partially Observable Risk-sensitive Reinforcement Learning with Hindsight ObservationTonghe Zhang, Yu Chen, Longbo HuangICML 2024
- Pessimism Meets Risk: Risk-Sensitive Offline Reinforcement LearningDake Zhang, Boxiang Lyu, Shuang Qiu, Mladen Kolar 等ICML 2024 · 被引用 4 次
- Cascaded Gaps: Towards Logarithmic Regret for Risk-Sensitive Reinforcement LearningYingjie Fei, Ruitu XuICML 2022 · 被引用 13 次
- Regret Bounds for Episodic Risk-Sensitive Linear Quadratic RegulatorWenhao Xu, Xuefeng Gao, Xuedong HeICLR 2025
