Learning Near Optimal Policies with Low Inherent Bellman Error
Andrea Zanette, Alessandro Lazaric, Mykel J. Kochenderfer, Emma Brunskill
摘要
We study the exploration problem with approximate linear action-value functions in episodic reinforcement learning under the notion of low inherent Bellman error, a condition normally employed to show convergence of approximate value iteration. First we relate this condition to other common frameworks and show that it is strictly more general than the low rank (or linear) MDP assumption of prior work. Second we provide an algorithm with a high probability regret bound where is the horizon, is the number of episodes, is the value if the inherent Bellman error and is the feature dimension at timestep . In addition, we show that the result is unimprovable beyond constants and logs by showing a matching lower bound. This has two important consequences: 1) it shows that exploration is possible using only batch assumptions with an algorithm that achieves the optimal statistical rate for the setting we consider, which is more general than prior work on low-rank MDPs 2) the lack of closedness (measured by the inherent Bellman error) is only amplified by despite working in the online setting. Finally, the algorithm reduces to the celebrated LinUCB when but with a different choice of the exploration parameter that allows handling misspecified contextual linear bandits. While computational tractability questions remain open for the MDP setting, this enriches the class of MDPs with a linear representation for the action-value function where statistically efficient reinforcement learning is possible.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper120
- Bellman Eluder Dimension: New Rich Classes of RL Problems, and Sample-Efficient AlgorithmsChi Jin, Qinghua Liu, Sobhan MiryoosefiNeurIPS 2021 · 被引用 264 次
- Policy Finetuning: Bridging Sample-Efficient Offline and Online Reinforcement LearningTengyang Xie, Nan Jiang, Huan Wang, Caiming Xiong 等NeurIPS 2021 · 被引用 207 次
- Reinforcement Learning with General Value Function Approximation: Provably Efficient Approach via Bounded Eluder DimensionRuosong Wang, Ruslan Salakhutdinov, Lin F. YangNeurIPS 2020 · 被引用 168 次
- Jump-Start Reinforcement LearningIkechukwu Uchendu, Ted Xiao, Yao Lu, Banghua Zhu 等ICML 2023 · 被引用 158 次
- Provably Efficient Reinforcement Learning for Discounted MDPs with Feature MappingDongruo Zhou, Jiafan He, Quanquan GuICML 2021 · 被引用 143 次
它引用的顶会 Paper2
- Learning with Good Feature Representations in Bandits and in RL with a Generative ModelTor Lattimore, Csaba Szepesvári, Gellért WeiszICML 2020 · 被引用 181 次
- Optimism in Reinforcement Learning with Generalized Linear Function ApproximationYining Wang, Ruosong Wang, Simon Shaolei Du, Akshay KrishnamurthyICLR 2021 · 被引用 54 次
相关 Paper
- Near-Optimal Representation Learning for Linear Bandits and Linear RLJiachen Hu, Xiaoyu Chen, Chi Jin, Lihong Li 等ICML 2021 · 被引用 60 次
- Reinforcement Learning in Linear MDPs: Constant Regret and Representation SelectionMatteo Papini, Andrea Tirinzoni, Aldo Pacchiano, Marcello Restelli 等NeurIPS 2021 · 被引用 26 次
- Provably Efficient Reward-Agnostic Navigation with Linear Value IterationAndrea Zanette, Alessandro Lazaric, Mykel J. Kochenderfer, Emma BrunskillNeurIPS 2020 · 被引用 68 次
- Corruption-Robust Algorithms with Uncertainty Weighting for Nonlinear Contextual Bandits and Markov Decision ProcessesChenlu Ye, Wei Xiong, Quanquan Gu, Tong ZhangICML 2023 · 被引用 40 次
- Nearly Minimax Optimal Reinforcement Learning with Linear Function ApproximationPihe Hu, Yu Chen, Longbo HuangICML 2022 · 被引用 38 次
