Linear Q-Learning Does Not Diverge in L2: Convergence Rates to a Bounded Set
Xinyu Liu, Zixuan Xie, Shangtong Zhang
摘要
Q-learning is one of the most fundamental reinforcement learning algorithms. It is widely believed that Q-learning with linear function approximation (i.e., linear Q-learning) suffers from possible divergence until the recent work Meyn (2024) which establishes the ultimate almost sure boundedness of the iterates of linear Q-learning. Building on this success, this paper further establishes the first L 2 convergence rate of linear Q-learning iterates (to a bounded set). Similar to Meyn (2024), we do not make any modification to the original linear Q-learning algorithm, do not make any Bellman completeness assumption, and do not make any near-optimality assumption on the behavior policy. All we need is an ϵ-softmax behavior policy with an adaptive temperature. The key to our analysis is the general result of stochastic approximations under Markovian noise with fast-changing transition functions. As a side product, we also use this general result to establish the L 2 convergence rate of tabular Q-learning with an ϵ-softmax behavior policy, for which we rely on a novel pseudo-contraction property of the weighted Bellman optimality operator.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper10
- A Finite-Time Analysis of Two Time-Scale Actor-Critic MethodsYue Wu, Weitong Zhang, Pan Xu, Quanquan GuNeurIPS 2020 · 被引用 189 次
- A Finite-Time Analysis of Q-Learning with Neural Network Function ApproximationPan Xu, Quanquan GuICML 2020 · 被引用 79 次
- Breaking the Deadly Triad with a Target NetworkShangtong Zhang, Hengshuai Yao, Shimon WhitesonICML 2021 · 被引用 61 次
- Provably Convergent Two-Timescale Off-Policy Actor-Critic with Function ApproximationShangtong Zhang, Bo Liu, Hengshuai Yao, Shimon WhitesonICML 2020 · 被引用 58 次
- On the Convergence and Sample Complexity Analysis of Deep Q-Networks with ε-Greedy ExplorationShuai Zhang, Hongkang Li, Meng Wang, Miao Liu 等NeurIPS 2023 · 被引用 57 次
相关 Paper
- The Mean-Squared Error of Double Q-LearningWentao Weng, Harsh Gupta, Niao He, Lei Ying 等NeurIPS 2020 · 被引用 19 次
- A new convergent variant of Q-learning with linear function approximationDiogo S. Carvalho, Francisco S. Melo, Pedro SantosNeurIPS 2020 · 被引用 39 次
- Regularized Q-LearningHan-Dong Lim, Donghwan LeeNeurIPS 2024 · 被引用 1 次
- Taming "data-hungry" reinforcement learning? Stability in continuous state-action spacesYaqi Duan, Martin J. WainwrightNeurIPS 2024 · 被引用 6 次
- Zap Q-Learning With Nonlinear Function ApproximationShuhang Chen, Adithya M. Devraj, Fan Lu, Ana Busic 等NeurIPS 2020 · 被引用 26 次
