Zap Q-Learning With Nonlinear Function Approximation
Shuhang Chen, Adithya M. Devraj, Fan Lu, Ana Busic, Sean P. Meyn
摘要
The Zap stochastic approximation (SA) algorithm was introduced recently as a means to accelerate convergence in reinforcement learning algorithms. While numerical results were impressive, stability (in the sense of boundedness of parameter estimates) was established in only a few special cases. This class of algorithms is generalized in this paper, and stability is established under very general conditions. This general result can be applied to a wide range of algorithms found in reinforcement learning. Two classes are considered in this paper: (i)The natural generalization of Watkins' algorithm is not always stable in function approximation settings. Parameter estimates may diverge to infinity even in the linear function approximation setting with a simple finite state-action MDP. Under mild conditions, the Zap SA algorithm provides a stable algorithm, even in the case of nonlinear function approximation. (ii) The GQ algorithm of Maei et. al. 2010 is designed to address the stability challenge. Analysis is provided to explain why the algorithm may be very slow to converge in practice. The new Zap GQ algorithm is stable even for nonlinear function approximation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Breaking the Deadly Triad with a Target NetworkShangtong Zhang, Hengshuai Yao, Shimon WhitesonICML 2021 · 被引用 61 次
- TD Convergence: An Optimization PerspectiveKavosh Asadi, Shoham Sabach, Yao Liu, Omer Gottesman 等NeurIPS 2023 · 被引用 17 次
- Why Target Networks Stabilise Temporal Difference MethodsMattie Fellows, Matthew J. A. Smith, Shimon WhitesonICML 2023 · 被引用 10 次
- Rank-One Modified Value IterationArman Sharifi Kolarijani, Tolga Ok, Peyman Mohajerin Esfahani, Mohamad Amin Sharifi KolarijaniICML 2025
相关 Paper
- Linear Q-Learning Does Not Diverge in L2: Convergence Rates to a Bounded SetXinyu Liu, Zixuan Xie, Shangtong ZhangICML 2025
- Finite-Sample Analysis of Contractive Stochastic Approximation Using Smooth Convex EnvelopesZaiwei Chen, Siva Theja Maguluri, Sanjay Shakkottai, Karthikeyan ShanmugamNeurIPS 2020 · 被引用 66 次
- A new convergent variant of Q-learning with linear function approximationDiogo S. Carvalho, Francisco S. Melo, Pedro SantosNeurIPS 2020 · 被引用 39 次
- Approaching Quartic Convergence Rates for Quasi-Stochastic Approximation with Application to Gradient-Free OptimizationCaio Kalil Lauand, Sean P. MeynNeurIPS 2022 · 被引用 13 次
- Regularization Guarantees Generalization in Bayesian Reinforcement Learning through Algorithmic StabilityAviv Tamar, Daniel Soudry, Ev ZisselmanAAAI 2022 · 被引用 9 次
