Temporal Difference Learning as Gradient Splitting
Rui Liu, Alex Olshevsky
摘要
Temporal difference learning with linear function approximation is a popular method to obtain a low-dimensional approximation of the value function of a policy in a Markov Decision Process. We give a new interpretation of this method in terms of a splitting of the gradient of an appropriately chosen function. As a consequence of this interpretation, convergence proofs for gradient descent can be applied almost verbatim to temporal difference learning. Beyond giving a new, fuller explanation of why temporal difference works, our interpretation also yields improved convergence times. We consider the setting with step-size, where previous comparable finite-time convergence time bounds for temporal difference learning had the multiplicative factor in front of the bound, with being the discount factor. We show that a minor variation on TD learning which estimates the mean of the value function separately has a convergence time where only multiplies an asymptotically negligible term.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Convergence of Actor-Critic with Multi-Layer Neural NetworksHaoxing Tian, Alex Olshevsky, Yannis PaschalidisNeurIPS 2023 · 被引用 13 次
- A General-Purpose Theorem for High-Probability Bounds of Stochastic Approximation with Polyak AveragingSajad Khodadadian, Martin ZubeldiaNeurIPS 2025 · 被引用 4 次
- On the Second-Order Convergence of Biased Policy Gradient AlgorithmsSiqiao Mu, Diego KlabjanICML 2024 · 被引用 4 次
- Towards Parameter-Free Temporal Difference LearningYunxiang LI, Mark Schmidt, Reza Babanezhad, Sharan VaswaniICML 2026 · 被引用 2 次
- Corruption-Tolerant Asynchronous Q-Learning with Near-Optimal RatesSreejeet Maity, Aritra MitraICML 2026 · 被引用 1 次
它引用的顶会 Paper1
相关 Paper
- TD Convergence: An Optimization PerspectiveKavosh Asadi, Shoham Sabach, Yao Liu, Omer Gottesman 等NeurIPS 2023 · 被引用 17 次
- Non-Asymptotic Analysis for Two Time-scale TDC with General Smooth Function ApproximationYue Wang, Shaofeng Zou, Yi ZhouNeurIPS 2021 · 被引用 12 次
- A Finite Sample Analysis of Distributional TD Learning with Linear Function ApproximationYang Peng, Kaicheng Jin, Liangyu Zhang, Zhihua ZhangNeurIPS 2025 · 被引用 6 次
- Loss Dynamics of Temporal Difference Reinforcement LearningBlake Bordelon, Paul Masset, Henry Kuo, Cengiz PehlevanNeurIPS 2023
- Backstepping Temporal Difference LearningHan-Dong Lim, Donghwan LeeICLR 2023
