Temporal Difference Learning as Gradient Splitting
Rui Liu, Alex Olshevsky
Abstract
Temporal difference learning with linear function approximation is a popular method to obtain a low-dimensional approximation of the value function of a policy in a Markov Decision Process. We give a new interpretation of this method in terms of a splitting of the gradient of an appropriately chosen function. As a consequence of this interpretation, convergence proofs for gradient descent can be applied almost verbatim to temporal difference learning. Beyond giving a new, fuller explanation of why temporal difference works, our interpretation also yields improved convergence times. We consider the setting with step-size, where previous comparable finite-time convergence time bounds for temporal difference learning had the multiplicative factor in front of the bound, with being the discount factor. We show that a minor variation on TD learning which estimates the mean of the value function separately has a convergence time where only multiplies an asymptotically negligible term.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9f0ac47b-7a9f-4760-a3b1-2c7eb145da7cCited by top-tier papers7
- Convergence of Actor-Critic with Multi-Layer Neural NetworksHaoxing Tian, Alex Olshevsky, Yannis PaschalidisNeurIPS 2023 · 13 citations
- A General-Purpose Theorem for High-Probability Bounds of Stochastic Approximation with Polyak AveragingSajad Khodadadian, Martin ZubeldiaNeurIPS 2025 · 4 citations
- On the Second-Order Convergence of Biased Policy Gradient AlgorithmsSiqiao Mu, Diego KlabjanICML 2024 · 4 citations
- Towards Parameter-Free Temporal Difference LearningYunxiang LI, Mark Schmidt, Reza Babanezhad, Sharan VaswaniICML 2026 · 2 citations
- Corruption-Tolerant Asynchronous Q-Learning with Near-Optimal RatesSreejeet Maity, Aritra MitraICML 2026 · 1 citation
Builds on1
Related papers
- TD Convergence: An Optimization PerspectiveKavosh Asadi, Shoham Sabach, Yao Liu, Omer Gottesman et al.NeurIPS 2023 · 17 citations
- Non-Asymptotic Analysis for Two Time-scale TDC with General Smooth Function ApproximationYue Wang, Shaofeng Zou, Yi ZhouNeurIPS 2021 · 12 citations
- A Finite Sample Analysis of Distributional TD Learning with Linear Function ApproximationYang Peng, Kaicheng Jin, Liangyu Zhang, Zhihua ZhangNeurIPS 2025 · 6 citations
- Loss Dynamics of Temporal Difference Reinforcement LearningBlake Bordelon, Paul Masset, Henry Kuo, Cengiz PehlevanNeurIPS 2023
- Backstepping Temporal Difference LearningHan-Dong Lim, Donghwan LeeICLR 2023
