Why Target Networks Stabilise Temporal Difference Methods
Mattie Fellows, Matthew J. A. Smith, Shimon Whiteson
Abstract
Integral to many recent successes in deep reinforcement learning has been a class of temporal difference methods that use infrequently updated target values for policy evaluation in a Markov Decision Process. At the same time, a complete theoretical explanation for the effectiveness of target networks remains elusive. In this work, we provide an analysis of this popular class of algorithms, to finally answer the question: "why do target networks stabilise TD learning"? To do so, we formalise the notion of a partially fitted policy evaluation method, which describes the use of target networks and bridges the gap between fitted methods and semigradient temporal difference algorithms. Using this framework we are able to uniquely characterise the so-called deadly triad-the use of TD updates with (nonlinear) function approximation and off-policy datawhich often leads to nonconvergent algorithms. This insight leads us to conclude that the use of target networks can mitigate the effects of poor conditioning in the Jacobian of the TD update. Furthermore, we show that under mild regularity conditions and a well tuned target network update frequency, convergence can be guaranteed even in the extremely challenging off-policy sampling and nonlinear function approximation setting.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ca1f9457-cb99-4ae3-aac7-c86dd70e3f4eCited by top-tier papers6
- Chunking the Critic: A Transformer-based Soft Actor-Critic with N-Step ReturnsDong Tian, Onur Celik, Gerhard NeumannICLR 2026 · 18 citations
- Target Networks and Over-parameterization Stabilize Off-policy Bootstrapping with Function ApproximationFengdi Che, Chenjun Xiao, Jincheng Mei, Bo Dai et al.ICML 2024 · 7 citations
- Bayesian Exploration NetworksMattie Fellows, Brandon Kaplowitz, Christian Schröder de Witt, Shimon WhitesonICML 2024 · 4 citations
- A Unifying View of Linear Function Approximation in Off-Policy RL Through Matrix Splitting and PreconditioningZechen Wu, Amy Greenwald, Ronald E. ParrNeurIPS 2025 · 4 citations
- Simplifying Deep Temporal Difference LearningMatteo Gallici, Mattie Fellows, Benjamin Ellis, Bartomeu Pou et al.ICLR 2025 · 1 citation
Builds on6
- What are the Statistical Limits of Offline RL with Linear Function Approximation?Ruosong Wang, Dean P. Foster, Sham M. KakadeICLR 2021 · 172 citations
- Breaking the Deadly Triad with a Target NetworkShangtong Zhang, Hengshuai Yao, Shimon WhitesonICML 2021 · 61 citations
- Gradient Temporal-Difference Learning with Regularized CorrectionsSina Ghiassian, Andrew Patterson, Shivam Garg, Dhawal Gupta et al.ICML 2020 · 49 citations
- Instabilities of Offline RL with Pre-Trained Neural RepresentationRuosong Wang, Yifan Wu, Ruslan Salakhutdinov, Sham M. KakadeICML 2021 · 46 citations
- A new convergent variant of Q-learning with linear function approximationDiogo S. Carvalho, Francisco S. Melo, Pedro SantosNeurIPS 2020 · 39 citations
Related papers
- The Pitfalls of Regularization in Off-Policy TD LearningGaurav Manek, J. Zico KolterNeurIPS 2022 · 7 citations
- Average-Reward Off-Policy Policy Evaluation with Function ApproximationShangtong Zhang, Yi Wan, Richard S. Sutton, Shimon WhitesonICML 2021 · 39 citations
- Revisiting a Design Choice in Gradient Temporal Difference LearningXiaochi Qian, Shangtong ZhangICLR 2025
- Emphatic Algorithms for Deep Reinforcement LearningRay Jiang, Tom Zahavy, Zhongwen Xu, Adam White et al.ICML 2021 · 22 citations
- Fixed-Horizon Temporal Difference Methods for Stable Reinforcement LearningKristopher De Asis, Alan Chan, Silviu Pitis, Richard S. Sutton et al.AAAI 2020 · 34 citations
