Breaking the Deadly Triad with a Target Network
Shangtong Zhang, Hengshuai Yao, Shimon Whiteson
Abstract
The deadly triad refers to the instability of a reinforcement learning algorithm when it employs off-policy learning, function approximation, and bootstrapping simultaneously. In this paper, we investigate the target network as a tool for breaking the deadly triad, providing theoretical support for the conventional wisdom that a target network stabilizes training. We first propose and analyze a novel target network update rule which augments the commonly used Polyak-averaging style update with two projections. We then apply the target network and ridge regularization in several divergent algorithms and show their convergence to regularized TD fixed points. Those algorithms are off-policy with linear function approximation and bootstrapping, spanning both policy evaluation and control, as well as both discounted and average-reward settings. In particular, we provide the first convergent linear -learning algorithms under nonrestrictive and changing behavior policies without bi-level optimization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7531eb78-ecaa-43f5-a83d-d926245a60f8Cited by top-tier papers20
- Understanding, Predicting and Better Resolving Q-Value Divergence in Offline-RLYang Yue, Rui Lu, Bingyi Kang, Shiji Song et al.NeurIPS 2023 · 28 citations
- Improving Deep Reinforcement Learning by Reducing the Chain Effect of Value and Policy ChurnHongyao Tang, Glen BersethNeurIPS 2024 · 23 citations
- On the Convergence of SARSA with Linear Function ApproximationShangtong Zhang, Remi Tachet des Combes, Romain LarocheICML 2023 · 19 citations
- TD Convergence: An Optimization PerspectiveKavosh Asadi, Shoham Sabach, Yao Liu, Omer Gottesman et al.NeurIPS 2023 · 17 citations
- Off-Policy Average Reward Actor-Critic with Deterministic Policy SearchNaman Saxena, Subhojyoti Khastagir, Shishir Kolathaya, Shalabh BhatnagarICML 2023 · 14 citations
Builds on10
- Reinforcement Learning in Feature Space: Matrix Bandit, Kernels, and Regret BoundLin Yang, Mengdi WangICML 2020 · 308 citations
- A Finite-Time Analysis of Two Time-Scale Actor-Critic MethodsYue Wu, Weitong Zhang, Pan Xu, Quanquan GuNeurIPS 2020 · 189 citations
- Learning and Planning in Average-Reward Markov Decision ProcessesYi Wan, Abhishek Naik, Richard S. SuttonICML 2021 · 82 citations
- A Finite-Time Analysis of Q-Learning with Neural Network Function ApproximationPan Xu, Quanquan GuICML 2020 · 79 citations
- Provably Convergent Two-Timescale Off-Policy Actor-Critic with Function ApproximationShangtong Zhang, Bo Liu, Hengshuai Yao, Shimon WhitesonICML 2020 · 58 citations
Related papers
- Why Target Networks Stabilise Temporal Difference MethodsMattie Fellows, Matthew J. A. Smith, Shimon WhitesonICML 2023 · 10 citations
- The Pitfalls of Regularization in Off-Policy TD LearningGaurav Manek, J. Zico KolterNeurIPS 2022 · 7 citations
- Average-Reward Off-Policy Policy Evaluation with Function ApproximationShangtong Zhang, Yi Wan, Richard S. Sutton, Shimon WhitesonICML 2021 · 39 citations
- Fixed-Horizon Temporal Difference Methods for Stable Reinforcement LearningKristopher De Asis, Alan Chan, Silviu Pitis, Richard S. Sutton et al.AAAI 2020 · 34 citations
- Revisiting a Design Choice in Gradient Temporal Difference LearningXiaochi Qian, Shangtong ZhangICLR 2025
