The Pitfalls of Regularization in Off-Policy TD Learning
Gaurav Manek, J. Zico Kolter
Abstract
Temporal Difference (TD) learning is ubiquitous in reinforcement learning, where it is often combined with off-policy sampling and function approximation. Unfortunately learning with this combination (known as the deadly triad ), exhibits instability and unbounded error. To account for this, modern RL methods often implicitly (or sometimes explicitly) assume that regularization is sufficient to mitigate the problem in practice; indeed, the standard deadly triad examples from the literature can be “fixed” via proper regularization. In this paper, we introduce a series of new counterexamples to show that the instability and unbounded error of TD methods is not solved by regularization. We demonstrate that, in the off-policy setting with linear function approximation, TD methods can fail to learn a non-trivial value function under any amount of regularization; we further show that regularization can induce divergence under common conditions; and we show that one of the most promising methods to mitigate this divergence (Emphatic TD algorithms) may also diverge under regularization. We further demonstrate such divergence when using neural networks as function approximators. Thus, we argue that the role of regularization in TD methods needs to be reconsidered, given that it is insufficient to prevent divergence and may itself introduce instability. There needs to be much more care in the practical and theoretical application of regularization to RL methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fc67cdcf-286b-4bd8-8c2a-02d9a8335217Cited by top-tier papers4
- TD Convergence: An Optimization PerspectiveKavosh Asadi, Shoham Sabach, Yao Liu, Omer Gottesman et al.NeurIPS 2023 · 17 citations
- The Statistical Benefits of Quantile Temporal-Difference Learning for Value EstimationMark Rowland, Yunhao Tang, Clare Lyle, Rémi Munos et al.ICML 2023 · 13 citations
- Target Networks and Over-parameterization Stabilize Off-policy Bootstrapping with Function ApproximationFengdi Che, Chenjun Xiao, Jincheng Mei, Bo Dai et al.ICML 2024 · 7 citations
- Regularized Q-LearningHan-Dong Lim, Donghwan LeeNeurIPS 2024 · 1 citation
Builds on8
- Revisiting Fundamentals of Experience ReplayWilliam Fedus, Prajit Ramachandran, Rishabh Agarwal, Yoshua Bengio et al.ICML 2020 · 303 citations
- DisCor: Corrective Feedback in Reinforcement Learning via Distribution CorrectionAviral Kumar, Abhishek Gupta, Sergey LevineNeurIPS 2020 · 124 citations
- DR3: Value-Based Deep Reinforcement Learning Requires Explicit RegularizationAviral Kumar, Rishabh Agarwal, Tengyu Ma, Aaron C. Courville et al.ICLR 2022 · 85 citations
- Breaking the Deadly Triad with a Target NetworkShangtong Zhang, Hengshuai Yao, Shimon WhitesonICML 2021 · 61 citations
- Provably Convergent Two-Timescale Off-Policy Actor-Critic with Function ApproximationShangtong Zhang, Bo Liu, Hengshuai Yao, Shimon WhitesonICML 2020 · 58 citations
Related papers
- Why Target Networks Stabilise Temporal Difference MethodsMattie Fellows, Matthew J. A. Smith, Shimon WhitesonICML 2023 · 10 citations
- Emphatic Algorithms for Deep Reinforcement LearningRay Jiang, Tom Zahavy, Zhongwen Xu, Adam White et al.ICML 2021 · 22 citations
- Revisiting a Design Choice in Gradient Temporal Difference LearningXiaochi Qian, Shangtong ZhangICLR 2025
- Average-Reward Off-Policy Policy Evaluation with Function ApproximationShangtong Zhang, Yi Wan, Richard S. Sutton, Shimon WhitesonICML 2021 · 39 citations
- Fixed-Horizon Temporal Difference Methods for Stable Reinforcement LearningKristopher De Asis, Alan Chan, Silviu Pitis, Richard S. Sutton et al.AAAI 2020 · 34 citations
