On Double Descent in Reinforcement Learning with LSTD and Random Features
David Brellmann, Eloïse Berthier, David Filliat, Goran Frehse
Abstract
Temporal Difference (TD) algorithms are widely used in Deep Reinforcement Learning (RL). Their performance is heavily influenced by the size of the neural network. While in supervised learning, the regime of over-parameterization and its benefits are well understood, the situation in RL is much less clear. In this paper, we present a theoretical analysis of the influence of network size and l 2 -regularization on performance. We identify the ratio between the number of parameters and the number of visited states as a crucial factor and define overparameterization as the regime when it is larger than one. Furthermore, we observe a double descent phenomenon, i.e., a sudden drop in performance around the parameter/state ratio of one. Leveraging random features and the lazy training regime, we study the regularized Least-Squared Temporal Difference (LSTD) algorithm in an asymptotic regime, as both the number of parameters and states go to infinity, maintaining a constant ratio. We derive deterministic limits of both the empirical and the true Mean-Squared Bellman Error (MSBE) that feature correction terms responsible for the double descent. Correction terms vanish when the l 2 -regularization is increased or the number of unvisited states goes to zero. Numerical experiments with synthetic and small real-world environments closely match the theoretical predictions. Lénaïc Chizat, Edouard Oyallon, and Francis Bach. On lazy training in differentiable programming. Advances in Neural Information Processing Systems, 32, 2019. Romain Couillet and Merouane Debbah. Random Matrix Methods for Wireless Communications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 02ce58a6-935b-449b-959b-c1bc81f54a6dCited by top-tier papers1
Ask how each one uses itBuilds on7
- Implicit Under-Parameterization Inhibits Data-Efficient Deep Reinforcement LearningAviral Kumar, Rishabh Agarwal, Dibya Ghosh, Sergey LevineICLR 2021 · 155 citations
- Understanding and Preventing Capacity Loss in Reinforcement LearningClare Lyle, Mark Rowland, Will DabneyICLR 2022 · 151 citations
- A random matrix analysis of random Fourier features: beyond the Gaussian kernel, a precise phase transition, and the corresponding double descentZhenyu Liao, Romain Couillet, Michael W. MahoneyNeurIPS 2020 · 102 citations
- Understanding and Leveraging Overparameterization in Recursive Value EstimationChenjun Xiao, Bo Dai, Jincheng Mei, Oscar A. Ramirez et al.ICLR 2022 · 17 citations
- Can Temporal-Difference and Q-Learning Learn Representation? A Mean-Field TheoryYufeng Zhang, Qi Cai, Zhuoran Yang, Yongxin Chen et al.NeurIPS 2020 · 12 citations
Related papers
- Double Trouble in Double Descent: Bias and Variance(s) in the Lazy RegimeStéphane d'Ascoli, Maria Refinetti, Giulio Biroli, Florent KrzakalaICML 2020 · 163 citations
- On the role of overparameterization in off-policy Temporal Difference learning with linear function approximationValentin ThomasNeurIPS 2022 · 2 citations
- On the Role of Optimization in Double Descent: A Least Squares StudyIlja Kuzborskij, Csaba Szepesvári, Omar Rivasplata, Amal Rannen-Triki et al.NeurIPS 2021 · 12 citations
- Exact expressions for double descent and implicit regularization via surrogate random designMichal Derezinski, Feynman T. Liang, Michael W. MahoneyNeurIPS 2020 · 81 citations
- DR3: Value-Based Deep Reinforcement Learning Requires Explicit RegularizationAviral Kumar, Rishabh Agarwal, Tengyu Ma, Aaron C. Courville et al.ICLR 2022 · 85 citations
