On the Role of Optimization in Double Descent: A Least Squares Study
Ilja Kuzborskij, Csaba Szepesvári, Omar Rivasplata, Amal Rannen-Triki, Razvan Pascanu
Abstract
Empirically it has been observed that the performance of deep neural networks steadily improves as we increase model size, contradicting the classical view on overfitting and generalization. Recently, the double descent phenomena has been proposed to reconcile this observation with theory, suggesting that the test error has a second descent when the model becomes sufficiently overparametrized, as the model size itself acts as an implicit regularizer. In this paper we add to the growing body of work in this space, providing a careful study of learning dynamics as a function of model size for the least squares scenario. We show an excess risk bound for the gradient descent solution of the least squares objective. The bound depends on the smallest non-zero eigenvalue of the covariance matrix of the input features, via a functional form that has the double descent behaviour. This gives a new perspective on the double descent curves reported in the literature. Our analysis of the excess risk allows to decouple the effect of optimisation and generalisation error. In particular, we find that in case of noiseless regression, double descent is explained solely by optimisation-related quantities, which was missed in studies focusing on the Moore-Penrose pseudoinverse solution. We believe that our derivation provides an alternative view compared to existing work, shedding some light on a possible cause of this phenomena, at least in the considered least squares setting. We empirically explore if our predictions hold for neural networks, in particular whether the covariance of intermediary hidden activations has a similar behaviour as the one predicted by our derivations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d2bee26b-12f0-489f-b47b-2091993f7c76Cited by top-tier papers5
- A U-turn on Double Descent: Rethinking Parameter Counting in Statistical LearningAlicia Curth, Alan Jeffares, Mihaela van der SchaarNeurIPS 2023 · 42 citations
- Phenomenology of Double Descent in Finite-Width Neural NetworksSidak Pal Singh, Aurélien Lucchi, Thomas Hofmann, Bernhard SchölkopfICLR 2022 · 12 citations
- Benign Overfitting in Deep Neural Networks under Lazy TrainingZhenyu Zhu, Fanghui Liu, Grigorios Chrysos, Francesco Locatello et al.ICML 2023 · 12 citations
- On the Double Descent of Random Features Models Trained with SGDFanghui Liu, Johan A. K. Suykens, Volkan CevherNeurIPS 2022 · 11 citations
- On the role of overparameterization in off-policy Temporal Difference learning with linear function approximationValentin ThomasNeurIPS 2022 · 2 citations
Builds on6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang et al.ICLR 2020 · 1,108 citations
- High-Performance Large-Scale Image Recognition Without NormalizationAndy Brock, Soham De, Samuel L. Smith, Karen SimonyanICML 2021 · 613 citations
- Batch Normalization Biases Residual Blocks Towards the Identity Function in Deep NetworksSoham De, Samuel L. SmithNeurIPS 2020 · 173 citations
- Understanding Double Descent Requires A Fine-Grained Bias-Variance DecompositionBen Adlam, Jeffrey PenningtonNeurIPS 2020 · 111 citations
Related papers
- The Neural Tangent Kernel in High Dimensions: Triple Descent and a Multi-Scale Theory of GeneralizationBen Adlam, Jeffrey PenningtonICML 2020 · 133 citations
- Generalization of Two-layer Neural Networks: An Asymptotic ViewpointJimmy Ba, Murat A. Erdogdu, Taiji Suzuki, Denny Wu et al.ICLR 2020 · 77 citations
- Exact expressions for double descent and implicit regularization via surrogate random designMichal Derezinski, Feynman T. Liang, Michael W. MahoneyNeurIPS 2020 · 81 citations
- Double Trouble in Double Descent: Bias and Variance(s) in the Lazy RegimeStéphane d'Ascoli, Maria Refinetti, Giulio Biroli, Florent KrzakalaICML 2020 · 163 citations
- Multi-scale Feature Learning Dynamics: Insights for Double DescentMohammad Pezeshki, Amartya Mitra, Yoshua Bengio, Guillaume LajoieICML 2022 · 33 citations
