Early Stopping in Deep Networks: Double Descent and How to Eliminate it
Reinhard Heckel, Fatih Furkan Yilmaz
摘要
Over-parameterized models, such as large deep networks, often exhibit a double descent phenomenon, where as a function of model size, error first decreases, increases, and decreases at last. This intriguing double descent behavior also occurs as a function of training epochs, and has been conjectured to arise because training epochs control the model complexity. In this paper, we show that such epoch-wise double descent arises for a different reason: It is caused by a superposition of two or more bias-variance tradeoffs that arise because different parts of the network are learned at different epochs, and eliminating this by proper scaling of stepsizes can significantly improve the early stopping performance. We show this analytically for i) linear regression, where differently scaled features give rise to a superposition of bias-variance tradeoffs, and for ii) a two-layer neural network, where the first and second layer each govern a bias-variance tradeoff. Inspired by this theory, we study two standard convolutional networks empirically, and show that eliminating epoch-wise double descent through adjusting stepsizes of different layers improves the early stopping performance significantly.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Data Determines Distributional Robustness in Contrastive Language Image Pre-training (CLIP)Alex Fang, Gabriel Ilharco, Mitchell Wortsman, Yuhao Wan 等ICML 2022 · 被引用 183 次
- Multi-scale Feature Learning Dynamics: Insights for Double DescentMohammad Pezeshki, Amartya Mitra, Yoshua Bengio, Guillaume LajoieICML 2022 · 被引用 33 次
- Defects of Convolutional Decoder Networks in Frequency RepresentationLing Tang, Wen Shen, Zhanpeng Zhou, Yuefeng Chen 等ICML 2023 · 被引用 17 次
- Post-Hoc Reversal: Are We Selecting Models Prematurely?Rishabh Ranjan, Saurabh Garg, Mrigank Raman, Carlos Guestrin 等NeurIPS 2024 · 被引用 8 次
- Towards Theoretical Analysis of Transformation Complexity of ReLU DNNsJie Ren, Mingjie Li, Meng Zhou, Shih-Han Chan 等ICML 2022 · 被引用 4 次
它引用的顶会 Paper6
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang 等ICLR 2020 · 被引用 1,108 次
- Rethinking Bias-Variance Trade-off for Generalization of Neural NetworksZitong Yang, Yaodong Yu, Chong You, Jacob Steinhardt 等ICML 2020 · 被引用 219 次
- Double Trouble in Double Descent: Bias and Variance(s) in the Lazy RegimeStéphane d'Ascoli, Maria Refinetti, Giulio Biroli, Florent KrzakalaICML 2020 · 被引用 163 次
- Optimal Regularization can Mitigate Double DescentPreetum Nakkiran, Prayaag Venkat, Sham M. Kakade, Tengyu MaICLR 2021 · 被引用 148 次
- Denoising and Regularization via Exploiting the Structural Bias of Convolutional GeneratorsReinhard Heckel, Mahdi SoltanolkotabiICLR 2020 · 被引用 91 次
相关 Paper
- Model, sample, and epoch-wise descents: exact solution of gradient flow in the random feature modelAntoine Bodin, Nicolas MacrisNeurIPS 2021 · 被引用 19 次
- Least Squares Regression Can Exhibit Under-Parameterized Double DescentXinyue Li, Rishi SonthaliaNeurIPS 2024 · 被引用 5 次
- Provable Benefits of Overparameterization in Model Compression: From Double Descent to Pruning Neural NetworksXiangyu Chang, Yingcong Li, Samet Oymak, Christos ThrampoulidisAAAI 2021 · 被引用 58 次
- On the Role of Optimization in Double Descent: A Least Squares StudyIlja Kuzborskij, Csaba Szepesvári, Omar Rivasplata, Amal Rannen-Triki 等NeurIPS 2021 · 被引用 12 次
- United We Stand: Using Epoch-Wise Agreement of Ensembles to Combat OverfitUri Stern, Daniel Shwartz, Daphna WeinshallAAAI 2024
