DSD²: Can We Dodge Sparse Double Descent and Compress the Neural Network Worry-Free?
Victor Quétu, Enzo Tartaglione
Abstract
Neoteric works have shown that modern deep learning models can exhibit a sparse double descent phenomenon. Indeed, as the sparsity of the model increases, the test performance first worsens since the model is overfitting the training data; then, the overfitting reduces, leading to an improvement in performance, and finally, the model begins to forget critical information, resulting in underfitting. Such a behavior prevents using traditional early stop criteria.
In this work, we have three key contributions. First, we propose a learning framework that avoids such a phenomenon and improves generalization. Second, we introduce an entropy measure providing more insights into the insurgence of this phenomenon and enabling the use of traditional stop criteria. Third, we provide a comprehensive quantitative analysis of contingent factors such as re-initialization methods, model width and depth, and dataset noise. The contributions are supported by empirical evidence in typical setups. Our code is available at https://github.com/VGCQ/DSD2.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang et al.ICLR 2020 · 1,108 citations
- Early-Learning Regularization Prevents Memorization of Noisy LabelsSheng Liu, Jonathan Niles-Weed, Narges Razavian, Carlos Fernandez-GrandaNeurIPS 2020 · 798 citations
- On-Device Training Under 256KB MemoryJi Lin, Ligeng Zhu, Wei-Ming Chen, Wei-Chen Wang et al.NeurIPS 2022 · 345 citations
Related papers
- Sparse Double Descent: Where Network Pruning Aggravates OverfittingZheng He, Zeke Xie, Quanzhi Zhu, Zengchang QinICML 2022 · 36 citations
- Double Trouble in Double Descent: Bias and Variance(s) in the Lazy RegimeStéphane d'Ascoli, Maria Refinetti, Giulio Biroli, Florent KrzakalaICML 2020 · 163 citations
- On the Role of Optimization in Double Descent: A Least Squares StudyIlja Kuzborskij, Csaba Szepesvári, Omar Rivasplata, Amal Rannen-Triki et al.NeurIPS 2021 · 12 citations
- Multi-scale Feature Learning Dynamics: Insights for Double DescentMohammad Pezeshki, Amartya Mitra, Yoshua Bengio, Guillaume LajoieICML 2022 · 33 citations
- The Neural Tangent Kernel in High Dimensions: Triple Descent and a Multi-Scale Theory of GeneralizationBen Adlam, Jeffrey PenningtonICML 2020 · 133 citations
