Sparse Double Descent: Where Network Pruning Aggravates Overfitting
Zheng He, Zeke Xie, Quanzhi Zhu, Zengchang Qin
摘要
People usually believe that network pruning not only reduces the computational cost of deep networks, but also prevents overfitting by decreasing model capacity. However, our work surprisingly discovers that network pruning sometimes even aggravates overfitting. We report an unexpected sparse double descent phenomenon that, as we increase model sparsity via network pruning, test performance first gets worse (due to overfitting), then gets better (due to relieved overfitting), and gets worse at last (due to forgetting useful information). While recent studies focused on the deep double descent with respect to model overparameterization, they failed to recognize that sparsity may also cause double descent. In this paper, we have three main contributions. First, we report the novel sparse double descent phenomenon through extensive experiments. Second, for this phenomenon, we propose a novel learning distance interpretation that the curve of learning distance of sparse models (from initialized parameters to final parameters) may correlate with the sparse double descent curve well and reflect generalization better than minima flatness. Third, in the context of sparse double descent, a winning ticket in the lottery ticket hypothesis surprisingly may not always win.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- A U-turn on Double Descent: Rethinking Parameter Counting in Statistical LearningAlicia Curth, Alan Jeffares, Mihaela van der SchaarNeurIPS 2023 · 被引用 42 次
- Deep Learning Through A Telescoping Lens: A Simple Model Provides Empirical Insights On Grokking, Gradient Boosting & BeyondAlan Jeffares, Alicia Curth, Mihaela van der SchaarNeurIPS 2024 · 被引用 11 次
- Dynamic Sparse Training via Balancing the Exploration-Exploitation Trade-offShaoyi Huang, Bowen Lei, Dongkuan Xu, Hongwu Peng 等DAC 2023 · 被引用 7 次
- DSD²: Can We Dodge Sparse Double Descent and Compress the Neural Network Worry-Free?Victor Quétu, Enzo TartaglioneAAAI 2024 · 被引用 6 次
- Calibrating the Rigged Lottery: Making All Tickets ReliableBowen Lei, Ruqi Zhang, Dongkuan Xu, Bani K. MallickICLR 2023 · 被引用 2 次
它引用的顶会 Paper22
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang 等ICLR 2020 · 被引用 1,108 次
- Linear Mode Connectivity and the Lottery Ticket HypothesisJonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICML 2020 · 被引用 750 次
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan 等ICLR 2020 · 被引用 705 次
- Comparing Rewinding and Fine-tuning in Neural Network PruningAlex Renda, Jonathan Frankle, Michael CarbinICLR 2020 · 被引用 437 次
相关 Paper
- Provable Benefits of Overparameterization in Model Compression: From Double Descent to Pruning Neural NetworksXiangyu Chang, Yingcong Li, Samet Oymak, Christos ThrampoulidisAAAI 2021 · 被引用 58 次
- Lottery Ticket Preserves Weight Correlation: Is It Desirable or Not?Ning Liu, Geng Yuan, Zhengping Che, Xuan Shen 等ICML 2021 · 被引用 34 次
- Plant 'n' Seek: Can You Find the Winning Ticket?Jonas Fischer, Rebekka BurkholzICLR 2022 · 被引用 21 次
- Dual Lottery Ticket HypothesisYue Bai, Huan Wang, Zhiqiang Tao, Kunpeng Li 等ICLR 2022 · 被引用 49 次
- Analyzing Lottery Ticket Hypothesis from PAC-Bayesian Theory PerspectiveKeitaro Sakamoto, Issei SatoNeurIPS 2022 · 被引用 11 次
