The Generalization-Stability Tradeoff In Neural Network Pruning
Brian R. Bartoldson, Ari S. Morcos, Adrian Barbu, Gordon Erlebacher
摘要
Pruning neural network parameters is often viewed as a means to compress models, but pruning has also been motivated by the desire to prevent overfitting. This motivation is particularly relevant given the perhaps surprising observation that a wide variety of pruning approaches increase test accuracy despite sometimes massive reductions in parameter counts. To better understand this phenomenon, we analyze the behavior of pruning over the course of training, finding that pruning's benefit to generalization increases with pruning's instability (defined as the drop in test accuracy immediately following pruning). We demonstrate that this "generalization-stability tradeoff" is present across a wide variety of pruning settings and propose a mechanism for its cause: pruning regularizes similarly to noise injection. Supporting this, we find less pruning stability leads to more model flatness and the benefits of pruning do not depend on permanent parameter removal. These results explain the compatibility of pruning-based generalization improvements and the high generalization recently observed in overparameterized networks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper27
- Chasing Sparsity in Vision Transformers: An End-to-End ExplorationTianlong Chen, Yu Cheng, Zhe Gan, Lu Yuan 等NeurIPS 2021 · 被引用 295 次
- Sparse Training via Boosting Pruning Plasticity with NeuroregenerationShiwei Liu, Tianlong Chen, Xiaohan Chen, Zahra Atashgahi 等NeurIPS 2021 · 被引用 145 次
- Coarsening the Granularity: Towards Structurally Sparse Lottery TicketsTianlong Chen, Xuxi Chen, Xiaolong Ma, Yanzhi Wang 等ICML 2022 · 被引用 42 次
- Pruning's Effect on Generalization Through the Lens of Training and RegularizationTian Jin, Michael Carbin, Daniel M. Roy, Jonathan Frankle 等NeurIPS 2022 · 被引用 40 次
- Sparse Double Descent: Where Network Pruning Aggravates OverfittingZheng He, Zeke Xie, Quanzhi Zhu, Zengchang QinICML 2022 · 被引用 36 次
它引用的顶会 Paper2
相关 Paper
- Pruning from ScratchYulong Wang, Xiaolu Zhang, Lingxi Xie, Jun Zhou 等AAAI 2020 · 被引用 219 次
- Provable Benefits of Overparameterization in Model Compression: From Double Descent to Pruning Neural NetworksXiangyu Chang, Yingcong Li, Samet Oymak, Christos ThrampoulidisAAAI 2021 · 被引用 58 次
- Picking Winning Tickets Before Training by Preserving Gradient FlowChaoqi Wang, Guodong Zhang, Roger B. GrosseICLR 2020 · 被引用 743 次
- Noise Injection Node Regularization for Robust LearningNoam Levi, Itay M. Bloch, Marat Freytsis, Tomer VolanskyICLR 2023 · 被引用 2 次
- Over-parameterized Model Optimization with Polyak-Łojasiewicz ConditionYixuan Chen, Yubin Shi, Mingzhi Dong, Xiaochen Yang 等ICLR 2023
