When to Prune? A Policy towards Early Structural Pruning
Maying Shen, Pavlo Molchanov, Hongxu Yin, José M. Álvarez
摘要
Pruning enables appealing reductions in network memory footprint and time complexity. Conventional post-training pruning techniques lean towards efficient inference while overlooking the heavy computation for training. Recent exploration of pre-training pruning at initialization hints on training cost reduction via pruning, but suffers noticeable performance degradation. We attempt to combine the benefits of both directions and propose a policy that prunes as early as possible during training without hurting performance. Instead of pruning at initialization, our method exploits initial dense training for few epochs to quickly guide the architecture, while constantly evaluating dominant sub-networks via neuron importance ranking. This unveils dominant sub-networks whose structures turn stable, allowing conventional pruning to be pushed earlier into the training. To do this early, we further introduce an Early Pruning Indicator (EPI) that relies on sub-network architectural similarity and quickly triggers pruning when the sub-network's architecture stabilizes. Through extensive experiments on ImageNet, we show that EPI empowers a quick tracking of early training epochs suitable for pruning, offering same efficacy as an otherwise “oracle” grid-search that scans through epochs and requires orders of magnitude more compute. Our method yields 1.4% top-l accuracy boost over state-of-the-art pruning counterparts, cuts down training cost on GPU by 2.4x, hence offers a new efficiency-accuracy boundary for network pruning during training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Structural Pruning via Latency-Saliency KnapsackMaying Shen, Hongxu Yin, Pavlo Molchanov, Lei Mao 等NeurIPS 2022 · 被引用 70 次
- APT: Adaptive Pruning and Tuning Pretrained Language Models for Efficient Training and InferenceBowen Zhao, Hannaneh Hajishirzi, Qingqing CaoICML 2024 · 被引用 31 次
- Adaptive Sharpness-Aware Pruning for Robust Sparse NetworksAnna Bair, Hongxu Yin, Maying Shen, Pavlo Molchanov 等ICLR 2024 · 被引用 19 次
- A Three-regime Model of Network PruningYefan Zhou, Yaoqing Yang, Arin Chang, Michael W. MahoneyICML 2023 · 被引用 15 次
- Towards Fairness-aware Adversarial Network PruningLei Zhang, Zhibo Wang, Xiaowei Dong, Yunhe Feng 等ICCV 2023 · 被引用 8 次
它引用的顶会 Paper12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Pruning neural networks without any data by iteratively conserving synaptic flowHidenori Tanaka, Daniel Kunin, Daniel L. K. Yamins, Surya GanguliNeurIPS 2020 · 被引用 884 次
- Linear Mode Connectivity and the Lottery Ticket HypothesisJonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICML 2020 · 被引用 750 次
- Picking Winning Tickets Before Training by Preserving Gradient FlowChaoqi Wang, Guodong Zhang, Roger B. GrosseICLR 2020 · 被引用 743 次
- The Early Phase of Neural Network TrainingJonathan Frankle, David J. Schwab, Ari S. MorcosICLR 2020 · 被引用 199 次
相关 Paper
- Towards Data-Agnostic Pruning At Initialization: What Makes a Good Sparse Mask?Hoang Pham, The-Anh Ta, Shiwei Liu, Lichuan Xiang 等NeurIPS 2023 · 被引用 16 次
- Winning the Lottery Ahead of Time: Efficient Early Network PruningJohn Rachwan, Daniel Zügner, Bertrand Charpentier, Simon Geisler 等ICML 2022 · 被引用 34 次
- Pruning from ScratchYulong Wang, Xiaolu Zhang, Lingxi Xie, Jun Zhou 等AAAI 2020 · 被引用 219 次
- Lottery Tickets on a Data Diet: Finding Initializations with Sparse Trainable NetworksMansheej Paul, Brett W. Larsen, Surya Ganguli, Jonathan Frankle 等NeurIPS 2022 · 被引用 27 次
- Good Subnetworks Provably Exist: Pruning via Greedy Forward SelectionMao Ye, Chengyue Gong, Lizhen Nie, Denny Zhou 等ICML 2020 · 被引用 123 次
