Network Pruning That Matters: A Case Study on Retraining Variants
Duong H. Le, Binh-Son Hua
Abstract
Network pruning is an effective method to reduce the computational expense of over-parameterized neural networks for deployment on low-resource systems. Recent state-of-the-art techniques for retraining pruned networks such as weight rewinding and learning rate rewinding have been shown to outperform the traditional fine-tuning technique in recovering the lost accuracy (Renda et al., 2020), but so far it is unclear what accounts for such performance. In this work, we conduct extensive experiments to verify and analyze the uncanny effectiveness of learning rate rewinding. We find that the reason behind the success of learning rate rewinding is the usage of a large learning rate. Similar phenomenon can be observed in other learning rate schedules that involve large learning rates, e.g., the 1-cycle learning rate schedule (Smith et al., 2019). By leveraging the right learning rate schedule in retraining, we demonstrate a counter-intuitive phenomenon in that randomly pruned networks could even achieve better performance than methodically pruned networks (fine-tuned with the conventional approach). Our results emphasize the cruciality of the learning rate schedule in pruned network retraining - a detail often overlooked by practitioners during the implementation of network pruning. One-sentence Summary: We study the effective of different retraining mechanisms while doing pruning
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 23eeec7d-ed83-4968-8714-4d056c68b1b2Cited by top-tier papers8
- Validating the Lottery Ticket Hypothesis with Inertial Manifold TheoryZeru Zhang, Jiayin Jin, Zijie Zhang, Yang Zhou et al.NeurIPS 2021 · 45 citations
- Sparse Double Descent: Where Network Pruning Aggravates OverfittingZheng He, Zeke Xie, Quanzhi Zhu, Zengchang QinICML 2022 · 36 citations
- Prune and Tune Ensembles: Low-Cost Ensemble Learning with Sparse Independent SubnetworksTim Whitaker, Darrell WhitleyAAAI 2022 · 28 citations
- Sparse Model Soups: A Recipe for Improved Pruning via Model AveragingMax Zimmer, Christoph Spiegel, Sebastian PokuttaICLR 2024 · 22 citations
- A Three-regime Model of Network PruningYefan Zhou, Yaoqing Yang, Arin Chang, Michael W. MahoneyICML 2023 · 15 citations
Builds on9
- Pruning neural networks without any data by iteratively conserving synaptic flowHidenori Tanaka, Daniel Kunin, Daniel L. K. Yamins, Surya GanguliNeurIPS 2020 · 884 citations
- Picking Winning Tickets Before Training by Preserving Gradient FlowChaoqi Wang, Guodong Zhang, Roger B. GrosseICLR 2020 · 743 citations
- Rigging the Lottery: Making All Tickets WinnersUtku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro et al.ICML 2020 · 723 citations
- Comparing Rewinding and Fine-tuning in Neural Network PruningAlex Renda, Jonathan Frankle, Michael CarbinICLR 2020 · 437 citations
- Soft Threshold Weight Reparameterization for Learnable SparsityAditya Kusupati, Vivek Ramanujan, Raghav Somani, Mitchell Wortsman et al.ICML 2020 · 266 citations
Related papers
- How I Learned to Stop Worrying and Love RetrainingMax Zimmer, Christoph Spiegel, Sebastian PokuttaICLR 2023
- Masks, Signs, And Learning Rate RewindingAdvait Harshal Gadhikar, Rebekka BurkholzICLR 2024 · 15 citations
- Good Subnetworks Provably Exist: Pruning via Greedy Forward SelectionMao Ye, Chengyue Gong, Lizhen Nie, Denny Zhou et al.ICML 2020 · 123 citations
- Learning with RetrospectionXiang Deng, Zhongfei ZhangAAAI 2021 · 20 citations
- Dynamic Model Pruning with FeedbackTao Lin, Sebastian U. Stich, Luis Barba, Daniil Dmitriev et al.ICLR 2020 · 229 citations
