How I Learned to Stop Worrying and Love Retraining
Max Zimmer, Christoph Spiegel, Sebastian Pokutta
Abstract
Many Neural Network Pruning approaches consist of several iterative training and pruning steps, seemingly losing a significant amount of their performance after pruning and then recovering it in the subsequent retraining phase. Recent works of Renda et al. (2020) and Le & Hua (2021) demonstrate the significance of the learning rate schedule during the retraining phase and propose specific heuristics for choosing such a schedule for IMP (Han et al., 2015). We place these findings in the context of the results of Li et al. (2020) regarding the training of models within a fixed training budget and demonstrate that, consequently, the retraining phase can be massively shortened using a simple linear learning rate schedule. Improving on existing retraining approaches, we additionally propose a method to adaptively select the initial value of the linear schedule. Going a step further, we propose similarly imposing a budget on the initial dense training phase and show that the resulting simple and efficient method is capable of outperforming significantly more complex or heavily parameterized state-of-the-art approaches that attempt to sparsify the network during training. These findings not only advance our understanding of the retraining phase, but more broadly question the belief that one should aim to avoid the need for retraining and reduce the negative effects of 'hard' pruning by incorporating the sparsification process into the standard training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1966f6bd-03ee-4979-a46b-d6d04ba2ea30Cited by top-tier papers3
- Estimating Canopy Height at ScaleJan Pauls, Max Zimmer, Una M. Kelly, Martin Schwartz et al.ICML 2024 · 27 citations
- Sparse Model Soups: A Recipe for Improved Pruning via Model AveragingMax Zimmer, Christoph Spiegel, Sebastian PokuttaICLR 2024 · 22 citations
- Capturing Temporal Dynamics in Large-Scale Canopy Tree Height EstimationJan Pauls, Max Zimmer, Berkant Turan, Sassan Saatchi et al.ICML 2025
Builds on11
- Pruning neural networks without any data by iteratively conserving synaptic flowHidenori Tanaka, Daniel Kunin, Daniel L. K. Yamins, Surya GanguliNeurIPS 2020 · 884 citations
- Rigging the Lottery: Making All Tickets WinnersUtku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro et al.ICML 2020 · 723 citations
- Comparing Rewinding and Fine-tuning in Neural Network PruningAlex Renda, Jonathan Frankle, Michael CarbinICLR 2020 · 437 citations
- Layer-adaptive Sparsity for the Magnitude-based PruningJaeho Lee, Sejun Park, Sangwoo Mo, Sungsoo Ahn et al.ICLR 2021 · 331 citations
- Soft Threshold Weight Reparameterization for Learnable SparsityAditya Kusupati, Vivek Ramanujan, Raghav Somani, Mitchell Wortsman et al.ICML 2020 · 266 citations
Related papers
- Network Pruning That Matters: A Case Study on Retraining VariantsDuong H. Le, Binh-Son HuaICLR 2021 · 45 citations
- Lottery Tickets on a Data Diet: Finding Initializations with Sparse Trainable NetworksMansheej Paul, Brett W. Larsen, Surya Ganguli, Jonathan Frankle et al.NeurIPS 2022 · 27 citations
- A Unified Framework for Soft Threshold PruningYanqi Chen, Zhengyu Ma, Wei Fang, Xiawu Zheng et al.ICLR 2023 · 6 citations
- Pruning from ScratchYulong Wang, Xiaolu Zhang, Lingxi Xie, Jun Zhou et al.AAAI 2020 · 219 citations
- Budgeted Training: Rethinking Deep Neural Network Training Under Resource ConstraintsMengtian Li, Ersin Yumer, Deva RamananICLR 2020 · 58 citations
