A Three-regime Model of Network Pruning
Yefan Zhou, Yaoqing Yang, Arin Chang, Michael W. Mahoney
Abstract
Recent work has highlighted the complex influence training hyperparameters, e.g., the number of training epochs, can have on the prunability of machine learning models. Perhaps surprisingly, a systematic approach to predict precisely how adjusting a specific hyperparameter will affect prunability remains elusive. To address this gap, we introduce a phenomenological model grounded in the statistical mechanics of learning. Our approach uses temperature-like and load-like parameters to model the impact of neural network (NN) training hyperparameters on pruning performance. A key empirical result we identify is a sharp transition phenomenon: depending on the value of a load-like parameter in the pruned model, increasing the value of a temperature-like parameter in the pre-pruned model may either enhance or impair subsequent pruning performance. Based on this transition, we build a three-regime model by taxonomizing the global structure of the pruned NN loss landscape. Our model reveals that the dichotomous effect of high temperature is associated with transitions between distinct types of global structures in the post-pruned model. Based on our results, we present three case-studies: 1) determining whether to increase or decrease a hyperparameter for improved pruning; 2) selecting the best model to prune from a family of models; and 3) tuning the hyperparameter of the Sharpness Aware Minimization method for better pruning performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 711397dd-a79a-4e98-b6bc-32a74c04e098Cited by top-tier papers7
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 794 citations
- AutoVP: An Automated Visual Prompting Framework and BenchmarkHsi-Ai Tsao, Lei Hsiung, Pin-Yu Chen, Si Liu et al.ICLR 2024 · 29 citations
- Temperature Balancing, Layer-wise Weight Analysis, and Neural Network TrainingYefan Zhou, Tianyu Pang, Keqin Liu, Charles H. Martin et al.NeurIPS 2023 · 29 citations
- Adaptive Sharpness-Aware Pruning for Robust Sparse NetworksAnna Bair, Hongxu Yin, Maying Shen, Pavlo Molchanov et al.ICLR 2024 · 19 citations
- MD tree: a model-diagnostic tree grown on loss landscapeYefan Zhou, Jianlong Chen, Qinxue Cao, Konstantin Schürholt et al.ICML 2024 · 2 citations
Builds on12
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- Linear Mode Connectivity and the Lottery Ticket HypothesisJonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICML 2020 · 750 citations
- Beyond neural scaling laws: beating power law scaling via data pruningBen Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli et al.NeurIPS 2022 · 720 citations
- Comparing Rewinding and Fine-tuning in Neural Network PruningAlex Renda, Jonathan Frankle, Michael CarbinICLR 2020 · 437 citations
- Layer-adaptive Sparsity for the Magnitude-based PruningJaeho Lee, Sejun Park, Sangwoo Mo, Sungsoo Ahn et al.ICLR 2021 · 331 citations
Related papers
- Taxonomizing local versus global structure in neural network loss landscapesYaoqing Yang, Liam Hodgkinson, Ryan Theisen, Joe Zou et al.NeurIPS 2021 · 51 citations
- The Generalization-Stability Tradeoff In Neural Network PruningBrian R. Bartoldson, Ari S. Morcos, Adrian Barbu, Gordon ErlebacherNeurIPS 2020 · 97 citations
- Probabilistic Neural Pruning via Sparsity Evolutionary Fokker-Planck-Kolmogorov EquationZhanfeng Mo, Haosen Shi, Sinno Jialin PanICLR 2025
- Pruning's Effect on Generalization Through the Lens of Training and RegularizationTian Jin, Michael Carbin, Daniel M. Roy, Jonathan Frankle et al.NeurIPS 2022 · 40 citations
- Dynamic Sparse Training: Find Efficient Sparse Network From Scratch With Trainable Masked LayersJunjie Liu, Zhe Xu, Runbin Shi, Ray C. C. Cheung et al.ICLR 2020 · 136 citations
