Winning the Lottery Ahead of Time: Efficient Early Network Pruning
John Rachwan, Daniel Zügner, Bertrand Charpentier, Simon Geisler, Morgane Ayle, Stephan Günnemann
Abstract
Pruning, the task of sparsifying deep neural networks, received increasing attention recently. Although state-of-the-art pruning methods extract highly sparse models, they neglect two main challenges: (1) the process of finding these sparse models is often very expensive; (2) unstructured pruning does not provide benefits in terms of GPU memory, training time, or carbon emissions. We propose Early Compression via Gradient Flow Preservation (EarlyCroP), which efficiently extracts state-of-the-art sparse models before or early in training addressing challenge (1), and can be applied in a structured manner addressing challenge (2). This enables us to train sparse networks on commodity GPUs whose dense versions would be too large, thereby saving costs and reducing hardware requirements. We empirically show that EarlyCroP outperforms a rich set of baselines for many tasks (incl. classification, regression) and domains (incl. computer vision, natural language processing, and reinforcment learning). Early-CroP leads to accuracy comparable to dense training while outperforming pruning baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d73d4476-2e7f-4d05-9ebb-3424d921e993Cited by top-tier papers7
- Shaving Weights with Occam's Razor: Bayesian Sparsification for Neural Networks using the Marginal LikelihoodRayen Dhahri, Alexander Immer, Bertrand Charpentier, Stephan Günnemann et al.NeurIPS 2024 · 10 citations
- Train Faster, Perform Better: Modular Adaptive Training in Over-Parameterized ModelsYubin Shi, Yixuan Chen, Mingzhi Dong, Xiaochen Yang et al.NeurIPS 2023 · 5 citations
- Pruning via Sparsity-indexed ODE: a Continuous Sparsity ViewpointZhanfeng Mo, Haosen Shi, Sinno Jialin PanICML 2023 · 5 citations
- Towards a Smaller Student: Capacity Dynamic Distillation for Efficient Image RetrievalYi Xie, Huaidong Zhang, Xuemiao Xu, Jianqing Zhu et al.CVPR 2023
- F^3OCUS - Federated Finetuning of Vision-Language Foundation Models with Optimal Client Layer Updating Strategy via Multi-objective Meta-HeuristicsPramit Saha, Felix Wagner, Divyanshu Mishra, Can Peng et al.CVPR 2025
Builds on10
- Pruning neural networks without any data by iteratively conserving synaptic flowHidenori Tanaka, Daniel Kunin, Daniel L. K. Yamins, Surya GanguliNeurIPS 2020 · 884 citations
- Linear Mode Connectivity and the Lottery Ticket HypothesisJonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICML 2020 · 750 citations
- Picking Winning Tickets Before Training by Preserving Gradient FlowChaoqi Wang, Guodong Zhang, Roger B. GrosseICLR 2020 · 743 citations
- Rigging the Lottery: Making All Tickets WinnersUtku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro et al.ICML 2020 · 723 citations
- The Lottery Ticket Hypothesis for Pre-trained BERT NetworksTianlong Chen, Jonathan Frankle, Shiyu Chang, Sijia Liu et al.NeurIPS 2020 · 428 citations
Related papers
- When to Prune? A Policy towards Early Structural PruningMaying Shen, Pavlo Molchanov, Hongxu Yin, José M. ÁlvarezCVPR 2022 · 46 citations
- Pruning from ScratchYulong Wang, Xiaolu Zhang, Lingxi Xie, Jun Zhou et al.AAAI 2020 · 219 citations
- PacTrain: Pruning and Adaptive Sparse Gradient Compression for Efficient Collective Communication in Distributed Deep LearningYisu Wang, Ruilong Wu, Xinjiao Li, Dirk KutscherDAC 2025 · 3 citations
- Growing Efficient Deep Networks by Structured Continuous SparsificationXin Yuan, Pedro Henrique Pamplona Savarese, Michael MaireICLR 2021 · 51 citations
- Effective Model Sparsification by Scheduled Grow-and-Prune MethodsXiaolong Ma, Minghai Qin, Fei Sun, Zejiang Hou et al.ICLR 2022 · 45 citations
