Winning the Lottery Ahead of Time: Efficient Early Network Pruning
John Rachwan, Daniel Zügner, Bertrand Charpentier, Simon Geisler, Morgane Ayle, Stephan Günnemann
摘要
Pruning, the task of sparsifying deep neural networks, received increasing attention recently. Although state-of-the-art pruning methods extract highly sparse models, they neglect two main challenges: (1) the process of finding these sparse models is often very expensive; (2) unstructured pruning does not provide benefits in terms of GPU memory, training time, or carbon emissions. We propose Early Compression via Gradient Flow Preservation (EarlyCroP), which efficiently extracts state-of-the-art sparse models before or early in training addressing challenge (1), and can be applied in a structured manner addressing challenge (2). This enables us to train sparse networks on commodity GPUs whose dense versions would be too large, thereby saving costs and reducing hardware requirements. We empirically show that EarlyCroP outperforms a rich set of baselines for many tasks (incl. classification, regression) and domains (incl. computer vision, natural language processing, and reinforcment learning). Early-CroP leads to accuracy comparable to dense training while outperforming pruning baselines.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Shaving Weights with Occam's Razor: Bayesian Sparsification for Neural Networks using the Marginal LikelihoodRayen Dhahri, Alexander Immer, Bertrand Charpentier, Stephan Günnemann 等NeurIPS 2024 · 被引用 10 次
- Train Faster, Perform Better: Modular Adaptive Training in Over-Parameterized ModelsYubin Shi, Yixuan Chen, Mingzhi Dong, Xiaochen Yang 等NeurIPS 2023 · 被引用 5 次
- Pruning via Sparsity-indexed ODE: a Continuous Sparsity ViewpointZhanfeng Mo, Haosen Shi, Sinno Jialin PanICML 2023 · 被引用 5 次
- Towards a Smaller Student: Capacity Dynamic Distillation for Efficient Image RetrievalYi Xie, Huaidong Zhang, Xuemiao Xu, Jianqing Zhu 等CVPR 2023
- F^3OCUS - Federated Finetuning of Vision-Language Foundation Models with Optimal Client Layer Updating Strategy via Multi-objective Meta-HeuristicsPramit Saha, Felix Wagner, Divyanshu Mishra, Can Peng 等CVPR 2025
它引用的顶会 Paper10
- Pruning neural networks without any data by iteratively conserving synaptic flowHidenori Tanaka, Daniel Kunin, Daniel L. K. Yamins, Surya GanguliNeurIPS 2020 · 被引用 884 次
- Linear Mode Connectivity and the Lottery Ticket HypothesisJonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICML 2020 · 被引用 750 次
- Picking Winning Tickets Before Training by Preserving Gradient FlowChaoqi Wang, Guodong Zhang, Roger B. GrosseICLR 2020 · 被引用 743 次
- Rigging the Lottery: Making All Tickets WinnersUtku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro 等ICML 2020 · 被引用 723 次
- The Lottery Ticket Hypothesis for Pre-trained BERT NetworksTianlong Chen, Jonathan Frankle, Shiyu Chang, Sijia Liu 等NeurIPS 2020 · 被引用 428 次
相关 Paper
- When to Prune? A Policy towards Early Structural PruningMaying Shen, Pavlo Molchanov, Hongxu Yin, José M. ÁlvarezCVPR 2022 · 被引用 46 次
- Pruning from ScratchYulong Wang, Xiaolu Zhang, Lingxi Xie, Jun Zhou 等AAAI 2020 · 被引用 219 次
- PacTrain: Pruning and Adaptive Sparse Gradient Compression for Efficient Collective Communication in Distributed Deep LearningYisu Wang, Ruilong Wu, Xinjiao Li, Dirk KutscherDAC 2025 · 被引用 3 次
- Growing Efficient Deep Networks by Structured Continuous SparsificationXin Yuan, Pedro Henrique Pamplona Savarese, Michael MaireICLR 2021 · 被引用 51 次
- Effective Model Sparsification by Scheduled Grow-and-Prune MethodsXiaolong Ma, Minghai Qin, Fei Sun, Zejiang Hou 等ICLR 2022 · 被引用 45 次
