Progressive Skeletonization: Trimming more fat from a network at initialization
Pau de Jorge, Amartya Sanyal, Harkirat S. Behl, Philip H. S. Torr, Grégory Rogez, Puneet K. Dokania
Abstract
Recent studies have shown that skeletonization (pruning parameters) of networks at initialization provides all the practical benefits of sparsity both at inference and training time, while only marginally degrading their performance. However, we observe that beyond a certain level of sparsity (approx 95%), these approaches fail to preserve the network performance, and to our surprise, in many cases perform even worse than trivial random pruning. To this end, we propose an objective to find a skeletonized network with maximum foresight connection sensitivity (FORCE) whereby the trainability, in terms of connection sensitivity, of a pruned network is taken into consideration. We then propose two approximate procedures to maximize our objective (1) Iterative SNIP: allows parameters that were unimportant at earlier stages of skeletonization to become important at later stages; and (2) FORCE: iterative process that allows exploration by allowing already pruned parameters to resurrect at later stages of skeletonization. Empirical analysis on a large suite of experiments show that our approach, while providing at least as good a performance as other recent approaches on moderate pruning levels, provide remarkably improved performance on higher pruning levels (could remove up to 99.5% parameters while keeping the networks trainable). INTRODUCTION The majority of pruning algorithms for Deep Neural Networks require training dense models and often fine-tuning sparse sub-networks in order to obtain their pruned counterparts. In Frankle & Carbin (2019), the authors provide empirical evidence to support the hypothesis that there exist sparse sub-networks that can be trained from scratch to achieve similar performance as the dense ones. However, their method to find such sub-networks requires training the full-sized model and intermediate sub-networks, making the process much more expensive. Recently, Lee et al. ( 2019 ) presented SNIP. Building upon almost a three decades old saliency criterion for pruning trained models (Mozer & Smolensky, 1989) , they are able to predict, at initialization, the importance each weight will have later in training. Pruning at initialization methods are much cheaper than conventional pruning methods. Moreover, while traditional pruning methods can help accelerate inference tasks, pruning at initialization may go one step further and provide the same benefits at train time Elsen et al. (2020) . Wang et al. (2020) (GRASP) noted that after applying the pruning mask, gradients are modified due to non-trivial interactions between weights. Thus, maximizing SNIP criterion before pruning might be sub-optimal. They present an approximation to maximize the gradient norm after pruning, where they treat pruning as a perturbation on the weight matrix and use the first order Taylor's approximation. While they show improved performance, their approximation involves computing a Hessian-vector product which is expensive both in terms of memory and computation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers37
- Pruning neural networks without any data by iteratively conserving synaptic flowHidenori Tanaka, Daniel Kunin, Daniel L. K. Yamins, Surya GanguliNeurIPS 2020 · 884 citations
- Pruning Neural Networks at Initialization: Why Are We Missing the Mark?Jonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICLR 2021 · 261 citations
- Do We Actually Need Dense Over-Parameterization? In-Time Over-Parameterization in Sparse TrainingShiwei Liu, Lu Yin, Decebal Constantin Mocanu, Mykola PechenizkiyICML 2021 · 146 citations
- Sparse Training via Boosting Pruning Plasticity with NeuroregenerationShiwei Liu, Tianlong Chen, Xiaohan Chen, Zahra Atashgahi et al.NeurIPS 2021 · 145 citations
- The Unreasonable Effectiveness of Random Pruning: Return of the Most Naive Baseline for Sparse TrainingShiwei Liu, Tianlong Chen, Xiaohan Chen, Li Shen et al.ICLR 2022 · 141 citations
Builds on7
- Pruning neural networks without any data by iteratively conserving synaptic flowHidenori Tanaka, Daniel Kunin, Daniel L. K. Yamins, Surya GanguliNeurIPS 2020 · 884 citations
- Picking Winning Tickets Before Training by Preserving Gradient FlowChaoqi Wang, Guodong Zhang, Roger B. GrosseICLR 2020 · 743 citations
- Rigging the Lottery: Making All Tickets WinnersUtku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro et al.ICML 2020 · 723 citations
- Soft Threshold Weight Reparameterization for Learnable SparsityAditya Kusupati, Vivek Ramanujan, Raghav Somani, Mitchell Wortsman et al.ICML 2020 · 266 citations
- Dynamic Model Pruning with FeedbackTao Lin, Sebastian U. Stich, Luis Barba, Daniil Dmitriev et al.ICLR 2020 · 229 citations
Related papers
- A Signal Propagation Perspective for Pruning Neural Networks at InitializationNamhoon Lee, Thalaiyasingam Ajanthan, Stephen Gould, Philip H. S. TorrICLR 2020 · 174 citations
- Prospect Pruning: Finding Trainable Weights at Initialization using Meta-GradientsMilad Alizadeh, Shyam A. Tailor, Luisa M. Zintgraf, Joost van Amersfoort et al.ICLR 2022 · 50 citations
- Robust Pruning at InitializationSoufiane Hayou, Jean-Francois Ton, Arnaud Doucet, Yee Whye TehICLR 2021 · 50 citations
- Towards Data-Agnostic Pruning At Initialization: What Makes a Good Sparse Mask?Hoang Pham, The-Anh Ta, Shiwei Liu, Lichuan Xiang et al.NeurIPS 2023 · 16 citations
- Learning Pruning-Friendly Networks via Frank-Wolfe: One-Shot, Any-Sparsity, And No RetrainingLu Miao, Xiaolong Luo, Tianlong Chen, Wuyang Chen et al.ICLR 2022 · 34 citations
