Powerpropagation: A sparsity inducing weight reparameterisation
Jonathan Schwarz, Siddhant M. Jayakumar, Razvan Pascanu, Peter E. Latham, Yee Whye Teh
Abstract
The training of sparse neural networks is becoming an increasingly important tool for reducing the computational footprint of models at training and evaluation, as well enabling the effective scaling up of models. Whereas much work over the years has been dedicated to specialised pruning techniques, little attention has been paid to the inherent effect of gradient based training on model sparsity. In this work, we introduce Powerpropagation, a new weight-parameterisation for neural networks that leads to inherently sparse models. Exploiting the behaviour of gradient descent, our method gives rise to weight updates exhibiting a "rich get richer" dynamic, leaving low-magnitude parameters largely unaffected by learning. Models trained in this manner exhibit similar performance, but have a distribution with markedly higher density at zero, allowing more parameters to be pruned safely. Powerpropagation is general, intuitive, cheap and straight-forward to implement and can readily be combined with various other techniques. To highlight its versatility, we explore it in two very different settings: Firstly, following a recent line of work, we investigate its effect on sparse training for resource-constrained settings. Here, we combine Powerpropagation with a traditional weight-pruning technique as well as recent state-of-the-art sparse-to-sparse algorithms, showing superior performance on the ImageNet benchmark. Secondly, we advocate the use of sparsity in overcoming catastrophic forgetting, where compressed representations allow accommodating a large number of tasks at fixed model capacity. In all cases our reparameterisation considerably increases the efficacy of the off-the-shelf methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 47b76e6c-7cb8-43b0-a935-5a4dee064917Cited by top-tier papers30
- More ConvNets in the 2020s: Scaling up Kernels Beyond 51x51 using SparsityShiwei Liu, Tianlong Chen, Xiaohan Chen, Xuxi Chen et al.ICLR 2023 · 87 citations
- The Emergence of Essential Sparsity in Large Pre-trained Models: The Weights that MatterAjay Jaiswal, Shiwei Liu, Tianlong Chen, Zhangyang WangNeurIPS 2023 · 57 citations
- Scaling Laws for Sparsely-Connected Foundation ModelsElias Frantar, Carlos Riquelme Ruiz, Neil Houlsby, Dan Alistarh et al.ICLR 2024 · 48 citations
- Dynamic Sparse Network for Time Series Classification: Learning What to "See"Qiao Xiao, Boqian Wu, Yu Zhang, Shiwei Liu et al.NeurIPS 2022 · 45 citations
- SPDY: Accurate Pruning with Speedup GuaranteesElias Frantar, Dan AlistarhICML 2022 · 45 citations
Builds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang et al.ICLR 2020 · 1,108 citations
- Rigging the Lottery: Making All Tickets WinnersUtku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro et al.ICML 2020 · 723 citations
- Gradient Projection Memory for Continual LearningGobinda Saha, Isha Garg, Kaushik RoyICLR 2021 · 409 citations
- Supermasks in SuperpositionMitchell Wortsman, Vivek Ramanujan, Rosanne Liu, Aniruddha Kembhavi et al.NeurIPS 2020 · 364 citations
Related papers
- Top-KAST: Top-K Always Sparse TrainingSiddhant M. Jayakumar, Razvan Pascanu, Jack W. Rae, Simon Osindero et al.NeurIPS 2020 · 116 citations
- Dynamic Model Pruning with FeedbackTao Lin, Sebastian U. Stich, Luis Barba, Daniil Dmitriev et al.ICLR 2020 · 229 citations
- ResRep: Lossless CNN Pruning via Decoupling Remembering and ForgettingXiaohan Ding, Tianxiang Hao, Jianchao Tan, Ji Liu et al.ICCV 2021 · 202 citations
- Dense for the Price of Sparse: Improved Performance of Sparsely Initialized Networks via a Subspace OffsetIlan Price, Jared TannerICML 2021 · 17 citations
- S: Sign-Sparse-Shift Reparametrization for Effective Training of Low-bit Shift NetworksXinlin Li, Bang Liu, Yaoliang Yu, Wulong Liu et al.NeurIPS 2021 · 12 citations
