Powerpropagation: A sparsity inducing weight reparameterisation
Jonathan Schwarz, Siddhant M. Jayakumar, Razvan Pascanu, Peter E. Latham, Yee Whye Teh
摘要
The training of sparse neural networks is becoming an increasingly important tool for reducing the computational footprint of models at training and evaluation, as well enabling the effective scaling up of models. Whereas much work over the years has been dedicated to specialised pruning techniques, little attention has been paid to the inherent effect of gradient based training on model sparsity. In this work, we introduce Powerpropagation, a new weight-parameterisation for neural networks that leads to inherently sparse models. Exploiting the behaviour of gradient descent, our method gives rise to weight updates exhibiting a "rich get richer" dynamic, leaving low-magnitude parameters largely unaffected by learning. Models trained in this manner exhibit similar performance, but have a distribution with markedly higher density at zero, allowing more parameters to be pruned safely. Powerpropagation is general, intuitive, cheap and straight-forward to implement and can readily be combined with various other techniques. To highlight its versatility, we explore it in two very different settings: Firstly, following a recent line of work, we investigate its effect on sparse training for resource-constrained settings. Here, we combine Powerpropagation with a traditional weight-pruning technique as well as recent state-of-the-art sparse-to-sparse algorithms, showing superior performance on the ImageNet benchmark. Secondly, we advocate the use of sparsity in overcoming catastrophic forgetting, where compressed representations allow accommodating a large number of tasks at fixed model capacity. In all cases our reparameterisation considerably increases the efficacy of the off-the-shelf methods.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper30
- More ConvNets in the 2020s: Scaling up Kernels Beyond 51x51 using SparsityShiwei Liu, Tianlong Chen, Xiaohan Chen, Xuxi Chen 等ICLR 2023 · 被引用 87 次
- The Emergence of Essential Sparsity in Large Pre-trained Models: The Weights that MatterAjay Jaiswal, Shiwei Liu, Tianlong Chen, Zhangyang WangNeurIPS 2023 · 被引用 57 次
- Scaling Laws for Sparsely-Connected Foundation ModelsElias Frantar, Carlos Riquelme Ruiz, Neil Houlsby, Dan Alistarh 等ICLR 2024 · 被引用 48 次
- Dynamic Sparse Network for Time Series Classification: Learning What to "See"Qiao Xiao, Boqian Wu, Yu Zhang, Shiwei Liu 等NeurIPS 2022 · 被引用 45 次
- SPDY: Accurate Pruning with Speedup GuaranteesElias Frantar, Dan AlistarhICML 2022 · 被引用 45 次
它引用的顶会 Paper10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang 等ICLR 2020 · 被引用 1,108 次
- Rigging the Lottery: Making All Tickets WinnersUtku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro 等ICML 2020 · 被引用 723 次
- Gradient Projection Memory for Continual LearningGobinda Saha, Isha Garg, Kaushik RoyICLR 2021 · 被引用 409 次
- Supermasks in SuperpositionMitchell Wortsman, Vivek Ramanujan, Rosanne Liu, Aniruddha Kembhavi 等NeurIPS 2020 · 被引用 364 次
相关 Paper
- Top-KAST: Top-K Always Sparse TrainingSiddhant M. Jayakumar, Razvan Pascanu, Jack W. Rae, Simon Osindero 等NeurIPS 2020 · 被引用 116 次
- Dynamic Model Pruning with FeedbackTao Lin, Sebastian U. Stich, Luis Barba, Daniil Dmitriev 等ICLR 2020 · 被引用 229 次
- ResRep: Lossless CNN Pruning via Decoupling Remembering and ForgettingXiaohan Ding, Tianxiang Hao, Jianchao Tan, Ji Liu 等ICCV 2021 · 被引用 202 次
- Dense for the Price of Sparse: Improved Performance of Sparsely Initialized Networks via a Subspace OffsetIlan Price, Jared TannerICML 2021 · 被引用 17 次
- S: Sign-Sparse-Shift Reparametrization for Effective Training of Low-bit Shift NetworksXinlin Li, Bang Liu, Yaoliang Yu, Wulong Liu 等NeurIPS 2021 · 被引用 12 次
