Top-KAST: Top-K Always Sparse Training
Siddhant M. Jayakumar, Razvan Pascanu, Jack W. Rae, Simon Osindero, Erich Elsen
Abstract
Sparse neural networks are becoming increasingly important as the field seeks to improve the performance of existing models by scaling them up, while simultaneously trying to reduce power consumption and computational footprint. Unfortunately, most existing methods for inducing performant sparse models still entail the instantiation of dense parameters, or dense gradients in the backward-pass, during training. For very large models this requirement can be prohibitive. In this work we propose Top-KAST, a method that preserves constant sparsity throughout training (in both the forward and backward-passes). We demonstrate the efficacy of our approach by showing that it performs comparably to or better than previous works when training models on the established ImageNet benchmark, whilst fully maintaining sparsity. In addition to our ImageNet results, we also demonstrate our approach in the domain of language modeling where the current best performing architectures tend to have tens of billions of parameters and scaling up does not yet seem to have saturated performance. Sparse versions of these architectures can be run with significantly fewer resources, making them more widely accessible and applicable. Furthermore, in addition to being effective, our approach is straightforward and can easily be implemented in a wide range of existing machine learning frameworks with only a few additional lines of code. We therefore hope that our contribution will help enable the broader community to explore the potential held by massive models, without incurring massive computational cost.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0dd40366-60cf-4d73-9307-38e76c1a6201Cited by top-tier papers44
- Chasing Sparsity in Vision Transformers: An End-to-End ExplorationTianlong Chen, Yu Cheng, Zhe Gan, Lu Yuan et al.NeurIPS 2021 · 295 citations
- Do We Actually Need Dense Over-Parameterization? In-Time Over-Parameterization in Sparse TrainingShiwei Liu, Lu Yin, Decebal Constantin Mocanu, Mykola PechenizkiyICML 2021 · 146 citations
- Sparse Training via Boosting Pruning Plasticity with NeuroregenerationShiwei Liu, Tianlong Chen, Xiaohan Chen, Zahra Atashgahi et al.NeurIPS 2021 · 145 citations
- The Unreasonable Effectiveness of Random Pruning: Return of the Most Naive Baseline for Sparse TrainingShiwei Liu, Tianlong Chen, Xiaohan Chen, Li Shen et al.ICLR 2022 · 141 citations
- Make Sharpness-Aware Minimization Stronger: A Sparsified Perturbation ApproachPeng Mi, Li Shen, Tianhe Ren, Yiyi Zhou et al.NeurIPS 2022 · 102 citations
Builds on2
Related papers
- Powerpropagation: A sparsity inducing weight reparameterisationJonathan Schwarz, Siddhant M. Jayakumar, Razvan Pascanu, Peter E. Latham et al.NeurIPS 2021 · 63 citations
- Picking Winning Tickets Before Training by Preserving Gradient FlowChaoqi Wang, Guodong Zhang, Roger B. GrosseICLR 2020 · 743 citations
- Winning the Lottery Ahead of Time: Efficient Early Network PruningJohn Rachwan, Daniel Zügner, Bertrand Charpentier, Simon Geisler et al.ICML 2022 · 34 citations
- SparseOpt: Addressing Normalization-induced Gradient Skew in Sparse TrainingAdnan Mohammed, Rohan Jain, Tom Jacobs, Ekansh Sharma et al.ICML 2026
- Dynamic Model Pruning with FeedbackTao Lin, Sebastian U. Stich, Luis Barba, Daniil Dmitriev et al.ICLR 2020 · 229 citations
