Fantastic Weights and How to Find Them: Where to Prune in Dynamic Sparse Training
Aleksandra Nowak, Bram Grooten, Decebal Constantin Mocanu, Jacek Tabor
Abstract
Dynamic Sparse Training (DST) is a rapidly evolving area of research that seeks to optimize the sparse initialization of a neural network by adapting its topology during training. It has been shown that under specific conditions, DST is able to outperform dense models. The key components of this framework are the pruning and growing criteria, which are repeatedly applied during the training process to adjust the network's sparse connectivity. While the growing criterion's impact on DST performance is relatively well studied, the influence of the pruning criterion remains overlooked. To address this issue, we design and perform an extensive empirical analysis of various pruning criteria to better understand their impact on the dynamics of DST solutions. Surprisingly, we find that most of the studied methods yield similar results. The differences become more significant in the low-density regime, where the best performance is predominantly given by the simplest technique: magnitude-based pruning. The code is provided at https://github.com/alooow/fantastic_weights_paper
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 42894ded-7104-4816-829a-d0635caa5a78Cited by top-tier papers7
- Advancing Dynamic Sparse Training by Exploring Optimization OpportunitiesJie Ji, Gen Li, Lu Yin, Minghai Qin et al.ICML 2024 · 10 citations
- Brain network science modelling of sparse neural networks enables Transformers and LLMs to perform as fully connectedYingtao Zhang, Diego Cerretti, Jialin Zhao, Wenjing Wu et al.NeurIPS 2025 · 5 citations
- A Recovery Guarantee for Sparse Neural NetworksSara Fridovich-Keil, Mert PilanciICLR 2026 · 1 citation
- SCAN: Bootstrapping Contrastive Pre-training for Data EfficiencyYangyang Guo, Mohan KankanhalliICCV 2025 · 1 citation
- Dynamic Sparse Training versus Dense Training: The Unexpected Winner in Image Corruption RobustnessBoqian Wu, Qiao Xiao, Shunxin Wang, Nicola Strisciuglio et al.ICLR 2025
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Pruning neural networks without any data by iteratively conserving synaptic flowHidenori Tanaka, Daniel Kunin, Daniel L. K. Yamins, Surya GanguliNeurIPS 2020 · 884 citations
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 794 citations
- Picking Winning Tickets Before Training by Preserving Gradient FlowChaoqi Wang, Guodong Zhang, Roger B. GrosseICLR 2020 · 743 citations
- Rigging the Lottery: Making All Tickets WinnersUtku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro et al.ICML 2020 · 723 citations
Related papers
- Gradient Flow in Sparse Neural Networks and How Lottery Tickets WinUtku Evci, Yani Ioannou, Cem Keskin, Yann N. DauphinAAAI 2022 · 106 citations
- SparseOpt: Addressing Normalization-induced Gradient Skew in Sparse TrainingAdnan Mohammed, Rohan Jain, Tom Jacobs, Ekansh Sharma et al.ICML 2026
- Dynamic Sparse Training: Find Efficient Sparse Network From Scratch With Trainable Masked LayersJunjie Liu, Zhe Xu, Runbin Shi, Ray C. C. Cheung et al.ICLR 2020 · 136 citations
- Do We Actually Need Dense Over-Parameterization? In-Time Over-Parameterization in Sparse TrainingShiwei Liu, Lu Yin, Decebal Constantin Mocanu, Mykola PechenizkiyICML 2021 · 146 citations
- NeurRev: Train Better Sparse Neural Network Practically via Neuron RevitalizationGen Li, Lu Yin, Jie Ji, Wei Niu et al.ICLR 2024 · 10 citations
