Advancing Dynamic Sparse Training by Exploring Optimization Opportunities
Jie Ji, Gen Li, Lu Yin, Minghai Qin, Geng Yuan, Linke Guo, Shiwei Liu, Xiaolong Ma
Abstract
Dynamic Sparse Training (DST) has been effectively addressing the substantial training resource requirements of increasingly large Deep Neural Networks (DNNs). Characterized by its dynamic "train-prune-grow" schedule during training, DST implicitly develops a bi-level structure for training the weights while discovering a subnetwork topology. However, such a structure is consistently overlooked by the current DST algorithms for further optimization opportunities, and these algorithms, on the other hand, solely optimize the weights while determining masks heuristically. In this paper, we extensively study DST algorithms and argue that the training scheme of DST naturally forms a bi-level problem in which the updating of weight and mask is interdependent. Based on this observation, we introduce a novel efficient training framework called BiDST, which for the first time, introduces bi-level optimization methodology into dynamic sparse training domain. Unlike traditional partialheuristic DST schemes, which suffer from suboptimal search efficiency for masks and miss the opportunity to fully explore the topological space of neural networks, BiDST excels at discovering excellent sparse patterns by optimizing mask and weight simultaneously, resulting in maximum 2.62% higher accuracy, 2.1× faster execution speed, and 25× reduced overhead. Code available at https://github.com/jjsrf/ BiDST-ICML2024 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1af7ddb5-68e6-459e-b64e-8936ac6b5173Cited by top-tier papers6
- A Single-Step, Sharpness-Aware Minimization is All You Need to Achieve Efficient and Accurate Sparse TrainingJie Ji, Gen Li, Jingjing Fu, Fatemeh Afghah et al.NeurIPS 2024 · 14 citations
- Brain network science modelling of sparse neural networks enables Transformers and LLMs to perform as fully connectedYingtao Zhang, Diego Cerretti, Jialin Zhao, Wenjing Wu et al.NeurIPS 2025 · 5 citations
- A Recovery Guarantee for Sparse Neural NetworksSara Fridovich-Keil, Mert PilanciICLR 2026 · 1 citation
- Sculpting Memory: Multi-Concept Forgetting in Diffusion Models via Dynamic Mask and Concept-Aware OptimizationGen Li, Yang Xiao, Jie Ji, Kaiyuan Deng et al.ICCV 2025 · 1 citation
- SparseOpt: Addressing Normalization-induced Gradient Skew in Sparse TrainingAdnan Mohammed, Rohan Jain, Tom Jacobs, Ekansh Sharma et al.ICML 2026
Builds on18
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Pruning neural networks without any data by iteratively conserving synaptic flowHidenori Tanaka, Daniel Kunin, Daniel L. K. Yamins, Surya GanguliNeurIPS 2020 · 884 citations
- Picking Winning Tickets Before Training by Preserving Gradient FlowChaoqi Wang, Guodong Zhang, Roger B. GrosseICLR 2020 · 743 citations
- Rigging the Lottery: Making All Tickets WinnersUtku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro et al.ICML 2020 · 723 citations
- Outlier Weighed Layerwise Sparsity (OWL): A Missing Secret Sauce for Pruning LLMs to High SparsityLu Yin, You Wu, Zhenyu Zhang, Cheng-Yu Hsieh et al.ICML 2024 · 183 citations
Related papers
- Dynamic Sparse Training: Find Efficient Sparse Network From Scratch With Trainable Masked LayersJunjie Liu, Zhe Xu, Runbin Shi, Ray C. C. Cheung et al.ICLR 2020 · 136 citations
- Federated Dynamic Sparse Training: Computing Less, Communicating Less, Yet Learning BetterSameer Bibikar, Haris Vikalo, Zhangyang Wang, Xiaohan ChenAAAI 2022 · 133 citations
- Do We Actually Need Dense Over-Parameterization? In-Time Over-Parameterization in Sparse TrainingShiwei Liu, Lu Yin, Decebal Constantin Mocanu, Mykola PechenizkiyICML 2021 · 146 citations
- Dynamic Sparse Training via Balancing the Exploration-Exploitation Trade-offShaoyi Huang, Bowen Lei, Dongkuan Xu, Hongwu Peng et al.DAC 2023 · 7 citations
- Fantastic Weights and How to Find Them: Where to Prune in Dynamic Sparse TrainingAleksandra Nowak, Bram Grooten, Decebal Constantin Mocanu, Jacek TaborNeurIPS 2023 · 23 citations
