Dynamic Sparse Training of Diagonally Sparse Networks
Abhishek Tyagi, Arjun Iyer, William H. Renninger, Christopher Kanan, Yuhao Zhu
Abstract
Recent advances in Dynamic Sparse Training (DST) have pushed the frontier of sparse neural network training in structured and unstructured contexts, matching dense-model performance while drastically reducing parameter counts to facilitate model scaling. However, unstructured sparsity often fails to translate into practical speedups on modern hardware. To address this shortcoming, we propose DynaDiag, a novel structured sparse-to-sparse DST method that performs at par with unstructured sparsity. Dyna-Diag enforces a diagonal sparsity pattern throughout training and preserves sparse computation in forward and backward passes. We further leverage the diagonal structure to accelerate computation via a custom CUDA kernel, rendering the method hardware-friendly. Empirical evaluations on diverse neural architectures demonstrate that our method maintains accuracy on par with unstructured counterparts while benefiting from tangible computational gains. Notably, with 90% sparse linear layers in ViTs, we observe up to a 3.13x speedup in online inference without sacrificing model performance and a 1.59x speedup in training on a GPU compared to equivalent unstructured layers. Our source code is available at https://github.com/ horizon-research/DynaDiag/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6255f848-451f-4a10-973a-166c93a6fbf4Cited by top-tier papers2
- VideoNSA: Native Sparse Attention Scales Video UnderstandingEnxin Song, Wenhao Chai, Shusheng Yang, Ethan Armand et al.ICLR 2026 · 11 citations
- SparseOpt: Addressing Normalization-induced Gradient Skew in Sparse TrainingAdnan Mohammed, Rohan Jain, Tom Jacobs, Ekansh Sharma et al.ICML 2026
Builds on22
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- Pruning neural networks without any data by iteratively conserving synaptic flowHidenori Tanaka, Daniel Kunin, Daniel L. K. Yamins, Surya GanguliNeurIPS 2020 · 884 citations
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 794 citations
- Picking Winning Tickets Before Training by Preserving Gradient FlowChaoqi Wang, Guodong Zhang, Roger B. GrosseICLR 2020 · 743 citations
Related papers
- Dynamic Sparse Training with Structured SparsityMike Lasby, Anna Golubeva, Utku Evci, Mihai Nica et al.ICLR 2024 · 37 citations
- Dynamic Sparsity Is Channel-Level Sparsity LearnerLu Yin, Gen Li, Meng Fang, Li Shen et al.NeurIPS 2023 · 29 citations
- Exposing and Exploiting Fine-Grained Block Structures for Fast and Accurate Sparse TrainingPeng Jiang, Lihan Hu, Shihui SongNeurIPS 2022 · 21 citations
- Network Expansion For Practical Training AccelerationNing Ding, Yehui Tang, Kai Han, Chao Xu et al.CVPR 2023
- Selfish Sparse RNN TrainingShiwei Liu, Decebal Constantin Mocanu, Yulong Pei, Mykola PechenizkiyICML 2021 · 43 citations
