Accelerated Sparse Neural Training: A Provable and Efficient Method to Find N: M Transposable Masks
Itay Hubara, Brian Chmiel, Moshe Island, Ron Banner, Joseph Naor, Daniel Soudry
Abstract
Unstructured pruning reduces the memory footprint in deep neural networks (DNNs). Recently, researchers proposed different types of structural pruning intending to reduce also the computation complexity. In this work, we first suggest a new measure called mask-diversity which correlates with the expected accuracy of the different types of structural pruning. We focus on the recently suggested N : M fine-grained block sparsity mask, in which for each block of M weights, we have at least N zeros. While N : M fine-grained block sparsity allows acceleration in actual modern hardware, it can be used only to accelerate the inference phase. In order to allow for similar accelerations in the training phase, we suggest a novel transposable fine-grained sparsity mask, where the same mask can be used for both forward and backward passes. Our transposable mask guarantees that both the weight matrix and its transpose follow the same sparsity pattern; thus, the matrix multiplication required for passing the error backward can also be accelerated. We formulate the problem of finding the optimal transposable-mask as a minimum-cost flow problem. Additionally, to speed up the minimum-cost flow computation, we also introduce a fast linear-time approximation that can be used when the masks dynamically change during training. Our experiments suggest a 2x speed-up in the matrix multiplications with no accuracy degradation over vision and language models. Finally, to solve the problem of switching between different structure constraints, we suggest a method to convert a pre-trained model with unstructured sparsity to an N : M fine-grained block sparsity model with little to no training. A reference implementation can be found at https: //github.com/papers-submission/structured_transposable_masks .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 169872e2-1911-4fc7-9c52-cb48f598db09Cited by top-tier papers65
- SparseGPT: Massive Language Models Can be Accurately Pruned in One-ShotElias Frantar, Dan AlistarhICML 2023 · 1,240 citations
- A Simple and Effective Pruning Approach for Large Language ModelsMingjie Sun, Zhuang Liu, Anna Bair, J. Zico KolterICLR 2024 · 794 citations
- Optimal Brain Compression: A Framework for Accurate Post-Training Quantization and PruningElias Frantar, Dan AlistarhNeurIPS 2022 · 440 citations
- Sparse Training via Boosting Pruning Plasticity with NeuroregenerationShiwei Liu, Tianlong Chen, Xiaohan Chen, Zahra Atashgahi et al.NeurIPS 2021 · 145 citations
- SpecInfer: Accelerating Large Language Model Serving with Tree-based Speculative Inference and VerificationXupeng Miao, Gabriele Oliaro, Zhihao Zhang, Xinhao Cheng et al.ASPLOS 2024 · 105 citations
Builds on6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Up or Down? Adaptive Rounding for Post-Training QuantizationMarkus Nagel, Rana Ali Amjad, Mart van Baalen, Christos Louizos et al.ICML 2020 · 816 citations
- Rigging the Lottery: Making All Tickets WinnersUtku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro et al.ICML 2020 · 723 citations
- Comparing Rewinding and Fine-tuning in Neural Network PruningAlex Renda, Jonathan Frankle, Michael CarbinICLR 2020 · 437 citations
- Sparse GPU kernels for deep learningTrevor Gale, Matei Zaharia, Cliff Young, Erich ElsenSC 2020 · 170 citations
Related papers
- Learning N: M Fine-grained Structured Sparse Neural Networks From ScratchAojun Zhou, Yukun Ma, Junnan Zhu, Jianbo Liu et al.ICLR 2021 · 301 citations
- Bi-directional Masks for Efficient N: M Sparse TrainingYuxin Zhang, Yiting Luo, Mingbao Lin, Yunshan Zhong et al.ICML 2023 · 23 citations
- BAME: Block-Aware Mask Evolution for Efficient N: M Sparse TrainingChenyi Yang, Wenjie Nie, Yuxin Zhang, Yuhang Wu et al.ICML 2025
- Minimum Variance Unbiased N: M Sparsity for the Neural GradientsBrian Chmiel, Itay Hubara, Ron Banner, Daniel SoudryICLR 2023
- HBP: Hierarchically Balanced Pruning and Accelerator Co-Design for Efficient DNN InferenceAo Ren, Yuhao Wang, Tao Zhang, Jiaxing Shi et al.DAC 2023 · 4 citations
