Towards Fully Sparse Training: Information Restoration with Spatial Similarity
Weixiang Xu, Xiangyu He, Ke Cheng, Peisong Wang, Jian Cheng
摘要
The 2:4 structured sparsity pattern released by NVIDIA Ampere architecture, requiring four consecutive values containing at least two zeros, enables doubling math throughput for matrix multiplications. Recent works mainly focus on inference speedup via 2:4 sparsity while training acceleration has been largely overwhelmed where backpropagation consumes around 70% of the training time. However, unlike inference, training speedup with structured pruning is nontrivial due to the need to maintain the fidelity of gradients and reduce the additional overhead of performing 2:4 sparsity online. For the first time, this article proposes fully sparse training (FST) where `fully' indicates that ALL matrix multiplications in forward/backward propagation are structurally pruned while maintaining accuracy. To this end, we begin with saliency analysis, investigating the sensitivity of different sparse objects to structured pruning. Based on the observation of spatial similarity among activations, we propose pruning activations with fixed 2:4 masks. Moreover, an Information Restoration block is proposed to retrieve the lost information, which can be implemented by efficient gradient-shift operation. Evaluation of accuracy and efficiency shows that we can achieve 2× training acceleration with negligible accuracy degradation on challenging large-scale classification and detection tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Make Sharpness-Aware Minimization Stronger: A Sparsified Perturbation ApproachPeng Mi, Li Shen, Tianhe Ren, Yiyi Zhou 等NeurIPS 2022 · 被引用 102 次
- Accelerating Transformer Pre-training with 2: 4 SparsityYuezhou Hu, Kang Zhao, Weiyu Huang, Jianfei Chen 等ICML 2024 · 被引用 19 次
- Towards Efficient Spiking Transformer: a Token Sparsification Framework for Training and Inference AccelerationZhengyang Zhuge, Peisong Wang, Xingting Yao, Jian ChengICML 2024 · 被引用 6 次
- Minimum Variance Unbiased N: M Sparsity for the Neural GradientsBrian Chmiel, Itay Hubara, Ron Banner, Daniel SoudryICLR 2023
它引用的顶会 Paper9
- Rigging the Lottery: Making All Tickets WinnersUtku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro 等ICML 2020 · 被引用 723 次
- Learning N: M Fine-grained Structured Sparse Neural Networks From ScratchAojun Zhou, Yukun Ma, Junnan Zhu, Jianbo Liu 等ICLR 2021 · 被引用 301 次
- Dynamic Model Pruning with FeedbackTao Lin, Sebastian U. Stich, Luis Barba, Daniil Dmitriev 等ICLR 2020 · 被引用 229 次
- Towards Accurate Post-training Network Quantization via Bit-Split and StitchingPeisong Wang, Qiang Chen, Xiangyu He, Jian ChengICML 2020 · 被引用 159 次
- Accelerated Sparse Neural Training: A Provable and Efficient Method to Find N: M Transposable MasksItay Hubara, Brian Chmiel, Moshe Island, Ron Banner 等NeurIPS 2021 · 被引用 148 次
相关 Paper
- S-STE: Continuous Pruning Function for Efficient 2: 4 Sparse Pre-trainingYuezhou Hu, Jun Zhu, Jianfei ChenNeurIPS 2024 · 被引用 14 次
- SAS: Structured Activation SparsificationYusuke Sekikawa, Shingo YashimaICLR 2024 · 被引用 1 次
- SparseTrain: Exploiting Dataflow Sparsity for Efficient Convolutional Neural Networks TrainingPengcheng Dai, Jianlei Yang, Xucheng Ye, Xingzhou Cheng 等DAC 2020 · 被引用 27 次
- Eureka: Efficient Tensor Cores for One-sided Unstructured Sparsity in DNN InferenceAshish Gondimalla, Mithuna Thottethodi, T. N. VijaykumarMICRO 2023 · 被引用 15 次
- Cascading structured pruning: enabling high data reuse for sparse DNN acceleratorsEdward Hanson, Shiyu Li, Hai Helen Li, Yiran ChenISCA 2022 · 被引用 30 次
