Sparse Weight Activation Training
Md Aamir Raihan, Tor M. Aamodt
摘要
Neural network training is computationally and memory intensive. Sparse training can reduce the burden on emerging hardware platforms designed to accelerate sparse computations, but it can also affect network convergence. In this work, we propose a novel CNN training algorithm called Sparse Weight Activation Training (SWAT). SWAT is more computation and memory-efficient than conventional training. SWAT modifies back-propagation based on the empirical insight that convergence during training tends to be robust to the elimination of (i) small magnitude weights during the forward pass and (ii) both small magnitude weights and activations during the backward pass. We evaluate SWAT on recent CNN architectures such as ResNet, VGG, DenseNet and WideResNet using CIFAR-10, CIFAR-100 and ImageNet datasets. For ResNet-50 on ImageNet SWAT reduces total floating-point operations (FLOPs) during training by 80% resulting in a 3.3× training speedup when run on a simulated sparse learning accelerator representative of emerging platforms while incurring only 1.63% reduction in validation accuracy. Moreover, SWAT reduces memory footprint during the backward pass by 23% to 50% for activations and 50% to 90% for weights. Code is available at https: //github.com/AamirRaihan/SWAT .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper33
- Chasing Sparsity in Vision Transformers: An End-to-End ExplorationTianlong Chen, Yu Cheng, Zhe Gan, Lu Yuan 等NeurIPS 2021 · 被引用 295 次
- Do We Actually Need Dense Over-Parameterization? In-Time Over-Parameterization in Sparse TrainingShiwei Liu, Lu Yin, Decebal Constantin Mocanu, Mykola PechenizkiyICML 2021 · 被引用 146 次
- Pixelated Butterfly: Simple and Efficient Sparse training for Neural Network ModelsBeidi Chen, Tri Dao, Kaizhao Liang, Jiaming Yang 等ICLR 2022 · 被引用 94 次
- ZeroFL: Efficient On-Device Training for Federated Learning with Local SparsityXinchi Qiu, Javier Fernández-Marqués, Pedro P. B. de Gusmao, Yan Gao 等ICLR 2022 · 被引用 87 次
- Deep Ensembling with No Overhead for either Training or Testing: The All-Round Blessings of Dynamic SparsityShiwei Liu, Tianlong Chen, Zahra Atashgahi, Xiaohan Chen 等ICLR 2022 · 被引用 62 次
它引用的顶会 Paper5
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang 等ICLR 2020 · 被引用 1,522 次
- Rigging the Lottery: Making All Tickets WinnersUtku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro 等ICML 2020 · 被引用 723 次
- Soft Threshold Weight Reparameterization for Learnable SparsityAditya Kusupati, Vivek Ramanujan, Raghav Somani, Mitchell Wortsman 等ICML 2020 · 被引用 266 次
- Dynamic Sparse Training: Find Efficient Sparse Network From Scratch With Trainable Masked LayersJunjie Liu, Zhe Xu, Runbin Shi, Ray C. C. Cheung 等ICLR 2020 · 被引用 136 次
- ReSprop: Reuse Sparsified BackpropagationNegar Goli, Tor M. AamodtCVPR 2020
相关 Paper
- SparseTrain: Exploiting Dataflow Sparsity for Efficient Convolutional Neural Networks TrainingPengcheng Dai, Jianlei Yang, Xucheng Ye, Xingzhou Cheng 等DAC 2020 · 被引用 27 次
- SparseOpt: Addressing Normalization-induced Gradient Skew in Sparse TrainingAdnan Mohammed, Rohan Jain, Tom Jacobs, Ekansh Sharma 等ICML 2026
- Stochastic Weight Averaging in Parallel: Large-Batch Training That Generalizes WellVipul Gupta, Santiago Akle Serrano, Dennis DeCosteICLR 2020 · 被引用 78 次
- SparseProp: Efficient Sparse Backpropagation for Faster Training of Neural Networks at the EdgeMahdi Nikdan, Tommaso Pegolotti, Eugenia Iofinova, Eldar Kurtic 等ICML 2023 · 被引用 14 次
- Inducing and Exploiting Activation Sparsity for Fast Inference on Deep Neural NetworksMark Kurtz, Justin Kopinsky, Rati Gelashvili, Alexander Matveev 等ICML 2020 · 被引用 163 次
