MEST: Accurate and Fast Memory-Economic Sparse Training Framework on the Edge
Geng Yuan, Xiaolong Ma, Wei Niu, Zhengang Li, Zhenglun Kong, Ning Liu, Yifan Gong, Zheng Zhan, Chaoyang He, Qing Jin, Siyue Wang, Minghai Qin
Abstract
Recently, a new trend of exploring sparsity for accelerating neural network training has emerged, embracing the paradigm of training on the edge. This paper proposes a novel Memory-Economic Sparse Training (MEST) framework targeting for accurate and fast execution on edge devices. The proposed MEST framework consists of enhancements by Elastic Mutation (EM) and Soft Memory Bound (&S) that ensure superior accuracy at high sparsity ratios. Different from the existing works for sparse training, this current work reveals the importance of sparsity schemes on the performance of sparse training in terms of accuracy as well as training speed on real edge devices. On top of that, the paper proposes to employ data efficiency for further acceleration of sparse training. Our results suggest that unforgettable examples can be identified in-situ even during the dynamic exploration of sparsity masks in the sparse training process, and therefore can be removed for further training speedup on edge devices. Comparing with state-of-the-art (SOTA) works on accuracy, our MEST increases Top-1 accuracy significantly on ImageNet when using the same unstructured sparsity scheme. Systematical evaluation on accuracy, training speed, and memory footprint are conducted, where the proposed MEST framework consistently outperforms representative SOTA works. Our codes are publicly available at: https://github.com/boone891214/MEST . Recently, a new trend of exploring sparsity for training acceleration of neural networks has emerged to embrace the promising training-on-the-edge paradigm. The first works in this direction use the pruning-at-initialization approach such as SNIP [9] and GraSP [10] that first obtains a fixed sparse model structure and then follows with a traditional training process. However, the whole process is still computation-and memory-intensive, and therefore not compatible with the end-to-end edge training paradigm. Such a sparse training methodology with the pre-fixed structure also faces the problem of compromised accuracy. † These authors contributed equally. 35th Conference on Neural Information Processing Systems (NeurIPS 2021).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f8609689-7748-4b88-9d58-da410a6c19d1Cited by top-tier papers49
- SparCL: Sparse Continual Learning on the EdgeZifeng Wang, Zheng Zhan, Yifan Gong, Geng Yuan et al.NeurIPS 2022 · 97 citations
- CHEX: CHannel EXploration for CNN Model CompressionZejiang Hou, Minghai Qin, Fei Sun, Xiaolong Ma et al.CVPR 2022 · 80 citations
- Dynamic Sparse No Training: Training-Free Fine-tuning for Sparse LLMsYuxin Zhang, Lirui Zhao, Mingbao Lin, Yunyun Sun et al.ICLR 2024 · 78 citations
- Plug-and-Play: An Efficient Post-training Pruning Method for Large Language ModelsYingtao Zhang, Haoli Bai, Haokun Lin, Jialin Zhao et al.ICLR 2024 · 72 citations
- F8Net: Fixed-Point 8-bit Only Multiplication for Network QuantizationQing Jin, Jian Ren, Richard Zhuang, Sumant Hanumante et al.ICLR 2022 · 57 citations
Builds on9
- Pruning neural networks without any data by iteratively conserving synaptic flowHidenori Tanaka, Daniel Kunin, Daniel L. K. Yamins, Surya GanguliNeurIPS 2020 · 884 citations
- Picking Winning Tickets Before Training by Preserving Gradient FlowChaoqi Wang, Guodong Zhang, Roger B. GrosseICLR 2020 · 743 citations
- Rigging the Lottery: Making All Tickets WinnersUtku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro et al.ICML 2020 · 723 citations
- Drawing Early-Bird Tickets: Toward More Efficient Training of Deep NetworksHaoran You, Chaojian Li, Pengfei Xu, Yonggan Fu et al.ICLR 2020 · 282 citations
- PatDNN: Achieving Real-Time DNN Execution on Mobile Devices with Pattern-based Weight PruningWei Niu, Xiaolong Ma, Sheng Lin, Shihao Wang et al.ASPLOS 2020 · 214 citations
Related papers
- Layer Freezing & Data Sieving: Missing Pieces of a Generic Framework for Sparse TrainingGeng Yuan, Yanyu Li, Sheng Li, Zhenglun Kong et al.NeurIPS 2022 · 27 citations
- Federated Dynamic Sparse Training: Computing Less, Communicating Less, Yet Learning BetterSameer Bibikar, Haris Vikalo, Zhangyang Wang, Xiaohan ChenAAAI 2022 · 133 citations
- Neurogenesis Dynamics-inspired Spiking Neural Network Training AccelerationShaoyi Huang, Haowen Fang, Kaleel Mahmood, Bowen Lei et al.DAC 2023 · 6 citations
- TinyTrain: Resource-Aware Task-Adaptive Sparse Training of DNNs at the Data-Scarce EdgeYoung D. Kwon, Rui Li, Stylianos I. Venieris, Jagmohan Chauhan et al.ICML 2024 · 25 citations
- MST-compression: Compressing and Accelerating Binary Neural Networks with Minimum Spanning TreeQuang Hieu Vo, Linh-Tam Tran, Sung-Ho Bae, Lok-Won Kim et al.ICCV 2023 · 2 citations
