NeurRev: Train Better Sparse Neural Network Practically via Neuron Revitalization
Gen Li, Lu Yin, Jie Ji, Wei Niu, Minghai Qin, Bin Ren, Linke Guo, Shiwei Liu, Xiaolong Ma
摘要
Dynamic Sparse Training (DST) employs a greedy search mechanism to identify an optimal sparse subnetwork by periodically pruning and growing network connections during training. To guarantee effectiveness, DST algorithms rely on high search frequency, which consequently, requires large learning rate and batch size to enforce stable neuron learning. Such settings demand extreme memory consumption, as well as generating significant system overheads that limit the wide deployment of deep learning-based applications on resource-constraint platforms. To reconcile such, we propose an Neuron Revitalizationframework for DST (NeurRev), based on an innovative finding that dormant neurons exist in the presence of weight sparsity and cannot be revitalized (i.e., activated for learning) even with a high sparse mask search frequency. These dormant neurons produce a large quantity of zeros during training, which contribute relatively little to the outputs of succeeding layers or to the final results. Different from most existing DST algorithms that spare no effort designing weight-growing criteria, NeurRev focuses on optimizing the long-neglected pruning part, which awakens dormant neurons by pruning and incurs no additional computation costs. As such, NerRev advances more effective neuron learning, which not only achieves outperformance accuracy in a variety of networks and datasets but also promotes low-cost dynamism at the system level. Systematical evaluations on training speed and system overhead are conducted on mobile devices, where the proposed NeurRev framework consistently outperforms representative state-of-the-arts. Code available in https://github.com/coulsonlee/NeurRev-ICLR2024 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- A Single-Step, Sharpness-Aware Minimization is All You Need to Achieve Efficient and Accurate Sparse TrainingJie Ji, Gen Li, Jingjing Fu, Fatemeh Afghah 等NeurIPS 2024 · 被引用 14 次
- Advancing Dynamic Sparse Training by Exploring Optimization OpportunitiesJie Ji, Gen Li, Lu Yin, Minghai Qin 等ICML 2024 · 被引用 10 次
- Brain network science modelling of sparse neural networks enables Transformers and LLMs to perform as fully connectedYingtao Zhang, Diego Cerretti, Jialin Zhao, Wenjing Wu 等NeurIPS 2025 · 被引用 5 次
- Forget-It-All: Multi-Concept Machine Unlearning via Concept-Aware Neuron MaskingKaiyuan Deng, Bo Hui, Gen Li, Jie Ji 等ICML 2026 · 被引用 1 次
- Sculpting Memory: Multi-Concept Forgetting in Diffusion Models via Dynamic Mask and Concept-Aware OptimizationGen Li, Yang Xiao, Jie Ji, Kaiyuan Deng 等ICCV 2025 · 被引用 1 次
它引用的顶会 Paper20
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le 等ICCV 2019 · 被引用 9,163 次
- Pruning neural networks without any data by iteratively conserving synaptic flowHidenori Tanaka, Daniel Kunin, Daniel L. K. Yamins, Surya GanguliNeurIPS 2020 · 被引用 884 次
- Linear Mode Connectivity and the Lottery Ticket HypothesisJonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICML 2020 · 被引用 750 次
- Picking Winning Tickets Before Training by Preserving Gradient FlowChaoqi Wang, Guodong Zhang, Roger B. GrosseICLR 2020 · 被引用 743 次
- Rigging the Lottery: Making All Tickets WinnersUtku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro 等ICML 2020 · 被引用 723 次
相关 Paper
- Selfish Sparse RNN TrainingShiwei Liu, Decebal Constantin Mocanu, Yulong Pei, Mykola PechenizkiyICML 2021 · 被引用 43 次
- Fantastic Weights and How to Find Them: Where to Prune in Dynamic Sparse TrainingAleksandra Nowak, Bram Grooten, Decebal Constantin Mocanu, Jacek TaborNeurIPS 2023 · 被引用 23 次
- Dynamic Model Pruning with FeedbackTao Lin, Sebastian U. Stich, Luis Barba, Daniil Dmitriev 等ICLR 2020 · 被引用 229 次
- The Dormant Neuron Phenomenon in Deep Reinforcement LearningGhada Sokar, Rishabh Agarwal, Pablo Samuel Castro, Utku EvciICML 2023 · 被引用 153 次
- Dynamic Sparse Training: Find Efficient Sparse Network From Scratch With Trainable Masked LayersJunjie Liu, Zhe Xu, Runbin Shi, Ray C. C. Cheung 等ICLR 2020 · 被引用 136 次
