Dynamic Sparsity Is Channel-Level Sparsity Learner
Lu Yin, Gen Li, Meng Fang, Li Shen, Tianjin Huang, Zhangyang Wang, Vlado Menkovski, Xiaolong Ma, Mykola Pechenizkiy, Shiwei Liu
摘要
Sparse training has received an upsurging interest in machine learning due to its tantalizing saving potential for the entire training process as well as inference. Dynamic sparse training (DST), as a leading sparse training approach, can train deep neural networks at high sparsity from scratch to match the performance of their dense counterparts. However, most if not all DST prior arts demonstrate their effectiveness on unstructured sparsity with highly irregular sparse patterns, which receives limited support in common hardware. This limitation hinders the usage of DST in practice. In this paper, we propose Channel-aware dynamic sparse (Chase), which for the first time seamlessly translates the promise of unstructured dynamic sparsity to GPU-friendly channel-level sparsity (not fine-grained N:M or group sparsity) during one end-to-end training process, without any ad-hoc operations. The resulting small sparse networks can be directly accelerated by commodity hardware, without using any particularly sparsity-aware hardware accelerators. This appealing outcome is partially motivated by a hidden phenomenon of dynamic sparsity: off-the-shelf unstructured DST implicitly involves biased parameter reallocation across channels, with a large fraction of channels (up to 60%) being sparser than others. By progressively identifying and removing these channels during training, our approach translates unstructured sparsity to channel-wise sparsity. Our experimental results demonstrate that Chase achieves 1.7× inference throughput speedup on common GPU devices without compromising accuracy with ResNet-50 on ImageNet. We release our codes in https://github.com/luuyin/chase . 37th Conference on Neural Information Processing Systems (NeurIPS 2023).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Dynamic Sparse Training with Structured SparsityMike Lasby, Anna Golubeva, Utku Evci, Mihai Nica 等ICLR 2024 · 被引用 37 次
- SpaFL: Communication-Efficient Federated Learning With Sparse Models And Low Computational OverheadMinsu Kim, Walid Saad, Mérouane Debbah, Choong Seon HongNeurIPS 2024 · 被引用 30 次
- Visual Prompting Upgrades Neural Network Sparsification: A Data-Model PerspectiveCan Jin, Tianjin Huang, Yihua Zhang, Mykola Pechenizkiy 等AAAI 2025 · 被引用 30 次
- A Single-Step, Sharpness-Aware Minimization is All You Need to Achieve Efficient and Accurate Sparse TrainingJie Ji, Gen Li, Jingjing Fu, Fatemeh Afghah 等NeurIPS 2024 · 被引用 14 次
- S-STE: Continuous Pruning Function for Efficient 2: 4 Sparse Pre-trainingYuezhou Hu, Jun Zhu, Jianfei ChenNeurIPS 2024 · 被引用 14 次
它引用的顶会 Paper30
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Swin Transformer V2: Scaling Up Capacity and ResolutionZe Liu, Han Hu, Yutong Lin, Zhuliang Yao 等CVPR 2022 · 被引用 2,138 次
- Pruning neural networks without any data by iteratively conserving synaptic flowHidenori Tanaka, Daniel Kunin, Daniel L. K. Yamins, Surya GanguliNeurIPS 2020 · 被引用 884 次
- Picking Winning Tickets Before Training by Preserving Gradient FlowChaoqi Wang, Guodong Zhang, Roger B. GrosseICLR 2020 · 被引用 743 次
相关 Paper
- Dynamic Sparse Training of Diagonally Sparse NetworksAbhishek Tyagi, Arjun Iyer, William H. Renninger, Christopher Kanan 等ICML 2025
- Exposing and Exploiting Fine-Grained Block Structures for Fast and Accurate Sparse TrainingPeng Jiang, Lihan Hu, Shihui SongNeurIPS 2022 · 被引用 21 次
- Learning N: M Fine-grained Structured Sparse Neural Networks From ScratchAojun Zhou, Yukun Ma, Junnan Zhu, Jianbo Liu 等ICLR 2021 · 被引用 301 次
- Get More at Once: Alternating Sparse Training with Gradient CorrectionLi Yang, Jian Meng, Jae-sun Seo, Deliang FanNeurIPS 2022 · 被引用 3 次
- DominoSearch: Find layer-wise fine-grained N: M sparse schemes from dense neural networksWei Sun, Aojun Zhou, Sander Stuijk, Rob G. J. Wijnhoven 等NeurIPS 2021 · 被引用 67 次
