Learning Pruning-Friendly Networks via Frank-Wolfe: One-Shot, Any-Sparsity, And No Retraining
Lu Miao, Xiaolong Luo, Tianlong Chen, Wuyang Chen, Dong Liu, Zhangyang Wang
摘要
We present a novel framework to train a large deep neural network (DNN) for only , which can then be pruned to to preserve competitive accuracy . Conventional methods often require (iterative) pruning followed by re-training, which not only incurs large overhead beyond the original DNN training but also can be sensitive to retraining hyperparameters. Our core idea is to re-cast the DNN training as an explicit process: that is formulated with an auxiliary -sparse polytope constraint, to encourage network weights to lie in a convex hull spanned by -sparse vectors, potentially resulting in more sparse weight matrices. We then leverage a stochastic Frank-Wolfe (SFW) algorithm to solve this new constrained optimization, which naturally leads to sparse weight updates each time. We further note an overlooked fact that existing DNN initializations were derived to enhance SGD training (e.g., avoid gradient explosion or collapse), but was unaligned with the challenges of training with SFW. We hence also present the first learning-based initialization scheme specifically for boosting SFW-based DNN training. Experiments on CIFAR-10 and Tiny-ImageNet datasets demonstrate that our new framework named consistently achieves the state-of-the-art performance on various benchmark DNNs over a wide range of pruning ratios. Moreover, SFW-pruning only needs to train once on the same model and dataset, for obtaining arbitrary ratios, while requiring neither iterative pruning nor retraining. All codes will be released to the public.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper12
- Pruner-Zero: Evolving Symbolic Pruning Metric From Scratch for Large Language ModelsPeijie Dong, Lujun Li, Zhenheng Tang, Xiang Liu 等ICML 2024 · 被引用 64 次
- Differentiable Transportation PruningYunqiang Li, Jan C. van Gemert, Torsten Hoefler, Bert Moons 等ICCV 2023 · 被引用 17 次
- Register Tiling for Unstructured Sparsity in Neural Network InferenceLucas Wilkinson, Kazem Cheshmi, Maryam Mehri DehnaviPLDI 2023 · 被引用 17 次
- OTOv2: Automatic, Generic, User-FriendlyTianyi Chen, Luming Liang, Tianyu Ding, Zhihui Zhu 等ICLR 2023 · 被引用 7 次
- UPSCALE: Unconstrained Channel PruningAlvin Wan, Hanxiang Hao, Kaushik Patnaik, Yueyang Xu 等ICML 2023 · 被引用 7 次
相关 Paper
- Only Train Once: A One-Shot Neural Network Training And Pruning FrameworkTianyi Chen, Bo Ji, Tianyu Ding, Biyi Fang 等NeurIPS 2021 · 被引用 135 次
- Auto- Train-Once: Controller Network Guided Automatic Network Pruning from ScratchXidong Wu, Shangqian Gao, Zeyu Zhang, Zhenzhen Li 等CVPR 2024 · 被引用 13 次
- Pruning from ScratchYulong Wang, Xiaolu Zhang, Lingxi Xie, Jun Zhou 等AAAI 2020 · 被引用 219 次
- Progressive Skeletonization: Trimming more fat from a network at initializationPau de Jorge, Amartya Sanyal, Harkirat S. Behl, Philip H. S. Torr 等ICLR 2021 · 被引用 110 次
- A Unified DNN Weight Pruning Framework Using Reweighted Optimization MethodsTianyun Zhang, Xiaolong Ma, Zheng Zhan, Shanglin Zhou 等DAC 2021 · 被引用 25 次
