Training Structured Neural Networks Through Manifold Identification and Variance Reduction
Zih-Syuan Huang, Ching-pei Lee
摘要
This paper proposes an algorithm, RMDA, for training neural networks (NNs) with a regularization term for promoting desired structures. RMDA does not incur computation additional to proximal SGD with momentum, and achieves variance reduction without requiring the objective function to be of the finite-sum form. Through the tool of manifold identification from nonlinear optimization, we prove that after a finite number of iterations, all iterates of RMDA possess a desired structure identical to that induced by the regularizer at the stationary point of asymptotic convergence, even in the presence of engineering tricks like data augmentation that complicate the training process. Experiments on training NNs with structured sparsity confirm that variance reduction is necessary for such an identification, and show that RMDA thus significantly outperforms existing methods for this task. For unstructured sparsity, RMDA also outperforms a state-of-the-art pruning method, validating the benefits of training structured NNs through regularization. Implementation of RMDA is available at https://www.github.com/zihsyuan1214/rmda .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper5
- Rigging the Lottery: Making All Tickets WinnersUtku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro 等ICML 2020 · 被引用 723 次
- Adaptive Proximal Gradient Methods for Structured Neural NetworksJihun Yun, Aurélie C. Lozano, Eunho YangNeurIPS 2021 · 被引用 34 次
- ProxSGD: Training Structured Neural Networks under Regularization and ConstraintsYang Yang, Yaxiong Yuan, Avraam Chatzimichailidis, Ruud J. G. van Sloun 等ICLR 2020 · 被引用 34 次
- Sparsifying Networks via Subdifferential InclusionSagar Verma, Jean-Christophe PesquetICML 2021 · 被引用 15 次
- Manifold Identification for Ultimately Communication-Efficient Distributed OptimizationYu-Sheng Li, Wei-Lin Chiang, Ching-Pei LeeICML 2020 · 被引用 6 次
相关 Paper
- Directional Pruning of Deep Neural NetworksShih-Kang Chao, Zhanyu Wang, Yue Xing, Guang ChengNeurIPS 2020 · 被引用 36 次
- Variance Reduction via Primal-Dual Accelerated Dual Averaging for Nonsmooth Convex Finite-SumsChaobing Song, Stephen J. Wright, Jelena DiakonikolasICML 2021 · 被引用 22 次
- Only Train Once: A One-Shot Neural Network Training And Pruning FrameworkTianyi Chen, Bo Ji, Tianyu Ding, Biyi Fang 等NeurIPS 2021 · 被引用 135 次
- Learning Compact Representations of Neural Networks using DiscriminAtive Masking (DAM)Jie Bu, Arka Daw, M. Maruf, Anuj KarpatneNeurIPS 2021 · 被引用 7 次
- Efficient Neural Network Training via Forward and Backward Propagation SparsificationXiao Zhou, Weizhong Zhang, Zonghao Chen, Shizhe Diao 等NeurIPS 2021 · 被引用 57 次
