Training Structured Neural Networks Through Manifold Identification and Variance Reduction
Zih-Syuan Huang, Ching-pei Lee
Abstract
This paper proposes an algorithm, RMDA, for training neural networks (NNs) with a regularization term for promoting desired structures. RMDA does not incur computation additional to proximal SGD with momentum, and achieves variance reduction without requiring the objective function to be of the finite-sum form. Through the tool of manifold identification from nonlinear optimization, we prove that after a finite number of iterations, all iterates of RMDA possess a desired structure identical to that induced by the regularizer at the stationary point of asymptotic convergence, even in the presence of engineering tricks like data augmentation that complicate the training process. Experiments on training NNs with structured sparsity confirm that variance reduction is necessary for such an identification, and show that RMDA thus significantly outperforms existing methods for this task. For unstructured sparsity, RMDA also outperforms a state-of-the-art pruning method, validating the benefits of training structured NNs through regularization. Implementation of RMDA is available at https://www.github.com/zihsyuan1214/rmda .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 240a043b-e69a-4f3a-8d9e-1076732eaef6Cited by top-tier papers1
Ask how each one uses itBuilds on5
- Rigging the Lottery: Making All Tickets WinnersUtku Evci, Trevor Gale, Jacob Menick, Pablo Samuel Castro et al.ICML 2020 · 723 citations
- Adaptive Proximal Gradient Methods for Structured Neural NetworksJihun Yun, Aurélie C. Lozano, Eunho YangNeurIPS 2021 · 34 citations
- ProxSGD: Training Structured Neural Networks under Regularization and ConstraintsYang Yang, Yaxiong Yuan, Avraam Chatzimichailidis, Ruud J. G. van Sloun et al.ICLR 2020 · 34 citations
- Sparsifying Networks via Subdifferential InclusionSagar Verma, Jean-Christophe PesquetICML 2021 · 15 citations
- Manifold Identification for Ultimately Communication-Efficient Distributed OptimizationYu-Sheng Li, Wei-Lin Chiang, Ching-Pei LeeICML 2020 · 6 citations
Related papers
- Directional Pruning of Deep Neural NetworksShih-Kang Chao, Zhanyu Wang, Yue Xing, Guang ChengNeurIPS 2020 · 36 citations
- Variance Reduction via Primal-Dual Accelerated Dual Averaging for Nonsmooth Convex Finite-SumsChaobing Song, Stephen J. Wright, Jelena DiakonikolasICML 2021 · 22 citations
- Only Train Once: A One-Shot Neural Network Training And Pruning FrameworkTianyi Chen, Bo Ji, Tianyu Ding, Biyi Fang et al.NeurIPS 2021 · 135 citations
- Learning Compact Representations of Neural Networks using DiscriminAtive Masking (DAM)Jie Bu, Arka Daw, M. Maruf, Anuj KarpatneNeurIPS 2021 · 7 citations
- Efficient Neural Network Training via Forward and Backward Propagation SparsificationXiao Zhou, Weizhong Zhang, Zonghao Chen, Shizhe Diao et al.NeurIPS 2021 · 57 citations
