ProxSGD: Training Structured Neural Networks under Regularization and Constraints
Yang Yang, Yaxiong Yuan, Avraam Chatzimichailidis, Ruud J. G. van Sloun, Lei Lei, Symeon Chatzinotas
Abstract
In this paper, we consider the problem of training structured neural networks (NN) with nonsmooth regularization (e.g. `1-norm) and constraints (e.g. interval constraints). We formulate training as a constrained nonsmooth nonconvex optimization problem, and propose a convergent proximal-type stochastic gradient descent (ProxSGD) algorithm. We show that under properly selected learning rates, with probability 1, every limit point of the sequence generated by the proposed Prox-SGD algorithm is a stationary point. Finally, to support the theoretical analysis and demonstrate the flexibility of ProxSGD, we show by extensive numerical tests how ProxSGD can be used to train either sparse or binary neural networks through an adequate selection of the regularization function and constraint set.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get f9ff5731-29d4-4fb9-b799-577362fd947cCited by top-tier papers7
- Adaptive Proximal Gradient Methods for Structured Neural NetworksJihun Yun, Aurélie C. Lozano, Eunho YangNeurIPS 2021 · 34 citations
- Compressed Decentralized Proximal Stochastic Gradient Method for Nonconvex Composite Problems with Heterogeneous DataYonggui Yan, Jie Chen, Pin-Yu Chen, Xiaodong Cui et al.ICML 2023 · 18 citations
- Training Structured Neural Networks Through Manifold Identification and Variance ReductionZih-Syuan Huang, Ching-pei LeeICLR 2022 · 10 citations
- Regularized Adaptive Momentum Dual Averaging with an Efficient Inexact Subproblem Solver for Training Structured Neural NetworkZih-Syuan Huang, Ching-pei LeeNeurIPS 2024
- Singularity-aware Optimization via Randomized Geometric Probing: Towards Stable Non-smooth OptimizationRuoran Xu, Borong She, Xiaobo Jin, Qiufeng WangICML 2026
Related papers
- Random Scaling and Momentum for Non-smooth Non-convex OptimizationQinzi Zhang, Ashok CutkoskyICML 2024 · 10 citations
- Local Regularizer Improves GeneralizationYikai Zhang, Hui Qu, Dimitris N. Metaxas, Chao ChenAAAI 2020 · 4 citations
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 402 citations
- Linear Regularizers Enforce the Strict Saddle PropertyMatthew Ubl, Matthew Hale, Kasra YazdaniAAAI 2023 · 3 citations
- When Expressivity Meets Trainability: Fewer than Neurons Can WorkJiawei Zhang, Yushun Zhang, Mingyi Hong, Ruoyu Sun et al.NeurIPS 2021 · 11 citations
