Adaptive Proximal Gradient Methods for Structured Neural Networks
Jihun Yun, Aurélie C. Lozano, Eunho Yang
Abstract
We consider the training of structured neural networks where the regularizer can be non-smooth and possibly non-convex. While popular machine learning libraries have resorted to stochastic (adaptive) subgradient approaches, the use of proximal gradient methods in the stochastic setting has been little explored and warrants further study, in particular regarding the incorporation of adaptivity. Towards this goal, we present a general framework of stochastic proximal gradient descent methods that allows for arbitrary positive preconditioners and lower semi-continuous regularizers. We derive two important instances of our framework: (i) the first proximal version of ADAM, one of the most popular adaptive SGD algorithm, and (ii) a revised version of PROXQUANT [1] for quantization-specific regularizers, which improves upon the original approach by incorporating the effect of preconditioners in the proximal mapping computations. We provide convergence guarantees for our framework and show that adaptive gradient methods can have faster convergence in terms of constant than vanilla SGD for sparse data. Lastly, we demonstrate the superiority of stochastic proximal methods compared to subgradient-based approaches via extensive experiments. Interestingly, our results indicate that the benefit of proximal approaches over sub-gradient counterparts is more pronounced for non-convex regularizers than for convex ones.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext baf6f2b3-1374-40ef-a654-4f213a0a1201Cited by top-tier papers7
- On the Convergence of Black-Box Variational InferenceKyurae Kim, Jisu Oh, Kaiwen Wu, Yi-An Ma et al.NeurIPS 2023 · 27 citations
- Training Structured Neural Networks Through Manifold Identification and Variance ReductionZih-Syuan Huang, Ching-pei LeeICLR 2022 · 10 citations
- TEDDY: Trimming Edges with Degree-based Discrimination StrategyHyunjin Seo, Jihun Yun, Eunho YangICLR 2024 · 3 citations
- Double Variance Reduction: A Smoothing Trick for Composite Optimization Problems without First-Order GradientHao Di, Haishan Ye, Yueling Zhang, Xiangyu Chang et al.ICML 2024 · 2 citations
- The Devil Is in Gradient Entanglement: Energy-Aware Gradient Coordinator for Robust Generalized Category DiscoveryHaiyang Zheng, Nan Pu, Yaqi Cai, Teng Long et al.CVPR 2026 · 1 citation
Builds on2
- Economy Statistical Recurrent Units For Inferring Nonlinear Granger CausalitySaurabh Khanna, Vincent Y. F. TanICLR 2020 · 93 citations
- ProxSGD: Training Structured Neural Networks under Regularization and ConstraintsYang Yang, Yaxiong Yuan, Avraam Chatzimichailidis, Ruud J. G. van Sloun et al.ICLR 2020 · 34 citations
Related papers
- ASGO: Adaptive Structured Gradient OptimizationKang An, Yuxing Liu, Rui Pan, Yi Ren et al.NeurIPS 2025 · 58 citations
- Amortized Proximal OptimizationJuhan Bae, Paul Vicol, Jeff Z. HaoChen, Roger B. GrosseNeurIPS 2022 · 15 citations
- Regularized Adaptive Momentum Dual Averaging with an Efficient Inexact Subproblem Solver for Training Structured Neural NetworkZih-Syuan Huang, Ching-pei LeeNeurIPS 2024
- Local Regularizer Improves GeneralizationYikai Zhang, Hui Qu, Dimitris N. Metaxas, Chao ChenAAAI 2020 · 4 citations
- Structured Inverse-Free Natural Gradient Descent: Memory-Efficient & Numerically-Stable KFACWu Lin, Felix Dangel, Runa Eschenhagen, Kirill Neklyudov et al.ICML 2024 · 7 citations
