spred: Solving L1 Penalty with SGD
Liu Ziyin, Zihao Wang
Abstract
We propose to minimize a generic differentiable objective with constraint using a simple reparametrization and straightforward stochastic gradient descent. Our proposal is the direct generalization of previous ideas that the penalty may be equivalent to a differentiable reparametrization with weight decay. We prove that the proposed method, spred, is an exact differentiable solver of and that the reparametrization trick is completely ``benign"for a generic nonconvex function. Practically, we demonstrate the usefulness of the method in (1) training sparse neural networks to perform gene selection tasks, which involves finding relevant features in a very high dimensional space, and (2) neural network compression task, to which previous attempts at applying the -penalty have been unsuccessful. Conceptually, our result bridges the gap between the sparsity in deep learning and conventional statistical learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dd333596-07f9-414e-955a-89e4f53d468aCited by top-tier papers1
Ask how each one uses itBuilds on4
- Pruning neural networks without any data by iteratively conserving synaptic flowHidenori Tanaka, Daniel Kunin, Daniel L. K. Yamins, Surya GanguliNeurIPS 2020 · 884 citations
- Soft Threshold Weight Reparameterization for Learnable SparsityAditya Kusupati, Vivek Ramanujan, Raghav Somani, Mitchell Wortsman et al.ICML 2020 · 266 citations
- Smooth Bilevel Programming for Sparse RegularizationClarice Poon, Gabriel PeyréNeurIPS 2021 · 23 citations
- SGD Can Converge to Local MaximaLiu Ziyin, Botao Li, James B. Simon, Masahito UedaICLR 2022 · 18 citations
Related papers
- Deep Weight Factorization: Sparse Learning Through the Lens of Artificial SymmetriesChris Kolb, Tobias Weber, Bernd Bischl, David RügamerICLR 2025
- Global Minimizers of ℓp-Regularized Objectives Yield the Sparsest ReLU Neural NetworksJulia B. Nakhleh, Robert D. NowakNeurIPS 2025
- Sparsifying Networks via Subdifferential InclusionSagar Verma, Jean-Christophe PesquetICML 2021 · 15 citations
- Global Optimality Beyond Two Layers: Training Deep ReLU Networks via Convex ProgramsTolga Ergen, Mert PilanciICML 2021 · 35 citations
- Efficient Neural Network Training via Forward and Backward Propagation SparsificationXiao Zhou, Weizhong Zhang, Zonghao Chen, Shizhe Diao et al.NeurIPS 2021 · 57 citations
