Theoretical Analysis of Sparse Optimization with Reparameterization, Weight Decay, and Adaptive Learning Rate
Huangyu Xu, Jingqin Yang, Qianqian Xu, Jiaye Teng
Abstract
Sparse optimization is a fundamental challenge in various practical applications. A popular approach to sparse optimization is ℓ p regularization. However, it may encounter optimization instability due to the unbounded gradients when 0 < p < 1. In this paper, we introduce a novel approach to sparse optimization termed ReWA, based on Reparameterization, Weight decay, and Adaptive learning rate. ReWA is closely connected to ℓ p -regularization, yet it unveils a distinct optimization landscape that helps mitigate instability issues. Experiments on CIFAR-10 and ImageNet with ResNets demonstrate that ReWA leads to significant sparsity improvements over the ℓ 1 -regularization approach while preserving test accuracy.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5e30efaf-fb1a-423f-a30a-4d20a8fe5e38Builds on20
- Pruning neural networks without any data by iteratively conserving synaptic flowHidenori Tanaka, Daniel Kunin, Daniel L. K. Yamins, Surya GanguliNeurIPS 2020 · 884 citations
- Sheared LLaMA: Accelerating Language Model Pre-training via Structured PruningMengzhou Xia, Tianyu Gao, Zhiyuan Zeng, Danqi ChenICLR 2024 · 453 citations
- A Function Space View of Bounded Norm Infinite Width ReLU Nets: The Multivariate CaseGreg Ongie, Rebecca Willett, Daniel Soudry, Nathan SrebroICLR 2020 · 172 citations
- Winning the Lottery with Continuous SparsificationPedro Savarese, Hugo Silva, Michael MaireNeurIPS 2020 · 162 citations
- Forget-free Continual Learning with Winning SubnetworksHaeyong Kang, Rusty John Lloyd Mina, Sultan Rizky Hikmawan Madjid, Jaehong Yoon et al.ICML 2022 · 159 citations
Related papers
- spred: Solving L1 Penalty with SGDLiu Ziyin, Zihao WangICML 2023 · 23 citations
- Rethinking Weight Decay for Robust Fine-Tuning of Foundation ModelsJunjiao Tian, Chengyue Huang, Zsolt KiraNeurIPS 2024 · 13 citations
- Initialization of Large Language Models via Reparameterization to Mitigate Loss SpikesKosuke Nishida, Kyosuke Nishida, Kuniko SaitoEMNLP 2024 · 2 citations
- Sparsity Outperforms Low-Rank Projections in Few-Shot AdaptationNairouz Mrabah, Nicolas Richet, Ismail Ben Ayed, Eric GrangerICCV 2025
- Global Minimizers of ℓp-Regularized Objectives Yield the Sparsest ReLU Neural NetworksJulia B. Nakhleh, Robert D. NowakNeurIPS 2025
