Theoretical Analysis of Sparse Optimization with Reparameterization, Weight Decay, and Adaptive Learning Rate
Huangyu Xu, Jingqin Yang, Qianqian Xu, Jiaye Teng
摘要
Sparse optimization is a fundamental challenge in various practical applications. A popular approach to sparse optimization is ℓ p regularization. However, it may encounter optimization instability due to the unbounded gradients when 0 < p < 1. In this paper, we introduce a novel approach to sparse optimization termed ReWA, based on Reparameterization, Weight decay, and Adaptive learning rate. ReWA is closely connected to ℓ p -regularization, yet it unveils a distinct optimization landscape that helps mitigate instability issues. Experiments on CIFAR-10 and ImageNet with ResNets demonstrate that ReWA leads to significant sparsity improvements over the ℓ 1 -regularization approach while preserving test accuracy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- Pruning neural networks without any data by iteratively conserving synaptic flowHidenori Tanaka, Daniel Kunin, Daniel L. K. Yamins, Surya GanguliNeurIPS 2020 · 被引用 884 次
- Sheared LLaMA: Accelerating Language Model Pre-training via Structured PruningMengzhou Xia, Tianyu Gao, Zhiyuan Zeng, Danqi ChenICLR 2024 · 被引用 453 次
- A Function Space View of Bounded Norm Infinite Width ReLU Nets: The Multivariate CaseGreg Ongie, Rebecca Willett, Daniel Soudry, Nathan SrebroICLR 2020 · 被引用 172 次
- Winning the Lottery with Continuous SparsificationPedro Savarese, Hugo Silva, Michael MaireNeurIPS 2020 · 被引用 162 次
- Forget-free Continual Learning with Winning SubnetworksHaeyong Kang, Rusty John Lloyd Mina, Sultan Rizky Hikmawan Madjid, Jaehong Yoon 等ICML 2022 · 被引用 159 次
相关 Paper
- spred: Solving L1 Penalty with SGDLiu Ziyin, Zihao WangICML 2023 · 被引用 23 次
- Rethinking Weight Decay for Robust Fine-Tuning of Foundation ModelsJunjiao Tian, Chengyue Huang, Zsolt KiraNeurIPS 2024 · 被引用 13 次
- Initialization of Large Language Models via Reparameterization to Mitigate Loss SpikesKosuke Nishida, Kyosuke Nishida, Kuniko SaitoEMNLP 2024 · 被引用 2 次
- Sparsity Outperforms Low-Rank Projections in Few-Shot AdaptationNairouz Mrabah, Nicolas Richet, Ismail Ben Ayed, Eric GrangerICCV 2025
- Global Minimizers of ℓp-Regularized Objectives Yield the Sparsest ReLU Neural NetworksJulia B. Nakhleh, Robert D. NowakNeurIPS 2025
