Implicit Sparse Regularization: The Impact of Depth and Early Stopping
Jiangyuan Li, Thanh Van Nguyen, Chinmay Hegde, Ka Wai Wong
Abstract
In this paper, we study the implicit bias of gradient descent for sparse regression. We extend results on regression with quadratic parametrization, which amounts to depth-2 diagonal linear networks, to more general depth-N networks, under more realistic settings of noise and correlated designs. We show that early stopping is crucial for gradient descent to converge to a sparse model, a phenomenon that we call implicit sparse regularization. This result is in sharp contrast to known results for noiseless and uncorrelated-design cases. We characterize the impact of depth and early stopping and show that for a general depth parameter N , gradient descent with early stopping achieves minimax optimal sparse recovery with sufficiently small initialization w 0 and step size η. In particular, we show that increasing depth enlarges the scale of working initialization and the early-stopping window so that this implicit sparse regularization effect is more likely to take place.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers18
- Robust Training under Label Noise by Over-parameterizationSheng Liu, Zhihui Zhu, Qing Qu, Chong YouICML 2022 · 152 citations
- On the Convergence of Gradient Flow on Multi-layer Linear ModelsHancheng Min, René Vidal, Enrique MalladaICML 2023 · 13 citations
- The Implicit Regularization of Momentum Gradient Descent in Overparametrized ModelsLi Wang, Zhiguo Fu, Yingcong Zhou, Zili YanAAAI 2023 · 9 citations
- Blessing of Depth in Linear Regression: Deeper Models Have Flatter Landscape Around the True SolutionJianhao Ma, Salar FattahiNeurIPS 2022 · 7 citations
- Implicit Regularization Leads to Benign Overfitting for Sparse Linear RegressionMo Zhou, Rong GeICML 2023 · 4 citations
Builds on3
- Towards Resolving the Implicit Bias of Gradient Descent for Matrix Factorization: Greedy Low-Rank LearningZhiyuan Li, Yuping Luo, Kaifeng LyuICLR 2021 · 155 citations
- Implicit Bias in Deep Linear Classification: Initialization Scale vs Training AccuracyEdward Moroshko, Blake E. Woodworth, Suriya Gunasekar, Jason D. Lee et al.NeurIPS 2020 · 98 citations
- The Implicit Bias of Depth: How Incremental Learning Drives GeneralizationDaniel Gissin, Shai Shalev-Shwartz, Amit DanielyICLR 2020 · 90 citations
Related papers
- Implicit Regularization for Group SparsityJiangyuan Li, Thanh Van Nguyen, Chinmay Hegde, Raymond K. W. WongICLR 2023 · 2 citations
- (S)GD over Diagonal Linear Networks: Implicit bias, Large Stepsizes and Edge of StabilityMathieu Even, Scott Pesme, Suriya Gunasekar, Nicolas FlammarionNeurIPS 2023 · 42 citations
- Implicit Bias of the Step Size in Linear Diagonal Neural NetworksMor Shpigel Nacson, Kavya Ravichandran, Nathan Srebro, Daniel SoudryICML 2022 · 57 citations
- Implicit Bias of (Stochastic) Gradient Descent for Rank-1 Linear Neural NetworkBochen Lyu, Zhanxing ZhuNeurIPS 2023 · 5 citations
- Benefits of Early Stopping in Gradient Descent for Overparameterized Logistic RegressionJingfeng Wu, Peter L. Bartlett, Matus Telgarsky, Bin YuICML 2025
