Implicit Sparse Regularization: The Impact of Depth and Early Stopping
Jiangyuan Li, Thanh Van Nguyen, Chinmay Hegde, Ka Wai Wong
摘要
In this paper, we study the implicit bias of gradient descent for sparse regression. We extend results on regression with quadratic parametrization, which amounts to depth-2 diagonal linear networks, to more general depth-N networks, under more realistic settings of noise and correlated designs. We show that early stopping is crucial for gradient descent to converge to a sparse model, a phenomenon that we call implicit sparse regularization. This result is in sharp contrast to known results for noiseless and uncorrelated-design cases. We characterize the impact of depth and early stopping and show that for a general depth parameter N , gradient descent with early stopping achieves minimax optimal sparse recovery with sufficiently small initialization w 0 and step size η. In particular, we show that increasing depth enlarges the scale of working initialization and the early-stopping window so that this implicit sparse regularization effect is more likely to take place.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper18
- Robust Training under Label Noise by Over-parameterizationSheng Liu, Zhihui Zhu, Qing Qu, Chong YouICML 2022 · 被引用 152 次
- On the Convergence of Gradient Flow on Multi-layer Linear ModelsHancheng Min, René Vidal, Enrique MalladaICML 2023 · 被引用 13 次
- The Implicit Regularization of Momentum Gradient Descent in Overparametrized ModelsLi Wang, Zhiguo Fu, Yingcong Zhou, Zili YanAAAI 2023 · 被引用 9 次
- Blessing of Depth in Linear Regression: Deeper Models Have Flatter Landscape Around the True SolutionJianhao Ma, Salar FattahiNeurIPS 2022 · 被引用 7 次
- Implicit Regularization Leads to Benign Overfitting for Sparse Linear RegressionMo Zhou, Rong GeICML 2023 · 被引用 4 次
它引用的顶会 Paper3
- Towards Resolving the Implicit Bias of Gradient Descent for Matrix Factorization: Greedy Low-Rank LearningZhiyuan Li, Yuping Luo, Kaifeng LyuICLR 2021 · 被引用 155 次
- Implicit Bias in Deep Linear Classification: Initialization Scale vs Training AccuracyEdward Moroshko, Blake E. Woodworth, Suriya Gunasekar, Jason D. Lee 等NeurIPS 2020 · 被引用 98 次
- The Implicit Bias of Depth: How Incremental Learning Drives GeneralizationDaniel Gissin, Shai Shalev-Shwartz, Amit DanielyICLR 2020 · 被引用 90 次
相关 Paper
- Implicit Regularization for Group SparsityJiangyuan Li, Thanh Van Nguyen, Chinmay Hegde, Raymond K. W. WongICLR 2023 · 被引用 2 次
- (S)GD over Diagonal Linear Networks: Implicit bias, Large Stepsizes and Edge of StabilityMathieu Even, Scott Pesme, Suriya Gunasekar, Nicolas FlammarionNeurIPS 2023 · 被引用 42 次
- Implicit Bias of the Step Size in Linear Diagonal Neural NetworksMor Shpigel Nacson, Kavya Ravichandran, Nathan Srebro, Daniel SoudryICML 2022 · 被引用 57 次
- Implicit Bias of (Stochastic) Gradient Descent for Rank-1 Linear Neural NetworkBochen Lyu, Zhanxing ZhuNeurIPS 2023 · 被引用 5 次
- Benefits of Early Stopping in Gradient Descent for Overparameterized Logistic RegressionJingfeng Wu, Peter L. Bartlett, Matus Telgarsky, Bin YuICML 2025
