Combining Explicit and Implicit Regularization for Efficient Learning in Deep Networks
Dan Zhao
摘要
Works on implicit regularization have studied gradient trajectories during the optimization process to explain why deep networks favor certain kinds of solutions over others. In deep linear networks, it has been shown that gradient descent implicitly regularizes toward low-rank solutions on matrix completion/factorization tasks. Adding depth not only improves performance on these tasks but also acts as an accelerative pre-conditioning that further enhances this bias towards lowrankedness. Inspired by this, we propose an explicit penalty to mirror this implicit bias which only takes effect with certain adaptive gradient optimizers (e.g. Adam). This combination can enable a degenerate single-layer network to achieve lowrank approximations with generalization error comparable to deep linear networks, making depth no longer necessary for learning. The single-layer network also performs competitively or out-performs various approaches for matrix completion over a range of parameter and data regimes despite its simplicity. Together with an optimizer's inductive bias, our findings suggest that explicit regularization can play a role in designing different, desirable forms of regularization and that a more nuanced understanding of this interplay may be necessary. 36th Conference on Neural Information Processing Systems (NeurIPS 2022).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Pretraining with Random Noise for Fast and Robust Learning without Weight TransportJeonghwan Cheon, Sang Wan Lee, Se-Bum PaikNeurIPS 2024 · 被引用 8 次
- Dependency Parsing is More Parameter-Efficient with NormalizationPaolo Gajo, Domenic Rosati, Hassan Sajjad, Alberto Barrón-CedeñoNeurIPS 2025
它引用的顶会 Paper7
- On the Variance of the Adaptive Learning Rate and BeyondLiyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen 等ICLR 2020 · 被引用 2,210 次
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan 等ICLR 2020 · 被引用 705 次
- Implicit Gradient RegularizationDavid G. T. Barrett, Benoit DherinICLR 2021 · 被引用 235 次
- On the Implicit Bias of Initialization Shape: Beyond Infinitesimal Mirror DescentShahar Azulay, Edward Moroshko, Mor Shpigel Nacson, Blake E. Woodworth 等ICML 2021 · 被引用 85 次
- Implicit Bias of Gradient Descent on Reparametrized Models: On Equivalence to Mirror DescentZhiyuan Li, Tianhao Wang, Jason D. Lee, Sanjeev AroraNeurIPS 2022 · 被引用 49 次
相关 Paper
- Implicit Regularization in Deep Learning May Not Be Explainable by NormsNoam Razin, Nadav CohenNeurIPS 2020 · 被引用 178 次
- Implicit Sparse Regularization: The Impact of Depth and Early StoppingJiangyuan Li, Thanh Van Nguyen, Chinmay Hegde, Ka Wai WongNeurIPS 2021 · 被引用 43 次
- Implicit Regularization with Polynomial Growth in Deep Tensor FactorizationKais Hariz, Hachem Kadri, Stéphane Ayache, Maher Moakher 等ICML 2022 · 被引用 4 次
- Implicit Bias in Leaky ReLU Networks Trained on High-Dimensional DataSpencer Frei, Gal Vardi, Peter L. Bartlett, Nathan Srebro 等ICLR 2023 · 被引用 5 次
- Implicit Bias and Loss of Plasticity in Matrix Completion: Depth Promotes Low-RanknessBaekrok Shin, Chulhee YunICLR 2026
