Implicit Regularization and Convergence for Weight Normalization
Xiaoxia Wu, Edgar Dobriban, Tongzheng Ren, Shanshan Wu, Zhiyuan Li, Suriya Gunasekar, Rachel A. Ward, Qiang Liu
摘要
Normalization methods such as batch [Ioffe and Szegedy, 2015] , weight [Salimans and Kingma, 2016] , instance [Ulyanov et al., 2016] , and layer normalization [Ba et al., 2016] have been widely used in modern machine learning. Here, we study the weight normalization (WN) method [Salimans and Kingma, 2016 ] and a variant called reparametrized projected gradient descent (rPGD) for overparametrized least squares regression. WN and rPGD reparametrize the weights with a scale g and a unit vector w and thus the objective function becomes non-convex. We show that this non-convex formulation has beneficial regularization effects compared to gradient descent on the original objective. These methods adaptively regularize the weights and converge close to the minimum 2 norm solution, even for initializations far from zero. For certain stepsizes of g and w, we show that they can converge close to the minimum norm solution. This is different from the behavior of gradient descent, which converges to the minimum norm solution only when started at a point in the range space of the feature matrix, and is thus more sensitive to initialization.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- The Role of Permutation Invariance in Linear Mode Connectivity of Neural NetworksRahim Entezari, Hanie Sedghi, Olga Saukh, Behnam NeyshaburICLR 2022 · 被引用 301 次
- Understanding the Generalization Benefit of Normalization Layers: Sharpness ReductionKaifeng Lyu, Zhiyuan Li, Sanjeev AroraNeurIPS 2022 · 被引用 111 次
- Implicit Regularization in Tensor FactorizationNoam Razin, Asaf Maman, Nadav CohenICML 2021 · 被引用 60 次
- The Implicit Bias of Adam on Separable DataChenyang Zhang, Difan Zou, Yuan CaoNeurIPS 2024 · 被引用 37 次
- Fast Mixing of Stochastic Gradient Descent with Normalization and Weight DecayZhiyuan Li, Tianhao Wang, Dingli YuNeurIPS 2022 · 被引用 19 次
它引用的顶会 Paper3
- Ridge Regression: Structure, Cross-Validation, and SketchingSifan Liu, Edgar DobribanICLR 2020 · 被引用 52 次
- Dropout: Explicit Forms and Capacity ControlRaman Arora, Peter L. Bartlett, Poorya Mianjy, Nathan SrebroICML 2021 · 被引用 43 次
- Optimization Theory for ReLU Neural Networks Trained with Normalization LayersYonatan Dukler, Quanquan Gu, Guido MontúfarICML 2020 · 被引用 30 次
相关 Paper
- On the Benefits of Weight Normalization for Overparameterized Matrix SensingYudong Wei, Liang Zhang, Bingcong Li, Niao HeICLR 2026 · 被引用 4 次
- AdamP: Slowing Down the Slowdown for Momentum Optimizers on Scale-invariant WeightsByeongho Heo, Sanghyuk Chun, Seong Joon Oh, Dongyoon Han 等ICLR 2021 · 被引用 165 次
- Investigating the Role of Weight Decay in Enhancing Nonconvex SGDTao Sun, Yuhao Huang, Li Shen, Kele Xu 等CVPR 2025
- Direction Matters: On the Implicit Bias of Stochastic Gradient Descent with Moderate Learning RateJingfeng Wu, Difan Zou, Vladimir Braverman, Quanquan GuICLR 2021 · 被引用 18 次
- Understanding the Disharmony between Weight Normalization Family and Weight DecayXiang Li, Shuo Chen, Jian YangAAAI 2020 · 被引用 18 次
