Implicit Bias of Gradient Descent for Non-Homogeneous Deep Networks
Yuhang Cai, Kangjie Zhou, Jingfeng Wu, Song Mei, Michael Lindsey, Peter L. Bartlett
摘要
We establish the asymptotic implicit bias of gradient descent (GD) for generic non-homogeneous deep networks under exponential loss. Specifically, we characterize three key properties of GD iterates starting from a sufficiently small empirical risk, where the threshold is determined by a measure of the network's non-homogeneity. First, we show that a normalized margin induced by the GD iterates increases nearly monotonically. Second, we prove that while the norm of the GD iterates diverges to infinity, the iterates themselves converge in direction. Finally, we establish that this directional limit satisfies the Karush-Kuhn-Tucker (KKT) conditions of a margin maximization problem. Prior works on implicit bias have focused exclusively on homogeneous networks; in contrast, our results apply to a broad class of nonhomogeneous networks satisfying a mild nearhomogeneity condition. In particular, our results apply to networks with residual connections and non-homogeneous activation functions, thereby resolving an open problem posed by Ji & Telgarsky (2020).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Hard labels sampled from sparse targets mislead rotation invariant algorithmsAvrajit Ghosh, Bin Yu, Manfred Warmuth, Peter BartlettICML 2026 · 被引用 1 次
- The Implicit Bias of Steepest Descent with Mini-batch Stochastic GradientJichu Li, Xuan Tang, Difan ZouICML 2026 · 被引用 1 次
- Strong Correlations Induce Cause Only Predictions in Transformer TrainingHaihan Zhang, Yimu Zhang, Cong FangICLR 2026
- Data Reconstruction: Identifiability and Optimization with Sample SplittingYujie Shen, Zihan Wang, Jian Qian, Qi LeiICML 2026
- The Implicit Bias of Depth: From Neural Collapse to Softmax CodesConnall Garrod, Jonathan Keating, Christos ThrampoulidisICML 2026
它引用的顶会 Paper7
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 被引用 402 次
- Directional convergence and alignment in deep learningZiwei Ji, Matus TelgarskyNeurIPS 2020 · 被引用 226 次
- Gradient Descent on Two-layer Nets: Margin Maximization and Simplicity BiasKaifeng Lyu, Zhiyuan Li, Runzhe Wang, Sanjeev AroraNeurIPS 2021 · 被引用 94 次
- Implicit Bias of Gradient Descent for Logistic Regression at the Edge of StabilityJingfeng Wu, Vladimir Braverman, Jason D. LeeNeurIPS 2023 · 被引用 46 次
- Large Stepsize Gradient Descent for Non-Homogeneous Two-Layer Networks: Margin Improvement and Fast OptimizationYuhang Cai, Jingfeng Wu, Song Mei, Michael Lindsey 等NeurIPS 2024 · 被引用 20 次
相关 Paper
- Implicit Bias of Gradient Descent for Two-layer ReLU and Leaky ReLU Networks on Nearly-orthogonal DataYiwen Kou, Zixiang Chen, Quanquan GuNeurIPS 2023 · 被引用 24 次
- Flavors of Margin: Implicit Bias of Steepest Descent in Homogeneous Neural NetworksNikolaos Tsilivis, Gal Vardi, Julia KempeICLR 2025
- The Asymmetric Maximum Margin Bias of Quasi-Homogeneous Neural NetworksDaniel Kunin, Atsushi Yamamura, Chao Ma, Surya GanguliICLR 2023 · 被引用 1 次
- Implicit Bias of Adversarial Training for Deep Neural NetworksBochen Lv, Zhanxing ZhuICLR 2022 · 被引用 8 次
- On Margin Maximization in Linear and ReLU NetworksGal Vardi, Ohad Shamir, Nati SrebroNeurIPS 2022 · 被引用 37 次
