Implicit Bias of Gradient Descent for Non-Homogeneous Deep Networks
Yuhang Cai, Kangjie Zhou, Jingfeng Wu, Song Mei, Michael Lindsey, Peter L. Bartlett
Abstract
We establish the asymptotic implicit bias of gradient descent (GD) for generic non-homogeneous deep networks under exponential loss. Specifically, we characterize three key properties of GD iterates starting from a sufficiently small empirical risk, where the threshold is determined by a measure of the network's non-homogeneity. First, we show that a normalized margin induced by the GD iterates increases nearly monotonically. Second, we prove that while the norm of the GD iterates diverges to infinity, the iterates themselves converge in direction. Finally, we establish that this directional limit satisfies the Karush-Kuhn-Tucker (KKT) conditions of a margin maximization problem. Prior works on implicit bias have focused exclusively on homogeneous networks; in contrast, our results apply to a broad class of nonhomogeneous networks satisfying a mild nearhomogeneity condition. In particular, our results apply to networks with residual connections and non-homogeneous activation functions, thereby resolving an open problem posed by Ji & Telgarsky (2020).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d356afe4-ea88-4f6f-9acd-32c114b61c0cCited by top-tier papers8
- Hard labels sampled from sparse targets mislead rotation invariant algorithmsAvrajit Ghosh, Bin Yu, Manfred Warmuth, Peter BartlettICML 2026 · 1 citation
- The Implicit Bias of Steepest Descent with Mini-batch Stochastic GradientJichu Li, Xuan Tang, Difan ZouICML 2026 · 1 citation
- Strong Correlations Induce Cause Only Predictions in Transformer TrainingHaihan Zhang, Yimu Zhang, Cong FangICLR 2026
- Data Reconstruction: Identifiability and Optimization with Sample SplittingYujie Shen, Zihan Wang, Jian Qian, Qi LeiICML 2026
- The Implicit Bias of Depth: From Neural Collapse to Softmax CodesConnall Garrod, Jonathan Keating, Christos ThrampoulidisICML 2026
Builds on7
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 402 citations
- Directional convergence and alignment in deep learningZiwei Ji, Matus TelgarskyNeurIPS 2020 · 226 citations
- Gradient Descent on Two-layer Nets: Margin Maximization and Simplicity BiasKaifeng Lyu, Zhiyuan Li, Runzhe Wang, Sanjeev AroraNeurIPS 2021 · 94 citations
- Implicit Bias of Gradient Descent for Logistic Regression at the Edge of StabilityJingfeng Wu, Vladimir Braverman, Jason D. LeeNeurIPS 2023 · 46 citations
- Large Stepsize Gradient Descent for Non-Homogeneous Two-Layer Networks: Margin Improvement and Fast OptimizationYuhang Cai, Jingfeng Wu, Song Mei, Michael Lindsey et al.NeurIPS 2024 · 20 citations
Related papers
- Implicit Bias of Gradient Descent for Two-layer ReLU and Leaky ReLU Networks on Nearly-orthogonal DataYiwen Kou, Zixiang Chen, Quanquan GuNeurIPS 2023 · 24 citations
- Flavors of Margin: Implicit Bias of Steepest Descent in Homogeneous Neural NetworksNikolaos Tsilivis, Gal Vardi, Julia KempeICLR 2025
- The Asymmetric Maximum Margin Bias of Quasi-Homogeneous Neural NetworksDaniel Kunin, Atsushi Yamamura, Chao Ma, Surya GanguliICLR 2023 · 1 citation
- Implicit Bias of Adversarial Training for Deep Neural NetworksBochen Lv, Zhanxing ZhuICLR 2022 · 8 citations
- On Margin Maximization in Linear and ReLU NetworksGal Vardi, Ohad Shamir, Nati SrebroNeurIPS 2022 · 37 citations
