Faster Directional Convergence of Linear Neural Networks under Spherically Symmetric Data
Dachao Lin, Ruoyu Sun, Zhihua Zhang
Abstract
In this paper, we study gradient methods for training deep linear neural networks with binary cross-entropy loss. In particular, we show global directional convergence guarantees from a polynomial rate to a linear rate for (deep) linear networks with spherically symmetric data distribution, which can be viewed as a specific zero-margin dataset. Our results do not require the assumptions in other works such as small initial loss, presumed convergence of weight direction, or overparameterization. We also characterize our findings in experiments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9af6e2bf-1dac-4e6a-a35d-b13c87ad383bCited by top-tier papers2
- NTK-SAP: Improving neural network pruning by aligning training dynamicsYite Wang, Dawei Li, Ruoyu SunICLR 2023 · 2 citations
- On Non-local Convergence Analysis of Deep Linear NetworksKun Chen, Dachao Lin, Zhihua ZhangICML 2022 · 1 citation
Builds on6
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 402 citations
- Directional convergence and alignment in deep learningZiwei Ji, Matus TelgarskyNeurIPS 2020 · 226 citations
- The Early Phase of Neural Network TrainingJonathan Frankle, David J. Schwab, Ari S. MorcosICLR 2020 · 199 citations
- Provable Benefit of Orthogonal Initialization in Optimizing Deep Linear NetworksWei Hu, Lechao Xiao, Jeffrey PenningtonICLR 2020 · 136 citations
- A unifying view on implicit bias in training linear neural networksChulhee Yun, Shankar Krishnan, Hossein MobahiICLR 2021 · 94 citations
Related papers
- Convergence Rates of Non-Convex Stochastic Gradient Descent Under a Generic Lojasiewicz Condition and Local SmoothnessKevin Scaman, Cédric Malherbe, Ludovic Dos SantosICML 2022 · 24 citations
- Towards Understanding Learning in Neural Networks with Linear TeachersRoei Sarussi, Alon Brutzkus, Amir GlobersonICML 2021 · 24 citations
- The Implicit Bias of Adam on Separable DataChenyang Zhang, Difan Zou, Yuan CaoNeurIPS 2024 · 37 citations
- On the Global Convergence of Training Deep Linear ResNetsDifan Zou, Philip M. Long, Quanquan GuICLR 2020 · 44 citations
- Gradient flow dynamics of shallow ReLU networks for square loss and orthogonal inputsEtienne Boursier, Loucas Pillaud-Vivien, Nicolas FlammarionNeurIPS 2022 · 92 citations
