Towards Understanding Learning in Neural Networks with Linear Teachers
Roei Sarussi, Alon Brutzkus, Amir Globerson
Abstract
Can a neural network minimizing cross-entropy learn linearly separable data? Despite progress in the theory of deep learning, this question remains unsolved. Here we prove that SGD globally optimizes this learning problem for a two-layer network with Leaky ReLU activations. The learned network can in principle be very complex. However, empirical evidence suggests that it often turns out to be approximately linear. We provide theoretical support for this phenomenon by proving that if network weights converge to two weight clusters, this will imply an approximately linear decision boundary. Finally, we show a condition on the optimization that leads to weight clustering. We provide empirical results that validate our theoretical analysis.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aba510d2-6103-4722-83ab-a32d06db9e8aCited by top-tier papers15
- Modality Competition: What Makes Joint Training of Multi-modal Network Fail in Deep Learning? (Provably)Yu Huang, Junyang Lin, Chang Zhou, Hongxia Yang et al.ICML 2022 · 168 citations
- Gradient Descent on Two-layer Nets: Margin Maximization and Simplicity BiasKaifeng Lyu, Zhiyuan Li, Runzhe Wang, Sanjeev AroraNeurIPS 2021 · 94 citations
- On Margin Maximization in Linear and ReLU NetworksGal Vardi, Ohad Shamir, Nati SrebroNeurIPS 2022 · 37 citations
- Implicit Regularization in Hierarchical Tensor Factorization and Deep Convolutional Neural NetworksNoam Razin, Asaf Maman, Nadav CohenICML 2022 · 34 citations
- The Double-Edged Sword of Implicit Bias: Generalization vs. Robustness in ReLU NetworksSpencer Frei, Gal Vardi, Peter L. Bartlett, Nati SrebroNeurIPS 2023 · 25 citations
Builds on6
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 402 citations
- Directional convergence and alignment in deep learningZiwei Ji, Matus TelgarskyNeurIPS 2020 · 226 citations
- Polylogarithmic width suffices for gradient descent to achieve arbitrarily small test error with shallow ReLU networksZiwei Ji, Matus TelgarskyICLR 2020 · 193 citations
- Learning Parities with Neural NetworksAmit Daniely, Eran MalachNeurIPS 2020 · 104 citations
- Implicit Bias in Deep Linear Classification: Initialization Scale vs Training AccuracyEdward Moroshko, Blake E. Woodworth, Suriya Gunasekar, Jason D. Lee et al.NeurIPS 2020 · 98 citations
Related papers
- Implicit Bias in Leaky ReLU Networks Trained on High-Dimensional DataSpencer Frei, Gal Vardi, Peter L. Bartlett, Nathan Srebro et al.ICLR 2023 · 5 citations
- The inductive bias of ReLU networks on orthogonally separable dataMary Phuong, Christoph H. LampertICLR 2021 · 53 citations
- Implicit Bias of Gradient Descent for Two-layer ReLU and Leaky ReLU Networks on Nearly-orthogonal DataYiwen Kou, Zixiang Chen, Quanquan GuNeurIPS 2023 · 24 citations
- Neural Collapse under Gradient Flow on Shallow ReLU Networks for Orthogonally Separable DataHancheng Min, Zhihui Zhu, René VidalNeurIPS 2025 · 3 citations
- Early Neuron Alignment in Two-layer ReLU Networks with Small InitializationHancheng Min, Enrique Mallada, René VidalICLR 2024 · 31 citations
