Benign Overfitting in Two-layer Convolutional Neural Networks
Yuan Cao, Zixiang Chen, Misha Belkin, Quanquan Gu
Abstract
Modern neural networks often have great expressive power and can be trained to overfit the training data, while still achieving a good test performance. This phenomenon is referred to as"benign overfitting". Recently, there emerges a line of works studying"benign overfitting"from the theoretical perspective. However, they are limited to linear models or kernel/random feature models, and there is still a lack of theoretical understanding about when and how benign overfitting occurs in neural networks. In this paper, we study the benign overfitting phenomenon in training a two-layer convolutional neural network (CNN). We show that when the signal-to-noise ratio satisfies a certain condition, a two-layer CNN trained by gradient descent can achieve arbitrarily small training and test loss. On the other hand, when this condition does not hold, overfitting becomes harmful and the obtained CNN can only achieve a constant level test loss. These together demonstrate a sharp phase transition between benign overfitting and harmful overfitting, driven by the signal-to-noise ratio. To the best of our knowledge, this is the first work that precisely characterizes the conditions under which benign overfitting can occur in training convolutional neural networks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0fc676d6-ac5f-47c6-b0ab-37a0ee1607c5Cited by top-tier papers73
- Emergent and Predictable Memorization in Large Language ModelsStella Biderman, USVSN Sai Prashanth, Lintang Sutawika, Hailey Schoelkopf et al.NeurIPS 2023 · 205 citations
- Towards Understanding the Mixture-of-Experts Layer in Deep LearningZixiang Chen, Yihe Deng, Yue Wu, Quanquan Gu et al.NeurIPS 2022 · 199 citations
- On the Role of Attention in Prompt-tuningSamet Oymak, Ankit Singh Rawat, Mahdi Soltanolkotabi, Christos ThrampoulidisICML 2023 · 67 citations
- Going Beyond Linear Mode Connectivity: The Layerwise Linear Feature ConnectivityZhanpeng Zhou, Yongyi Yang, Xiaojiang Yang, Junchi Yan et al.NeurIPS 2023 · 56 citations
- Benign, Tempered, or Catastrophic: Toward a Refined Taxonomy of OverfittingNeil Mallinar, James B. Simon, Amirhesam Abedsoltan, Parthe Pandit et al.NeurIPS 2022 · 53 citations
Builds on9
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 402 citations
- Polylogarithmic width suffices for gradient descent to achieve arbitrarily small test error with shallow ReLU networksZiwei Ji, Matus TelgarskyICLR 2020 · 193 citations
- Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep LearningZeyuan Allen-Zhu, Yuanzhi LiICLR 2023 · 151 citations
- The Neural Tangent Kernel in High Dimensions: Triple Descent and a Multi-Scale Theory of GeneralizationBen Adlam, Jeffrey PenningtonICML 2020 · 133 citations
- A random matrix analysis of random Fourier features: beyond the Gaussian kernel, a precise phase transition, and the corresponding double descentZhenyu Liao, Romain Couillet, Michael W. MahoneyNeurIPS 2020 · 102 citations
Related papers
- Benign Overfitting in Two-layer ReLU Convolutional Neural NetworksYiwen Kou, Zixiang Chen, Yuanzhou Chen, Quanquan GuICML 2023 · 32 citations
- Benign overfitting in leaky ReLU networks with moderate input dimensionKedar Karhadkar, Erin George, Michael Murray, Guido F. Montúfar et al.NeurIPS 2024 · 5 citations
- Benign Overfitting in Single-Head AttentionRoey Magen, Shuning Shang, Zhiwei Xu, Spencer Frei et al.NeurIPS 2025 · 12 citations
- Directional Convergence, Benign Overfitting of Gradient Descent in leaky ReLU two-layer Neural NetworksIchiro HashimotoICLR 2026
- From Tempered to Benign Overfitting in ReLU Neural NetworksGuy Kornowski, Gilad Yehudai, Ohad ShamirNeurIPS 2023 · 18 citations
