Benign Overfitting in Two-layer ReLU Convolutional Neural Networks
Yiwen Kou, Zixiang Chen, Yuanzhou Chen, Quanquan Gu
摘要
Modern deep learning models with great expressive power can be trained to overfit the training data but still generalize well. This phenomenon is referred to as benign overfitting. Recently, a few studies have attempted to theoretically understand benign overfitting in neural networks. However, these works are either limited to neural networks with smooth activation functions or to the neural tangent kernel regime. How and when benign overfitting can occur in ReLU neural networks remains an open problem. In this work, we seek to answer this question by establishing algorithm-dependent risk bounds for learning two-layer ReLU convolutional neural networks with label-flipping noise. We show that, under mild conditions, the neural network trained by gradient descent can achieve near-zero training loss and Bayes optimal test risk. Our result also reveals a sharp transition between benign and harmful overfitting under different conditions on data distribution in terms of test risk. Experiments on synthetic data back up our theory.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper32
- Federated Learning from Vision-Language Foundation Models: Theoretical Analysis and MethodBikang Pan, Wei Huang, Ye ShiNeurIPS 2024 · 被引用 28 次
- Diffusion Models Demand Contrastive Guidance for Adversarial Purification to AdvanceMingyuan Bai, Wei Huang, Tenghui Li, Andong Wang 等ICML 2024 · 被引用 18 次
- Stable Minima Cannot Overfit in Univariate ReLU Networks: Generalization by Large Step SizesDan Qiao, Kaiqi Zhang, Esha Singh, Daniel Soudry 等NeurIPS 2024 · 被引用 15 次
- Minimum-Norm Interpolation Under Covariate ShiftNeil Mallinar, Austin Zane, Spencer Frei, Bin YuICML 2024 · 被引用 13 次
- Benign Overfitting in Single-Head AttentionRoey Magen, Shuning Shang, Zhiwei Xu, Spencer Frei 等NeurIPS 2025 · 被引用 12 次
它引用的顶会 Paper7
- Towards Understanding Ensemble, Knowledge Distillation and Self-Distillation in Deep LearningZeyuan Allen-Zhu, Yuanzhi LiICLR 2023 · 被引用 151 次
- The Neural Tangent Kernel in High Dimensions: Triple Descent and a Multi-Scale Theory of GeneralizationBen Adlam, Jeffrey PenningtonICML 2020 · 被引用 133 次
- Benign Overfitting in Two-layer Convolutional Neural NetworksYuan Cao, Zixiang Chen, Misha Belkin, Quanquan GuNeurIPS 2022 · 被引用 121 次
- Feature Purification: How Adversarial Training Performs Robust Deep LearningZeyuan Allen-Zhu, Yuanzhi LiFOCS 2021 · 被引用 83 次
- Risk Bounds for Over-parameterized Maximum Margin Classification on Sub-Gaussian MixturesYuan Cao, Quanquan Gu, Mikhail BelkinNeurIPS 2021 · 被引用 57 次
相关 Paper
- Benign Overfitting in Two-Layer ReLU Convolutional Neural Networks for XOR DataXuran Meng, Difan Zou, Yuan CaoICML 2024 · 被引用 11 次
- Benign Overfitting in Deep Neural Networks under Lazy TrainingZhenyu Zhu, Fanghui Liu, Grigorios Chrysos, Francesco Locatello 等ICML 2023 · 被引用 12 次
- Training shallow ReLU networks on noisy data using hinge loss: when do we overfit and is it benign?Erin George, Michael Murray, William Swartworth, Deanna NeedellNeurIPS 2023 · 被引用 9 次
- Benign, Tempered, or Catastrophic: Toward a Refined Taxonomy of OverfittingNeil Mallinar, James B. Simon, Amirhesam Abedsoltan, Parthe Pandit 等NeurIPS 2022 · 被引用 53 次
- Directional Convergence, Benign Overfitting of Gradient Descent in leaky ReLU two-layer Neural NetworksIchiro HashimotoICLR 2026
