Simplicity Bias and Optimization Threshold in Two-Layer ReLU Networks
Etienne Boursier, Nicolas Flammarion
摘要
Understanding generalization of overparametrized models remains a fundamental challenge in machine learning. The literature mostly studies generalization from an interpolation point of view, taking convergence towards a global minimum of the training loss for granted. This interpolation paradigm does not seem valid for complex tasks such as in-context learning or diffusion. It has instead been empirically observed that the trained models go from global minima to spurious local minima of the training loss as the number of training samples becomes larger than some level we call optimization threshold. This paper explores theoretically this phenomenon in the context of two-layer ReLU networks. We demonstrate that, despite overparametrization, networks might converge towards simpler solutions rather than interpolating training data, which leads to a drastic improvement on the test loss. Our analysis relies on the so called early alignment phase, during which neurons align toward specific directions. This directional alignment leads to a simplicity bias, wherein the network approximates the ground truth model without converging to the global minimum of the training loss. Our results suggest this bias, resulting in an optimization threshold from which interpolation is not reached anymore, is beneficial and enhances the generalization of trained models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- On the Closed-Form of Flow Matching: Generalization Does Not Arise from Target StochasticityQuentin Bertrand, Anne Gagneux, Mathurin Massias, Rémi EmonetNeurIPS 2025 · 被引用 49 次
- Saddle-to-Saddle Dynamics Explains A Simplicity Bias Across Neural Network ArchitecturesYedi Zhang, Andrew M. Saxe, Peter E. LathamICLR 2026 · 被引用 15 次
- A Theoretical Framework for Grokking: Interpolation followed by Riemannian Norm MinimisationEtienne Boursier, Scott Pesme, Radu-Alexandru DragomirNeurIPS 2025 · 被引用 12 次
- Convergence of the Gradient Flow for Shallow ReLU Networks on Weakly Interacting DataLéo Dana, Loucas Pillaud-Vivien, Francis BachNeurIPS 2025 · 被引用 1 次
- Noise Stability of Transformer ModelsThemistoklis Haris, Zihan Zhang, Yuichi YoshidaICLR 2026
它引用的顶会 Paper17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar 等ICLR 2021 · 被引用 1,270 次
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang 等ICLR 2020 · 被引用 1,108 次
- Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?Sewon Min, Xinxi Lyu, Ari Holtzman, Mikel Artetxe 等EMNLP 2022 · 被引用 634 次
- The Pitfalls of Simplicity Bias in Neural NetworksHarshay Shah, Kaustav Tamuly, Aditi Raghunathan, Prateek Jain 等NeurIPS 2020 · 被引用 503 次
相关 Paper
- Sharpness Minimization Algorithms Do Not Only Minimize Sharpness To Achieve Better GeneralizationKaiyue Wen, Zhiyuan Li, Tengyu MaNeurIPS 2023 · 被引用 53 次
- Generalization Below the Edge of Stability: The Role of Data GeometryTongtong Liang, Alexander Cloninger, Rahul Parhi, Yu-Xiang WangICLR 2026 · 被引用 4 次
- Annihilation of Spurious Minima in Two-Layer ReLU NetworksYossi Arjevani, Michael FieldNeurIPS 2022 · 被引用 14 次
- Generalization Error Bounds of Gradient Descent for Learning Over-Parameterized Deep ReLU NetworksYuan Cao, Quanquan GuAAAI 2020 · 被引用 168 次
- Over-Alignment vs Over-Fitting: The Role of Feature Learning Strength in GeneralizationTaesun Yeom, Taehyeok Ha, Jaeho LeeICML 2026
