The Pitfalls of Simplicity Bias in Neural Networks
Harshay Shah, Kaustav Tamuly, Aditi Raghunathan, Prateek Jain, Praneeth Netrapalli
Abstract
Several works have proposed Simplicity Bias (SB)-the tendency of standard training procedures such as Stochastic Gradient Descent (SGD) to find simple models-to justify why neural networks generalize well [1, 49, 74] . However, the precise notion of simplicity remains vague. Furthermore, previous settings [67, 24] that use SB to justify why neural networks generalize well do not simultaneously capture the non-robustness of neural networks-a widely observed phenomenon in practice [71, 36] . We attempt to reconcile SB and the superior standard generalization of neural networks with the non-robustness observed in practice by designing datasets that (a) incorporate a precise notion of simplicity, (b) comprise multiple predictive features with varying levels of simplicity, and (c) capture the non-robustness of neural networks trained on real data. Through theoretical analysis and targeted experiments on these datasets, we make four observations: (i) SB of SGD and variants can be extreme: neural networks can exclusively rely on the simplest feature and remain invariant to all predictive complex features. (ii) The extreme aspect of SB could explain why seemingly benign distribution shifts and small adversarial perturbations significantly degrade model performance. (iii) Contrary to conventional wisdom, SB can also hurt generalization on the same data distribution, as SB persists even when the simplest feature has less predictive power than the more complex features. (iv) Common approaches to improve generalization and robustness-ensembles and adversarial training-can fail in mitigating SB and its pitfalls. Given the role of SB in training neural networks, we hope that the proposed datasets and methods serve as an effective testbed to evaluate novel algorithmic approaches aimed at avoiding the pitfalls of SB.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c1919fca-1f99-45da-9981-2a7649a423a1Cited by top-tier papers181
- Domain Generalization using Causal MatchingDivyat Mahajan, Shruti Tople, Amit SharmaICML 2021 · 399 citations
- Gradient Starvation: A Learning Proclivity in Neural NetworksMohammad Pezeshki, Sékou-Oumar Kaba, Yoshua Bengio, Aaron C. Courville et al.NeurIPS 2021 · 378 citations
- Fishr: Invariant Gradient Variances for Out-of-Distribution GeneralizationAlexandre Ramé, Corentin Dancette, Matthieu CordICML 2022 · 262 citations
- On Feature Learning in the Presence of Spurious CorrelationsPavel Izmailov, Polina Kirichenko, Nate Gruver, Andrew Gordon WilsonNeurIPS 2022 · 208 citations
- Intriguing Properties of Contrastive LossesTing Chen, Calvin Luo, Lala LiNeurIPS 2021 · 206 citations
Builds on9
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha et al.S&P 2016 · 3,275 citations
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural NetworksWeilin Xu, David Evans, Yanjun QiNDSS 2018 · 1,633 citations
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 1,352 citations
- What shapes feature representations? Exploring datasets, architectures, and trainingKatherine L. Hermann, Andrew K. LampinenNeurIPS 2020 · 186 citations
Related papers
- Simplicity Bias in 1-Hidden Layer Neural NetworksDepen Morwani, Jatin Batra, Prateek Jain, Praneeth NetrapalliNeurIPS 2023 · 33 citations
- Evading the Simplicity Bias: Training a Diverse Set of Models Discovers Solutions with Superior OOD GeneralizationDamien Teney, Ehsan Abbasnejad, Simon Lucey, Anton van den HengelCVPR 2022 · 32 citations
- The Rich and the Simple: On the Implicit Bias of Adam and SGDBhavya Vasudeva, Jung Hoon Lee, Vatsal Sharan, Mahdi SoltanolkotabiNeurIPS 2025 · 14 citations
- Simplicity Bias in Overparameterized Machine LearningYakir BerchenkoAAAI 2024 · 7 citations
- Feature Averaging: An Implicit Bias of Gradient Descent Leading to Non-Robustness in Neural NetworksBinghui Li, Zhixuan Pan, Kaifeng Lyu, Jian LiICLR 2025
