Simplicity Bias in Overparameterized Machine Learning
Yakir Berchenko
摘要
A thorough theoretical understanding of the surprising generalization ability of deep networks (and other overparameterized models) is still lacking. Here we demonstrate that simplicity bias is a major phenomenon to be reckoned with in overparameterized machine learning. In addition to explaining the outcome of simplicity bias, we also study its source: following concrete rigorous examples, we argue that (i) simplicity bias can explain generalization in overparameterized learning models such as neural networks; (ii) simplicity bias and excellent generalization are optimizer-independent, as our example shows, and although the optimizer affects training, it is not the driving force behind simplicity bias; (iii) simplicity bias in pre-training models, and subsequent posteriors, is universal and stems from the subtle fact that uniformly-at-random constructed priors are not uniformly-at-random sampled ; and (iv) in neural network models, the biasing mechanism in wide (and shallow) networks is different from the biasing mechanism in deep (and narrow) networks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- How Uniform Random Weights Induce Non-uniform Bias: Typical Interpolating Neural Networks Generalize with Narrow TeachersGon Buzaglo, Itamar Harel, Mor Shpigel Nacson, Alon Brutzkus 等ICML 2024 · 被引用 11 次
- Do Neural Networks Need Gradient Descent to Generalize? A Theoretical StudyYotam Alexander, Yonatan Slutzky, Yuval Ran-Milo, Nadav CohenNeurIPS 2025 · 被引用 3 次
- SimBa: Simplicity Bias for Scaling Up Parameters in Deep Reinforcement LearningHojoon Lee, Dongyoon Hwang, Donghu Kim, Hyunseung Kim 等ICLR 2025
- Revisiting the Volume HypothesisAri Pakman, Lior Kreimer, Yakir BerchenkoICML 2026
- Network Sparsity Unlocks the Scaling Potential of Deep Reinforcement LearningGuozheng Ma, Lu Li, Zilin Wang, Li Shen 等ICML 2025
它引用的顶会 Paper2
相关 Paper
- Bias of Stochastic Gradient Descent or the Architecture: Disentangling the Effects of Overparameterization of Neural NetworksAmit Peleg, Matthias HeinICML 2024
- Neural networks trained with SGD learn distributions of increasing complexityMaria Refinetti, Alessandro Ingrosso, Sebastian GoldtICML 2023 · 被引用 58 次
- Neural Redshift: Random Networks are not Random FunctionsDamien Teney, Armand Mihai Nicolicioiu, Valentin Hartmann, Ehsan AbbasnejadCVPR 2024 · 被引用 7 次
- Sharpness Minimization Algorithms Do Not Only Minimize Sharpness To Achieve Better GeneralizationKaiyue Wen, Zhiyuan Li, Tengyu MaNeurIPS 2023 · 被引用 53 次
- Generalization Error Bounds of Gradient Descent for Learning Over-Parameterized Deep ReLU NetworksYuan Cao, Quanquan GuAAAI 2020 · 被引用 168 次
