How Uniform Random Weights Induce Non-uniform Bias: Typical Interpolating Neural Networks Generalize with Narrow Teachers
Gon Buzaglo, Itamar Harel, Mor Shpigel Nacson, Alon Brutzkus, Nathan Srebro, Daniel Soudry
摘要
Background. A main theoretical puzzle is why over-parameterized Neural Networks (NNs) generalize well when trained to zero loss (i.e., so they interpolate the data). Usually, the NN is trained with Stochastic Gradient Descent (SGD) or one of its variants. However, recent empirical work examined the generalization of a random NN that interpolates the data: the NN was sampled from a seemingly uniform prior over the parameters, conditioned on that the NN perfectly classifies the training set. Interestingly, such a NN sample typically generalized as well as SGD-trained NNs. Contributions. We prove that such a random NN interpolator typically generalizes well if there exists an underlying narrow ``teacher NN'' that agrees with the labels. Specifically, we show that such a `flat' prior over the NN parameterization induces a rich prior over the NN functions, due to the redundancy in the NN structure. In particular, this creates a bias towards simpler functions, which require less relevant parameters to represent -- enabling learning with a sample complexity approximately proportional to the complexity of the teacher (roughly, the number of non-redundant parameters), rather than the student's.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Stable Minima Cannot Overfit in Univariate ReLU Networks: Generalization by Large Step SizesDan Qiao, Kaiqi Zhang, Esha Singh, Daniel Soudry 等NeurIPS 2024 · 被引用 15 次
- Provable Tempered Overfitting of Minimal Nets and Typical NetsItamar Harel, William Hoza, Gal Vardi, Itay Evron 等NeurIPS 2024 · 被引用 7 次
- Mitigating the Curse of Detail: Scaling Arguments for Feature Learning and Sample ComplexityNoa Rubin, Orit Davidovich, Zohar RingelICLR 2026 · 被引用 5 次
- Do Neural Networks Need Gradient Descent to Generalize? A Theoretical StudyYotam Alexander, Yonatan Slutzky, Yuval Ran-Milo, Nadav CohenNeurIPS 2025 · 被引用 3 次
- Unsupervised Translation of Emergent CommunicationIdo Levy, Orr Paradise, Boaz Carmeli, Ron Meir 等AAAI 2025 · 被引用 3 次
它引用的顶会 Paper10
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 被引用 402 次
- Proving the Lottery Ticket Hypothesis: Pruning is All You NeedEran Malach, Gilad Yehudai, Shai Shalev-Shwartz, Ohad ShamirICML 2020 · 被引用 327 次
- Nonuniform-to-Uniform Quantization: Towards Accurate Quantization via Generalized Straight-Through EstimationZechun Liu, Kwang-Ting Cheng, Dong Huang, Eric P. Xing 等CVPR 2022 · 被引用 108 次
- Implicit Bias in Deep Linear Classification: Initialization Scale vs Training AccuracyEdward Moroshko, Blake E. Woodworth, Suriya Gunasekar, Jason D. Lee 等NeurIPS 2020 · 被引用 98 次
- PAC-Bayes Compression Bounds So Tight That They Can Explain GeneralizationSanae Lotfi, Marc Finzi, Sanyam Kapoor, Andres Potapczynski 等NeurIPS 2022 · 被引用 98 次
相关 Paper
- Bias of Stochastic Gradient Descent or the Architecture: Disentangling the Effects of Overparameterization of Neural NetworksAmit Peleg, Matthias HeinICML 2024
- Simplicity Bias in Overparameterized Machine LearningYakir BerchenkoAAAI 2024 · 被引用 7 次
- Landscape Connectivity and Dropout Stability of SGD Solutions for Over-parameterized Neural NetworksAlexander Shevchenko, Marco MondelliICML 2020 · 被引用 41 次
- Stochastic Collapse: How Gradient Noise Attracts SGD Dynamics Towards Simpler SubnetworksFeng Chen, Daniel Kunin, Atsushi Yamamura, Surya GanguliNeurIPS 2023 · 被引用 52 次
- Generalizablity of Memorization Neural NetworkLijia Yu, Xiao-Shan Gao, Lijun Zhang, Yibo MiaoNeurIPS 2024 · 被引用 5 次
