Redundant representations help generalization in wide neural networks
Diego Doimo, Aldo Glielmo, Sebastian Goldt, Alessandro Laio
摘要
Deep neural networks (DNNs) defy the classical bias-variance trade-off; adding parameters to a DNN that interpolates its training data will typically improve its generalization performance. Explaining the mechanism behind this ‘benign overfitting’ in deep networks remains an outstanding challenge. Here, we study the last hidden layer representations of various state-of-the-art convolutional neural networks and find that if the last hidden representation is wide enough, its neurons tend to split into groups that carry identical information and differ from each other only by statistically independent noise. The number of these groups increases linearly with the width of the layer, but only if the width is above a critical value. We show that redundant neurons appear only when the training is regularized and the training error is zero.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Learning Curves for Noisy Heterogeneous Feature-Subsampled Ridge EnsemblesBenjamin S. Ruben, Cengiz PehlevanNeurIPS 2023 · 被引用 1 次
- OrthoSolver: A Neural Proper Orthogonal Decomposition Solver For PDEsYing Pang, Jingyuan Wang, Jiahao Ji, Fanhao MuICLR 2026
它引用的顶会 Paper9
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang 等ICLR 2020 · 被引用 1,108 次
- Revisiting ResNets: Improved Training and Scaling StrategiesIrwan Bello, William Fedus, Xianzhi Du, Ekin Dogus Cubuk 等NeurIPS 2021 · 被引用 378 次
- A Constructive Prediction of the Generalization Error Across ScalesJonathan S. Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, Nir ShavitICLR 2020 · 被引用 265 次
- Rethinking Bias-Variance Trade-off for Generalization of Neural NetworksZitong Yang, Yaodong Yu, Chong You, Jacob Steinhardt 等ICML 2020 · 被引用 219 次
- Double Trouble in Double Descent: Bias and Variance(s) in the Lazy RegimeStéphane d'Ascoli, Maria Refinetti, Giulio Biroli, Florent KrzakalaICML 2020 · 被引用 163 次
相关 Paper
- Frivolous Units: Wider Networks Are Not Really That WideStephen Casper, Xavier Boix, Vanessa D'Amario, Ling Guo 等AAAI 2021 · 被引用 20 次
- The Curious Case of Benign MemorizationSotiris Anagnostidis, Gregor Bachmann, Lorenzo Noci, Thomas HofmannICLR 2023 · 被引用 1 次
- Strong inductive biases provably prevent harmless interpolationMichael Aerni, Marco Milanta, Konstantin Donhauser, Fanny YangICLR 2023
- Intraclass clustering: an implicit learning ability that regularizes DNNsSimon Carbonnelle, Christophe De VleeschouwerICLR 2021 · 被引用 2 次
- Batch normalization is sufficient for universal function approximation in CNNsRebekka BurkholzICLR 2024 · 被引用 8 次
