Redundant representations help generalization in wide neural networks
Diego Doimo, Aldo Glielmo, Sebastian Goldt, Alessandro Laio
Abstract
Deep neural networks (DNNs) defy the classical bias-variance trade-off; adding parameters to a DNN that interpolates its training data will typically improve its generalization performance. Explaining the mechanism behind this ‘benign overfitting’ in deep networks remains an outstanding challenge. Here, we study the last hidden layer representations of various state-of-the-art convolutional neural networks and find that if the last hidden representation is wide enough, its neurons tend to split into groups that carry identical information and differ from each other only by statistically independent noise. The number of these groups increases linearly with the width of the layer, but only if the width is above a critical value. We show that redundant neurons appear only when the training is regularized and the training error is zero.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0968ad7a-4733-44cd-b2c7-76af3222f8ddCited by top-tier papers2
- Learning Curves for Noisy Heterogeneous Feature-Subsampled Ridge EnsemblesBenjamin S. Ruben, Cengiz PehlevanNeurIPS 2023 · 1 citation
- OrthoSolver: A Neural Proper Orthogonal Decomposition Solver For PDEsYing Pang, Jingyuan Wang, Jiahao Ji, Fanhao MuICLR 2026
Builds on9
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang et al.ICLR 2020 · 1,108 citations
- Revisiting ResNets: Improved Training and Scaling StrategiesIrwan Bello, William Fedus, Xianzhi Du, Ekin Dogus Cubuk et al.NeurIPS 2021 · 378 citations
- A Constructive Prediction of the Generalization Error Across ScalesJonathan S. Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, Nir ShavitICLR 2020 · 265 citations
- Rethinking Bias-Variance Trade-off for Generalization of Neural NetworksZitong Yang, Yaodong Yu, Chong You, Jacob Steinhardt et al.ICML 2020 · 219 citations
- Double Trouble in Double Descent: Bias and Variance(s) in the Lazy RegimeStéphane d'Ascoli, Maria Refinetti, Giulio Biroli, Florent KrzakalaICML 2020 · 163 citations
Related papers
- Frivolous Units: Wider Networks Are Not Really That WideStephen Casper, Xavier Boix, Vanessa D'Amario, Ling Guo et al.AAAI 2021 · 20 citations
- The Curious Case of Benign MemorizationSotiris Anagnostidis, Gregor Bachmann, Lorenzo Noci, Thomas HofmannICLR 2023 · 1 citation
- Strong inductive biases provably prevent harmless interpolationMichael Aerni, Marco Milanta, Konstantin Donhauser, Fanny YangICLR 2023
- Intraclass clustering: an implicit learning ability that regularizes DNNsSimon Carbonnelle, Christophe De VleeschouwerICLR 2021 · 2 citations
- Batch normalization is sufficient for universal function approximation in CNNsRebekka BurkholzICLR 2024 · 8 citations
