Frivolous Units: Wider Networks Are Not Really That Wide
Stephen Casper, Xavier Boix, Vanessa D'Amario, Ling Guo, Martin Schrimpf, Kasper Vinken, Gabriel Kreiman
Abstract
A remarkable characteristic of overparameterized deep neural networks (DNNs) is that their accuracy does not degrade when the network width is increased. Recent evidence suggests that developing compressible representations allows the complexity of large networks to be adjusted for the learning task at hand. However, these representations are poorly understood. A promising strand of research inspired from biology involves studying representations at the unit level as it offers a more granular interpretation of the neural mechanisms. In order to better understand what facilitates increases in width without decreases in accuracy, we ask: Are there mechanisms at the unit level by which networks control their effective complexity? If so, how do these depend on the architecture, dataset, and hyperparameters? We identify two distinct types of “frivolous” units that proliferate when the network’s width increases: prunable units which can be dropped out of the network without significant change to the output and redundant units whose activities can be expressed as a linear combination of others. These units imply complexity constraints as the function the network computes could be expressed without them. We also identify how the development of these units can be influenced by architecture and a number of training factors. Together, these results help to explain why the accuracy of DNNs does not degrade when width is increased and highlight the importance of frivolous units toward understanding implicit regularization in DNNs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 261ff61b-9de6-48aa-ae6a-3d0315a2fba0Cited by top-tier papers2
- Manipulating dropout reveals an optimal balance of efficiency and robustness in biological and machine visual systemsJacob S. Prince, Gabriel Fajardo, George A. Alvarez, Talia KonkleICLR 2024 · 2 citations
- Budgeted Training for Vision TransformerZhuofan Xia, Xuran Pan, Xuan Jin, Yuan He et al.ICLR 2023
Builds on5
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang et al.ICLR 2020 · 1,108 citations
- Compositional Explanations of NeuronsJesse Mu, Jacob AndreasNeurIPS 2020 · 229 citations
- Double Trouble in Double Descent: Bias and Variance(s) in the Lazy RegimeStéphane d'Ascoli, Maria Refinetti, Giulio Biroli, Florent KrzakalaICML 2020 · 163 citations
- On the Number of Linear Regions of Convolutional Neural NetworksHuan Xiong, Lei Huang, Mengyang Yu, Li Liu et al.ICML 2020 · 80 citations
- Compression based bound for non-compressed network: unified generalization error analysis of large compressible deep neural networkTaiji Suzuki, Hiroshi Abe, Tomoaki NishimuraICLR 2020 · 57 citations
Related papers
- Redundant representations help generalization in wide neural networksDiego Doimo, Aldo Glielmo, Sebastian Goldt, Alessandro LaioNeurIPS 2022 · 13 citations
- Nonlinear Advantage: Trained Networks Might Not Be As Complex as You ThinkChristian H. X. Ali Mehmeti-Göpel, Jan DisselhoffICML 2023 · 6 citations
- Pruning from ScratchYulong Wang, Xiaolu Zhang, Lingxi Xie, Jun Zhou et al.AAAI 2020 · 219 citations
- Composable Sparse Subnetworks via Maximum-Entropy PrincipleFrancesco Caso, Samuele Fonio, Simone Monaco, Nicola Saccomanno et al.ICLR 2026
- Selectivity considered harmful: evaluating the causal impact of class selectivity in DNNsMatthew L. Leavitt, Ari S. MorcosICLR 2021 · 34 citations
