On Measuring Excess Capacity in Neural Networks
Florian Graf, Sebastian Zeng, Bastian Rieck, Marc Niethammer, Roland Kwitt
Abstract
We study the excess capacity of deep networks in the context of supervised classification. That is, given a capacity measure of the underlying hypothesis class - in our case, empirical Rademacher complexity - to what extent can we (a priori) constrain this class while retaining an empirical error on a par with the unconstrained regime? To assess excess capacity in modern architectures (such as residual networks), we extend and unify prior Rademacher complexity bounds to accommodate function composition and addition, as well as the structure of convolutions. The capacity-driving terms in our bounds are the Lipschitz constants of the layers and an (2, 1) group norm distance to the initializations of the convolution weights. Experiments on benchmark datasets of varying task difficulty indicate that (1) there is a substantial amount of excess capacity per task, and (2) capacity can be kept at a surprisingly similar level across tasks. Overall, this suggests a notion of compressibility with respect to weight norms, complementary to classic compression via weight pruning. Source code is available at https://github.com/rkwitt/excess_capacity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 02345b4b-e55f-4ed1-9b61-e1a1ae640571Cited by top-tier papers13
- On the Effectiveness of Lipschitz-Driven Rehearsal in Continual LearningLorenzo Bonicelli, Matteo Boschini, Angelo Porrello, Concetto Spampinato et al.NeurIPS 2022 · 64 citations
- Norm-based Generalization Bounds for Sparse Neural NetworksTomer Galanti, Mengjia Xu, Liane Galanti, Tomaso A. PoggioNeurIPS 2023 · 19 citations
- Transformed Low-Rank Parameterization Can Help Robust Generalization for Tensor Neural NetworksAndong Wang, Chao Li, Mingyuan Bai, Zhong Jin et al.NeurIPS 2023 · 12 citations
- Generalization Analysis of Deep Non-linear Matrix CompletionAntoine Ledent, Rodrigo AlvesICML 2024 · 5 citations
- On the Expressivity and Sample Complexity of Node-Individualized Graph Neural NetworksPaolo Pellizzoni, Till Hendrik Schulz, Dexiong Chen, Karsten M. BorgwardtNeurIPS 2024 · 5 citations
Builds on4
- A Geometric Analysis of Neural Collapse with Unconstrained FeaturesZhihui Zhu, Tianyu Ding, Jinxin Zhou, Xiao Li et al.NeurIPS 2021 · 303 citations
- Generalization bounds for deep convolutional neural networksPhilip M. Long, Hanie SedghiICLR 2020 · 102 citations
- Compression based bound for non-compressed network: unified generalization error analysis of large compressible deep neural networkTaiji Suzuki, Hiroshi Abe, Tomoaki NishimuraICLR 2020 · 57 citations
- Measuring Generalization with Optimal TransportChing-Yao Chuang, Youssef Mroueh, Kristjan H. Greenewald, Antonio Torralba et al.NeurIPS 2021 · 33 citations
Related papers
- Norm-Based Generalisation Bounds for Deep Multi-Class Convolutional Neural NetworksAntoine Ledent, Waleed Mustafa, Yunwen Lei, Marius KloftAAAI 2021 · 24 citations
- Reproducing Kernel Banach Space Models for Neural Networks with Application to Rademacher Complexity AnalysisAlistair Shilton, Sunil Gupta, Santu Rana, Svetha VenkateshNeurIPS 2025
- Dropout: Explicit Forms and Capacity ControlRaman Arora, Peter L. Bartlett, Poorya Mianjy, Nathan SrebroICML 2021 · 43 citations
- Distance-Based Regularisation of Deep Networks for Fine-TuningHenry Gouk, Timothy M. Hospedales, Massimiliano PontilICLR 2021 · 65 citations
- On Contraction of Sequential and Offset Rademacher ComplexitiesAdam Block, Alexander Rakhlin, Mark SellkeICML 2026
