On Measuring Excess Capacity in Neural Networks
Florian Graf, Sebastian Zeng, Bastian Rieck, Marc Niethammer, Roland Kwitt
摘要
We study the excess capacity of deep networks in the context of supervised classification. That is, given a capacity measure of the underlying hypothesis class - in our case, empirical Rademacher complexity - to what extent can we (a priori) constrain this class while retaining an empirical error on a par with the unconstrained regime? To assess excess capacity in modern architectures (such as residual networks), we extend and unify prior Rademacher complexity bounds to accommodate function composition and addition, as well as the structure of convolutions. The capacity-driving terms in our bounds are the Lipschitz constants of the layers and an (2, 1) group norm distance to the initializations of the convolution weights. Experiments on benchmark datasets of varying task difficulty indicate that (1) there is a substantial amount of excess capacity per task, and (2) capacity can be kept at a surprisingly similar level across tasks. Overall, this suggests a notion of compressibility with respect to weight norms, complementary to classic compression via weight pruning. Source code is available at https://github.com/rkwitt/excess_capacity.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- On the Effectiveness of Lipschitz-Driven Rehearsal in Continual LearningLorenzo Bonicelli, Matteo Boschini, Angelo Porrello, Concetto Spampinato 等NeurIPS 2022 · 被引用 64 次
- Norm-based Generalization Bounds for Sparse Neural NetworksTomer Galanti, Mengjia Xu, Liane Galanti, Tomaso A. PoggioNeurIPS 2023 · 被引用 19 次
- Transformed Low-Rank Parameterization Can Help Robust Generalization for Tensor Neural NetworksAndong Wang, Chao Li, Mingyuan Bai, Zhong Jin 等NeurIPS 2023 · 被引用 12 次
- Generalization Analysis of Deep Non-linear Matrix CompletionAntoine Ledent, Rodrigo AlvesICML 2024 · 被引用 5 次
- On the Expressivity and Sample Complexity of Node-Individualized Graph Neural NetworksPaolo Pellizzoni, Till Hendrik Schulz, Dexiong Chen, Karsten M. BorgwardtNeurIPS 2024 · 被引用 5 次
它引用的顶会 Paper4
- A Geometric Analysis of Neural Collapse with Unconstrained FeaturesZhihui Zhu, Tianyu Ding, Jinxin Zhou, Xiao Li 等NeurIPS 2021 · 被引用 303 次
- Generalization bounds for deep convolutional neural networksPhilip M. Long, Hanie SedghiICLR 2020 · 被引用 102 次
- Compression based bound for non-compressed network: unified generalization error analysis of large compressible deep neural networkTaiji Suzuki, Hiroshi Abe, Tomoaki NishimuraICLR 2020 · 被引用 57 次
- Measuring Generalization with Optimal TransportChing-Yao Chuang, Youssef Mroueh, Kristjan H. Greenewald, Antonio Torralba 等NeurIPS 2021 · 被引用 33 次
相关 Paper
- Norm-Based Generalisation Bounds for Deep Multi-Class Convolutional Neural NetworksAntoine Ledent, Waleed Mustafa, Yunwen Lei, Marius KloftAAAI 2021 · 被引用 24 次
- Reproducing Kernel Banach Space Models for Neural Networks with Application to Rademacher Complexity AnalysisAlistair Shilton, Sunil Gupta, Santu Rana, Svetha VenkateshNeurIPS 2025
- Dropout: Explicit Forms and Capacity ControlRaman Arora, Peter L. Bartlett, Poorya Mianjy, Nathan SrebroICML 2021 · 被引用 43 次
- Distance-Based Regularisation of Deep Networks for Fine-TuningHenry Gouk, Timothy M. Hospedales, Massimiliano PontilICLR 2021 · 被引用 65 次
- On Contraction of Sequential and Offset Rademacher ComplexitiesAdam Block, Alexander Rakhlin, Mark SellkeICML 2026
