Generalization Bounds for Rank-sparse Neural Networks
Antoine Ledent, Rodrigo Alves, Yunwen Lei
Abstract
It has been recently observed in much of the literature that neural networks exhibit a bottleneck rank property: for larger depths, the activation and weights of neural networks trained with gradient-based methods tend to be of approximately low rank. In fact, the rank of the activations of each layer converges to a fixed value referred to as the ``bottleneck rank'', which is the minimum rank required to represent the training data. This perspective is in line with the observation that regularizing linear networks (without activations) with weight decay is equivalent to minimizing the Schatten quasi norm of the neural network. In this paper we investigate the implications of this phenomenon for generalization. More specifically, we prove generalization bounds for neural networks which exploit the approximate low rank structure of the weight matrices if present. The final results rely on the Schatten quasi norms of the weight matrices: for small , the bounds exhibit a sample complexity where and are the width and depth of the neural network respectively and where is the rank of the weight matrices. As increases, the bound behaves more like a norm-based bound instead.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 76430127-a9cc-45ee-a41e-62a0fafb1e7fCited by top-tier papers2
- A Refined Generalization Analysis for Extreme Multi-class Supervised Contrastive Representation LearningMinh Hieu Nong, Antoine LedentICML 2026
- Sharper Generalization Guarantees for Asynchronous SGD: Beyond Lipschitzness, Smoothness and Data HomogeneityYufeng Xie, Yunwen LeiICML 2026
Builds on27
- Provable Guarantees for Self-Supervised Deep Learning with Spectral Contrastive LossJeff Z. HaoChen, Colin Wei, Adrien Gaidon, Tengyu MaNeurIPS 2021 · 425 citations
- Inductive Biases and Variable Creation in Self-Attention MechanismsBenjamin L. Edelman, Surbhi Goel, Sham M. Kakade, Cyril ZhangICML 2022 · 154 citations
- PAC-Bayes Compression Bounds So Tight That They Can Explain GeneralizationSanae Lotfi, Marc Finzi, Sanyam Kapoor, Andres Potapczynski et al.NeurIPS 2022 · 98 citations
- Distance-Based Regularisation of Deep Networks for Fine-TuningHenry Gouk, Timothy M. Hospedales, Massimiliano PontilICLR 2021 · 65 citations
- Feature learning in deep classifiers through Intermediate Neural CollapseAkshay Rangamani, Marius Lindegaard, Tomer Galanti, Tomaso A. PoggioICML 2023 · 64 citations
Related papers
- Koopman-based generalization bound: New aspect for full-rank weightsYuka Hashimoto, Sho Sonoda, Isao Ishikawa, Atsushi Nitanda et al.ICLR 2024 · 6 citations
- Bottleneck Structure in Learned Features: Low-Dimension vs Regularity TradeoffArthur JacotNeurIPS 2023 · 20 citations
- Which Frequencies do CNNs Need? Emergent Bottleneck Structure in Feature LearningYuxiao Wen, Arthur JacotICML 2024 · 9 citations
- Generalization Analysis of Deep Non-linear Matrix CompletionAntoine Ledent, Rodrigo AlvesICML 2024 · 5 citations
- Batch normalization provably avoids ranks collapse for randomly initialised deep networksHadi Daneshmand, Jonas Moritz Kohler, Francis R. Bach, Thomas Hofmann et al.NeurIPS 2020 · 73 citations
