Lune

NeurIPS2025Top-tier venue

Generalization Bounds for Rank-sparse Neural Networks

Antoine Ledent, Rodrigo Alves, Yunwen Lei

2025Year
4Citations
2Top-tier citations

Abstract

It has been recently observed in much of the literature that neural networks exhibit a bottleneck rank property: for larger depths, the activation and weights of neural networks trained with gradient-based methods tend to be of approximately low rank. In fact, the rank of the activations of each layer converges to a fixed value referred to as the ``bottleneck rank'', which is the minimum rank required to represent the training data. This perspective is in line with the observation that regularizing linear networks (without activations) with weight decay is equivalent to minimizing the Schatten pp quasi norm of the neural network. In this paper we investigate the implications of this phenomenon for generalization. More specifically, we prove generalization bounds for neural networks which exploit the approximate low rank structure of the weight matrices if present. The final results rely on the Schatten pp quasi norms of the weight matrices: for small pp, the bounds exhibit a sample complexity O~(WrL2) \widetilde{O}(WrL^2) where WW and LL are the width and depth of the neural network respectively and where rr is the rank of the weight matrices. As pp increases, the bound behaves more like a norm-based bound instead.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 76430127-a9cc-45ee-a41e-62a0fafb1e7f

Cited by top-tier papers2

Ask how each one uses it

Builds on27

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines