Lune

NeurIPS2025顶会

Generalization Bounds for Rank-sparse Neural Networks

Antoine Ledent, Rodrigo Alves, Yunwen Lei

2025年份
4被引次数
2顶会引用

摘要

It has been recently observed in much of the literature that neural networks exhibit a bottleneck rank property: for larger depths, the activation and weights of neural networks trained with gradient-based methods tend to be of approximately low rank. In fact, the rank of the activations of each layer converges to a fixed value referred to as the ``bottleneck rank'', which is the minimum rank required to represent the training data. This perspective is in line with the observation that regularizing linear networks (without activations) with weight decay is equivalent to minimizing the Schatten pp quasi norm of the neural network. In this paper we investigate the implications of this phenomenon for generalization. More specifically, we prove generalization bounds for neural networks which exploit the approximate low rank structure of the weight matrices if present. The final results rely on the Schatten pp quasi norms of the weight matrices: for small pp, the bounds exhibit a sample complexity O~(WrL2) \widetilde{O}(WrL^2) where WW and LL are the width and depth of the neural network respectively and where rr is the rank of the weight matrices. As pp increases, the bound behaves more like a norm-based bound instead.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper2

问问它们各自怎么用它

它引用的顶会 Paper27

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖