Bottleneck Structure in Learned Features: Low-Dimension vs Regularity Tradeoff
Arthur Jacot
Abstract
Previous work has shown that DNNs with large depth and -regularization are biased towards learning low-dimensional representations of the inputs, which can be interpreted as minimizing a notion of rank of the learned function , conjectured to be the Bottleneck rank. We compute finite depth corrections to this result, revealing a measure of regularity which bounds the pseudo-determinant of the Jacobian and is subadditive under composition and addition. This formalizes a balance between learning low-dimensional representations and minimizing complexity/irregularity in the feature maps, allowing the network to learn the `right' inner dimension. Finally, we prove the conjectured bottleneck structure in the learned features as : for large depths, almost all hidden representations are approximately -dimensional, and almost all weight matrices have singular values close to 1 while the others are . Interestingly, the use of large learning rates is required to guarantee an order NTK which in turns guarantees infinite depth convergence of the representations of almost all layers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 46a83d40-070f-4896-a1aa-28161b502cdbCited by top-tier papers11
- Implicit bias of SGD in L2-regularized linear DNNs: One-way jumps from high to low rankZihan Wang, Arthur JacotICLR 2024 · 27 citations
- Layer-wise linear mode connectivityLinara Adilova, Maksym Andriushchenko, Michael Kamp, Asja Fischer et al.ICLR 2024 · 22 citations
- Mixed Dynamics In Linear Networks: Unifying the Lazy and Active RegimesZhenfeng Tu, Santiago Aranguri, Arthur JacotNeurIPS 2024 · 18 citations
- Neural collapse vs. low-rank bias: Is deep neural collapse really optimal?Peter Súkeník, Christoph H. Lampert, Marco MondelliNeurIPS 2024 · 14 citations
- Which Frequencies do CNNs Need? Emergent Bottleneck Structure in Feature LearningYuxiao Wen, Arthur JacotICML 2024 · 9 citations
Builds on8
- Label Noise SGD Provably Prefers Flat Global MinimizersAlex Damian, Tengyu Ma, Jason D. LeeNeurIPS 2021 · 155 citations
- What Happens after SGD Reaches Zero Loss? --A Mathematical FrameworkZhiyuan Li, Tianhao Wang, Sanjeev AroraICLR 2022 · 121 citations
- High-dimensional limit theorems for SGD: Effective dynamics and critical scalingGérard Ben Arous, Reza Gheissari, Aukosh JagannathNeurIPS 2022 · 94 citations
- The asymptotic spectrum of the Hessian of DNN throughout trainingArthur Jacot, Franck Gabriel, Clément HonglerICLR 2020 · 39 citations
- Training invariances and the low-rank phenomenon: beyond linear networksThien Le, Stefanie JegelkaICLR 2022 · 39 citations
Related papers
- Generalization Bounds for Rank-sparse Neural NetworksAntoine Ledent, Rodrigo Alves, Yunwen LeiNeurIPS 2025 · 4 citations
- Tight Bounds on the Smallest Eigenvalue of the Neural Tangent Kernel for Deep ReLU NetworksQuynh Nguyen, Marco Mondelli, Guido F. MontúfarICML 2021 · 98 citations
- The Persistence of Neural Collapse Despite Low-Rank BiasConnall Garrod, Jonathan P. KeatingNeurIPS 2025 · 2 citations
- Feature Learning in -regularized DNNs: Attraction/Repulsion and SparsityArthur Jacot, Eugene A. Golikov, Clément Hongler, Franck GabrielNeurIPS 2022 · 22 citations
- Implicit Bias of Large Depth Networks: a Notion of Rank for Nonlinear FunctionsArthur JacotICLR 2023 · 2 citations
