Bottleneck Structure in Learned Features: Low-Dimension vs Regularity Tradeoff
Arthur Jacot
摘要
Previous work has shown that DNNs with large depth and -regularization are biased towards learning low-dimensional representations of the inputs, which can be interpreted as minimizing a notion of rank of the learned function , conjectured to be the Bottleneck rank. We compute finite depth corrections to this result, revealing a measure of regularity which bounds the pseudo-determinant of the Jacobian and is subadditive under composition and addition. This formalizes a balance between learning low-dimensional representations and minimizing complexity/irregularity in the feature maps, allowing the network to learn the `right' inner dimension. Finally, we prove the conjectured bottleneck structure in the learned features as : for large depths, almost all hidden representations are approximately -dimensional, and almost all weight matrices have singular values close to 1 while the others are . Interestingly, the use of large learning rates is required to guarantee an order NTK which in turns guarantees infinite depth convergence of the representations of almost all layers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Implicit bias of SGD in L2-regularized linear DNNs: One-way jumps from high to low rankZihan Wang, Arthur JacotICLR 2024 · 被引用 27 次
- Layer-wise linear mode connectivityLinara Adilova, Maksym Andriushchenko, Michael Kamp, Asja Fischer 等ICLR 2024 · 被引用 22 次
- Mixed Dynamics In Linear Networks: Unifying the Lazy and Active RegimesZhenfeng Tu, Santiago Aranguri, Arthur JacotNeurIPS 2024 · 被引用 18 次
- Neural collapse vs. low-rank bias: Is deep neural collapse really optimal?Peter Súkeník, Christoph H. Lampert, Marco MondelliNeurIPS 2024 · 被引用 14 次
- Which Frequencies do CNNs Need? Emergent Bottleneck Structure in Feature LearningYuxiao Wen, Arthur JacotICML 2024 · 被引用 9 次
它引用的顶会 Paper8
- Label Noise SGD Provably Prefers Flat Global MinimizersAlex Damian, Tengyu Ma, Jason D. LeeNeurIPS 2021 · 被引用 155 次
- What Happens after SGD Reaches Zero Loss? --A Mathematical FrameworkZhiyuan Li, Tianhao Wang, Sanjeev AroraICLR 2022 · 被引用 121 次
- High-dimensional limit theorems for SGD: Effective dynamics and critical scalingGérard Ben Arous, Reza Gheissari, Aukosh JagannathNeurIPS 2022 · 被引用 94 次
- The asymptotic spectrum of the Hessian of DNN throughout trainingArthur Jacot, Franck Gabriel, Clément HonglerICLR 2020 · 被引用 39 次
- Training invariances and the low-rank phenomenon: beyond linear networksThien Le, Stefanie JegelkaICLR 2022 · 被引用 39 次
相关 Paper
- Generalization Bounds for Rank-sparse Neural NetworksAntoine Ledent, Rodrigo Alves, Yunwen LeiNeurIPS 2025 · 被引用 4 次
- Tight Bounds on the Smallest Eigenvalue of the Neural Tangent Kernel for Deep ReLU NetworksQuynh Nguyen, Marco Mondelli, Guido F. MontúfarICML 2021 · 被引用 98 次
- The Persistence of Neural Collapse Despite Low-Rank BiasConnall Garrod, Jonathan P. KeatingNeurIPS 2025 · 被引用 2 次
- Feature Learning in -regularized DNNs: Attraction/Repulsion and SparsityArthur Jacot, Eugene A. Golikov, Clément Hongler, Franck GabrielNeurIPS 2022 · 被引用 22 次
- Implicit Bias of Large Depth Networks: a Notion of Rank for Nonlinear FunctionsArthur JacotICLR 2023 · 被引用 2 次
