Which Frequencies do CNNs Need? Emergent Bottleneck Structure in Feature Learning
Yuxiao Wen, Arthur Jacot
Abstract
We describe the emergence of a Convolution Bottleneck (CBN) structure in CNNs, where the network uses its first few layers to transform the input representation into a representation that is supported only along a few frequencies and channels, before using the last few layers to map back to the outputs. We define the CBN rank, which describes the number and type of frequencies that are kept inside the bottleneck, and partially prove that the parameter norm required to represent a function scales as depth times the CBN rank . We also show that the parameter norm depends at next order on the regularity of . We show that any network with almost optimal parameter norm will exhibit a CBN structure in both the weights and - under the assumption that the network is stable under large learning rate - the activations, which motivates the common practice of down-sampling; and we verify that the CBN results still hold with down-sampling. Finally we use the CBN structure to interpret the functions learned by CNNs on a number of tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Neural collapse vs. low-rank bias: Is deep neural collapse really optimal?Peter Súkeník, Christoph H. Lampert, Marco MondelliNeurIPS 2024 · 14 citations
- Saddle-To-Saddle Dynamics in Deep ReLU Networks: Low-Rank Bias in the First Saddle EscapeIoannis Bantzis, James B. Simon, Arthur JacotICLR 2026 · 4 citations
- Geometric Inductive Biases of Deep Networks: The Role of Data and ArchitectureSajad Movahedi, Antonio Orvieto, Seyed-Mohsen Moosavi-DezfooliICLR 2025
- Wide Neural Networks Trained with Weight Decay Provably Exhibit Neural CollapseArthur Jacot, Peter Súkeník, Zihan Wang, Marco MondelliICLR 2025
- When Diffusion Models Memorize: Inductive Biases in Probability Flow of Minimum-Norm Shallow Neural NetsChen Zeno, Hila Manor, Greg Ongie, Nir Weinberger et al.ICML 2025
Builds on11
- Towards Resolving the Implicit Bias of Gradient Descent for Matrix Factorization: Greedy Low-Rank LearningZhiyuan Li, Yuping Luo, Kaifeng LyuICLR 2021 · 155 citations
- Label Noise SGD Provably Prefers Flat Global MinimizersAlex Damian, Tengyu Ma, Jason D. LeeNeurIPS 2021 · 155 citations
- What Happens after SGD Reaches Zero Loss? --A Mathematical FrameworkZhiyuan Li, Tianhao Wang, Sanjeev AroraICLR 2022 · 121 citations
- Representation Costs of Linear Neural Networks: Analysis and DesignZhen Dai, Mina Karzand, Nathan SrebroNeurIPS 2021 · 34 citations
- Learning with convolution and pooling operations in kernel methodsTheodor Misiakiewicz, Song MeiNeurIPS 2022 · 30 citations
Related papers
- Generalization Bounds for Rank-sparse Neural NetworksAntoine Ledent, Rodrigo Alves, Yunwen LeiNeurIPS 2025 · 4 citations
- Bottleneck Structure in Learned Features: Low-Dimension vs Regularity TradeoffArthur JacotNeurIPS 2023 · 20 citations
- The Hessian perspective into the Nature of Convolutional Neural NetworksSidak Pal Singh, Thomas Hofmann, Bernhard SchölkopfICML 2023 · 12 citations
- Batch normalization is sufficient for universal function approximation in CNNsRebekka BurkholzICLR 2024 · 8 citations
- Linear CNNs Discover the Statistical Structure of the Dataset Using Only the Most Dominant FrequenciesHannah Pinson, Joeri Lenaerts, Vincent GinisICML 2023 · 8 citations
