Linear CNNs Discover the Statistical Structure of the Dataset Using Only the Most Dominant Frequencies
Hannah Pinson, Joeri Lenaerts, Vincent Ginis
摘要
We here present a stepping stone towards a deeper understanding of convolutional neural networks (CNNs) in the form of a theory of learning in linear CNNs. Through analyzing the gradient descent equations, we discover that the evolution of the network during training is determined by the interplay between the dataset structure and the convolutional network structure. We show that linear CNNs discover the statistical structure of the dataset with non-linear, ordered, stage-like transitions, and that the speed of discovery changes depending on the relationship between the dataset and the convolutional network structure. Moreover, we find that this interplay lies at the heart of what we call the "dominant frequency bias", where linear CNNs arrive at these discoveries using only the dominant frequencies of the different structural parts present in the dataset. We furthermore provide experiments that show how our theory relates to deep, non-linear CNNs used in practice. Our findings shed new light on the inner working of CNNs, and can help explain their shortcut learning and their tendency to rely on texture instead of shape. In addition to a neural network's pre-defined architecture, the parameters of the network obtain an implicit structure during training. For example, it has been shown that weight matrices can exhibit structural patterns, such as clusters and branches (Voss et al., 2021; Casper et al., 2022) . On the other hand, the input dataset also has an implicit structure arising from patterns and relationships between the samples. E.g., in a classification task, dogs are more visually similar
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Towards Combating Frequency Simplicity-biased Learning for Domain GeneralizationXilin He, Jingyu Hu, Qinliang Lin, Cheng Luo 等NeurIPS 2024 · 被引用 16 次
- Convolutions and More as Einsum: A Tensor Network Perspective with Advances for Second-Order MethodsFelix DangelNeurIPS 2024 · 被引用 5 次
- A Solvable Attention for Neural Scaling LawsBochen Lyu, Di Wang, Zhanxing ZhuICLR 2025
- Domain Adaptive Object Detection via Dynamic Causal RefinementZeyu Ma, Jiaqi Huang, Yitong Qin, Ziqiang Zheng 等ICML 2026
- FreqDebias: Towards Generalizable Deepfake Detection via Consistency-Driven Frequency DebiasingHossein Kashiani, Niloufar Alipour Talemi, Fatemeh AfghahCVPR 2025
它引用的顶会 Paper5
- Frequency Bias in Neural Networks for Input of Non-Uniform DensityRonen Basri, Meirav Galun, Amnon Geifman, David W. Jacobs 等ICML 2020 · 被引用 229 次
- Neural Networks as Kernel Learners: The Silent Alignment EffectAlexander B. Atanasov, Blake Bordelon, Cengiz PehlevanICLR 2022 · 被引用 110 次
- Exact learning dynamics of deep linear networks with prior knowledgeLukas Braun, Clémentine C. J. Dominé, James Fitzgerald, Andrew M. SaxeNeurIPS 2022 · 被引用 75 次
- Neural networks trained with SGD learn distributions of increasing complexityMaria Refinetti, Alessandro Ingrosso, Sebastian GoldtICML 2023 · 被引用 58 次
- The dynamics of representation learning in shallow, non-linear autoencodersMaria Refinetti, Sebastian GoldtICML 2022 · 被引用 25 次
相关 Paper
- Which Layer is Learning Faster? A Systematic Exploration of Layer-wise Convergence Rate for Deep Neural NetworksYixiong Chen, Alan L. Yuille, Zongwei ZhouICLR 2023
- Shape or Texture: Understanding Discriminative Features in CNNsMd. Amirul Islam, Matthew Kowal, Patrick Esser, Sen Jia 等ICLR 2021 · 被引用 86 次
- What do neural networks learn in image classification? A frequency shortcut perspectiveShunxin Wang, Raymond N. J. Veldhuis, Christoph Brune, Nicola StrisciuglioICCV 2023 · 被引用 51 次
- Deep Frequency Principle Towards Understanding Why Deeper Learning Is FasterZhiqin John Xu, Hanxu ZhouAAAI 2021 · 被引用 67 次
- Implicit Bias of Linear Equivariant NetworksHannah Lawrence, Bobak Toussi Kiani, Kristian G. Georgiev, Andrew K. DienesICML 2022 · 被引用 18 次
