The Hessian perspective into the Nature of Convolutional Neural Networks
Sidak Pal Singh, Thomas Hofmann, Bernhard Schölkopf
摘要
While Convolutional Neural Networks (CNNs) have long been investigated and applied, as well as theorized, we aim to provide a slightly different perspective into their nature -- through the perspective of their Hessian maps. The reason is that the loss Hessian captures the pairwise interaction of parameters and therefore forms a natural ground to probe how the architectural aspects of CNN get manifested in its structure and properties. We develop a framework relying on Toeplitz representation of CNNs, and then utilize it to reveal the Hessian structure and, in particular, its rank. We prove tight upper bounds (with linear activations), which closely follow the empirical trend of the Hessian rank and hold in practice in more general settings. Overall, our work generalizes and establishes the key insight that, even in CNNs, the Hessian rank grows as the square root of the number of parameters.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Neglected Hessian component explains mysteries in sharpness regularizationYann N. Dauphin, Atish Agarwala, Hossein MobahiNeurIPS 2024 · 被引用 16 次
- Theoretical Characterisation of the Gauss Newton Conditioning in Neural NetworksJim Zhao, Sidak Pal Singh, Aurélien LucchiNeurIPS 2024 · 被引用 7 次
- Convolutions and More as Einsum: A Tensor Network Perspective with Advances for Second-Order MethodsFelix DangelNeurIPS 2024 · 被引用 5 次
- Avoiding spurious sharpness minimization broadens applicability of SAMSidak Pal Singh, Hossein Mobahi, Atish Agarwala, Yann N. DauphinICML 2025
- Logits are All We Need to Adapt Closed ModelsGaurush Hiranandani, Haolun Wu, Subhojyoti Mukherjee, Sanmi KoyejoICML 2025
它引用的顶会 Paper10
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer 等NeurIPS 2021 · 被引用 3,862 次
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan 等ICLR 2020 · 被引用 705 次
相关 Paper
- Analytic Insights into Structure and Rank of Neural Network Hessian MapsSidak Pal Singh, Gregor Bachmann, Thomas HofmannNeurIPS 2021 · 被引用 60 次
- Which Frequencies do CNNs Need? Emergent Bottleneck Structure in Feature LearningYuxiao Wen, Arthur JacotICML 2024 · 被引用 9 次
- Investigating the Overlooked Hessian Structure: From CNNs to LLMsQian-Yuan Tang, Yufei Gu, Yunfeng Cai, Mingming Sun 等ICML 2025
- Linear CNNs Discover the Statistical Structure of the Dataset Using Only the Most Dominant FrequenciesHannah Pinson, Joeri Lenaerts, Vincent GinisICML 2023 · 被引用 8 次
- On Lipschitz Regularization of Convolutional Layers using Toeplitz Matrix TheoryAlexandre Araujo, Benjamin Négrevergne, Yann Chevaleyre, Jamal AtifAAAI 2021 · 被引用 31 次
