The Hessian perspective into the Nature of Convolutional Neural Networks
Sidak Pal Singh, Thomas Hofmann, Bernhard Schölkopf
Abstract
While Convolutional Neural Networks (CNNs) have long been investigated and applied, as well as theorized, we aim to provide a slightly different perspective into their nature -- through the perspective of their Hessian maps. The reason is that the loss Hessian captures the pairwise interaction of parameters and therefore forms a natural ground to probe how the architectural aspects of CNN get manifested in its structure and properties. We develop a framework relying on Toeplitz representation of CNNs, and then utilize it to reveal the Hessian structure and, in particular, its rank. We prove tight upper bounds (with linear activations), which closely follow the empirical trend of the Hessian rank and hold in practice in more general settings. Overall, our work generalizes and establishes the key insight that, even in CNNs, the Hessian rank grows as the square root of the number of parameters.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Neglected Hessian component explains mysteries in sharpness regularizationYann N. Dauphin, Atish Agarwala, Hossein MobahiNeurIPS 2024 · 16 citations
- Theoretical Characterisation of the Gauss Newton Conditioning in Neural NetworksJim Zhao, Sidak Pal Singh, Aurélien LucchiNeurIPS 2024 · 7 citations
- Convolutions and More as Einsum: A Tensor Network Perspective with Advances for Second-Order MethodsFelix DangelNeurIPS 2024 · 5 citations
- Avoiding spurious sharpness minimization broadens applicability of SAMSidak Pal Singh, Hossein Mobahi, Atish Agarwala, Yann N. DauphinICML 2025
- Logits are All We Need to Adapt Closed ModelsGaurush Hiranandani, Haolun Wu, Subhojyoti Mukherjee, Sanmi KoyejoICML 2025
Builds on10
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan et al.ICLR 2020 · 705 citations
Related papers
- Analytic Insights into Structure and Rank of Neural Network Hessian MapsSidak Pal Singh, Gregor Bachmann, Thomas HofmannNeurIPS 2021 · 60 citations
- Which Frequencies do CNNs Need? Emergent Bottleneck Structure in Feature LearningYuxiao Wen, Arthur JacotICML 2024 · 9 citations
- Investigating the Overlooked Hessian Structure: From CNNs to LLMsQian-Yuan Tang, Yufei Gu, Yunfeng Cai, Mingming Sun et al.ICML 2025
- Linear CNNs Discover the Statistical Structure of the Dataset Using Only the Most Dominant FrequenciesHannah Pinson, Joeri Lenaerts, Vincent GinisICML 2023 · 8 citations
- On Lipschitz Regularization of Convolutional Layers using Toeplitz Matrix TheoryAlexandre Araujo, Benjamin Négrevergne, Yann Chevaleyre, Jamal AtifAAAI 2021 · 31 citations
