Does the Data Induce Capacity Control in Deep Learning?
Rubing Yang, Jialin Mao, Pratik Chaudhari
Abstract
We show that the input correlation matrix of typical classification datasets has an eigenspectrum where, after a sharp initial drop, a large number of small eigenvalues are distributed uniformly over an exponentially large range. This structure is mirrored in a network trained on this data: we show that the Hessian and the Fisher Information Matrix (FIM) have eigenvalues that are spread uniformly over exponentially large ranges. We call such eigenspectra "sloppy" because sets of weights corresponding to small eigenvalues can be changed by large magnitudes without affecting the loss. Networks trained on atypical datasets with non-sloppy inputs do not share these traits and deep networks trained on such datasets generalize poorly. Inspired by this, we study the hypothesis that sloppiness of inputs aids generalization in deep networks. We show that if the Hessian is sloppy, we can compute non-vacuous PAC-Bayes generalization bounds analytically. By exploiting our empirical observation that training predominantly takes place in the non-sloppy subspace of the FIM, we develop data-distribution dependent PAC-Bayes priors that lead to accurate generalization bounds using numerical optimization. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Robust Fine-Tuning of Deep Neural Networks with Hessian-based Generalization GuaranteesHaotian Ju, Dongyue Li, Hongyang R. ZhangICML 2022 · 41 citations
- Measuring and Controlling Solution Degeneracy across Task-Trained Recurrent Neural NetworksAnn Huang, Satpreet Harcharan Singh, Flavio Martinelli, Kanaka RajanNeurIPS 2025 · 22 citations
- Spectral Bias Outside the Training Set for Deep Networks in the Kernel RegimeBenjamin Bowman, Guido F. MontúfarNeurIPS 2022 · 17 citations
- A Picture of the Space of Typical Learnable TasksRahul Ramesh, Jialin Mao, Itay Griniasty, Rubing Yang et al.ICML 2023 · 7 citations
- Low-Rank Curvature for Zeroth-Order Optimization in LLM Fine-tuningHyunseok Seung, Jaewoo Lee, Hyunsuk KoAAAI 2026 · 2 citations
Related papers
- Unifying Low Dimensional Spectra in Deep LearningConnall Garrod, Jonathan KeatingICML 2026 · 12 citations
- Investigating the Overlooked Hessian Structure: From CNNs to LLMsQian-Yuan Tang, Yufei Gu, Yunfeng Cai, Mingming Sun et al.ICML 2025
- Catastrophic Fisher Explosion: Early Phase Fisher Matrix Impacts GeneralizationStanislaw Jastrzebski, Devansh Arpit, Oliver Åstrand, Giancarlo Kerg et al.ICML 2021 · 78 citations
- Analytic Characterization of the Hessian in Shallow ReLU Models: A Tale of SymmetryYossi Arjevani, Michael FieldNeurIPS 2020 · 22 citations
- What training reveals about neural network complexityAndreas Loukas, Marinos Poiitis, Stefanie JegelkaNeurIPS 2021 · 12 citations
