Does the Data Induce Capacity Control in Deep Learning?
Rubing Yang, Jialin Mao, Pratik Chaudhari
摘要
We show that the input correlation matrix of typical classification datasets has an eigenspectrum where, after a sharp initial drop, a large number of small eigenvalues are distributed uniformly over an exponentially large range. This structure is mirrored in a network trained on this data: we show that the Hessian and the Fisher Information Matrix (FIM) have eigenvalues that are spread uniformly over exponentially large ranges. We call such eigenspectra "sloppy" because sets of weights corresponding to small eigenvalues can be changed by large magnitudes without affecting the loss. Networks trained on atypical datasets with non-sloppy inputs do not share these traits and deep networks trained on such datasets generalize poorly. Inspired by this, we study the hypothesis that sloppiness of inputs aids generalization in deep networks. We show that if the Hessian is sloppy, we can compute non-vacuous PAC-Bayes generalization bounds analytically. By exploiting our empirical observation that training predominantly takes place in the non-sloppy subspace of the FIM, we develop data-distribution dependent PAC-Bayes priors that lead to accurate generalization bounds using numerical optimization. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Robust Fine-Tuning of Deep Neural Networks with Hessian-based Generalization GuaranteesHaotian Ju, Dongyue Li, Hongyang R. ZhangICML 2022 · 被引用 41 次
- Measuring and Controlling Solution Degeneracy across Task-Trained Recurrent Neural NetworksAnn Huang, Satpreet Harcharan Singh, Flavio Martinelli, Kanaka RajanNeurIPS 2025 · 被引用 22 次
- Spectral Bias Outside the Training Set for Deep Networks in the Kernel RegimeBenjamin Bowman, Guido F. MontúfarNeurIPS 2022 · 被引用 17 次
- A Picture of the Space of Typical Learnable TasksRahul Ramesh, Jialin Mao, Itay Griniasty, Rubing Yang 等ICML 2023 · 被引用 7 次
- Low-Rank Curvature for Zeroth-Order Optimization in LLM Fine-tuningHyunseok Seung, Jaewoo Lee, Hyunsuk KoAAAI 2026 · 被引用 2 次
相关 Paper
- Unifying Low Dimensional Spectra in Deep LearningConnall Garrod, Jonathan KeatingICML 2026 · 被引用 12 次
- Investigating the Overlooked Hessian Structure: From CNNs to LLMsQian-Yuan Tang, Yufei Gu, Yunfeng Cai, Mingming Sun 等ICML 2025
- Catastrophic Fisher Explosion: Early Phase Fisher Matrix Impacts GeneralizationStanislaw Jastrzebski, Devansh Arpit, Oliver Åstrand, Giancarlo Kerg 等ICML 2021 · 被引用 78 次
- Analytic Characterization of the Hessian in Shallow ReLU Models: A Tale of SymmetryYossi Arjevani, Michael FieldNeurIPS 2020 · 被引用 22 次
- What training reveals about neural network complexityAndreas Loukas, Marinos Poiitis, Stefanie JegelkaNeurIPS 2021 · 被引用 12 次
