PAC-Bayes Compression Bounds So Tight That They Can Explain Generalization
Sanae Lotfi, Marc Finzi, Sanyam Kapoor, Andres Potapczynski, Micah Goldblum, Andrew Gordon Wilson
Abstract
While there has been progress in developing non-vacuous generalization bounds for deep neural networks, these bounds tend to be uninformative about why deep learning works. In this paper, we develop a compression approach based on quantizing neural network parameters in a linear subspace, profoundly improving on previous results to provide state-of-the-art generalization bounds on a variety of tasks, including transfer learning. We use these tight bounds to better understand the role of model size, equivariance, and the implicit biases of optimization, for generalization in deep learning. Notably, we find large models can be compressed to a much greater extent than previously known, encapsulating Occam's razor. We also argue for data-independent bounds in explaining generalization.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e46a2e33-3bea-47ab-ad0b-5f1c167d8ca2Cited by top-tier papers35
- How Does Information Bottleneck Help Deep Learning?Kenji Kawaguchi, Zhun Deng, Xu Ji, Jiaoyang HuangICML 2023 · 117 citations
- Bayesian Model Selection, the Marginal Likelihood, and GeneralizationSanae Lotfi, Pavel Izmailov, Gregory W. Benton, Micah Goldblum et al.ICML 2022 · 83 citations
- Non-Vacuous Generalization Bounds for Large Language ModelsSanae Lotfi, Marc Anton Finzi, Yilun Kuang, Tim G. J. Rudner et al.ICML 2024 · 49 citations
- Simplifying Neural Network Training Under Class ImbalanceRavid Shwartz-Ziv, Micah Goldblum, Yucen Lily Li, C. Bayan Bruss et al.NeurIPS 2023 · 45 citations
- Unlocking Tokens as Data Points for Generalization Bounds on Larger Language ModelsSanae Lotfi, Yilun Kuang, Marc Finzi, Brandon Amos et al.NeurIPS 2024 · 29 citations
Builds on11
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang et al.ICLR 2020 · 1,108 citations
- What Are Bayesian Neural Network Posteriors Really Like?Pavel Izmailov, Sharad Vikram, Matthew D. Hoffman, Andrew Gordon WilsonICML 2021 · 458 citations
- Generalizing Convolutional Neural Networks for Equivariance to Lie Groups on Arbitrary Continuous DataMarc Finzi, Samuel Stanton, Pavel Izmailov, Andrew Gordon WilsonICML 2020 · 372 citations
- On the Noisy Gradient Descent that Generalizes as SGDJingfeng Wu, Wenqing Hu, Haoyi Xiong, Jun Huan et al.ICML 2020 · 125 citations
- On the Generalization Benefit of Noise in Stochastic Gradient DescentSamuel L. Smith, Erich Elsen, Soham DeICML 2020 · 122 citations
Related papers
- Compression based bound for non-compressed network: unified generalization error analysis of large compressible deep neural networkTaiji Suzuki, Hiroshi Abe, Tomoaki NishimuraICLR 2020 · 57 citations
- In-Context Learning and Occam's RazorEric Elmoznino, Tom Marty, Tejas Kasetty, Léo Gagnon et al.ICML 2025
- PAC-Bayes Information BottleneckZifeng Wang, Shao-Lun Huang, Ercan Engin Kuruoglu, Jimeng Sun et al.ICLR 2022 · 42 citations
- Generalization Bounds via Meta-Learned Model Representations: PAC-Bayes and Sample Compression HypernetworksBenjamin Leblanc, Mathieu Bazinet, Nathaniel D'Amours, Alexandre Drouin et al.ICML 2025
- VC dimension of partially quantized neural networks in the overparametrized regimeYutong Wang, Clayton ScottICLR 2022 · 1 citation
