Slicing Mutual Information Generalization Bounds for Neural Networks
Kimia Nadjahi, Kristjan H. Greenewald, Rickard Brüel Gabrielsson, Justin Solomon
摘要
The ability of machine learning (ML) algorithms to generalize well to unseen data has been studied through the lens of information theory, by bounding the generalization error with the input-output mutual information (MI), i.e., the MI between the training data and the learned hypothesis. Yet, these bounds have limited practicality for modern ML applications (e.g., deep learning), due to the difficulty of evaluating MI in high dimensions. Motivated by recent findings on the compressibility of neural networks, we consider algorithms that operate by slicing the parameter space, i.e., trained on random lower-dimensional subspaces. We introduce new, tighter information-theoretic generalization bounds tailored for such algorithms, demonstrating that slicing improves generalization. Our bounds offer significant computational and statistical advantages over standard MI bounds, as they rely on scalable alternative measures of dependence, i.e., disintegrated mutual information and -sliced mutual information. Then, we extend our analysis to algorithms whose parameters do not need to exactly lie on random subspaces, by leveraging rate-distortion theory. This strategy yields generalization bounds that incorporate a distortion term measuring model compressibility under slicing, thereby tightening existing bounds without compromising performance or requiring model compression. Building on this, we propose a regularization scheme enabling practitioners to control generalization through compressibility. Finally, we empirically validate our results and achieve the computation of non-vacuous information-theoretic generalization bounds for neural networks, a task that was previously out of reach.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Tighter CMI-Based Generalization Bounds via Stochastic Projection and QuantizationMilad Sefidgaran, Kimia Nadjahi, Abdellatif ZaidiNeurIPS 2025 · 被引用 2 次
- Curse of Slicing: Why Sliced Mutual Information is a Deceptive Measure of Statistical DependenceAlexander Semenenko, Ivan Butakov, Ivan Oseledets, Alexey FrolovICLR 2026 · 被引用 1 次
- Compress then Serve: Serving Thousands of LoRA Adapters with Little OverheadRickard Brüel Gabrielsson, Jiacheng Zhu, Onkar Bhardwaj, Leshem Choshen 等ICML 2025
它引用的顶会 Paper12
- Once-for-All: Train One Network and Specialize it for Efficient DeploymentHan Cai, Chuang Gan, Tianzhe Wang, Zhekai Zhang 等ICLR 2020 · 被引用 1,522 次
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan 等ICLR 2020 · 被引用 705 次
- Relative Flatness and GeneralizationHenning Petzka, Michael Kamp, Linara Adilova, Cristian Sminchisescu 等NeurIPS 2021 · 被引用 114 次
- PAC-Bayes Compression Bounds So Tight That They Can Explain GeneralizationSanae Lotfi, Marc Finzi, Sanyam Kapoor, Andres Potapczynski 等NeurIPS 2022 · 被引用 98 次
- Hausdorff Dimension, Heavy Tails, and Generalization in Neural NetworksUmut Simsekli, Ozan Sener, George Deligiannidis, Murat A. ErdogduNeurIPS 2020 · 被引用 79 次
相关 Paper
- How Does Information Bottleneck Help Deep Learning?Kenji Kawaguchi, Zhun Deng, Xu Ji, Jiaoyang HuangICML 2023 · 被引用 117 次
- -Sliced Mutual Information: A Quantitative Study of Scalability with DimensionZiv Goldfeld, Kristjan H. Greenewald, Theshani Nuradha, Galen ReevesNeurIPS 2022 · 被引用 15 次
- Sliced Mutual Information: A Scalable Measure of Statistical DependenceZiv Goldfeld, Kristjan H. GreenewaldNeurIPS 2021 · 被引用 48 次
- Information-theoretic generalization bounds for black-box learning algorithmsHrayr Harutyunyan, Maxim Raginsky, Greg Ver Steeg, Aram GalstyanNeurIPS 2021 · 被引用 61 次
- Information Bottleneck Analysis of Deep Neural Networks via Lossy CompressionIvan Butakov, Aleksander Tolmachev, Sofia Malanchuk, Anna Neopryatnaya 等ICLR 2024 · 被引用 20 次
