PAC-Bayes Information Bottleneck
Zifeng Wang, Shao-Lun Huang, Ercan Engin Kuruoglu, Jimeng Sun, Xi Chen, Yefeng Zheng
Abstract
Understanding the source of the superior generalization ability of NNs remains one of the most important problems in ML research. There have been a series of theoretical works trying to derive non-vacuous bounds for NNs. Recently, the compression of information stored in weights (IIW) is proved to play a key role in NNs generalization based on the PAC-Bayes theorem. However, no solution of IIW has ever been provided, which builds a barrier for further investigation of the IIW's property and its potential in practical deep learning. In this paper, we propose an algorithm for the efficient approximation of IIW. Then, we build an IIW-based information bottleneck on the trade-off between accuracy and information complexity of NNs, namely PIB. From PIB, we can empirically identify the fitting to compressing phase transition during NNs' training and the concrete connection between the IIW compression and the generalization. Besides, we verify that IIW is able to explain NNs in broad cases, e.g., varying batch sizes, over-parameterization, and noisy labels. Moreover, we propose an MCMC-based algorithm to sample from the optimal weight posterior characterized by PIB, which fulfills the potential of IIW in enhancing NNs in practice.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 208d2d28-49c7-43ee-bac9-2bdec6436678Cited by top-tier papers15
- Leveraging Inpainting for Single-Image Shadow RemovalXiaoguang Li, Qing Guo, Rabab Abdelfattah, Di Lin et al.ICCV 2023 · 40 citations
- Gradient-based Parameter Selection for Efficient Fine-TuningZhi Zhang, Qizhe Zhang, Zijun Gao, Renrui Zhang et al.CVPR 2024 · 19 citations
- Cauchy-Schwarz Divergence Information Bottleneck for RegressionShujian Yu, Xi Yu, Sigurd Løkse, Robert Jenssen et al.ICLR 2024 · 16 citations
- Improving Generalization in Federated Learning with Model-Data Mutual Information Regularization: A Posterior Inference ApproachHao Zhang, Chenglin Li, Nuowen Kan, Ziyang Zheng et al.NeurIPS 2024 · 13 citations
- Lessons from Generalization Error Analysis of Federated Learning: You May Communicate Less Often!Milad Sefidgaran, Romain Chor, Abdellatif Zaidi, Yijun WanICML 2024 · 11 citations
Builds on4
- Graph Information BottleneckTailin Wu, Hongyu Ren, Pan Li, Jure LeskovecNeurIPS 2020 · 366 citations
- Information Theoretic Counterfactual Learning from Missing-Not-At-Random FeedbackZifeng Wang, Xi Chen, Rui Wen, Shao-Lun Huang et al.NeurIPS 2020 · 95 citations
- Less Is Better: Unweighted Data Subsampling via Influence FunctionZifeng Wang, Hong Zhu, Zhenhua Dong, Xiuqiang He et al.AAAI 2020 · 61 citations
- Disentangled Information BottleneckZiqi Pan, Li Niu, Jianfu Zhang, Liqing ZhangAAAI 2021 · 55 citations
Related papers
- Information Bottleneck Analysis of Deep Neural Networks via Lossy CompressionIvan Butakov, Aleksander Tolmachev, Sofia Malanchuk, Anna Neopryatnaya et al.ICLR 2024 · 20 citations
- How Does Information Bottleneck Help Deep Learning?Kenji Kawaguchi, Zhun Deng, Xu Ji, Jiaoyang HuangICML 2023 · 117 citations
- Weight matrices compression based on PDB model in deep neural networksXiaoling Wu, Junpeng Zhu, Zeng LiICML 2025
- PAC-Bayes Compression Bounds So Tight That They Can Explain GeneralizationSanae Lotfi, Marc Finzi, Sanyam Kapoor, Andres Potapczynski et al.NeurIPS 2022 · 98 citations
- Compression based bound for non-compressed network: unified generalization error analysis of large compressible deep neural networkTaiji Suzuki, Hiroshi Abe, Tomoaki NishimuraICLR 2020 · 57 citations
