PAC-Bayes Information Bottleneck
Zifeng Wang, Shao-Lun Huang, Ercan Engin Kuruoglu, Jimeng Sun, Xi Chen, Yefeng Zheng
摘要
Understanding the source of the superior generalization ability of NNs remains one of the most important problems in ML research. There have been a series of theoretical works trying to derive non-vacuous bounds for NNs. Recently, the compression of information stored in weights (IIW) is proved to play a key role in NNs generalization based on the PAC-Bayes theorem. However, no solution of IIW has ever been provided, which builds a barrier for further investigation of the IIW's property and its potential in practical deep learning. In this paper, we propose an algorithm for the efficient approximation of IIW. Then, we build an IIW-based information bottleneck on the trade-off between accuracy and information complexity of NNs, namely PIB. From PIB, we can empirically identify the fitting to compressing phase transition during NNs' training and the concrete connection between the IIW compression and the generalization. Besides, we verify that IIW is able to explain NNs in broad cases, e.g., varying batch sizes, over-parameterization, and noisy labels. Moreover, we propose an MCMC-based algorithm to sample from the optimal weight posterior characterized by PIB, which fulfills the potential of IIW in enhancing NNs in practice.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Leveraging Inpainting for Single-Image Shadow RemovalXiaoguang Li, Qing Guo, Rabab Abdelfattah, Di Lin 等ICCV 2023 · 被引用 40 次
- Gradient-based Parameter Selection for Efficient Fine-TuningZhi Zhang, Qizhe Zhang, Zijun Gao, Renrui Zhang 等CVPR 2024 · 被引用 19 次
- Cauchy-Schwarz Divergence Information Bottleneck for RegressionShujian Yu, Xi Yu, Sigurd Løkse, Robert Jenssen 等ICLR 2024 · 被引用 16 次
- Improving Generalization in Federated Learning with Model-Data Mutual Information Regularization: A Posterior Inference ApproachHao Zhang, Chenglin Li, Nuowen Kan, Ziyang Zheng 等NeurIPS 2024 · 被引用 13 次
- Lessons from Generalization Error Analysis of Federated Learning: You May Communicate Less Often!Milad Sefidgaran, Romain Chor, Abdellatif Zaidi, Yijun WanICML 2024 · 被引用 11 次
它引用的顶会 Paper4
- Graph Information BottleneckTailin Wu, Hongyu Ren, Pan Li, Jure LeskovecNeurIPS 2020 · 被引用 366 次
- Information Theoretic Counterfactual Learning from Missing-Not-At-Random FeedbackZifeng Wang, Xi Chen, Rui Wen, Shao-Lun Huang 等NeurIPS 2020 · 被引用 95 次
- Less Is Better: Unweighted Data Subsampling via Influence FunctionZifeng Wang, Hong Zhu, Zhenhua Dong, Xiuqiang He 等AAAI 2020 · 被引用 61 次
- Disentangled Information BottleneckZiqi Pan, Li Niu, Jianfu Zhang, Liqing ZhangAAAI 2021 · 被引用 55 次
相关 Paper
- Information Bottleneck Analysis of Deep Neural Networks via Lossy CompressionIvan Butakov, Aleksander Tolmachev, Sofia Malanchuk, Anna Neopryatnaya 等ICLR 2024 · 被引用 20 次
- How Does Information Bottleneck Help Deep Learning?Kenji Kawaguchi, Zhun Deng, Xu Ji, Jiaoyang HuangICML 2023 · 被引用 117 次
- Weight matrices compression based on PDB model in deep neural networksXiaoling Wu, Junpeng Zhu, Zeng LiICML 2025
- PAC-Bayes Compression Bounds So Tight That They Can Explain GeneralizationSanae Lotfi, Marc Finzi, Sanyam Kapoor, Andres Potapczynski 等NeurIPS 2022 · 被引用 98 次
- Compression based bound for non-compressed network: unified generalization error analysis of large compressible deep neural networkTaiji Suzuki, Hiroshi Abe, Tomoaki NishimuraICLR 2020 · 被引用 57 次
