Information-Theoretic Understanding of Population Risk Improvement with Model Compression
Yuheng Bu, Weihao Gao, Shaofeng Zou, Venugopal V. Veeravalli
摘要
We show that model compression can improve the population risk of a pre-trained model, by studying the tradeoff between the decrease in the generalization error and the increase in the empirical risk with model compression. We first prove that model compression reduces an information-theoretic bound on the generalization error; this allows for an interpretation of model compression as a regularization technique to avoid overfitting. We then characterize the increase in empirical risk with model compression using rate distortion theory. These results imply that the population risk could be improved by model compression if the decrease in generalization error exceeds the increase in empirical risk. We show through a linear regression example that such a decrease in population risk due to model compression is indeed possible. Our theoretical results further suggest that the Hessian-weighted K-means clustering compression approach can be improved by regularizing the distance between the clustering centers. We provide experiments with neural networks to support our theoretical assertions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- An Exact Characterization of the Generalization Error for the Gibbs AlgorithmGholamali Aminian, Yuheng Bu, Laura Toni, Miguel R. D. Rodrigues 等NeurIPS 2021 · 被引用 75 次
- Tighter Expected Generalization Error Bounds via Wasserstein DistanceBorja Rodríguez Gálvez, Germán Bassi, Ragnar Thobaben, Mikael SkoglundNeurIPS 2021 · 被引用 52 次
- Rate-Distortion Analysis of Minimum Excess Risk in Bayesian LearningHassan Hafez-Kolahi, Behrad Moniri, Shohreh Kasaei, Mahdieh Soleymani BaghshahICML 2021 · 被引用 12 次
- Towards the Fundamental Limits of Knowledge Transfer over Finite DomainsQingyue Zhao, Banghua ZhuICLR 2024 · 被引用 5 次
- Knowledge Distillation as Semiparametric InferenceTri Dao, Govinda M. Kamath, Vasilis Syrgkanis, Lester MackeyICLR 2021 · 被引用 4 次
相关 Paper
- Compression based bound for non-compressed network: unified generalization error analysis of large compressible deep neural networkTaiji Suzuki, Hiroshi Abe, Tomoaki NishimuraICLR 2020 · 被引用 57 次
- Slicing Mutual Information Generalization Bounds for Neural NetworksKimia Nadjahi, Kristjan H. Greenewald, Rickard Brüel Gabrielsson, Justin SolomonICML 2024 · 被引用 5 次
- Cauchy-Schwarz Divergence Information Bottleneck for RegressionShujian Yu, Xi Yu, Sigurd Løkse, Robert Jenssen 等ICLR 2024 · 被引用 16 次
- Measuring Model Complexity of Neural Networks with Curve Activation FunctionsXia Hu, Weiqing Liu, Jiang Bian, Jian PeiKDD 2020 · 被引用 25 次
- Weight matrices compression based on PDB model in deep neural networksXiaoling Wu, Junpeng Zhu, Zeng LiICML 2025
