Information-Theoretic Understanding of Population Risk Improvement with Model Compression
Yuheng Bu, Weihao Gao, Shaofeng Zou, Venugopal V. Veeravalli
Abstract
We show that model compression can improve the population risk of a pre-trained model, by studying the tradeoff between the decrease in the generalization error and the increase in the empirical risk with model compression. We first prove that model compression reduces an information-theoretic bound on the generalization error; this allows for an interpretation of model compression as a regularization technique to avoid overfitting. We then characterize the increase in empirical risk with model compression using rate distortion theory. These results imply that the population risk could be improved by model compression if the decrease in generalization error exceeds the increase in empirical risk. We show through a linear regression example that such a decrease in population risk due to model compression is indeed possible. Our theoretical results further suggest that the Hessian-weighted K-means clustering compression approach can be improved by regularizing the distance between the clustering centers. We provide experiments with neural networks to support our theoretical assertions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- An Exact Characterization of the Generalization Error for the Gibbs AlgorithmGholamali Aminian, Yuheng Bu, Laura Toni, Miguel R. D. Rodrigues et al.NeurIPS 2021 · 75 citations
- Tighter Expected Generalization Error Bounds via Wasserstein DistanceBorja Rodríguez Gálvez, Germán Bassi, Ragnar Thobaben, Mikael SkoglundNeurIPS 2021 · 52 citations
- Rate-Distortion Analysis of Minimum Excess Risk in Bayesian LearningHassan Hafez-Kolahi, Behrad Moniri, Shohreh Kasaei, Mahdieh Soleymani BaghshahICML 2021 · 12 citations
- Towards the Fundamental Limits of Knowledge Transfer over Finite DomainsQingyue Zhao, Banghua ZhuICLR 2024 · 5 citations
- Knowledge Distillation as Semiparametric InferenceTri Dao, Govinda M. Kamath, Vasilis Syrgkanis, Lester MackeyICLR 2021 · 4 citations
Related papers
- Compression based bound for non-compressed network: unified generalization error analysis of large compressible deep neural networkTaiji Suzuki, Hiroshi Abe, Tomoaki NishimuraICLR 2020 · 57 citations
- Slicing Mutual Information Generalization Bounds for Neural NetworksKimia Nadjahi, Kristjan H. Greenewald, Rickard Brüel Gabrielsson, Justin SolomonICML 2024 · 5 citations
- Cauchy-Schwarz Divergence Information Bottleneck for RegressionShujian Yu, Xi Yu, Sigurd Løkse, Robert Jenssen et al.ICLR 2024 · 16 citations
- Measuring Model Complexity of Neural Networks with Curve Activation FunctionsXia Hu, Weiqing Liu, Jiang Bian, Jian PeiKDD 2020 · 25 citations
- Weight matrices compression based on PDB model in deep neural networksXiaoling Wu, Junpeng Zhu, Zeng LiICML 2025
