A Bayesian Optimization Framework for Neural Network Compression
Xingchen Ma, Amal Rannen Triki, Maxim Berman, Christos Sagonas, Jacques Calì, Matthew B. Blaschko
Abstract
Neural network compression is an important step for deploying neural networks where speed is of high importance, or on devices with limited memory. It is necessary to tune compression parameters in order to achieve the desired trade-off between size and performance. This is often done by optimizing the loss on a validation set of data, which should be large enough to approximate the true risk and therefore yield sufficient generalization ability. However, using a full validation set can be computationally expensive. In this work, we develop a general Bayesian optimization framework for optimizing functions that are computed based on U-statistics. We propagate Gaussian uncertainties from the statistics through the Bayesian optimization framework yielding a method that gives a probabilistic approximation certificate of the result. We then apply this to parameter selection in neural network compression. Compression objectives that can be written as U-statistics are typically based on empirical risk and knowledge distillation for deep discriminative models. We demonstrate our method on VGG and ResNet models, and the resulting system can find optimal compression parameters for relatively high-dimensional parametrizations in a matter of minutes on a standard desktop machine, orders of magnitude faster than competing methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eecac771-2fc5-4034-ba6d-c304bac52b2aCited by top-tier papers1
Ask how each one uses itRelated papers
- Bayesian Knowledge Distillation: A Bayesian Perspective of Distillation with Uncertainty QuantificationLuyang Fang, Yongkai Chen, Wenxuan Zhong, Ping MaICML 2024 · 10 citations
- Robust Model Compression Using Deep HypothesesOmri Armstrong, Ran Gilad-BachrachAAAI 2021 · 2 citations
- Distribution-Aware Tensor Decomposition for Compression of Convolutional Neural NetworksAlper Kalle, Théo Rudkiewicz, Mohamed Ouerfelli, Mohamed TamaazoustiNeurIPS 2025 · 2 citations
- Automatic Neural Network Compression by Sparsity-Quantization Joint Learning: A Constrained Optimization-Based ApproachHaichuan Yang, Shupeng Gui, Yuhao Zhu, Ji LiuCVPR 2020
- Low-Rank Compression of Neural Nets: Learning the Rank of Each LayerYerlan Idelbayev, Miguel Á. Carreira-PerpiñánCVPR 2020
