Model Preserving Compression for Neural Networks
Jerry Chee, Megan Flynn, Anil Damle, Christopher De Sa
Abstract
After training complex deep learning models, a common task is to compress the model to reduce compute and storage demands. When compressing, it is desirable to preserve the original model's per-example decisions (e.g., to go beyond top-1 accuracy or preserve robustness), maintain the network's structure, automatically determine per-layer compression levels, and eliminate the need for fine tuning. No existing compression methods simultaneously satisfy these criteria-we introduce a principled approach that does by leveraging interpolative decompositions. Our approach simultaneously selects and eliminates channels (analogously, neurons), then constructs an interpolation matrix that propagates a correction into the next layer, preserving the network's structure. Consequently, our method achieves good performance even without fine tuning and admits theoretical analysis. Our theoretical generalization bound for a one layer network lends itself naturally to a heuristic that allows our method to automatically choose per-layer sizes for deep networks. We demonstrate the efficacy of our approach with strong empirical performance on a variety of tasks, models, and datasets-from simple one-hiddenlayer networks to deep networks on ImageNet. * equal contribution Preprint. Under review.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 278aef28-ba04-402e-ae0a-963f9ce9f7e1Cited by top-tier papers1
Ask how each one uses itBuilds on10
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Linear Mode Connectivity and the Lottery Ticket HypothesisJonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICML 2020 · 750 citations
- Neuron-level Structured Pruning using Polarization RegularizerTao Zhuang, Zhixuan Zhang, Yuheng Huang, Xiaoyi Zeng et al.NeurIPS 2020 · 168 citations
- Provable Filter Pruning for Efficient Neural NetworksLucas Liebenwein, Cenk Baykal, Harry Lang, Dan Feldman et al.ICLR 2020 · 161 citations
- DeepHoyer: Learning Sparser Neural Network with Differentiable Scale-Invariant Sparsity MeasuresHuanrui Yang, Wei Wen, Hai LiICLR 2020 · 109 citations
Related papers
- Compressing Neural Networks: Towards Determining the Optimal Layer-wise DecompositionLucas Liebenwein, Alaa Maalouf, Dan Feldman, Daniela RusNeurIPS 2021 · 60 citations
- Towards Compact CNNs via Collaborative CompressionYuchao Li, Shaohui Lin, Jianzhuang Liu, Qixiang Ye et al.CVPR 2021
- How Informative is the Approximation Error from Tensor Decomposition for Neural Network Compression?Jetze Schuurmans, Kim Batselier, Julian F. P. KooijICLR 2023
- Data-Efficient Structured Pruning via Submodular OptimizationMarwa El Halabi, Suraj Srinivas, Simon Lacoste-JulienNeurIPS 2022 · 31 citations
- OPQ: Compressing Deep Neural Networks with One-shot Pruning-QuantizationPeng Hu, Xi Peng, Hongyuan Zhu, Mohamed M. Sabry Aly et al.AAAI 2021 · 79 citations
