Structured Multi-Hashing for Model Compression
Elad Eban, Yair Movshovitz-Attias, Hao Wu, Mark Sandler, Andrew Poon, Yerlan Idelbayev, Miguel Á. Carreira-Perpiñán
摘要
Despite the success of deep neural networks (DNNs), state-of-the-art models are too large to deploy on lowresource devices or common server configurations in which multiple models are held in memory. Model compression methods address this limitation by reducing the memory footprint, latency, or energy consumption of a model with minimal impact on accuracy. We focus on the task of reducing the number of learnable variables in the model. In this work we combine ideas from weight hashing and dimensionality reductions resulting in a simple and powerful structured multi-hashing method based on matrix products that allows direct control of model size of any deep network and is trained end-to-end. We demonstrate the strength of our approach by compressing models from the ResNet, EfficientNet, and Mo-bileNet architecture families. Our method allows us to drastically decrease the number of variables while maintaining high accuracy. For instance, by applying our approach to EfficentNet-B4 (16M parameters) we reduce it to to the size of B0 (5M parameters), while gaining over 3% in accuracy over B0 baseline. On the commonly used benchmark CIFAR10 we reduce the ResNet32 model by 75% with no loss in quality, and are able to do a 10x compression while still achieving above 90% accuracy. * The author contribute equally to this paper. Elad and Yair contributed equally to the paper. They jointly proposed the idea of structured-multi-hashing. Yair was the main contributor to the manuscript. Elad wrote most of the code and ran EfficentNet experiments. Hao contributed to coding and experiments. Yerlan ran CIFAR and ResNet experiments and simplified some aspects of the structured hashing. Miguel advised Yerlan on issues about optimization and deep net compression. Mark, and Andrew helped with MobileNet, and ResNet experiments. † Worked performed while at Google Research.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- GDP: Stabilized Neural Network Pruning via Gates with Differentiable PolarizationYi Guo, Huan Yuan, Jianchao Tan, Zhangyang Wang 等ICCV 2021 · 被引用 52 次
- T-Basis: a Compact Representation for Neural NetworksAnton Obukhov, Maxim V. Rakhuba, Stamatios Georgoulis, Menelaos Kanakis 等ICML 2020 · 被引用 32 次
- BiPer: Binary Neural Networks Using a Periodic FunctionEdwin Vargas, Claudia V. Correa P., Carlos Hinojosa, Henry ArguelloCVPR 2024 · 被引用 10 次
- Network Memory Footprint Compression Through Jointly Learnable Codebooks and MappingsEdouard Yvinec, Arnaud Dapogny, Kevin BaillyICLR 2024 · 被引用 2 次
- Massively Scaling Heteroscedastic ClassifiersMark Collier, Rodolphe Jenatton, Basil Mustafa, Neil Houlsby 等ICLR 2023 · 被引用 1 次
它引用的顶会 Paper1
相关 Paper
- Asymmetric Deep Hashing for Efficient Hash Code CompressionShu Zhao, Dayan Wu, Wanqian Zhang, Yu Zhou 等ACM MM 2020 · 被引用 18 次
- Hardware-Aware Compression with Random Operation Access Specific Tile (ROAST) HashingAditya Desai, Keren Zhou, Anshumali ShrivastavaICML 2023 · 被引用 5 次
- And the Bit Goes Down: Revisiting the Quantization of Neural NetworksPierre Stock, Armand Joulin, Rémi Gribonval, Benjamin Graham 等ICLR 2020 · 被引用 157 次
- Binary Neural Network Hashing for Image RetrievalWanqian Zhang, Dayan Wu, Yu Zhou, Bo Li 等SIGIR 2021 · 被引用 19 次
- Neural Epitome Search for Architecture-Agnostic Network CompressionDaquan Zhou, Xiaojie Jin, Qibin Hou, Kaixin Wang 等ICLR 2020 · 被引用 13 次
