Training with Quantization Noise for Extreme Model Compression
Pierre Stock, Angela Fan, Benjamin Graham, Edouard Grave, Rémi Gribonval, Hervé Jégou, Armand Joulin
Abstract
We tackle the problem of producing compact models, maximizing their accuracy for a given model size. A standard solution is to train networks with Quantization Aware Training (Jacob et al., 2018) , where the weights are quantized during training and the gradients approximated with the Straight-Through Estimator (Bengio et al., 2013) . In this paper, we extend this approach to work beyond int8 fixedpoint quantization with extreme compression methods where the approximations introduced by STE are severe, such as Product Quantization. Our proposal is to only quantize a different random subset of weights during each forward, allowing for unbiased gradients to flow through the other weights. Controlling the amount of noise and its form allows for extreme compression rates while maintaining the performance of the original model. As a result we establish new state-of-the-art compromises between accuracy and model size both in natural language processing and image classification. For example, applying our method to state-of-the-art Transformer and ConvNet architectures, we can achieve 82.5% accuracy on MNLI by compressing RoBERTa to 14 MB and 80.0% top-1 accuracy on ImageNet by compressing an EfficientNet-B3 to 3.3 MB. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1af4388d-cd04-47fc-a729-99331e5fa582Cited by top-tier papers57
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Going deeper with Image TransformersHugo Touvron, Matthieu Cord, Alexandre Sablayrolles, Gabriel Synnaeve et al.ICCV 2021 · 1,279 citations
- TransReID: Transformer-based Object Re-IdentificationShuting He, Hao Luo, Pichao Wang, Fan Wang et al.ICCV 2021 · 1,172 citations
- GPT3.int8(): 8-bit Matrix Multiplication for Transformers at ScaleTim Dettmers, Mike Lewis, Younes Belkada, Luke ZettlemoyerNeurIPS 2022 · 1,012 citations
- Movement Pruning: Adaptive Sparsity by Fine-TuningVictor Sanh, Thomas Wolf, Alexander M. RushNeurIPS 2020 · 656 citations
Builds on4
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Compressive Transformers for Long-Range Sequence ModellingJack W. Rae, Anna Potapenko, Siddhant M. Jayakumar, Chloe Hillier et al.ICLR 2020 · 833 citations
- Reducing Transformer Depth on Demand with Structured DropoutAngela Fan, Edouard Grave, Armand JoulinICLR 2020 · 695 citations
- And the Bit Goes Down: Revisiting the Quantization of Neural NetworksPierre Stock, Armand Joulin, Rémi Gribonval, Benjamin Graham et al.ICLR 2020 · 157 citations
Related papers
- Gradient Regularization for Quantization RobustnessMilad Alizadeh, Arash Behboodi, Mart van Baalen, Christos Louizos et al.ICLR 2020 · 8 citations
- Adaptive Loss-Aware Quantization for Multi-Bit NetworksZhongnan Qu, Zimu Zhou, Yun Cheng, Lothar ThieleCVPR 2020
- NIPQ: Noise proxy-based Integrated Pseudo-QuantizationJuncheol Shin, Junhyuk So, Sein Park, Seungyeop Kang et al.CVPR 2023
- Bit-shrinking: Limiting Instantaneous Sharpness for Improving Post-training QuantizationChen Lin, Bo Peng, Zheyang Li, Wenming Tan et al.CVPR 2023
- Robust Training of Neural Networks at Arbitrary Precision and SparsityChengxi Ye, Grace Chu, Yanfeng Liu, Yichi Zhang et al.ICLR 2026 · 2 citations
