Permute, Quantize, and Fine-Tune: Efficient Compression of Neural Networks
Julieta Martinez, Jashan Shewakramani, Ting-Wei Liu, Ioan Andrei Barsan, Wenyuan Zeng, Raquel Urtasun
摘要
Compressing large neural networks is an important step for their deployment in resource-constrained computational platforms. In this context, vector quantization is an appealing framework that expresses multiple parameters using a single code, and has recently achieved state-of-the-art network compression on a range of core vision and natural language processing tasks. Key to the success of vector quantization is deciding which parameter groups should be compressed together. Previous work has relied on heuristics that group the spatial dimension of individual convolutional filters, but a general solution remains unaddressed. This is desirable for pointwise convolutions (which dominate modern architectures), linear layers (which have no notion of spatial dimension), and convolutions (when more than one filter is compressed to the same codeword). In this paper we make the observation that the weights of two adjacent layers can be permuted while expressing the same function. We then establish a connection to rate-distortion theory and search for permutations that result in networks that are easier to compress. Finally, we rely on an annealed quantization algorithm to better compress the network and achieve higher final accuracy. We show results on image classification, object detection, and segmentation, reducing the gap with the uncompressed model by 40 to 70% w.r.t. the current state of the art. All our experiments can be reproduced using the code at https://github.com/uber-research/ permute-quantize-finetune .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Compressing LLMs: The Truth is Rarely Pure and Never SimpleAjay Kumar Jaiswal, Zhe Gan, Xianzhi Du, Bowen Zhang 等ICLR 2024 · 被引用 61 次
- Generalizing Weather Forecast to Fine-grained Temporal Scales via Physics-AI Hybrid ModelingWanghan Xu, Fenghua Ling, Wenlong Zhang, Tao Han 等NeurIPS 2024 · 被引用 31 次
- VQ4DiT: Efficient Post-Training Vector Quantization for Diffusion TransformersJuncan Deng, Shuaiting Li, Zeyu Wang, Hong Gu 等AAAI 2025 · 被引用 12 次
- MVQ: Towards Efficient DNN Compression and Acceleration with Masked Vector QuantizationShuaiting Li, Chengxuan Wang, Juncan Deng, Zeyu Wang 等ASPLOS 2025 · 被引用 5 次
- SSVQ: Unleashing the Potential of Vector Quantization with Sign-SplittingShuaiting Li, Juncan Deng, Chengxuan Wang, Kedong Xu 等ICCV 2025 · 被引用 2 次
它引用的顶会 Paper2
相关 Paper
- NVTC: Nonlinear Vector Transform CodingRunsen Feng, Zongyu Guo, Weiping Li, Zhibo ChenCVPR 2023
- Qinco2: Vector Compression and Search with Improved Implicit Neural CodebooksThéophane Vallaeys, Matthew J. Muckley, Jakob Verbeek, Matthijs DouzeICLR 2025
- OPQ: Compressing Deep Neural Networks with One-shot Pruning-QuantizationPeng Hu, Xi Peng, Hongyuan Zhu, Mohamed M. Sabry Aly 等AAAI 2021 · 被引用 79 次
- FSNet: Compression of Deep Convolutional Neural Networks by Filter SummaryYingzhen Yang, Jiahui Yu, Nebojsa Jojic, Jun Huan 等ICLR 2020 · 被引用 19 次
- And the Bit Goes Down: Revisiting the Quantization of Neural NetworksPierre Stock, Armand Joulin, Rémi Gribonval, Benjamin Graham 等ICLR 2020 · 被引用 157 次
