OMPQ: Orthogonal Mixed Precision Quantization
Yuexiao Ma, Taisong Jin, Xiawu Zheng, Yan Wang, Huixia Li, Yongjian Wu, Guannan Jiang, Wei Zhang, Rongrong Ji
摘要
To bridge the ever-increasing gap between deep neural networks' complexity and hardware capability, network quantization has attracted more and more research attention. The latest trend of mixed precision quantization takes advantage of hardware's multiple bit-width arithmetic operations to unleash the full potential of network quantization. However, existing approaches rely heavily on an extremely time-consuming search process and various relaxations when seeking the optimal bit configuration. To address this issue, we propose to optimize a proxy metric of network orthogonality that can be efficiently solved with linear programming, which proves to be highly correlated with quantized model accuracy and bit-width. Our approach significantly reduces the search time and the required data amount by orders of magnitude, but without a compromise on quantization accuracy. Specifically, we achieve 72.08% Top-1 accuracy on ResNet-18 with 6.7Mb parameters, which does not require any searching iterations. Given the high efficiency and low data dependency of our algorithm, we use it for the post-training quantization, which achieves 71.27% Top-1 accuracy on MobileNetV2 with only 1.5Mb parameters.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper27
- F8Net: Fixed-Point 8-bit Only Multiplication for Network QuantizationQing Jin, Jian Ren, Richard Zhuang, Sumant Hanumante 等ICLR 2022 · 被引用 57 次
- Discovering Sparsity Allocation for Layer-wise Pruning of Large Language ModelsLujun Li, Peijie Dong, Zhenheng Tang, Xiang Liu 等NeurIPS 2024 · 被引用 51 次
- EMQ: Evolving Training-free Proxies for Automated Mixed Precision QuantizationPeijie Dong, Lujun Li, Zimian Wei, Xin Niu 等ICCV 2023 · 被引用 51 次
- Adaptive Layer Sparsity for Large Language Models via Activation Correlation AssessmentWei Li, Lujun Li, Mark G. Lee, Shengjie SunNeurIPS 2024 · 被引用 39 次
- Quantized Visual Geometry Grounded TransformerWeilun Feng, Haotong Qin, Mingqiang Wu, Chuanguang Yang 等ICLR 2026 · 被引用 17 次
它引用的顶会 Paper6
- Data-Free Quantization Through Weight Equalization and Bias CorrectionMarkus Nagel, Mart van Baalen, Tijmen Blankevoort, Max WellingICCV 2019 · 被引用 622 次
- BRECQ: Pushing the Limit of Post-Training Quantization by Block ReconstructionYuhang Li, Ruihao Gong, Xu Tan, Yang Yang 等ICLR 2021 · 被引用 619 次
- Differentiable Soft Quantization: Bridging Full-Precision and Low-Bit Neural NetworksRuihao Gong, Xianglong Liu, Shenghu Jiang, Tianxiang Li 等ICCV 2019 · 被引用 540 次
- HAWQ-V3: Dyadic Neural Network QuantizationZhewei Yao, Zhen Dong, Zhangcheng Zheng, Amir Gholami 等ICML 2021 · 被引用 240 次
- BinaryDuo: Reducing Gradient Mismatch in Binary Activation Network by Coupling Binary ActivationsHyungjun Kim, Kyungsu Kim, Jinseok Kim, Jae-Joon KimICLR 2020 · 被引用 51 次
相关 Paper
- Leveraging Inter-Layer Dependency for Post -Training QuantizationChangbao Wang, Dandan Zheng, Yuanliu Liu, Liang LiNeurIPS 2022 · 被引用 27 次
- Efficient and Generalizable Mixed-Precision Quantization via Topological EntropyNan Li, Yonghui Su, Lianbo MaNeurIPS 2025 · 被引用 4 次
- FracBits: Mixed Precision Quantization via Fractional Bit-WidthsLinjie Yang, Qing JinAAAI 2021 · 被引用 95 次
- InfoQ: Mixed-Precision Quantization via Global Information FlowMehmet Emre Akbulut, Hazem Hesham Yousef Shalby, Fabrizio Pittorino, Manuel RoveriAAAI 2026 · 被引用 2 次
- One-Shot Model for Mixed-Precision QuantizationIvan Koryakovskiy, Alexandra Yakovleva, Valentin Buchnev, Temur Isaev 等CVPR 2023
