Activation and Weight Distribution Balancing for Optimal Post-Training Quantization in Learned Image Compression
Jie Yu, Songping Mai, Peng Zhang, Yucheng Jiang, Jian Cheng
摘要
Recently, Learned Image Compression (LIC) models have garnered significant attention due to their superior performance in comparison to traditional image codecs. However, the growing complexity of these deep learning-based models results in high memory consumption and computational load, which limits their practical deployment. Quantization has emerged as a promising technique to reduce both the storage requirements and computational overhead. Despite its success in high-level vision tasks like image recognition and object detection, quantization techniques applied to LIC models remain underexplored. In this work, we identify the unique challenges of quantizing LIC models, specifically focusing on the impact of latent distribution ranges in high-bitrate. We observe that the activation layers of high-bitrate models exhibit a wider distribution range, which causes significant performance degradation after quantization. Furthermore, we explore the limitations of existing LIC quantization schemes, such as per-channel quantization for activation layers, which result in poor hardware acceleration performance and increased data storage overhead. To address these challenges, we propose the activation and weight distribution balancing post-training quantization (AWDB-PTQ) method for LIC models, which uses a coarse-to-fine strategy to optimize balancing coefficients. In addition, we employ per-tensor activation quantization and symmetric uniform quantization to better facilitate hardware acceleration. Experimental results demonstrate that our proposed method outperforms existing methods in terms of both compression performance and computational efficiency. Our code and data are available at: https://github.com/jie-yu16/AWDB-PTQ.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper1
问问它们各自怎么用它相关 Paper
- QuEST: Low-Bit Diffusion Model Quantization via Efficient Selective FinetuningHaoxuan Wang, Yuzhang Shang, Zhihang Yuan, Junyi Wu 等ICCV 2025 · 被引用 3 次
- ABQ-LLM: Arbitrary-Bit Quantized Inference Acceleration for Large Language ModelsChao Zeng, Songwei Liu, Yusheng Xie, Hong Liu 等AAAI 2025 · 被引用 24 次
- DynaQuant: Dynamic Mixed-Precision Quantization for Learned Image CompressionYouneng Bao, Yulong Cheng, Yiping Liu, Yichen Yang 等AAAI 2026
- MLWQ: Efficient Small Language Model Deployment via Multi-Level Weight QuantizationChun Hu, Junhui He, Shangyu Wu, Yuxin He 等EMNLP 2025 · 被引用 1 次
- Learnable Companding Quantization for Accurate Low-Bit Neural NetworksKohei YamamotoCVPR 2021
