MPQ-DM: Mixed Precision Quantization for Extremely Low Bit Diffusion Models
Weilun Feng, Haotong Qin, Chuanguang Yang, Zhulin An, Libo Huang, Boyu Diao, Fei Wang, Renshuai Tao, Yongjun Xu, Michele Magno
Abstract
Diffusion models have received wide attention in generation tasks. However, the expensive computation cost prevents the application of diffusion models in resource-constrained scenarios. Quantization emerges as a practical solution that significantly saves storage and computation by reducing the bit-width of parameters. However, the existing quantization methods for diffusion models still cause severe degradation in performance, especially under extremely low bit-widths (2-4 bit). The primary decrease in performance comes from the significant discretization of activation values at low bit quantization. Too few activation candidates are unfriendly for outlier significant weight channel quantization, and the discretized features prevent stable learning over different time steps of the diffusion model. This paper presents MPQ-DM, a Mixed-Precision Quantization method for Diffusion Models. The proposed MPQ-DM mainly relies on two techniques: (1) To mitigate the quantization error caused by outlier severe weight channels, we propose an Outlier-Driven Mixed Quantization (OMQ) technique that uses Kurtosis to quantify outlier salient channels and apply optimized intra-layer mixed-precision bit-width allocation to recover accuracy performance within target efficiency. (2) To robustly learn representations crossing time steps, we construct a Time-Smoothed Relation Distillation (TRD) scheme between the quantized diffusion model and its full-precision counterpart, transferring discrete and continuous latent to a unified relation space to reduce the representation inconsistency. Comprehensive experiments demonstrate that MPQ-DM achieves significant accuracy gains under extremely low bit-widths compared with SOTA quantization methods. MPQ-DM achieves a 58% FID decrease under W2A4 setting compared with baseline, while all other methods even collapse.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4d1b5521-c941-40b2-8673-d39f4b89079fCited by top-tier papers15
- Quantized Visual Geometry Grounded TransformerWeilun Feng, Haotong Qin, Mingqiang Wu, Chuanguang Yang et al.ICLR 2026 · 17 citations
- QuantSparse: Comprehensively Compressing Video Diffusion Transformer with Model Quantization and Attention SparsificationWeilun Feng, Chuanguang Yang, Haotong Qin, Mingqiang Wu et al.ICLR 2026 · 8 citations
- PersonaLive! Expressive Portrait Image Animation for Live StreamingZhiyuan Li, Chi-Man Pun, Chen Fang, Jue Wang et al.CVPR 2026 · 6 citations
- RobuQ: Pushing DiTs to W1.58A2 via Robust Activation QuantizationKaicheng Yang, Xun Zhang, Haotong Qin, Yucheng Lin et al.ICML 2026 · 5 citations
- Test-Time Iterative Error Correction for Efficient Diffusion ModelsYunshan Zhong, Weiqi Yan, Yuxin ZhangICLR 2026 · 4 citations
Builds on31
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
- Learned Step Size quantizationSteven K. Esser, Jeffrey L. McKinstry, Deepika Bablani, Rathinakumar Appuswamy et al.ICLR 2020 · 1,037 citations
Related papers
- TCAQ-DM: Timestep-Channel Adaptive Quantization for Diffusion ModelsHaocheng Huang, Jiaxin Chen, Jinyang Guo, Ruiyi Zhan et al.AAAI 2025 · 4 citations
- Optimizing Quantized Diffusion Models via Distillation with Cross-Timestep Error CorrectionYanxi Li, Chengbin DuAAAI 2025 · 2 citations
- DilateQuant: Accurate and Efficient Quantization-Aware Training for Diffusion Models via Weight DilationXuewen Liu, Zhikai Li, Minghao Jiang, Mengjuan Chen et al.ACM MM 2025
- PTQD: Accurate Post-Training Quantization for Diffusion ModelsYefei He, Luping Liu, Jing Liu, Weijia Wu et al.NeurIPS 2023 · 219 citations
- BiDM: Pushing the Limit of Quantization for Diffusion ModelsXingyu Zheng, Xianglong Liu, Yichen Bian, Xudong Ma et al.NeurIPS 2024 · 12 citations
