Qua2SeDiMo: Quantifiable Quantization Sensitivity of Diffusion Models
Keith G. Mills, Mohammad Salameh, Ruichen Chen, Negar Hassanpour, Wei Lu, Di Niu
Abstract
Diffusion Models (DM) have democratized AI image generation through an iterative denoising process. Quantization is a major technique to alleviate the inference cost and reduce the size of DM denoiser networks. However, as denoisers evolve from variants of convolutional U-Nets toward newer Transformer architectures, it is of growing importance to understand the quantization sensitivity of different weight layers, operations and architecture types to performance. In this work, we address this challenge with Qua 2 SeDiMo, a mixedprecision Post-Training Quantization framework that generates explainable insights on the cost-effectiveness of various model weight quantization methods for different denoiser operation types and block structures. We leverage these insights to make high-quality mixed-precision quantization decisions for a myriad of diffusion models ranging from foundational U-Nets to state-of-the-art Transformers. As a result, Qua 2 SeDiMo can construct 3.4-bit, 3.9-bit, 3.65-bit and 3.7bit weight quantization on PixArt-α, PixArt-Σ, Hunyuan-DiT and SDXL, respectively. We further pair our weightquantization configurations with 6-bit activation quantization and outperform existing approaches in terms of quantitative metrics and generative image quality.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 77f772ff-cae7-48d7-8a01-3331f903a139Cited by top-tier papers2
- Re-ttention: Ultra Sparse Visual Generation via Attention Statistical ReshapeRuichen Chen, Keith G. Mills, Liyao Jiang, Chao Gao et al.NeurIPS 2025 · 10 citations
- Error Propagation Mechanisms and Compensation Strategies for Quantized Diffusion ModelsSongwei Liu, Chao Zeng, Chenqian Yan, Xurui Peng et al.ICML 2026 · 4 citations
Builds on26
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- How Attentive are Graph Attention Networks?Shaked Brody, Uri Alon, Eran YahavICLR 2022 · 1,717 citations
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu et al.ICML 2023 · 1,493 citations
Related papers
- VETA-DiT: Variance-Equalized and Temporally Adaptive Quantization for Efficient 4-bit Diffusion TransformersQinkai Xu, Yijin Liu, Yang Chen, Lin F. Yang et al.NeurIPS 2025 · 3 citations
- PTQD: Accurate Post-Training Quantization for Diffusion ModelsYefei He, Luping Liu, Jing Liu, Weijia Wu et al.NeurIPS 2023 · 219 citations
- RobuQ: Pushing DiTs to W1.58A2 via Robust Activation QuantizationKaicheng Yang, Xun Zhang, Haotong Qin, Yucheng Lin et al.ICML 2026 · 5 citations
- SQ-DM: Accelerating Diffusion Models with Aggressive Quantization and Temporal SparsityZichen Fan, Steve Dai, Rangharajan Venkatesan, Dennis Sylvester et al.DAC 2025 · 3 citations
- Quant-dLLM: Post-Training Extreme Low-Bit Quantization for Diffusion Large Language ModelsTianao Zhang, Zhiteng Li, Xianglong Yan, Haotong Qin et al.ICLR 2026 · 11 citations
