TCAQ-DM: Timestep-Channel Adaptive Quantization for Diffusion Models
Haocheng Huang, Jiaxin Chen, Jinyang Guo, Ruiyi Zhan, Yunhong Wang
摘要
Diffusion models have achieved remarkable success in the image and video generation tasks. Nevertheless, they often require a large amount of memory and time overhead during inference, due to the complex network architecture and considerable number of timesteps for iterative diffusion. Recently, the post-training quantization (PTQ) technique has proved a promising way to reduce the inference cost by quantizing the float-point operations to low-bit ones. However, most of them fail to tackle with the large variations in the distribution of activations across distinct channels and timesteps, as well as the inconsistent of input between quantization and inference on diffusion models, thus leaving much room for improvement. To address the above issues, we propose a novel method dubbed Timestep-Channel Adaptive Quantization for Diffusion Models (TCAQ-DM). Specifically, we develop a timestep-channel joint reparameterization (TCR) module to balance the activation range along both the timesteps and channels, facilitating the successive reconstruction procedure. Subsequently, we employ a dynamically adaptive quantization (DAQ) module that mitigate the quantization error by selecting an optimal quantizer for each post-Softmax layers according to their specific types of distributions. Moreover, we present a progressively aligned reconstruction (PAR) strategy to mitigate the bias caused by the input mismatch. Extensive experiments on various benchmarks and distinct diffusion models demonstrate that the proposed method substantially outperforms the state-of-the-art approaches in most cases, especially yielding comparable FID metrics to the full precision model on CIFAR-10 in the W6A6 setting, while enabling generating available images in the W4A4 settings.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- FPSAttention: Training-Aware FP8 and Sparsity Co-Design for Fast Video DiffusionAkide Liu, Zeyu Zhang, Zhexin Li, Xuehai Bai 等NeurIPS 2025 · 被引用 19 次
- Beyond Uniformity: Sample and Frequency Meta Weighting for Post-Training Quantization of Diffusion ModelsVan Cuong Pham, Anh Hoang, Cuong Nguyen, Trung Le 等ICLR 2026
- APHQ-ViT: Post-Training Quantization with Average Perturbation Hessian Based Reconstruction for Vision TransformersZhuguanyu Wu, Jiayi Zhang, Jiaxin Chen, Jinyang Guo 等CVPR 2025
- DilateQuant: Accurate and Efficient Quantization-Aware Training for Diffusion Models via Weight DilationXuewen Liu, Zhikai Li, Minghao Jiang, Mengjuan Chen 等ACM MM 2025
它引用的顶会 Paper31
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan 等NeurIPS 2022 · 被引用 2,948 次
- Up or Down? Adaptive Rounding for Post-Training QuantizationMarkus Nagel, Rana Ali Amjad, Mart van Baalen, Christos Louizos 等ICML 2020 · 被引用 816 次
相关 Paper
- Q-Diffusion: Quantizing Diffusion ModelsXiuyu Li, Yijiang Liu, Long Lian, Huanrui Yang 等ICCV 2023 · 被引用 279 次
- Gradient-Aligned Calibration for Post-Training Quantization of Diffusion ModelsDung Anh Hoang, Cuong Pham, Trung Le, Jianfei Cai 等ICLR 2026 · 被引用 1 次
- TFMQ-DM: Temporal Feature Maintenance Quantization for Diffusion ModelsYushi Huang, Ruihao Gong, Jing Liu, Tianlong Chen 等CVPR 2024
- Towards Accurate Post-Training Quantization for Diffusion ModelsChangyuan Wang, Ziwei Wang, Xiuwei Xu, Yansong Tang 等CVPR 2024 · 被引用 8 次
- Temporal Dynamic Quantization for Diffusion ModelsJunhyuk So, Jungwon Lee, Daehyun Ahn, Hyungjun Kim 等NeurIPS 2023 · 被引用 109 次
