SQ-DM: Accelerating Diffusion Models with Aggressive Quantization and Temporal Sparsity
Zichen Fan, Steve Dai, Rangharajan Venkatesan, Dennis Sylvester, Brucek Khailany
摘要
Diffusion models have gained significant popularity in image generation tasks. However, generating high-quality content remains notably slow because it requires running model inference over many time steps. To accelerate these models, we propose to aggressively quantize both weights and activations, while simultaneously promoting significant activation sparsity. We further observe that the stated sparsity pattern varies among different channels and evolves across time steps. To support this quantization and sparsity scheme, we present a novel diffusion model accelerator featuring a heterogeneous mixed-precision dense-sparse architecture, channel-last address mapping, and a time-step-aware sparsity detector for efficient handling of the sparsity pattern. Our 4-bit quantization technique demonstrates superior generation quality compared to existing -bit methods. Our custom accelerator achieves speed-up and 51.5% energy reduction compared to traditional dense accelerators.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Test-Time Iterative Error Correction for Efficient Diffusion ModelsYunshan Zhong, Weiqi Yan, Yuxin ZhangICLR 2026 · 被引用 4 次
- DDiT: Dynamic Patch Scheduling for Efficient Diffusion TransformersDahye Kim, Deepti Ghadiyaram, Raghudeep GaddeCVPR 2026 · 被引用 3 次
它引用的顶会 Paper13
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 被引用 3,959 次
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan 等NeurIPS 2022 · 被引用 2,948 次
- SIGMA: A Sparse and Irregular GEMM Accelerator with Flexible Interconnects for DNN TrainingEric Qin, Ananda Samajdar, Hyoukjun Kwon, Vineet Nadella 等HPCA 2020 · 被引用 490 次
- Q-Diffusion: Quantizing Diffusion ModelsXiuyu Li, Yijiang Liu, Long Lian, Huanrui Yang 等ICCV 2023 · 被引用 279 次
- Structural Pruning for Diffusion ModelsGongfan Fang, Xinyin Ma, Xinchao WangNeurIPS 2023 · 被引用 257 次
相关 Paper
- MHDiff: Memory- and Hardware-Efficient Diffusion Acceleration via Focal Pixel Aware QuantizationChunyu Qi, Xuhang Wang, Ruiyang Chen, Yuanzheng Yao 等DAC 2025
- QuEST: Low-Bit Diffusion Model Quantization via Efficient Selective FinetuningHaoxuan Wang, Yuzhang Shang, Zhihang Yuan, Junyi Wu 等ICCV 2025 · 被引用 3 次
- Ditto: Accelerating Diffusion Model via Temporal Value SimilaritySungbin Kim, Hyunwuk Lee, Wonho Cho, Mincheol Park 等HPCA 2025 · 被引用 9 次
- RT-Lynx: Putting the GEMM Sparsity In a Right Way for Diffusion ModelsXing Cong, Hanlin Tang, Kan Liu, Lan Tao 等ICML 2026 · 被引用 1 次
- Outlier-Aware Post-Training Quantization for Discrete Graph Diffusion ModelsZheng Gong, Ying SunICML 2025
