BiDM: Pushing the Limit of Quantization for Diffusion Models
Xingyu Zheng, Xianglong Liu, Yichen Bian, Xudong Ma, Yulun Zhang, Jiakai Wang, Jinyang Guo, Haotong Qin
摘要
Diffusion models (DMs) have been significantly developed and widely used in various applications due to their excellent generative qualities. However, the expensive computation and massive parameters of DMs hinder their practical use in resource-constrained scenarios. As one of the effective compression approaches, quantization allows DMs to achieve storage saving and inference acceleration by reducing bit-width while maintaining generation performance. However, as the most extreme quantization form, 1-bit binarization causes the generation performance of DMs to face severe degradation or even collapse. This paper proposes a novel method, namely BiDM, for fully binarizing weights and activations of DMs, pushing quantization to the 1-bit limit. From a temporal perspective, we introduce the Timestep-friendly Binary Structure (TBS), which uses learnable activation binarizers and cross-timestep feature connections to address the highly timestep-correlated activation features of DMs. From a spatial perspective, we propose Space Patched Distillation (SPD) to address the difficulty of matching binary features during distillation, focusing on the spatial locality of image generation tasks and noise estimation networks. As the first work to fully binarize DMs, the W1A1 BiDM on the LDM-4 model for LSUN-Bedrooms 256256 achieves a remarkable FID of 22.74, significantly outperforming the current state-of-the-art general binarization methods with an FID of 59.44 and invalid generative samples, and achieves up to excellent 28.0 times storage and 52.7 times OPs savings. The code is available at https://github.com/Xingyu-Zheng/BiDM .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- QuantSparse: Comprehensively Compressing Video Diffusion Transformer with Model Quantization and Attention SparsificationWeilun Feng, Chuanguang Yang, Haotong Qin, Mingqiang Wu 等ICLR 2026 · 被引用 8 次
- First-Order Error Matters: Accurate Compensation for Quantized Large Language ModelsXingyu Zheng, Haotong Qin, Yuye Li, Haoran Chu 等AAAI 2026 · 被引用 2 次
- S2Q-VDiT: Accurate Quantized Video Diffusion Transformer with Salient Data and Sparse Token DistillationWeilun Feng, Haotong Qin, Chuanguang Yang, Xiangqi Li 等NeurIPS 2025 · 被引用 1 次
它引用的顶会 Paper33
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- DPM-Solver: A Fast ODE Solver for Diffusion Probabilistic Model Sampling in Around 10 StepsCheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen 等NeurIPS 2022 · 被引用 2,653 次
相关 Paper
- BinaryDM: Accurate Weight Binarization for Efficient Diffusion ModelsXingyu Zheng, Xianglong Liu, Haotong Qin, Xudong Ma 等ICLR 2025
- Q-DM: An Efficient Low-bit Quantized Diffusion ModelYanjing Li, Sheng Xu, Xianbin Cao, Xiao Sun 等NeurIPS 2023 · 被引用 70 次
- MPQ-DM: Mixed Precision Quantization for Extremely Low Bit Diffusion ModelsWeilun Feng, Haotong Qin, Chuanguang Yang, Zhulin An 等AAAI 2025 · 被引用 19 次
- BiBERT: Accurate Fully Binarized BERTHaotong Qin, Yifu Ding, Mingyuan Zhang, Qinghua Yan 等ICLR 2022 · 被引用 121 次
- Optimizing Quantized Diffusion Models via Distillation with Cross-Timestep Error CorrectionYanxi Li, Chengbin DuAAAI 2025 · 被引用 2 次
