BitsFusion: 1.99 bits Weight Quantization of Diffusion Model
Yang Sui, Yanyu Li, Anil Kag, Yerlan Idelbayev, Junli Cao, Ju Hu, Dhritiman Sagar, Bo Yuan, Sergey Tulyakov, Jian Ren
Abstract
Diffusion-based image generation models have achieved great success in recent years by showing the capability of synthesizing high-quality content. However, these models contain a huge number of parameters, resulting in a significantly large model size. Saving and transferring them is a major bottleneck for various applications, especially those running on resource-constrained devices. In this work, we develop a novel weight quantization method that quantizes the UNet from Stable Diffusion v1.5 to 1.99 bits, achieving a model with 7.9X smaller size while exhibiting even better generation quality than the original one. Our approach includes several novel techniques, such as assigning optimal bits to each layer, initializing the quantized model for better performance, and improving the training strategy to dramatically reduce quantization error. Furthermore, we extensively evaluate our quantized model across various benchmark datasets and through human evaluation to demonstrate its superior generation quality.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers18
- DeltaQuant: 4-bit Video Diffusion Models with Spatiotemporal Delta SmoothingXingyang Li, Samuel Tesfai, Zhekai Zhang, Haocheng Xi et al.CVPR 2026 · 7 citations
- SANA-Sprint: One-Step Diffusion with Continuous-Time Consistency DistillationJunsong Chen, Shuchen Xue, Yuyang Zhao, Jincheng Yu et al.ICCV 2025 · 6 citations
- Are Image-to-Video Models Good Zero-Shot Image Editors?Zechuan Zhang, Zhenyuan Chen, Zongxin Yang, Yi YangCVPR 2026 · 4 citations
- QuEST: Low-Bit Diffusion Model Quantization via Efficient Selective FinetuningHaoxuan Wang, Yuzhang Shang, Zhihang Yuan, Junyi Wu et al.ICCV 2025 · 3 citations
- Diffusion on Demand: Selective Caching and Modulation for Efficient GenerationHee Min Choi, Hyoa Kang, Dokwan Oh, Nam Ik ChoNeurIPS 2025 · 2 citations
Builds on49
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
Related papers
- MPQ-DM: Mixed Precision Quantization for Extremely Low Bit Diffusion ModelsWeilun Feng, Haotong Qin, Chuanguang Yang, Zhulin An et al.AAAI 2025 · 19 citations
- Temporal Dynamic Quantization for Diffusion ModelsJunhyuk So, Jungwon Lee, Daehyun Ahn, Hyungjun Kim et al.NeurIPS 2023 · 109 citations
- Q-Diffusion: Quantizing Diffusion ModelsXiuyu Li, Yijiang Liu, Long Lian, Huanrui Yang et al.ICCV 2023 · 279 citations
- Text Embedding Knows How to Quantize Text-Guided Diffusion ModelsHongjae Lee, Myungjun Son, Dongjea Kang, Seung-Won JungICCV 2025 · 2 citations
- Qua2SeDiMo: Quantifiable Quantization Sensitivity of Diffusion ModelsKeith G. Mills, Mohammad Salameh, Ruichen Chen, Negar Hassanpour et al.AAAI 2025
