Memory-Efficient Generative Models via Product Quantization
Jie Shao, Hanxiao Zhang, Hao Yu, Jianxin Wu
摘要
The rapid progress in generative models has significantly enhanced the quality of image generation. However, as these models grow larger, deploying and fine-tuning them becomes increasingly challenging. While conventional quantization techniques help reduce model size, they struggle to achieve high compression rates without significant performance loss. As a result, memory footprint remains a critical challenge for generative models. In this work, we explore the extreme compression of generative models through codebook quantization, drastically reducing model size while maintaining performance. We extend product quantization for model compression, significantly increasing codebook capacity, which is crucial for preserving the generative quality of diffusion models. We also introduce a codebook compression method for memory efficiency. To further minimize performance degradation, we develop EM calibration with re-initialization that optimizes both assignments and centroids. By compressing the model to as low as 1 bit (achieving a 13× reduction in model size), we obtain a highly compact generative model with remarkable image quality. Extensive experiments on Ima-geNet demonstrate the superiority of our method over existing techniques. Furthermore, we validate its effectiveness across various generation, language and 3D tasks, highlighting its broad applicability and robust performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper33
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 被引用 11,724 次
相关 Paper
- BitsFusion: 1.99 bits Weight Quantization of Diffusion ModelYang Sui, Yanyu Li, Anil Kag, Yerlan Idelbayev 等NeurIPS 2024 · 被引用 48 次
- QuEST: Low-Bit Diffusion Model Quantization via Efficient Selective FinetuningHaoxuan Wang, Yuzhang Shang, Zhihang Yuan, Junyi Wu 等ICCV 2025 · 被引用 3 次
- Q-Diffusion: Quantizing Diffusion ModelsXiuyu Li, Yijiang Liu, Long Lian, Huanrui Yang 等ICCV 2023 · 被引用 279 次
- Text Embedding Knows How to Quantize Text-Guided Diffusion ModelsHongjae Lee, Myungjun Son, Dongjea Kang, Seung-Won JungICCV 2025 · 被引用 2 次
- CodeGEMM: A Codebook-Centric Approach to Efficient GEMM in Quantized LLMsGunho Park, Jeongin Bae, Byeongwook Kim, Baeseong Park 等NeurIPS 2025 · 被引用 3 次
