QVGen: Pushing the Limit of Quantized Video Generative Models
Yushi Huang, Ruihao Gong, Jing Liu, Yifu Ding, Chengtao Lv, Haotong Qin, Jun Zhang
摘要
Video diffusion models (DMs) have enabled high-quality video synthesis. Yet, their substantial computational and memory demands pose serious challenges to real-world deployment, even on high-end GPUs. As a commonly adopted solution, quantization has proven notable success in reducing cost for image DMs, while its direct application to video DMs remains ineffective. In this paper, we present QVGen, a novel quantization-aware training (QAT) framework tailored for high-performance and inference-efficient video DMs under extremely low-bit quantization (e.g., -bit or below). We begin with a theoretical analysis demonstrating that reducing the gradient norm is essential to facilitate convergence for QAT. To this end, we introduce auxiliary modules () to mitigate large quantization errors, leading to significantly enhanced convergence. To eliminate the inference overhead of , we propose a rank-decay strategy that progressively eliminates . Specifically, we repeatedly employ singular value decomposition (SVD) and a proposed rank-based regularization to identify and decay low-contributing components. This strategy retains performance while zeroing out additional inference overhead. Extensive experiments across state-of-the-art (SOTA) video DMs, with parameter sizes ranging from , show that QVGen is the first to reach full-precision comparable quality under -bit settings. Moreover, it significantly outperforms existing methods. For instance, our -bit CogVideoX-2B achieves improvements of in Dynamic Degree and in Scene Consistency on VBench. Code and models are available at https://github.com/ModelTC/QVGen.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- PAROAttention: Pattern-Aware ReOrdering for Efficient Sparse and Quantized Attention in Visual Generation ModelsTianchen Zhao, Ke Hong, Xinhao Yang, Xuefeng Xiao 等NeurIPS 2025 · 被引用 19 次
- LinVideo: A Post-Training Framework towards O(n) Attention in Efficient Video GenerationYushi Huang, Xingtong Ge, Ruihao Gong, Chengtao Lv 等CVPR 2026 · 被引用 9 次
- Just-in-Time: Training-Free Spatial Acceleration for Diffusion TransformersWenhao Sun, Ji Li, Zhaoqiang LiuCVPR 2026 · 被引用 5 次
它引用的顶会 Paper56
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
相关 Paper
- Quant VideoGen: Auto-Regressive Long Video Generation via 2-Bit KV-Cache QuantizationHaocheng Xi, Shuo Yang, Yilong Zhao, Muyang Li 等ICML 2026 · 被引用 13 次
- Q-VDiT: Towards Accurate Quantization and Distillation of Video-Generation Diffusion TransformersWeilun Feng, Chuanguang Yang, Haotong Qin, Xiangqi Li 等ICML 2025
- DVD-Quant: Data-free Video Diffusion Transformers QuantizationZhiteng Li, Hanxuan Li, Junyi Wu, Kai Liu 等ICLR 2026 · 被引用 13 次
- EfficientDM: Efficient Quantization-Aware Fine-Tuning of Low-Bit Diffusion ModelsYefei He, Jing Liu, Weijia Wu, Hong Zhou 等ICLR 2024 · 被引用 78 次
- VETA-DiT: Variance-Equalized and Temporally Adaptive Quantization for Efficient 4-bit Diffusion TransformersQinkai Xu, Yijin Liu, Yang Chen, Lin F. Yang 等NeurIPS 2025 · 被引用 3 次
