QVGen: Pushing the Limit of Quantized Video Generative Models
Yushi Huang, Ruihao Gong, Jing Liu, Yifu Ding, Chengtao Lv, Haotong Qin, Jun Zhang
Abstract
Video diffusion models (DMs) have enabled high-quality video synthesis. Yet, their substantial computational and memory demands pose serious challenges to real-world deployment, even on high-end GPUs. As a commonly adopted solution, quantization has proven notable success in reducing cost for image DMs, while its direct application to video DMs remains ineffective. In this paper, we present QVGen, a novel quantization-aware training (QAT) framework tailored for high-performance and inference-efficient video DMs under extremely low-bit quantization (e.g., -bit or below). We begin with a theoretical analysis demonstrating that reducing the gradient norm is essential to facilitate convergence for QAT. To this end, we introduce auxiliary modules () to mitigate large quantization errors, leading to significantly enhanced convergence. To eliminate the inference overhead of , we propose a rank-decay strategy that progressively eliminates . Specifically, we repeatedly employ singular value decomposition (SVD) and a proposed rank-based regularization to identify and decay low-contributing components. This strategy retains performance while zeroing out additional inference overhead. Extensive experiments across state-of-the-art (SOTA) video DMs, with parameter sizes ranging from , show that QVGen is the first to reach full-precision comparable quality under -bit settings. Moreover, it significantly outperforms existing methods. For instance, our -bit CogVideoX-2B achieves improvements of in Dynamic Degree and in Scene Consistency on VBench. Code and models are available at https://github.com/ModelTC/QVGen.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b4a8431b-d5fc-4ceb-bed8-6e81a4c76873Cited by top-tier papers3
- PAROAttention: Pattern-Aware ReOrdering for Efficient Sparse and Quantized Attention in Visual Generation ModelsTianchen Zhao, Ke Hong, Xinhao Yang, Xuefeng Xiao et al.NeurIPS 2025 · 19 citations
- LinVideo: A Post-Training Framework towards O(n) Attention in Efficient Video GenerationYushi Huang, Xingtong Ge, Ruihao Gong, Chengtao Lv et al.CVPR 2026 · 9 citations
- Just-in-Time: Training-Free Spatial Acceleration for Diffusion TransformersWenhao Sun, Ji Li, Zhaoqiang LiuCVPR 2026 · 5 citations
Builds on56
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
Related papers
- Quant VideoGen: Auto-Regressive Long Video Generation via 2-Bit KV-Cache QuantizationHaocheng Xi, Shuo Yang, Yilong Zhao, Muyang Li et al.ICML 2026 · 13 citations
- Q-VDiT: Towards Accurate Quantization and Distillation of Video-Generation Diffusion TransformersWeilun Feng, Chuanguang Yang, Haotong Qin, Xiangqi Li et al.ICML 2025
- DVD-Quant: Data-free Video Diffusion Transformers QuantizationZhiteng Li, Hanxuan Li, Junyi Wu, Kai Liu et al.ICLR 2026 · 13 citations
- EfficientDM: Efficient Quantization-Aware Fine-Tuning of Low-Bit Diffusion ModelsYefei He, Jing Liu, Weijia Wu, Hong Zhou et al.ICLR 2024 · 78 citations
- VETA-DiT: Variance-Equalized and Temporally Adaptive Quantization for Efficient 4-bit Diffusion TransformersQinkai Xu, Yijin Liu, Yang Chen, Lin F. Yang et al.NeurIPS 2025 · 3 citations
