DVD-Quant: Data-free Video Diffusion Transformers Quantization
Zhiteng Li, Hanxuan Li, Junyi Wu, Kai Liu, Haotong Qin, Linghe Kong, Guihai Chen, Yulun Zhang, Xiaokang Yang
Abstract
Diffusion Transformers (DiTs) have emerged as the state-of-the-art architecture for video generation, yet their computational and memory demands hinder practical deployment. While post-training quantization (PTQ) presents a promising approach to accelerate Video DiT models, existing methods suffer from two critical limitations: (1) dependence on computation-heavy and inflexible calibration procedures, and (2) considerable performance deterioration after quantization. To address these challenges, we propose DVD-Quant, a novel Data-free quantization framework for Video DiTs. Our approach integrates three key innovations: (1) Bounded-init Grid Refinement (BGR) and (2) Auto-scaling Rotated Quantization (ARQ) for calibration data-free quantization error reduction, as well as (3) -Guided Bit Switching (-GBS) for adaptive bit-width allocation. Extensive experiments across multiple video generation benchmarks demonstrate that DVD-Quant achieves an approximately 2 speedup over full-precision baselines on advanced DiT models while maintaining visual fidelity. Notably, DVD-Quant is the first to enable W4A4 PTQ for Video DiTs without compromising video quality. Code and models will be available at https://github.com/lhxcs/DVD-Quant.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a7cee276-1070-4ebc-a7a7-ac4ec534041cCited by top-tier papers2
- Plug-and-Play Fidelity Optimization for Diffusion Transformer Acceleration via Cumulative Error MinimizationTong Shao, Yusen Fu, Guoying Sun, Jingde Kong et al.ICLR 2026 · 2 citations
- InfVSR: Toward Consistency-Driven Streaming Generative Video Super-ResolutionZiqing Zhang, Kai Liu, Zheng Chen, Xi Li et al.ICML 2026
Builds on23
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- SmoothQuant: Accurate and Efficient Post-Training Quantization for Large Language ModelsGuangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu et al.ICML 2023 · 1,493 citations
- QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMsSaleh Ashkboos, Amirkeivan Mohtashami, Maximilian L. Croci, Bo Li et al.NeurIPS 2024 · 723 citations
- OmniQuant: Omnidirectionally Calibrated Quantization for Large Language ModelsWenqi Shao, Mengzhao Chen, Zhaoyang Zhang, Peng Xu et al.ICLR 2024 · 395 citations
Related papers
- Q-DiT: Accurate Post-Training Quantization for Diffusion TransformersLei Chen, Yuan Meng, Chen Tang, Xinzhu Ma et al.CVPR 2025
- VQ4DiT: Efficient Post-Training Vector Quantization for Diffusion TransformersJuncan Deng, Shuaiting Li, Zeyu Wang, Hong Gu et al.AAAI 2025 · 12 citations
- VETA-DiT: Variance-Equalized and Temporally Adaptive Quantization for Efficient 4-bit Diffusion TransformersQinkai Xu, Yijin Liu, Yang Chen, Lin F. Yang et al.NeurIPS 2025 · 3 citations
- ViDiT-Q: Efficient and Accurate Quantization of Diffusion Transformers for Image and Video GenerationTianchen Zhao, Tongcheng Fang, Haofeng Huang, Rui Wan et al.ICLR 2025
- Q-DiT4SR: Exploration of Detail-Preserving Diffusion Transformer Quantization for Real-World Image Super-ResolutionXun Zhang, Kaicheng Yang, Hongliang Lu, Haotong Qin et al.ICML 2026 · 2 citations
