Quant VideoGen: Auto-Regressive Long Video Generation via 2-Bit KV-Cache Quantization
Haocheng Xi, Shuo Yang, Yilong Zhao, Muyang Li, Han Cai, Xingyang Li, Yujun Lin, Zhuoyang Zhang, Jintao Zhang, Xiuyu Li, Zhiying Xu, Jun Wu
Abstract
Despite rapid progress in auto-regressive video diffusion, we identify an emerging system–algorithm bottleneck that limits both deployability and generation quality: KV-cache memory. In auto-regressive video generation models, the KV-cache grows with generation history and quickly dominates GPU memory (often ≥30 GB), preventing deployment on widely available hardware. More critically, memory-bounded KV budgets constrain the effective working memory, directly degrading long-horizon consistency in identity, layout, and motion. To address this challenge, we present Quant VideoGen (QVG), a training-free KV-cache quantization framework for auto-regressive video diffusion models. QVG exploits video’s inherent spatiotemporal redundancy via Semantic-Aware Smoothing, producing low-magnitude, quantization-friendly residuals. Building on this, QVG introduces Progressive Residual Quantization, a coarse-to-fine multi-stage scheme that further reduces quantization error while enabling a smooth quality–memory trade-off. Across LongCat-Video, HY-WorldPlay, and Self-Forcing, QVG establishes a new Pareto frontier between quality and memory efficiency, reducing KV memory by up to 7.0× with less than 4% end-to-end latency overhead, while delivering significantly better generation quality than existing baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on16
- Diffusion Forcing: Next-token Prediction Meets Full-Sequence DiffusionBoyuan Chen, Diego Marti Monso, Yilun Du, Max Simchowitz et al.NeurIPS 2024 · 751 citations
- Self Forcing: Bridging the Train-Test Gap in Autoregressive Video DiffusionXun Huang, Zhengqi Li, Guande He, Mingyuan Zhou et al.NeurIPS 2025 · 628 citations
- FreeNoise: Tuning-Free Longer Video Diffusion via Noise ReschedulingHaonan Qiu, Menghan Xia, Yong Zhang, Yingqing He et al.ICLR 2024 · 171 citations
- WorldMem: Long-term Consistent World Simulation with MemoryZeqi Xiao, Yushi Lan, Yifan Zhou, Wenqi Ouyang et al.NeurIPS 2025 · 165 citations
- Frame Context Packing and Drift Prevention in Next-Frame-Prediction Video Diffusion ModelsLvmin Zhang, Shengqu Cai, Muyang Li, Gordon Wetzstein et al.NeurIPS 2025 · 132 citations
Related papers
- FAST-AR: Fast Autoregressive Video Diffusion and World Models with Temporal Cache Compression and Sparse AttentionDvir Samuel, Issar Tzachor, Matan Levy, Michael Green et al.ICML 2026 · 7 citations
- Accelerating Autoregressive Video Diffusion via History-Guided Cache and Residual CorrectionKepan Nan, Wangbo Zhao, Penghao Zhou, Jun Li et al.CVPR 2026
- QVGen: Pushing the Limit of Quantized Video Generative ModelsYushi Huang, Ruihao Gong, Jing Liu, Yifu Ding et al.ICLR 2026 · 22 citations
- Deep Forcing: Training-Free Long Video Generation with Deep Sink and Participative CompressionJung Yi, Wooseok Jang, Paul Cho, Jisu Nam et al.ICML 2026
- Flow Caching for Autoregressive Video GenerationYuexiao Ma, Xuzhe Zheng, Jing Xu, Xiwei Xu et al.ICLR 2026 · 20 citations
