Lune

CVPR2026顶会

DeltaQuant: 4-bit Video Diffusion Models with Spatiotemporal Delta Smoothing

Xingyang Li, Samuel Tesfai, Zhekai Zhang, Haocheng Xi, Shuo Yang, Lvmin Zhang, Yufei Sun, Kelly Peng, Maneesh Agrawala, Ion Stoica, Kurt Keutzer, Jun-Yan Zhu

出版方
2026年份
7被引次数

摘要

Prompt: The camera follows a white vintage SUV with a black roof rack speeding up a steep dirt road surrounded by redwoods on a mountain slope. Dust kicks up as the sunlight creates a warm glow, emphasizing the rugged, serene landscape. * Equal contribution tokens) to 4 bits while keeping core tokens in FP8. This decomposition substantially reduces quantization error with minimal overhead. For weight quantization, DeltaQuant incorporates SVDQuant's low-rank decomposition to further reduce quantization error. We also implement an efficient kernel that translates DeltaQuant's computational benefits into real-world speedups. Extensive experiments on Wan2.2 I2V, Wan2.2 T2V, and LTX-Video T2V demonstrate that DeltaQuant maintains high generation fidelity. On Wan2.2, it compresses model size by 2.9× and reduces memory footprint by 2.3×. DeltaQuant is compatible with efficient attention mechanisms and few-step distillation. When integrated with these techniques, it achieves an additional 3.0× acceleration, for a total 111.8× end-to-end speedup.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper31

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖