Video-SVD: Efficient Video Diffusion via Orthogonal Basis Composition
Zhang Wan, Yu Li, Tianze Huang, Haochen Li, Juan Cao, Sheng Tang
摘要
Video Diffusion Transformers (VDiTs) represent the state of the art in video generation but remain constrained by the quadratic complexity of dense self-attention. To address this attention bottleneck, we analyze the pre-softmax matrix () and reveal two key properties: (1) video attention exhibits an effective low-dimensional structure with rapid singular-value decay, and (2) real motion induces hybrid spatio-temporal patterns rather than rigid ``spatial vs. temporal'' layouts. Guided by these observations, we propose Video-SVD, a training-free and plug-and-play acceleration method that does not modify the original network parameters. Video-SVD learns checkpoint-adaptive orthogonal bases offline and, at inference time, replaces expensive dense attention computation with lightweight online subspace projection and basis composition. To preserve high fidelity, Video-SVD further employs layer-shared dual-stream residual modules to recover fine-grained content details and positional information. Across HunyuanVideo and Wan2.1 backbones, Video-SVD achieves significant end-to-end speedups while maintaining high visual quality, reaching 1.92 on HunyuanVideo, 1.75 on Wan2.1-1.3B, and 1.79 on Wan2.1-14B.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Consistency ModelsYang Song, Prafulla Dhariwal, Mark Chen, Ilya SutskeverICML 2023 · 被引用 1,720 次
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar 等ICLR 2021 · 被引用 1,270 次
- Sparse VideoGen2: Accelerate Video Generation with Sparse Attention via Semantic-Aware PermutationShuo Yang, Haocheng Xi, Yilong Zhao, Muyang Li 等NeurIPS 2025 · 被引用 114 次
- DeepCache: Accelerating Diffusion Models for FreeXinyin Ma, Gongfan Fang, Xinchao WangCVPR 2024 · 被引用 87 次
相关 Paper
- Sparse Video-Gen: Accelerating Video Diffusion Transformers with Spatial-Temporal SparsityHaocheng Xi, Shuo Yang, Yilong Zhao, Chenfeng Xu 等ICML 2025
- Attention Sparsity is Input-Stable: Training-Free Sparse Attention for Video Generation via Offline Sparsity Profiling and Online QK Co-ClusteringJiayi Luo, Jiayu Chen, Jiankun Wang, Cong Wang 等ICML 2026 · 被引用 5 次
- Faster Video Diffusion with Trainable Sparse AttentionPeiyuan Zhang, Yongqi Chen, Haofeng Huang, Will Lin 等NeurIPS 2025 · 被引用 6 次
- Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion TransformersPengtao Chen, Xianfang Zeng, Maosen Zhao, Mingzhu Shen 等AAAI 2026
- VMoBA: Mixture-of-Block Attention for Video Diffusion ModelsJianzong Wu, Liang Hou, Haotian Yang, Ye Tian 等ICLR 2026 · 被引用 36 次
