DSA: Efficient Inference For Video Generation Models via Distributed Sparse Attention
Shenggui Li, Runyu Lu, qiaoling chen, Haiyan Yin, Yueming Lyu, Yonggang Wen, Ivor Tsang, Tianwei Zhang
Abstract
Diffusion Transformer models have driven the rapid advances in video generation, achieving state-of-the-art quality and flexibility. However, their attention mechanism remains a major performance bottleneck, as its dense computation scales quadratically with the sequence length. To overcome this limitation and reduce the generation latency, we propose DSA, a novel attention mechanism that integrates sparse attention with distributed inference for diffusion-based video generation. By leveraging carefully-designed parallelism strategies and scheduling, DSA significantly reduces redundant computation while preserving global context. Extensive experiments on benchmark datasets demonstrate that, when deployed on 8 GPUs, DSA achieves up to 1.43× inference speedup than the existing distributed method and 10.79× faster than single-GPU inference.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aa732c1d-3345-4e8c-9117-57e20e58adf4Builds on27
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- Minute-Long Videos with Dual ParallelismsZeqing Wang, Bowen Zheng, Xingyi Yang, Zhenxiong Tan et al.AAAI 2026
- PipeDiT: Accelerating Diffusion Transformers in Video Generation with Task Pipelining and Model DecouplingSijie Wang, Qiang Wang, Shaohuai ShiAAAI 2026
- Astraea: A Token-wise Acceleration Framework for Video Diffusion TransformersHaosong Liu, Yuge Cheng, Wenxuan Miao, Zihan Liu et al.ICLR 2026 · 9 citations
- SLA: Beyond Sparsity in Diffusion Transformers via Fine-Tunable Sparse–Linear AttentionJintao Zhang, Haoxu Wang, Kai Jiang, Shuo Yang et al.ICLR 2026 · 57 citations
- TimeRipples: Accelerating vDiTs by Understanding the Spatio-Temporal Correlations in Latent SpaceWenxuan Mao, Yulin Sun, Aiyue Chen, Jing Lin et al.CVPR 2026
