Dynamic Sparsity in Large-Scale Video DiT Training
Xin Tan, Yuetao Chen, Yimin Jiang, Xing Chen, Kun Yan, Nan Duan, Yibo Zhu, Daxin Jiang, Hong Xu
摘要
Diffusion Transformers (DiTs) have shown remarkable performance in generating high-quality videos. However, the quadratic complexity of 3D full attention remains a bottleneck in scaling DiT training, especially with high-definition, lengthy videos, where it can consume up to 95% of processing time and demand specialized context parallelism.
This paper introduces DSV to accelerate video DiT training by leveraging the dynamic attention sparsity we empirically observe. DSV uses a two-stage algorithm to capture the dynamic sparsity patterns via low-rank based approximation of the original query and key. It employs custom kernels to efficiently identify critical key-value pairs and compute the sparse attention. To accommodate the new sparsity dimension, DSV adopts a hybrid sparsity-aware context parallelism that re-balances the skewed workload across attention heads and blocks due to sparsity heterogeneity. DSV achieves up to 3.02× higher training throughput, scaling to 128 GPUs and 520k token lengths, without quality loss.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Radial Attention: 𝒪(n log n) Sparse Attention with Energy Decay for Long Video GenerationXingyang Li, Muyang Li, Tianle Cai, Haocheng Xi 等NeurIPS 2025 · 被引用 66 次
- Hierarchical Fine-grained Preference Optimization for Physically Plausible Video GenerationHarold Haodong Chen, Haojian Huang, Qifeng Chen, Harry Yang 等NeurIPS 2025 · 被引用 25 次
- Training-Free and Adaptive Sparse Attention for Efficient Long Video GenerationYifei Xia, Suhan Ling, Fangcheng Fu, Yujie Wang 等ICCV 2025 · 被引用 6 次
- AdaCluster: Adaptive Query-Key Clustering for Sparse Attention in Video GenerationHaoyue Tan, Shengnan Wang, Yulin Qiao, Juncheng Zhang 等CVPR 2026 · 被引用 5 次
- VecAttention: Vector-wise Sparse Attention for Accelerating Long Context InferenceAnmin Liu, Ruixuan Yang, Huiqiang Jiang, Bin Lin 等CVPR 2026 · 被引用 4 次
它引用的顶会 Paper25
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
相关 Paper
- Sparse Video-Gen: Accelerating Video Diffusion Transformers with Spatial-Temporal SparsityHaocheng Xi, Shuo Yang, Yilong Zhao, Chenfeng Xu 等ICML 2025
- DFSAttn: Dynamic Fine-grained Sparse Attention for Efficient Video GenerationJie Hu, Zixiang Gao, Yutong He, Kun YuanICML 2026
- SLA: Beyond Sparsity in Diffusion Transformers via Fine-Tunable Sparse–Linear AttentionJintao Zhang, Haoxu Wang, Kai Jiang, Shuo Yang 等ICLR 2026 · 被引用 57 次
- DSA: Efficient Inference For Video Generation Models via Distributed Sparse AttentionShenggui Li, Runyu Lu, qiaoling chen, Haiyan Yin 等ICLR 2026
- Mixture of Distributions Matters: Dynamic Sparse Attention for Efficient Video Diffusion TransformersYuxi Liu, Yipeng Hu, Zekun Zhang, Kunze Jiang 等ICML 2026 · 被引用 5 次
