PipeDiT: Accelerating Diffusion Transformers in Video Generation with Task Pipelining and Model Decoupling
Sijie Wang, Qiang Wang, Shaohuai Shi
Abstract
Video generation has been advancing rapidly, and diffusion transformer (DiT) based models have demonstrated remarkable capabilities. However, their practical deployment is often hindered by slow inference speeds and high memory consumption. In this paper, we propose a novel pipelining framework named PipeDiT to accelerate video generation, which is equipped with three main innovations. First, we design a pipelining algorithm (PipeSP) for sequence parallelism (SP) to enable the computation of latent generation and communication among multiple GPUs to be pipelined, thus reducing inference latency. Second, we propose DeDiVAE to decouple the diffusion module and the variational autoencoder (VAE) module into two GPU groups, the executions of which can also be pipelined to reduce memory consumption and inference latency. Third, to better utilize the GPU resources in the VAE group, we propose an attention co-processing (Aco) method to further reduce the overall video generation latency. We integrate our PipeDiT into both OpenSoraPlan and Hun-yuanVideo, two state-of-the-art open-source video generation frameworks, and conduct extensive experiments on two 8-GPU systems. Experimental results show that, under many common resolution and timestep configurations, our PipeDiT achieves 1.06× to 4.02× speedups over OpenSoraPlan and HunyuanVideo.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4ac51a7c-3f2a-414a-b253-f5e892a4441bBuilds on8
- SnapFusion: Text-to-Image Diffusion Model on Mobile Devices within Two SecondsYanyu Li, Huan Wang, Qing Jin, Ju Hu et al.NeurIPS 2023 · 300 citations
- Parallel Sampling of Diffusion ModelsAndy Shih, Suneel Belkhale, Stefano Ermon, Dorsa Sadigh et al.NeurIPS 2023 · 144 citations
- DeepCache: Accelerating Diffusion Models for FreeXinyin Ma, Gongfan Fang, Xinchao WangCVPR 2024 · 87 citations
- LoongServe: Efficiently Serving Long-Context Large Language Models with Elastic Sequence ParallelismBingyang Wu, Shengyu Liu, Yinmin Zhong, Peng Sun et al.SOSP 2024 · 32 citations
- FlexSP: Accelerating Large Language Model Training via Flexible Sequence ParallelismYujie Wang, Shiju Wang, Shenhan Zhu, Fangcheng Fu et al.ASPLOS 2025 · 8 citations
Related papers
- DSA: Efficient Inference For Video Generation Models via Distributed Sparse AttentionShenggui Li, Runyu Lu, qiaoling chen, Haiyan Yin et al.ICLR 2026
- Minute-Long Videos with Dual ParallelismsZeqing Wang, Bowen Zheng, Xingyi Yang, Zhenxiong Tan et al.AAAI 2026
- Astraea: A Token-wise Acceleration Framework for Video Diffusion TransformersHaosong Liu, Yuge Cheng, Wenxuan Miao, Zihan Liu et al.ICLR 2026 · 9 citations
- Sparse-vDiT: Unleashing the Power of Sparse Attention to Accelerate Video Diffusion TransformersPengtao Chen, Xianfang Zeng, Maosen Zhao, Mingzhu Shen et al.AAAI 2026
- SLA: Beyond Sparsity in Diffusion Transformers via Fine-Tunable Sparse–Linear AttentionJintao Zhang, Haoxu Wang, Kai Jiang, Shuo Yang et al.ICLR 2026 · 57 citations
