Minute-Long Videos with Dual Parallelisms
Zeqing Wang, Bowen Zheng, Xingyi Yang, Zhenxiong Tan, Yuecong Xu, Xinchao Wang
摘要
Diffusion Transformer (DiT)-based video diffusion models generate high-quality videos at scale but incur prohibitive processing latency and memory costs for long videos. To address this, we propose a novel distributed inference strategy, termed DualParal. The core idea is that, instead of generating an entire video on a single GPU, we parallelize both temporal frames and model layers across GPUs. However, a naive implementation of this division faces a key limitation: since diffusion models require synchronized noise levels across frames, this implementation leads to the serialization of original parallelisms. We leverage a block-wise denoising scheme to handle this. Namely, we process a sequence of frame blocks through the pipeline with progressively decreasing noise levels. Each GPU handles a specific block and layer subset while passing previous results to the next GPU, enabling asynchronous computation and communication. To further optimize performance, we incorporate two key enhancements. Firstly, a feature cache is implemented on each GPU to store and reuse features from the prior block as context, minimizing inter-GPU communication and redundant computation. Secondly, we employ a coordinated noise initialization strategy, ensuring globally consistent temporal dynamics by sharing initial noise patterns across GPUs without extra resource costs. Together, these enable fast, artifact-free, and infinitely long video generation. Applied to the latest diffusion transformer video generator, our method efficiently produces 1,025-frame videos with up to 6.54× lower latency and 1.48× lower memory cost on 8×RTX 4090 GPUs. * Corresponding Author Preprint. Under review. hidden [11, 17] or input [25, 12] sequences using a full model replica on each device. However, they incurs high memory overhead due to the entire model on every device [25, 12] . In contrast, pipeline parallelism [6] mitigates memory usage by partitioning the model across devices as a device pipeline [9, 16, 24] . Therefore, an ideal solution would combine the sequence parallelism with pipeline parallelism to maximize speed and minimize memory usage. However, naively combining sequence and pipeline parallelism is fundamentally conflicting. The core issue stems from the inherent synchronization property of video diffusion models: all input tokens must pass through an entire layer together before any can move on. In pipeline parallelism, this means the full input must finish processing on one device (e.g., Device 1) before passing to the next (e.g., Device 2). This requirement directly contradicts sequence parallelism, which splits the input across devices. As a result, all distributed parts must be gathered back onto a single device for serialized processing on specific model layers. Only then can all parts enter the next pipeline stage, i.e. next device. This repeated gathering serializes computation and negates the benefits of sequence parallelism, reintroducing a serial bottleneck and significant communication overhead. To address this conflict, we propose a novel distributed inference strategy, termed DualParal. At a high level, DualParal divides both the video sequence and model into chunks and applies parallel processing across both.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper16
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- FreeNoise: Tuning-Free Longer Video Diffusion via Noise ReschedulingHaonan Qiu, Menghan Xia, Yong Zhang, Yingqing He 等ICLR 2024 · 被引用 171 次
- TeraPipe: Token-Level Pipeline Parallelism for Training Large-Scale Language ModelsZhuohan Li, Siyuan Zhuang, Shiyuan Guo, Danyang Zhuo 等ICML 2021 · 被引用 160 次
- FIFO-Diffusion: Generating Infinite Videos from Text without TrainingJihwan Kim, Junoh Kang, Jinyoung Choi, Bohyung HanNeurIPS 2024 · 被引用 156 次
相关 Paper
- DSA: Efficient Inference For Video Generation Models via Distributed Sparse AttentionShenggui Li, Runyu Lu, qiaoling chen, Haiyan Yin 等ICLR 2026
- PipeDiT: Accelerating Diffusion Transformers in Video Generation with Task Pipelining and Model DecouplingSijie Wang, Qiang Wang, Shaohuai ShiAAAI 2026
- Accelerating Diffusion via Hybrid Data-Pipeline Parallelism Based on Conditional Guidance SchedulingEuisoo Jung, Byunghyun Kim, Hyunjin Kim, Seonghye Cho 等CVPR 2026
- PipeFusion: Patch-level Pipeline Parallelism for Diffusion Transformers InferenceJiarui Fang, Jinzhe Pan, Aoyu Li, Xibo Sun 等NeurIPS 2025 · 被引用 36 次
- AsyncDiff: Parallelizing Diffusion Models by Asynchronous DenoisingZigeng Chen, Xinyin Ma, Gongfan Fang, Zhenxiong Tan 等NeurIPS 2024 · 被引用 33 次
