DSP: Dynamic Sequence Parallelism for Multi-Dimensional Transformers
Xuanlei Zhao, Shenggan Cheng, Chang Chen, Zangwei Zheng, Ziming Liu, Zheming Yang, Yang You
Abstract
Scaling multi-dimensional transformers to long sequences is important across various domains. The challenges of large memory requirements and slow speed of such sequences require sequence parallelism. All existing approaches fall under the category of embedded sequence parallelism, which are limited to shard along a single sequence dimension, thereby introducing significant communication overhead. However, multidimensional transformers involve independent calculation across multiple sequence dimensions. To this end, we propose Dynamic Sequence Parallelism (DSP) as a novel abstraction of sequence parallelism. DSP dynamically switches the parallel dimension according to the computation stage with efficient resharding strategy. DSP offers significant reductions in communication costs, adaptability across modules, and ease of use with minimal constraints. Experiments demonstrate DSP's superiority over state-of-the-art sequence parallelism methods by remarkable throughput improvements ranging from 32.2% to 10×, with at least 50% communication volume reduction.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- BWCache: Accelerating Video Diffusion Transformers through Block-Wise CachingHanshuai Cui, Zhiqing Tang, Zhifei Xu, Zhi Yao et al.ICLR 2026 · 11 citations
- LeMiCa: Lexicographic Minimax Path Caching for Efficient Diffusion-Based Video GenerationHuanlin Gao, Ping Chen, Fuyuan Shi, Chao Tan et al.NeurIPS 2025 · 9 citations
- MeanCache: From Instantaneous to Average Velocity for Accelerating Flow Matching InferenceHuanlin Gao, Ping Chen, Fuyuan Shi, Ruijia Wu et al.ICLR 2026 · 7 citations
- PreciseCache: Precise Feature Caching for Efficient and High-fidelity Video GenerationJiangshan Wang, Kang Zhao, Jiayi Guo, Jiayu Wang et al.ICLR 2026 · 6 citations
- Adaptive Caching for Faster Video Generation With Diffusion TransformersKumara Kahatapitiya, Haozhe Liu, Sen He, Ding Liu et al.ICCV 2025 · 5 citations
Builds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
Related papers
- Sequence Parallelism: Long Sequence Training from System PerspectiveShenggui Li, Fuzhao Xue, Chaitanya Baranwal, Yongbin Li et al.ACL 2023 · 29 citations
- ParaDySe: A Parallel Strategy Switching Framework for Dynamic Sequences in Transformer-based Large Language ModelsZhixin Ou, Peng Liang, Linbo Qiao, Jianchen Han et al.AAAI 2026
- StarTrail: Concentric Ring Sequence Parallelism for Efficient Near-Infinite-Context Transformer Model TrainingZiming Liu, Shaoyu Wang, Shenggan Cheng, Zhongkai Zhao et al.NeurIPS 2025 · 4 citations
- HelixPipe: Efficient Distributed Training of Long Sequence Transformers with Attention Parallel Pipeline ParallelismGeng Zhang, Shenggan Cheng, Xuanlei Zhao, Ziming Liu et al.PPoPP 2026 · 3 citations
- Untied Ulysses: Memory-Efficient Context Parallelism via Headwise ChunkingRavi Ghadia, Maksim Abraham, Sergei Vorobyov, Max RyabininICML 2026
