FlexSP: Accelerating Large Language Model Training via Flexible Sequence Parallelism
Yujie Wang, Shiju Wang, Shenhan Zhu, Fangcheng Fu, Xinyi Liu, Xuefeng Xiao, Huixia Li, Jiashi Li, Faming Wu, Bin Cui
Abstract
Extending the context length (i.e., the maximum supported sequence length) of LLMs is of paramount significance. To facilitate long context training of LLMs, sequence parallelism has emerged as an essential technique, which scatters each input sequence across multiple devices and necessitates communication to process the sequence. In essence, existing sequence parallelism methods assume homogeneous sequence lengths (i.e., all input sequences are equal in length) and therefore leverages a single, static scattering strategy for all input sequences. However, in reality, the sequence lengths in LLM training corpora exhibit substantial variability, often following a long-tail distribution, which leads to workload heterogeneity.
In this paper, we show that employing a single, static strategy results in inefficiency and resource under-utilization, highlighting the need for adaptive approaches to handle the heterogeneous workloads across sequences. To address this, we propose a heterogeneity-adaptive sequence parallelism method. For each training step, our approach captures the variability in sequence lengths and assigns the optimal combination of scattering strategies based on workload characteristics. We model this problem as a linear programming optimization and design an efficient and effective solver to find the optimal solution. Furthermore, we implement our
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers9
- Hydraulis: Balancing Large Transformer Model Training via Co-designing Parallel Strategies and Data AssignmentHaoyang Li, Fangcheng Fu, Sheng Lin, Hao Ge et al.SIGMOD 2026 · 7 citations
- Hetu v2: A General and Scalable Deep Learning System with Hierarchical and Heterogeneous Single Program Multiple Data AnnotationsHaoyang Li, Fangcheng Fu, Hao Ge, Sheng Lin et al.OSDI 2026 · 7 citations
- Malleus: Straggler-Resilient Hybrid Parallel Training of Large-scale Models via Malleable Data and Model ParallelizationHaoyang Li, Fangcheng Fu, Hao Ge, Sheng Lin et al.SIGMOD 2025 · 6 citations
- ResiHP: Taming LLM Training Failures with Dynamic Hybrid ParallelismTenghui Ma, Jihu Guo, Wei Gao, Sitian Lu et al.HPDC 2026
- PipeDiT: Accelerating Diffusion Transformers in Video Generation with Task Pipelining and Model DecouplingSijie Wang, Qiang Wang, Shaohuai ShiAAAI 2026
Builds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie et al.NeurIPS 2020 · 3,159 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
Related papers
- Skrull: Towards Efficient Long Context Fine-tuning through Dynamic Data SchedulingHongtao Xu, Wenting Shen, Yuanxin Wei, Ang Wang et al.NeurIPS 2025 · 3 citations
- Enabling Parallelism Hot Switching for Efficient Training of Large Language ModelsHao Ge, Fangcheng Fu, Haoyang Li, Xuanyu Wang et al.SOSP 2024 · 4 citations
- Efficient Long Context Fine-tuning with Chunk FlowXiulong Yuan, Hongtao Xu, Wenting Shen, Ang Wang et al.ICML 2025
- ByteScale: Communication-Efficient Scaling of LLM Training with a 2048K Context Length on 16384 GPUsHao Ge, Junda Feng, Qi Huang, Fangcheng Fu et al.SIGCOMM 2025 · 4 citations
- ParaDySe: A Parallel Strategy Switching Framework for Dynamic Sequences in Transformer-based Large Language ModelsZhixin Ou, Peng Liang, Linbo Qiao, Jianchen Han et al.AAAI 2026
