AutoSP: Unlocking Long-Context LLM Training Via Compiler-Based Sequence Parallelism
Ahan Gupta, Zhihao Wang, Neel Dani, Masahiro Tanaka, Olatunji Ruwase, Minjia Zhang
Abstract
Large-language-models (LLMs) demonstrate enormous utility in long-context tasks which require processing prompts that consist of tens to hundreds of thousands of tokens. However, existing LLM training libraries do not provide easy to use abstractions to optimize for long-context training, instead focusing on optimizations for models with large parameter counts through ZeRO-3/FSDP, Tensor and Pipeline parallelism. This forces users to rewrite LLM training libraries to incorporate compositions of various complex long-context optimizations, such as sequence-parallelism, to training pipelines; a process that requires in-depth expertise, reducing developer productivity. To tackle these challenges, we introduce AutoSP: the first automated solution to automatically optimize LLM training for longer-contexts. AutoSP compiles models and applies a targeted set of optimizations: automated sequence parallelism, and long-context aware activation-checkpointing, to drastically enhance LLM trainability at negligible cost to throughput. Our evaluation demonstrates AutoSP's capability on both NVIDIA and AMD hardware, increasing training contexts by upto 2.7 and 2.5 respectively over competitive hand-written baseline at negligible cost to runtime performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ea356fbe-4b79-4422-a358-2a8f26c875a6Builds on9
- GShard: Scaling Giant Models with Conditional Computation and Automatic ShardingDmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen et al.ICLR 2021 · 1,954 citations
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
- PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph CompilationJason Ansel, Edward Z. Yang, Horace He, Natalia Gimelshein et al.ASPLOS 2024 · 693 citations
- Scalable Multi-Hop Relational Reasoning for Knowledge-Aware Question AnsweringYanlin Feng, Xinyue Chen, Bill Yuchen Lin, Peifeng Wang et al.EMNLP 2020 · 207 citations
- DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language ModelsDamai Dai, Chengqi Deng, Chenggang Zhao, R. X. Xu et al.ACL 2024 · 171 citations
Related papers
- Long-Context Attention Benchmark: From Kernel Efficiency to Distributed Context ParallelismTao Bu, Qiangang Wang, Bowen Zeng, Hanwen Sun et al.ICLR 2026
- BurstEngine: An efficient distributed framework for training transformers On extremely Long sequences of over 1M tokensAo Sun, Weilin Zhao, Xu Han, Cheng Yang et al.SC 2025 · 1 citation
- Mario: Near Zero-cost Activation Checkpointing in Pipeline ParallelismWeijian Liu, Mingzhen Li, Guangming Tan, Weile JiaPPoPP 2025 · 4 citations
- Adapting Language Models to Compress ContextsAlexis Chevalier, Alexander Wettig, Anirudh Ajith, Danqi ChenEMNLP 2023 · 34 citations
- Out of the Memory Barrier: A Highly Memory-Efficient Training System for LLMs with Million-Token ContextsWenhao Li, Daohai Yu, Gen Luo, Yuxin Zhang et al.ICLR 2026 · 5 citations
