Untied Ulysses: Memory-Efficient Context Parallelism via Headwise Chunking
Ravi Ghadia, Maksim Abraham, Sergei Vorobyov, Max Ryabinin
摘要
Efficiently processing long sequences with Transformer models usually requires splitting the computations across accelerators via context parallelism. The dominant approaches in this family of methods, such as Ring Attention or Deep-Speed Ulysses, enable scaling over the context dimension but do not focus on memory efficiency, which limits the sequence lengths they can support. More advanced techniques, such as Fully Pipelined Distributed Transformer or activation offloading, can further extend the possible context length at the cost of training throughput. In this paper, we present UPipe, a simple yet effective context parallelism technique that performs fine-grained chunking at the attention head level. This technique significantly reduces the activation memory usage of self-attention, breaking the activation memory barrier and unlocking much longer context lengths. Our approach reduces intermediate tensor memory usage in the attention layer by as much as 87.5% for 32B Transformers, while matching previous context parallelism techniques in training speed. UPipe can support the context length of 5M tokens when training Llama3-8B on a single 8×H100 node, improving upon prior methods by over 25%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper4
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precisionJay Shah, Ganesh Bikshandi, Ying Zhang, Vijay Thakkar 等NeurIPS 2024 · 被引用 727 次
- Sequence Parallelism: Long Sequence Training from System PerspectiveShenggui Li, Fuzhao Xue, Chaitanya Baranwal, Yongbin Li 等ACL 2023 · 被引用 29 次
- M-LongDoc: A Benchmark For Multimodal Super-Long Document Understanding And A Retrieval-Aware Tuning FrameworkYew Ken Chia, Liying Cheng, Hou Pong Chan, Maojia Song 等EMNLP 2025 · 被引用 3 次
相关 Paper
- HelixPipe: Efficient Distributed Training of Long Sequence Transformers with Attention Parallel Pipeline ParallelismGeng Zhang, Shenggan Cheng, Xuanlei Zhao, Ziming Liu 等PPoPP 2026 · 被引用 3 次
- SlimPipe: Memory-Thrifty and Efficient Pipeline Parallelism for Long-Context LLM TrainingZhouyang Li, Yuliang Liu, Wei Zhang, Tailing Yuan 等SC 2025 · 被引用 5 次
- Mini-Sequence Transformers: Optimizing Intermediate Memory for Long Sequences TrainingCheng Luo, Jiawei Zhao, Zhuoming Chen, Beidi Chen 等NeurIPS 2024 · 被引用 6 次
- MEMO: Fine-grained Tensor Management For Ultra-long Context LLM TrainingPinxue Zhao, Hailin Zhang, Fangcheng Fu, Xiaonan Nie 等SIGMOD 2025 · 被引用 4 次
- RingX: Scalable Parallel Attention for Long-Context Learning on HPCJunqi Yin, Mijanur Palash, Mallikarjun Shankar, Feiyi WangSC 2025 · 被引用 1 次
