DCP: Addressing Input Dynamism In Long-Context Training via Dynamic Context Parallelism
Chenyu Jiang, Zhenkun Cai, Ye Tian, Zhen Jia, Yida Wang, Chuan Wu
摘要
Context parallelism has emerged as a key technique to support long-context training, a growing trend in generative AI for modern large models. However, existing context parallel methods rely on static parallelization configurations that overlook the dynamic nature of training data, specifically, the variability in sequence lengths and token relationships (i.e., attention patterns) across samples. As a result, these methods often suffer from unnecessary communication overhead and imbalanced computation. In this paper, we present DCP, a dynamic context parallel training framework that introduces fine-grained blockwise partitioning of both data and computation. By enabling flexible mapping of data and computation blocks to devices, DCP can adapt to varying sequence characteristics, effectively reducing communication and improving memory and computation balance. Micro-benchmarks demonstrate that DCP accelerates attention by 1.19x 2.45x under causal masks and 2.15x 3.77x under sparse attention patterns. Additionally, we observe up to 0.94x 1.16x end-to-end training speed-up for causal masks, and 1.00x 1.46x for sparse masks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Hetu v2: A General and Scalable Deep Learning System with Hierarchical and Heterogeneous Single Program Multiple Data AnnotationsHaoyang Li, Fangcheng Fu, Hao Ge, Sheng Lin 等OSDI 2026 · 被引用 7 次
- SAS: Sparse Attention Synthesizer for Efficient Language Model InferenceYuan Zhou, Shaojie Xiang, Lingfan Yu, Zhenyu Song 等EuroSys 2026
它引用的顶会 Paper11
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
相关 Paper
- UltraAttn: Efficiently Parallelizing Attention through Hierarchical Context-TilingHaoyu Yang, Zan Zong, Yuyang Jin, Kinman Lei 等SC 2025 · 被引用 1 次
- Long-Context Attention Benchmark: From Kernel Efficiency to Distributed Context ParallelismTao Bu, Qiangang Wang, Bowen Zeng, Hanwen Sun 等ICLR 2026
- Zeppelin: Balancing Variable-length Workloads in Data Parallel Large Model TrainingChang Chen, Tiancheng Chen, Jiangfei Duan, Qianchao Zhu 等EuroSys 2026
- ByteScale: Communication-Efficient Scaling of LLM Training with a 2048K Context Length on 16384 GPUsHao Ge, Junda Feng, Qi Huang, Fangcheng Fu 等SIGCOMM 2025 · 被引用 4 次
- Enabling Parallelism Hot Switching for Efficient Training of Large Language ModelsHao Ge, Fangcheng Fu, Haoyang Li, Xuanyu Wang 等SOSP 2024 · 被引用 4 次
