Efficient Long Context Fine-tuning with Chunk Flow
Xiulong Yuan, Hongtao Xu, Wenting Shen, Ang Wang, Xiafei Qiu, Jie Zhang, Yuqiong Liu, Bowen Yu, Junyang Lin, Mingzhen Li, Weile Jia, Yong Li, Wei Lin
Abstract
Long context fine-tuning of large language models(LLMs) involves training on datasets that are predominantly composed of short sequences and a small proportion of longer sequences. However, existing approaches overlook this long-tail distribution and employ training strategies designed specifically for long sequences. Moreover, these approaches also fail to address the challenges posed by variable sequence lengths during distributed training, such as load imbalance in data parallelism and severe pipeline bubbles in pipeline parallelism. These issues lead to suboptimal training performance and poor GPU resource utilization. To tackle these problems, we propose a chunk-centric training method named ChunkFlow. ChunkFlow reorganizes input sequences into uniformly sized chunks by consolidating short sequences and splitting longer ones. This approach achieves optimal computational efficiency and balance among training inputs. Additionally, Chunk-Flow incorporates a state-aware chunk scheduling mechanism to ensure that the peak memory usage during training is primarily determined by the chunk size rather than the maximum sequence length in the dataset. Integrating this scheduling mechanism with existing pipeline scheduling algorithms further enhances the performance of distributed training. Experimental results demonstrate that, compared with Megatron-LM, Chunk-Flow can be up to 4.53x faster in the long context fine-tuning of LLMs. Furthermore, we believe that ChunkFlow serves as an effective solution for a broader range of scenarios, such as long context continual pre-training, where datasets contain variable-length sequences.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9926888d-888f-474b-9510-d52cb036bd90Cited by top-tier papers4
- Future-Gain Guided Test-Time Learning for Large Language ModelsLangYu Bian, Jinwu Hu, Zitian Zhang, Dongjin Yang et al.ICML 2026 · 12 citations
- Skrull: Towards Efficient Long Context Fine-tuning through Dynamic Data SchedulingHongtao Xu, Wenting Shen, Yuanxin Wei, Ang Wang et al.NeurIPS 2025 · 3 citations
- AReaL-DTA: Dynamic Tree Attention for Efficient Reinforcement Learning of Large Language ModelsJiarui Zhang, Yuchen Yang, Ran Yan, Zhiyu Mei et al.ICML 2026
- AdaHC: Accelerating Multi-Token Prediction with Adaptive Head Chunking with Pipeline ParallelismYan Wang, Chang Si, Kaiming Yang, Zhipeng Zhang et al.ICML 2026
Builds on10
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
- Efficient large-scale language model training on GPU clusters using megatron-LMDeepak Narayanan, Mohammad Shoeybi, Jared Casper, Patrick LeGresley et al.SC 2021 · 576 citations
- LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation DatasetLianmin Zheng, Wei-Lin Chiang, Ying Sheng, Tianle Li et al.ICLR 2024 · 419 citations
- DAPPLE: a pipelined data parallel approach for training large modelsShiqing Fan, Yi Rong, Chen Meng, Zongyan Cao et al.PPoPP 2021 · 224 citations
- Data Engineering for Scaling Language Models to 128K ContextYao Fu, Rameswar Panda, Xinyao Niu, Xiang Yue et al.ICML 2024 · 204 citations
Related papers
- HelixPipe: Efficient Distributed Training of Long Sequence Transformers with Attention Parallel Pipeline ParallelismGeng Zhang, Shenggan Cheng, Xuanlei Zhao, Ziming Liu et al.PPoPP 2026 · 3 citations
- Dataset Decomposition: Faster LLM Training with Variable Sequence Length CurriculumHadi Pouransari, Chun-Liang Li, Jen-Hao Rick Chang, Pavan Kumar Anasosalu Vasu et al.NeurIPS 2024 · 40 citations
- SlimPipe: Memory-Thrifty and Efficient Pipeline Parallelism for Long-Context LLM TrainingZhouyang Li, Yuliang Liu, Wei Zhang, Tailing Yuan et al.SC 2025 · 5 citations
- Memorize Step by Step: Efficient Long-Context Prefilling with Incremental Memory and Decremental ChunkZhiyuan Zeng, Qipeng Guo, Xiaoran Liu, Zhangyue Yin et al.EMNLP 2024
- MEMO: Fine-grained Tensor Management For Ultra-long Context LLM TrainingPinxue Zhao, Hailin Zhang, Fangcheng Fu, Xiaonan Nie et al.SIGMOD 2025 · 4 citations
