Skrull: Towards Efficient Long Context Fine-tuning through Dynamic Data Scheduling
Hongtao Xu, Wenting Shen, Yuanxin Wei, Ang Wang, Guo Runfan, Tianxing Wang, Yong Li, Mingzhen Li, Weile Jia
摘要
Long-context supervised fine-tuning (Long-SFT) plays a vital role in enhancing the performance of large language models (LLMs) on long-context tasks. To smoothly adapt LLMs to long-context scenarios, this process typically entails training on mixed datasets containing both long and short sequences. However, this heterogeneous sequence length distribution poses significant challenges for existing training systems, as they fail to simultaneously achieve high training efficiency for both long and short sequences, resulting in sub-optimal end-to-end system performance in Long-SFT. In this paper, we present a novel perspective on data scheduling to address the challenges posed by the heterogeneous data distributions in Long-SFT. We propose Skrull, a dynamic data scheduler specifically designed for efficient long-SFT. Through dynamic data scheduling, Skrull balances the computation requirements of long and short sequences, improving overall training efficiency. Furthermore, we formulate the scheduling process as a joint optimization problem and thoroughly analyze the trade-offs involved. Based on those analysis, Skrull employs a lightweight scheduling algorithm to achieve near-zero cost online scheduling in Long-SFT. Finally, we implement Skrull upon DeepSpeed, a stateof-the-art distributed training system for LLMs. Experimental results demonstrate that Skrull outperforms DeepSpeed by 3.76x on average (up to 7.54x) in real-world long-SFT scenarios.
However, this heterogeneous data distribution in Long-SFT poses significant challenges for existing distributed LLM training frameworks [14,20,15], exhibiting sub-optimal efficiency. For instance, the heterogeneous data distribution poses a dilemma for parallelism and memory-reduction strategies. Specifically, long sequences necessitate context parallelism and other memory-reduction approaches 39th Conference on Neural Information Processing Systems (NeurIPS 2025). due to their tremendous memory requirements. However, those approaches compromise the training efficiency for short ones due to the overheads like unnecessary communication and GPU underutilization. Moreover, the wide sequence length distribution in long-SFT worsen the mismatch of computation characteristics in Attention module, which exhibit quadratic computational complexity and linear memory consumption [7,6], leading to another dilemma for load balance problem.
To tackle the above challenges, we propose Skrull, a dynamic data scheduler dedicated for Long-SFT scenarios. Skrull efficiently handle the unique data distributions in Long-SFT scenario through two main components: Distributed-Aware Context Parallelism (DACP) and Global Data Scheduling (GDS). DACP selectively shards sequences and schedules them across different workers to minimize the performance degradation while maintains the ability of handling long sequence. GDS enlarge the scope of scheduling and improve the GPU utilization during training. The two components collaborate with each other at different scheduling granularities. Furthermore, to achieve the optimal performance, we formulate the scheduling process as a joint optimization problem and design a lightweight heuristic algorithm to solve it at runtime. Experimental results demonstrate that Skrull improves the end-to-end training performance by 3.76x on average (up to 7.54x) compared to DeepSpeed, a state-of-the-art distributed LLM training framework.
Our key contributions are summarized as follows:
• We provide a new perspective of data scheduling to address the heterogeneous sequence length distribution.
• We propose a new context parallelism called DACP based on fine-grained data scheduling, which maintaining both the processing capabilities for long sequences and efficiency for short sequences, enabling efficient training on heterogeneous data distribution in long-SFT scenario.
• We implement coarse-grained global data scheduling (GDS) and further formulate GDS and DACP as a joint optimization problem through performance modeling.
• We design a lightweight heuristic algorithm and achieve performance gains by 3.76x on average (with a peak improvement of 7.54×) in real-world datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra 等NeurIPS 2022 · 被引用 5,493 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 被引用 852 次
- LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation DatasetLianmin Zheng, Wei-Lin Chiang, Ying Sheng, Tianle Li 等ICLR 2024 · 被引用 419 次
相关 Paper
- Efficient Long Context Fine-tuning with Chunk FlowXiulong Yuan, Hongtao Xu, Wenting Shen, Ang Wang 等ICML 2025
- FlexSP: Accelerating Large Language Model Training via Flexible Sequence ParallelismYujie Wang, Shiju Wang, Shenhan Zhu, Fangcheng Fu 等ASPLOS 2025 · 被引用 8 次
- ByteScale: Communication-Efficient Scaling of LLM Training with a 2048K Context Length on 16384 GPUsHao Ge, Junda Feng, Qi Huang, Fangcheng Fu 等SIGCOMM 2025 · 被引用 4 次
- Dataset Decomposition: Faster LLM Training with Variable Sequence Length CurriculumHadi Pouransari, Chun-Liang Li, Jen-Hao Rick Chang, Pavan Kumar Anasosalu Vasu 等NeurIPS 2024 · 被引用 40 次
- Zeppelin: Balancing Variable-length Workloads in Data Parallel Large Model TrainingChang Chen, Tiancheng Chen, Jiangfei Duan, Qianchao Zhu 等EuroSys 2026
