Lune

NeurIPS2025Top-tier venue

Skrull: Towards Efficient Long Context Fine-tuning through Dynamic Data Scheduling

Hongtao Xu, Wenting Shen, Yuanxin Wei, Ang Wang, Guo Runfan, Tianxing Wang, Yong Li, Mingzhen Li, Weile Jia

2025Year
3Citations

Abstract

Long-context supervised fine-tuning (Long-SFT) plays a vital role in enhancing the performance of large language models (LLMs) on long-context tasks. To smoothly adapt LLMs to long-context scenarios, this process typically entails training on mixed datasets containing both long and short sequences. However, this heterogeneous sequence length distribution poses significant challenges for existing training systems, as they fail to simultaneously achieve high training efficiency for both long and short sequences, resulting in sub-optimal end-to-end system performance in Long-SFT. In this paper, we present a novel perspective on data scheduling to address the challenges posed by the heterogeneous data distributions in Long-SFT. We propose Skrull, a dynamic data scheduler specifically designed for efficient long-SFT. Through dynamic data scheduling, Skrull balances the computation requirements of long and short sequences, improving overall training efficiency. Furthermore, we formulate the scheduling process as a joint optimization problem and thoroughly analyze the trade-offs involved. Based on those analysis, Skrull employs a lightweight scheduling algorithm to achieve near-zero cost online scheduling in Long-SFT. Finally, we implement Skrull upon DeepSpeed, a stateof-the-art distributed training system for LLMs. Experimental results demonstrate that Skrull outperforms DeepSpeed by 3.76x on average (up to 7.54x) in real-world long-SFT scenarios.

However, this heterogeneous data distribution in Long-SFT poses significant challenges for existing distributed LLM training frameworks [14,20,15], exhibiting sub-optimal efficiency. For instance, the heterogeneous data distribution poses a dilemma for parallelism and memory-reduction strategies. Specifically, long sequences necessitate context parallelism and other memory-reduction approaches 39th Conference on Neural Information Processing Systems (NeurIPS 2025). due to their tremendous memory requirements. However, those approaches compromise the training efficiency for short ones due to the overheads like unnecessary communication and GPU underutilization. Moreover, the wide sequence length distribution in long-SFT worsen the mismatch of computation characteristics in Attention module, which exhibit quadratic computational complexity and linear memory consumption [7,6], leading to another dilemma for load balance problem.

To tackle the above challenges, we propose Skrull, a dynamic data scheduler dedicated for Long-SFT scenarios. Skrull efficiently handle the unique data distributions in Long-SFT scenario through two main components: Distributed-Aware Context Parallelism (DACP) and Global Data Scheduling (GDS). DACP selectively shards sequences and schedules them across different workers to minimize the performance degradation while maintains the ability of handling long sequence. GDS enlarge the scope of scheduling and improve the GPU utilization during training. The two components collaborate with each other at different scheduling granularities. Furthermore, to achieve the optimal performance, we formulate the scheduling process as a joint optimization problem and design a lightweight heuristic algorithm to solve it at runtime. Experimental results demonstrate that Skrull improves the end-to-end training performance by 3.76x on average (up to 7.54x) compared to DeepSpeed, a state-of-the-art distributed LLM training framework.

Our key contributions are summarized as follows:

• We provide a new perspective of data scheduling to address the heterogeneous sequence length distribution.

• We propose a new context parallelism called DACP based on fine-grained data scheduling, which maintaining both the processing capabilities for long sequences and efficiency for short sequences, enabling efficient training on heterogeneous data distribution in long-SFT scenario.

• We implement coarse-grained global data scheduling (GDS) and further formulate GDS and DACP as a joint optimization problem through performance modeling.

• We design a lightweight heuristic algorithm and achieve performance gains by 3.76x on average (with a peak improvement of 7.54×) in real-world datasets.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

Builds on9

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines