Skrull: Towards Efficient Long Context Fine-tuning through Dynamic Data Scheduling
Hongtao Xu, Wenting Shen, Yuanxin Wei, Ang Wang, Guo Runfan, Tianxing Wang, Yong Li, Mingzhen Li, Weile Jia
Abstract
Long-context supervised fine-tuning (Long-SFT) plays a vital role in enhancing the performance of large language models (LLMs) on long-context tasks. To smoothly adapt LLMs to long-context scenarios, this process typically entails training on mixed datasets containing both long and short sequences. However, this heterogeneous sequence length distribution poses significant challenges for existing training systems, as they fail to simultaneously achieve high training efficiency for both long and short sequences, resulting in sub-optimal end-to-end system performance in Long-SFT. In this paper, we present a novel perspective on data scheduling to address the challenges posed by the heterogeneous data distributions in Long-SFT. We propose Skrull, a dynamic data scheduler specifically designed for efficient long-SFT. Through dynamic data scheduling, Skrull balances the computation requirements of long and short sequences, improving overall training efficiency. Furthermore, we formulate the scheduling process as a joint optimization problem and thoroughly analyze the trade-offs involved. Based on those analysis, Skrull employs a lightweight scheduling algorithm to achieve near-zero cost online scheduling in Long-SFT. Finally, we implement Skrull upon DeepSpeed, a stateof-the-art distributed training system for LLMs. Experimental results demonstrate that Skrull outperforms DeepSpeed by 3.76x on average (up to 7.54x) in real-world long-SFT scenarios.
However, this heterogeneous data distribution in Long-SFT poses significant challenges for existing distributed LLM training frameworks [14,20,15], exhibiting sub-optimal efficiency. For instance, the heterogeneous data distribution poses a dilemma for parallelism and memory-reduction strategies. Specifically, long sequences necessitate context parallelism and other memory-reduction approaches 39th Conference on Neural Information Processing Systems (NeurIPS 2025). due to their tremendous memory requirements. However, those approaches compromise the training efficiency for short ones due to the overheads like unnecessary communication and GPU underutilization. Moreover, the wide sequence length distribution in long-SFT worsen the mismatch of computation characteristics in Attention module, which exhibit quadratic computational complexity and linear memory consumption [7,6], leading to another dilemma for load balance problem.
To tackle the above challenges, we propose Skrull, a dynamic data scheduler dedicated for Long-SFT scenarios. Skrull efficiently handle the unique data distributions in Long-SFT scenario through two main components: Distributed-Aware Context Parallelism (DACP) and Global Data Scheduling (GDS). DACP selectively shards sequences and schedules them across different workers to minimize the performance degradation while maintains the ability of handling long sequence. GDS enlarge the scope of scheduling and improve the GPU utilization during training. The two components collaborate with each other at different scheduling granularities. Furthermore, to achieve the optimal performance, we formulate the scheduling process as a joint optimization problem and design a lightweight heuristic algorithm to solve it at runtime. Experimental results demonstrate that Skrull improves the end-to-end training performance by 3.76x on average (up to 7.54x) compared to DeepSpeed, a state-of-the-art distributed LLM training framework.
Our key contributions are summarized as follows:
• We provide a new perspective of data scheduling to address the heterogeneous sequence length distribution.
• We propose a new context parallelism called DACP based on fine-grained data scheduling, which maintaining both the processing capabilities for long sequences and efficiency for short sequences, enabling efficient training on heterogeneous data distribution in long-SFT scenario.
• We implement coarse-grained global data scheduling (GDS) and further formulate GDS and DACP as a joint optimization problem through performance modeling.
• We design a lightweight heuristic algorithm and achieve performance gains by 3.76x on average (with a peak improvement of 7.54×) in real-world datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on9
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 852 citations
- LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation DatasetLianmin Zheng, Wei-Lin Chiang, Ying Sheng, Tianle Li et al.ICLR 2024 · 419 citations
Related papers
- Efficient Long Context Fine-tuning with Chunk FlowXiulong Yuan, Hongtao Xu, Wenting Shen, Ang Wang et al.ICML 2025
- FlexSP: Accelerating Large Language Model Training via Flexible Sequence ParallelismYujie Wang, Shiju Wang, Shenhan Zhu, Fangcheng Fu et al.ASPLOS 2025 · 8 citations
- ByteScale: Communication-Efficient Scaling of LLM Training with a 2048K Context Length on 16384 GPUsHao Ge, Junda Feng, Qi Huang, Fangcheng Fu et al.SIGCOMM 2025 · 4 citations
- Dataset Decomposition: Faster LLM Training with Variable Sequence Length CurriculumHadi Pouransari, Chun-Liang Li, Jen-Hao Rick Chang, Pavan Kumar Anasosalu Vasu et al.NeurIPS 2024 · 40 citations
- Zeppelin: Balancing Variable-length Workloads in Data Parallel Large Model TrainingChang Chen, Tiancheng Chen, Jiangfei Duan, Qianchao Zhu et al.EuroSys 2026
