Demystifying Data Organization for Enhanced LLM Training
Yalun Dai, Yangyu Huang, Tongshen Yang, Yonghan Wang, Xin Zhang, Wenshan Wu, Qihao Zhao, Hao Li, Yuanyuan Gao, Kim-Hui Yap, Scarlett Li
摘要
Large Language Models (LLMs) have revolutionized various fields, yet their training efficiency is heavily reliant on effective data curation. While data selection has been widely studied, the strategic data organization for enhanced training remains an underexplored area, particularly since current LLMs are often trained for only one or a few epochs. This paper systematically explores the influence of data organization on LLM training by reusing pre-computed sample-level scores originally generated for data efficiency, thereby incurring minimal additional computational overhead. We identify and formalize four key guidelines for optimizing data organization: Boundary Sharpening, Cyclic Scheduling, Curriculum Continuity, and Local Diversity. Guided by them, we introduce two novel data ordering methods termed STR and SAW. Extensive experiments across different model scales and data sizes, encompassing both pre-training and SFT stages, validate the effectiveness of our summarized guidelines. They also demonstrate the robustness of our proposed data ordering methods in enhancing the stability and performance of LLM training. Github Link: https://github.com/microsoft/data-efficacy/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- WizardCoder: Empowering Code Large Language Models with Evol-InstructZiyang Luo, Can Xu, Pu Zhao, Qingfeng Sun 等ICLR 2024 · 被引用 945 次
- Data Selection for Language Models via Importance ResamplingSang Michael Xie, Shibani Santurkar, Tengyu Ma, Percy LiangNeurIPS 2023 · 被引用 383 次
- QuRating: Selecting High-Quality Data for Training Language ModelsAlexander Wettig, Aatmik Gupta, Saumya Malik, Danqi ChenICML 2024 · 被引用 138 次
相关 Paper
- D: Dynamic Directional Graph-Constrained Data Scheduling for LLM TrainingYuanjian Xu, Jianing Hao, Guang Zhang, Zhong LiICML 2026
- Difficulty Is Not Enough: Curriculum Learning for LLMs Fine-tuning Must Consider UtilityZishang Jiang, Jinyi Han, Tingyun Li, Xinyi Wang 等AAAI 2026 · 被引用 1 次
- How Learning Rate Decay Wastes Your Best Data in Curriculum-Based LLM PretrainingKairong Luo, Zhenbo Sun, Haodong Wen, Xinyu Shi 等ICLR 2026 · 被引用 12 次
- Rethinking Data Curation in LLM Training: Online Reweighting Offers Better Generalization than Offline MethodsWanru Zhao, Yihong Chen, Yuzhi Tang, Wentao Ma 等ICLR 2026 · 被引用 4 次
- Understanding Data Temporality Impact on Large Language Models Pre-trainingRomain Fabre, Hippolyte Pilchen, Franck SIGNE TALLA, Patrick Perez 等ICML 2026
