Efficient Pretraining Data Selection for Language Models via Multi-Actor Collaboration
Tianyi Bai, Ling Yang, Zhen Hao Wong, Fupeng Sun, Xinlin Zhuang, Jiahui Peng, Chi Zhang, Lijun Wu, Jiantao Qiu, Wentao Zhang, Binhang Yuan, Conghui He
Abstract
Efficient data selection is crucial to accelerate the pretraining of language model (LMs). While various methods have been proposed to enhance data efficiency, limited research has addressed the inherent conflicts between these approaches to achieve optimal data selection for LM pretraining. To tackle this problem, we propose a multi-actor collaborative data selection mechanism: each data selection method independently prioritizes data based on its criterion and updates its prioritization rules using the current state of the model, functioning as an independent actor for data selection; and a console is designed to adjust the impacts of different actors at various stages and dynamically integrate information from all actors throughout the LM pretraining process. We conduct extensive empirical studies to evaluate our multi-actor framework. The experimental results demonstrate that our approach significantly improves data efficiency, accelerates convergence in LM pretraining, and achieves an average relative performance gain up to 10.5% across multiple language model benchmarks compared to the state-of-the-art methods. Code and checkpoints are publicly released at https: //github.com/Relaxed-System-Lab/ multi-actor-data-selection.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a32a27cd-8c26-4962-af9a-2fe472490a7fCited by top-tier papers3
- Selective Learning for Deep Time Series ForecastingYisong Fu, Zezhi Shao, Chengqing Yu, Yujie Li et al.NeurIPS 2025 · 10 citations
- LRAgent: Efficient KV Cache Sharing for Multi-LoRA LLM AgentsHyesung Jeon, Hyeongju Ha, jae-joon kimICML 2026 · 5 citations
- D: Dynamic Directional Graph-Constrained Data Scheduling for LLM TrainingYuanjian Xu, Jianing Hao, Guang Zhang, Zhong LiICML 2026
Builds on16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong et al.ICML 2022 · 1,173 citations
- LESS: Selecting Influential Data for Targeted Instruction TuningMengzhou Xia, Sadhika Malladi, Suchin Gururangan, Sanjeev Arora et al.ICML 2024 · 460 citations
Related papers
- FIRE: Flexible Integration of Data Quality Ratings for Effective PretrainingLiangyu Xu, Xuemiao Zhang, Feiyu Duan, Sirui Wang et al.EMNLP 2025
- Group-Level Data Selection for Efficient PretrainingZichun Yu, Fei Peng, Jie Lei, Arnold Overwijk et al.NeurIPS 2025 · 13 citations
- Predictive Data Selection: The Data That Predicts Is the Data That TeachesKaShun Shum, Yuzhen Huang, Hongjian Zou, Qi Ding et al.ICML 2025
- MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence ModelsZichun Yu, Spandan Das, Chenyan XiongNeurIPS 2024 · 117 citations
- LLM Data Selection and Utilization via Dynamic Bi-level OptimizationYang Yu, Kai Han, Hang Zhou, Yehui Tang et al.ICML 2025
