Data Selection via Optimal Control for Language Models
Yuxian Gu, Li Dong, Hongning Wang, Yaru Hao, Qingxiu Dong, Furu Wei, Minlie Huang
摘要
This work investigates the selection of high-quality pre-training data from massive corpora to enhance LMs' capabilities for downstream usage. We formulate data selection as a generalized Optimal Control problem, which can be solved theoretically by Pontryagin's Maximum Principle (PMP), yielding a set of necessary conditions that characterize the relationship between optimal data selection and LM training dynamics. Based on these theoretical results, we introduce PMP-based Data Selection (PDS), a framework that approximates optimal data selection by solving the PMP conditions. In our experiments, we adopt PDS to select data from CommmonCrawl and show that the PDS-selected corpus accelerates the learning of LMs and constantly boosts their performance on a wide range of downstream tasks across various model sizes. Moreover, the benefits of PDS extend to 400B models trained on 10T tokens, as evidenced by the extrapolation of the test loss curves according to the Scaling Laws. PDS also improves data utilization when the pre-training data is limited, by reducing the data demand by 1.8 times, which helps mitigate the quick exhaustion of available web-crawled corpora. Our code, model, and data can be found at https: //github.com/microsoft/LMOps/tree/main/data_selection .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper24
- DataRater: Meta-Learned Dataset CurationDan Andrei Calian, Gregory Farquhar, Iurii Kemaev, Luisa M. Zintgraf 等NeurIPS 2025 · 被引用 17 次
- Adaptive Defense against Harmful Fine-Tuning for Large Language Models via Bayesian Data SchedulerZixuan Hu, Li Shen, Zhenyi Wang, Yongxian Wei 等NeurIPS 2025 · 被引用 16 次
- Group-Level Data Selection for Efficient PretrainingZichun Yu, Fei Peng, Jie Lei, Arnold Overwijk 等NeurIPS 2025 · 被引用 13 次
- How Learning Rate Decay Wastes Your Best Data in Curriculum-Based LLM PretrainingKairong Luo, Zhenbo Sun, Haodong Wen, Xinyu Shi 等ICLR 2026 · 被引用 12 次
- Influence-Preserving Proxies for Gradient-Based Data Selection in LLM FineTuningSirui Chen, Yunzhe Qi, Mengting Ai, Yifan Sun 等ICLR 2026 · 被引用 9 次
它引用的顶会 Paper22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- ZeRO: memory optimizations toward training trillion parameter modelsSamyam Rajbhandari, Jeff Rasley, Olatunji Ruwase, Yuxiong HeSC 2020 · 被引用 852 次
- Are Emergent Abilities of Large Language Models a Mirage?Rylan Schaeffer, Brando Miranda, Sanmi KoyejoNeurIPS 2023 · 被引用 796 次
相关 Paper
- Efficient Pretraining Data Selection for Language Models via Multi-Actor CollaborationTianyi Bai, Ling Yang, Zhen Hao Wong, Fupeng Sun 等ACL 2025
- OPUS: Towards Efficient and Principled Data Selection in Large Language Model Pre-training in Every IterationShaobo Wang, Xuan Ouyang, Tianyi Xu, Yuzheng Hu 等ICML 2026 · 被引用 12 次
- Instruction Pre-Training: Language Models are Supervised Multitask LearnersDaixuan Cheng, Yuxian Gu, Shaohan Huang, Junyu Bi 等EMNLP 2024 · 被引用 13 次
- Improving Pretraining Data Using Perplexity CorrelationsTristan Thrush, Christopher Potts, Tatsunori HashimotoICLR 2025
- Predictive Data Selection: The Data That Predicts Is the Data That TeachesKaShun Shum, Yuzhen Huang, Hongjian Zou, Qi Ding 等ICML 2025
