Joint Selection for Large-Scale Pre-Training Data via Policy Gradient-based Mask Learning
Ziqing Fan, Yuqiao Xian, Yan Sun, Ke Shen, Li Shen
Abstract
A fine-grained data recipe is crucial for pre-training large language models (LLMs), as it can significantly enhance training efficiency and model performance. One important ingredient in the recipe is to select samples based on scores produced by defined rules, LLM judgment, or statistical information in embeddings, which can be roughly categorized into quality and diversity metrics. Due to the high computational cost when applied to trillion-scale token pre-training datasets such as FineWeb and DCLM, these two or more types of metrics are rarely considered jointly in a single selection process. However, in our empirical study, selecting samples based on quality metrics exhibit severe diminishing returns during long-term pre-training, while selecting on diversity metrics removes too many valuable high-quality samples, both of which limit pre-trained LLMs' capabilities. Therefore, we introduce DATAMASK, a novel and efficient joint learning framework designed for large-scale pre-training data selection that can simultaneously optimize multiple types of metrics in a unified process, with this study focusing specifically on quality and diversity metrics. DATAMASK approaches the selection process as a mask learning problem, involving iterative sampling of data masks, computation of policy gradients based on predefined objectives with sampled masks, and updating of mask sampling logits. Through policy gradient-based optimization and various acceleration enhancements, it significantly reduces selection time by 98.9% compared to greedy algorithm, enabling our study to explore joint learning within trillion-scale tokens. With DATAMASK, we select a subset of about 10% from the 15 trillion-token FineWeb dataset, termed FineWeb-Mask. Evaluated across 12 diverse tasks, this high-quality and diverse subset achieves significant improvements of 3.2% on a 1.5B dense model and 1.9% on a 7B MoE model after pre-training with hundreds of billions of tokens, demonstrating its effectiveness. Source code is available at: https://github.com/ByteDance-Seed/DATAMASK.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on7
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- QuRating: Selecting High-Quality Data for Training Language ModelsAlexander Wettig, Aatmik Gupta, Saumya Malik, Danqi ChenICML 2024 · 138 citations
- Combinatorial Optimization for Panoptic Segmentation: A Fully Differentiable ApproachAhmed Abbas, Paul SwobodaNeurIPS 2021 · 16 citations
- GneissWeb: Preparing High Quality Data for LLMs at ScaleHajar Emami Gohari, Swanand Ravindra Kadhe, Yousaf Shah, Constantin M Adam et al.ICLR 2026 · 7 citations
Related papers
- MASS: Mathematical Data Selection via Skill Graphs for Pretraining Large Language ModelsJiazheng Li, Lu Yu, Qing Cui, Zhiqiang Zhang et al.ICML 2025
- FIRE: Flexible Integration of Data Quality Ratings for Effective PretrainingLiangyu Xu, Xuemiao Zhang, Feiyu Duan, Sirui Wang et al.EMNLP 2025
- Predictive Data Selection: The Data That Predicts Is the Data That TeachesKaShun Shum, Yuzhen Huang, Hongjian Zou, Qi Ding et al.ICML 2025
- Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language ModelsXinlin Zhuang, Jiahui Peng, Ren Ma, Yinfan Wang et al.ACL 2025 · 15 citations
- OPUS: Towards Efficient and Principled Data Selection in Large Language Model Pre-training in Every IterationShaobo Wang, Xuan Ouyang, Tianyi Xu, Yuzheng Hu et al.ICML 2026 · 12 citations
