DataMan: Data Manager for Pre-training Large Language Models
Ru Peng, Kexin Yang, Yawen Zeng, Junyang Lin, Dayiheng Liu, Junbo Zhao
摘要
The performance emergence of large language models (LLMs) driven by data scaling laws makes the selection of pre-training data increasingly important. However, existing methods rely on limited heuristics and human intuition, lacking comprehensive and clear guidelines. To address this, we are inspired by "reverse thinking" -prompting LLMs to self-identify which criteria benefit its performance. As its pre-training capabilities are related to perplexity (PPL), we derive 14 quality criteria from the causes of text perplexity anomalies and introduce 15 common application domains to support domain mixing. In this paper, we train a Data Manager (DataMan) to learn quality ratings and domain recognition from pointwise rating, and use it to annotate a 447B token pre-training corpus with 14 quality ratings and domain type. Our experiments validate our approach, using DataMan to select 30B tokens to train a 1.3B-parameter language model, demonstrating significant improvements in in-context learning (ICL), perplexity, and instructionfollowing ability over the state-of-the-art baseline. The best-performing model, based on the Overall Score l=5 surpasses a model trained with 50% more data using uniform sampling. We continue pre-training with high-rated, domain-specific data annotated by DataMan to enhance domain-specific ICL performance and thus verify DataMan's domain mixing ability. Our findings emphasize the importance of quality ranking, the complementary nature of quality criteria, and their low correlation with perplexity, analyzing misalignment between PPL and ICL performance. We also thoroughly analyzed our pre-training dataset, examining its composition, the distribution of quality ratings, and the original document sources.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language ModelsXinlin Zhuang, Jiahui Peng, Ren Ma, Yinfan Wang 等ACL 2025 · 被引用 15 次
- ADEPT: Continual Pretraining via Adaptive Expansion and Dynamic Decoupled TuningJinyang Zhang, Yue Fang, Hongxin Ding, Weibin Liao 等ICLR 2026 · 被引用 5 次
- MuRating: A High Quality Data Selecting Approach to Multilingual Large Language Model PretrainingZhixun Chen, Ping Guo, Wenhan Han, Yifan Zhang 等NeurIPS 2025 · 被引用 4 次
- RePro: Training Language Models to Faithfully Recycle the Web for PretrainingZichun Yu, Chenyan XiongICML 2026 · 被引用 2 次
- DeMix: Debugging Training Data with Mixed Data Error Types by Investigating Influence VectorsJiale Deng, Yanyan Shen, Xiaogang Shi, Junjun ChaiKDD 2026
它引用的顶会 Paper18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
相关 Paper
- QuRating: Selecting High-Quality Data for Training Language ModelsAlexander Wettig, Aatmik Gupta, Saumya Malik, Danqi ChenICML 2024 · 被引用 138 次
- Improving Pretraining Data Using Perplexity CorrelationsTristan Thrush, Christopher Potts, Tatsunori HashimotoICLR 2025
- DecorateLM: Data Engineering through Corpus Rating, Tagging, and Editing with Language ModelsRanchi Zhao, Zhen Leng Thai, Yifan Zhang, Shengding Hu 等EMNLP 2024
- CMR Scaling Law: Predicting Critical Mixture Ratios for Continual Pre-training of Language ModelsJiawei Gu, Zacc Yang, Chuanghao Ding, Rui Zhao 等EMNLP 2024 · 被引用 2 次
- RegMix: Data Mixture as Regression for Language Model Pre-trainingQian Liu, Xiaosen Zheng, Niklas Muennighoff, Guangtao Zeng 等ICLR 2025
