QuRating: Selecting High-Quality Data for Training Language Models
Alexander Wettig, Aatmik Gupta, Saumya Malik, Danqi Chen
Abstract
Selecting high-quality pre-training data is important for creating capable language models, but existing methods rely on simple heuristics. We introduce QuRating, a method for selecting pretraining data that can capture human intuitions about data quality. In this paper, we investigate four qualities-writing style, required expertise, facts & trivia, and educational value-and find that LLMs are able to discern these qualities, especially when making pairwise judgments of texts. We train a QuRater model to learn scalar ratings from pairwise judgments, and use it to annotate a 260B training corpus with quality ratings for each of the four criteria. In our experiments, we select 30B tokens according to the different quality ratings and train 1.3B-parameter language models on the selected data. We find that it is important to balance quality and diversity. When we sample using quality ratings as logits over documents, our models obtain lower perplexity and stronger in-context learning performance than baselines. Our best model is based on educational value and performs similarly to a model trained with uniform sampling for 50% more steps. Beyond data selection, we use the quality ratings to construct a training curriculum which improves performance without changing the training dataset. We extensively analyze the quality ratings and discuss their characteristics, biases, and wider implications. 2 2 0 2 4 Writing Style Wiki-en -de -ru Book StackEx. Github ArXiv 4 2 0 2 4 Facts & Trivia 4 2 0 2 4 Educational Value 4 2 0 2 4 Required Expertise Figure 4. Distribution of quality ratings, normalized for each criterion to have zero mean and unit standard deviation across the corpus. protein, gene and energy, climate, species are rated highly on required expertise, educational value, and facts & trivia. Meanwhile, the book, author cluster tends to obtain high ratings in writing style. However, almost all clusters encompass a wide range of quality ratings. Comparison to perplexity filtering. We compare sequencelevel log-likelihood scores from Llama-2-7b (Touvron et al., 2023b) with the quality ratings across 1M training sequences and visualize the relationship in Figure 7 in the appendix. We observe that documents with low quality ratings have a wide range of likelihoods, and the Spearman correlation coefficient varies between 0.50 for writing style to -0.02 for required expertise. Therefore, QuRating is meaningfully different from selecting texts based on perplexity scores from a strong LLM (Marion et al., 2023). Data Inspection We study raw documents from each of the domains and clusters discussed in Section 6.1. We select training examples at the 5th, 30th, 70th and 95th percentile for each criterion, and feature random extracts in Appendix F without any cherrypicking. While this is a minute sliver of the training data, the documents still exhibit clear qualitative differences and we invite the reader to inspect them in the appendix.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d4dd8c34-9460-4e48-a1fb-b23fea88ccc4Cited by top-tier papers74
- RewardBench 2: Advancing Reward Model EvaluationSaumya Malik, Valentina Pyatkin, Sander Land, Jacob Morrison et al.ICLR 2026 · 139 citations
- MATES: Model-Aware Data Selection for Efficient Pretraining with Data Influence ModelsZichun Yu, Spandan Das, Chenyan XiongNeurIPS 2024 · 117 citations
- GREATS: Online Selection of High-Quality Data for LLM Training in Every IterationJiachen T. Wang, Tong Wu, Dawn Song, Prateek Mittal et al.NeurIPS 2024 · 91 citations
- Scaling Laws for Optimal Data MixturesMustafa Shukor, Louis Béthune, Dan Busbridge, David Grangier et al.NeurIPS 2025 · 54 citations
- Quality-Driven Curation of Remote Sensing Vision-Language Data via Learned Scoring ModelsDilxat Muhtar, Enzhuo Zhang, Zhenshi Li, Feng Gu et al.NeurIPS 2025 · 16 citations
Builds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
Related papers
- DataMan: Data Manager for Pre-training Large Language ModelsRu Peng, Kexin Yang, Yawen Zeng, Junyang Lin et al.ICLR 2025
- CritiQ: Mining Data Quality Criteria from Human PreferencesHonglin Guo, Kai Lv, Qipeng Guo, Tianyi Liang et al.ACL 2025
- MuRating: A High Quality Data Selecting Approach to Multilingual Large Language Model PretrainingZhixun Chen, Ping Guo, Wenhan Han, Yifan Zhang et al.NeurIPS 2025 · 4 citations
- Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language ModelsXinlin Zhuang, Jiahui Peng, Ren Ma, Yinfan Wang et al.ACL 2025 · 15 citations
- DecorateLM: Data Engineering through Corpus Rating, Tagging, and Editing with Language ModelsRanchi Zhao, Zhen Leng Thai, Yifan Zhang, Shengding Hu et al.EMNLP 2024
