ScalingFilter: Assessing Data Quality through Inverse Utilization of Scaling Laws
Ruihang Li, Yixuan Wei, Miaosen Zhang, Nenghai Yu, Han Hu, Houwen Peng
Abstract
High-quality data is crucial for the pre-training performance of large language models. Unfortunately, existing quality filtering methods rely on a known high-quality dataset as reference, which can introduce potential bias and compromise diversity. In this paper, we propose ScalingFilter, a novel approach that evaluates text quality based on the perplexity difference between two language models trained on the same data, thereby eliminating the influence of the reference dataset in the filtering process. An theoretical analysis shows that ScalingFilter is equivalent to an inverse utilization of scaling laws. Through training models with 1.3B parameters on the same data source processed by various quality filters, we find ScalingFilter can improve zero-shot performance of pre-trained models in downstream tasks. To assess the bias introduced by quality filtering, we introduce semantic diversity, a metric of utilizing text embedding models for semantic representations. Extensive experiments reveal that semantic diversity is a reliable indicator of dataset diversity, and ScalingFilter achieves an optimal balance between downstream performance and semantic diversity. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 94e9a805-60a3-427d-bdc4-7a5bc4821f40Cited by top-tier papers2
- Predictable Scale (Part II) - Farseer: A Refined Scaling Law in LLMsHouyi Li, Wenzhen Zheng, Qiufeng Wang, Zhenyu Ding et al.NeurIPS 2025 · 4 citations
- Difficulty Is Not Enough: Curriculum Learning for LLMs Fine-tuning Must Consider UtilityZishang Jiang, Jinyi Han, Tingyun Li, Xinyi Wang et al.AAAI 2026 · 1 citation
Builds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong et al.ICML 2022 · 1,173 citations
- Unsupervised Cross-lingual Representation Learning at ScaleAlexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary et al.ACL 2020 · 539 citations
Related papers
- Your UnEmbedding Matrix is Secretly a Feature Lens for Text EmbeddingsSonghao Wu, Zhongxin Chen, Yuxuan Liu, Heng Cui et al.KDD 2026 · 1 citation
- Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language ModelsXinlin Zhuang, Jiahui Peng, Ren Ma, Yinfan Wang et al.ACL 2025 · 15 citations
- Analyzing Similarity Metrics for Data Selection for Language Model PretrainingDylan Sam, Ayan Chakrabarti, Afshin Rostamizadeh, Srikumar Ramalingam et al.NeurIPS 2025 · 3 citations
- Removing Noise, not Finding Gold: Quality Filtering for Large-Scale PretrainingThiziri Nait Saada, Louis Béthune, Michal Klein, David Grangier et al.ICML 2026
- Prior-based Noisy Text Data Filtering: Fast and Strong Alternative For PerplexityYeongbin Seo, Gayoung Kim, Jaehyung Kim, Jinyoung YeoICLR 2026
