Why Less is More (Sometimes): A Theory of Data Curation
Elvis Dohmatob, Mohammad Pezeshki, Reyhane Askari Hemmat
摘要
This paper introduces a theoretical framework to resolve a central paradox in modern machine learning: When is it better to use less data? This question has become critical as classical scaling laws suggesting more is more'' (Sun et al., 2025) are challenged by methods like LIMO (less is more'') and s1 (Ye et al., 2025; Muenighoff et al., 2025), which achieve superior performance with small, aggressively curated datasets. Here, we study data curation strategies where an imperfect oracle selects the training examples according to their difficulty and correctness. Our results provide exact scaling law curves for test error under both label-agnostic and label-aware curation rules, revealing when and why keeping only a subset of data can improve generalization. In contrast to classical scaling laws, we show that under certain conditions, small curated datasets can outperform full datasets, and we provide analytical conditions for this by deriving precise phase transition curves tied to data size and quality. We validate these theoretical claims with empirical results on ImageNet, confirming our predictions about when curation improves accuracy and can even mitigate model collapse. Furthermore, our framework provides a principled explanation for the contradictory curation strategies recently observed in LLM mathematical reasoning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- TarGATE: Target-Aware Data Selection via Token-Attenuation GatesXiandi Luo, Shiwei Li, Haozhao Wang, Yihao Ouyang 等ICML 2026
- Asymptotic Theory of Iterated Empirical Risk Minimization, with Applications to Active LearningHugo Cui, Yue LuICML 2026
- When Sample Selection Bias Precipitates Model CollapseXinbao Qiao, Xianglong Du, Wei Liu, Jingqi Zhang 等ICML 2026
- Difficulty-Diversity Collaborative Filtering for Data-Efficient LLM Fine-TuningLong P. Hoang, Wenxuan Zhang, Wei LuICLR 2026
它引用的顶会 Paper18
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Beyond neural scaling laws: beating power law scaling via data pruningBen Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli 等NeurIPS 2022 · 被引用 720 次
- Self-Consuming Generative Models Go MADSina Alemohammad, Josue Casco-Rodriguez, Lorenzo Luzi, Ahmed Imtiaz Humayun 等ICLR 2024 · 被引用 279 次
- Does Writing with Language Models Reduce Content Diversity?Vishakh Padmakumar, He HeICLR 2024 · 被引用 173 次
- RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-FoldAmrith Setlur, Saurabh Garg, Xinyang Geng, Naman Garg 等NeurIPS 2024 · 被引用 143 次
相关 Paper
- Escaping Collapse: The Strength of Weak Data for Large Language Model TrainingKareem Amin, Sara Babakniya, Alex Bie, Weiwei Kong 等NeurIPS 2025 · 被引用 17 次
- Improving Data Efficiency via Curating LLM-Driven Rating SystemsJinlong Pang, Jiaheng Wei, Ankit Shah, Zhaowei Zhu 等ICLR 2025
- Rethinking Data Curation in LLM Training: Online Reweighting Offers Better Generalization than Offline MethodsWanru Zhao, Yihong Chen, Yuzhi Tang, Wentao Ma 等ICLR 2026 · 被引用 4 次
- What Do Learning Dynamics Reveal About Generalization in LLM Mathematical Reasoning?Katie Kang, Amrith Setlur, Dibya Ghosh, Jacob Steinhardt 等ICML 2025
- Small Data, Big Decisions: Model Selection in the Small-Data RegimeJörg Bornschein, Francesco Visin, Simon OsinderoICML 2020 · 被引用 48 次
