Model Performance Scaling with Multiple Data Sources
Tatsunori Hashimoto
Abstract
Real-world machine learning systems are often trained using a mix of data sources with varying cost and quality. Understanding how the size and composition of a training dataset affect model performance is critical for advancing our understanding of generalization, as well as designing more effective data collection policies. We show that there is a simple scaling law that predicts the loss incurred by a model even under varying dataset composition. Our work expands recent observations of scaling laws for log-linear generalization error in the i.i.d setting and uses this to cast model performance prediction as a learning problem. Using the theory of optimal experimental design, we derive a simple rational function approximation to generalization error that can be fitted using a few model training runs. Our approach can achieve highly accurate (r 2 ≈ .9) predictions of model performance under substantial extrapolation in two different standard supervised learning tasks and is accurate (r 2 ≈ .83) on more challenging machine translation and question answering tasks where many baselines achieve worse-thanrandom performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3cfc98ba-a0d3-4f0d-9e0a-9bf571dbb743Cited by top-tier papers15
- Mandoline: Model Evaluation under Distribution ShiftMayee F. Chen, Karan Goel, Nimit Sharad Sohoni, Fait Poms et al.ICML 2021 · 84 citations
- DOGE: Domain Reweighting with Generalization EstimationSimin Fan, Matteo Pagliardini, Martin JaggiICML 2024 · 79 citations
- Scaling laws for learning with real and surrogate dataAyush Jain, Andrea Montanari, Eren SasogluNeurIPS 2024 · 30 citations
- Performance Scaling via Optimal Transport: Enabling Data Selection from Partially Revealed SourcesFeiyang Kang, Hoang Anh Just, Anit Kumar Sahu, Ruoxi JiaNeurIPS 2023 · 21 citations
- ATLAS: Adaptive Transfer Scaling Laws for Multilingual Pretraining, Finetuning, and Decoding the Curse of MultilingualityShayne Longpre, Sneha Kudugunta, Niklas Muennighoff, I-Hung Hsu et al.ICLR 2026 · 20 citations
Builds on3
- A Constructive Prediction of the Generalization Error Across ScalesJonathan S. Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, Nir ShavitICLR 2020 · 265 citations
- Data Valuation using Reinforcement LearningJinsung Yoon, Sercan Ömer Arik, Tomas PfisterICML 2020 · 236 citations
- A Distributional Framework For Data ValuationAmirata Ghorbani, Michael P. Kim, James ZouICML 2020 · 152 citations
Related papers
- How Much More Data Do I Need? Estimating Requirements for Downstream TasksRafid Mahmood, James Lucas, David Acuna, Daiqing Li et al.CVPR 2022 · 21 citations
- Scaling Laws for the Value of Individual Data Points in Machine LearningIan Connick Covert, Wenlong Ji, Tatsunori Hashimoto, James ZouICML 2024 · 12 citations
- Parameters vs FLOPs: Scaling Laws for Optimal Sparsity for Mixture-of-Experts Language ModelsSamira Abnar, Harshay Shah, Dan Busbridge, Alaaeldin El-Nouby et al.ICML 2025
- InfoLaw: Information Scaling Laws for Large Language Models with Quality-Weighted Mixture Data and RepetitionWeidong Zhou, Fengze Liu, LIU, Ping Guo et al.ICML 2026 · 1 citation
- Scaling Laws for Downstream Task Performance in Machine TranslationBerivan Isik, Natalia Ponomareva, Hussein Hazimeh, Dimitris Paparas et al.ICLR 2025
