Lune

ICML2021Top-tier venue

Model Performance Scaling with Multiple Data Sources

Tatsunori Hashimoto

2021Year
38Citations
15Top-tier citations

Abstract

Real-world machine learning systems are often trained using a mix of data sources with varying cost and quality. Understanding how the size and composition of a training dataset affect model performance is critical for advancing our understanding of generalization, as well as designing more effective data collection policies. We show that there is a simple scaling law that predicts the loss incurred by a model even under varying dataset composition. Our work expands recent observations of scaling laws for log-linear generalization error in the i.i.d setting and uses this to cast model performance prediction as a learning problem. Using the theory of optimal experimental design, we derive a simple rational function approximation to generalization error that can be fitted using a few model training runs. Our approach can achieve highly accurate (r 2 ≈ .9) predictions of model performance under substantial extrapolation in two different standard supervised learning tasks and is accurate (r 2 ≈ .83) on more challenging machine translation and question answering tasks where many baselines achieve worse-thanrandom performance.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 3cfc98ba-a0d3-4f0d-9e0a-9bf571dbb743

Cited by top-tier papers15

Ask how each one uses it

Builds on3

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines