DataRater: Meta-Learned Dataset Curation
Dan Andrei Calian, Gregory Farquhar, Iurii Kemaev, Luisa M. Zintgraf, Matteo Hessel, Jeremy Shar, Junhyuk Oh, András György, Tom Schaul, Jeff Dean, Hado Philip van Hasselt, David Silver
Abstract
The quality of foundation models depends heavily on their training data. Consequently, great efforts have been put into dataset curation. Yet most approaches rely on manual tuning of coarse-grained mixtures of large buckets of data, or filtering by hand-crafted heuristics. An approach that is ultimately more scalable (let alone more satisfying) is to learn which data is actually valuable for training. This type of meta-learning could allow more sophisticated, fine-grained, and effective curation. Our proposed DataRater is an instance of this idea. It estimates the value of training on any particular data point. This is done by meta-learning using `meta-gradients', with the objective of improving training efficiency on held out data. In extensive experiments across a range of model scales and datasets, we find that using our DataRater to filter data is highly effective, resulting in significantly improved compute efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 97c80568-e14b-4264-96c1-ee4d902d0f9eBuilds on22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- LIMA: Less Is More for AlignmentChunting Zhou, Pengfei Liu, Puxin Xu, Srinivasan Iyer et al.NeurIPS 2023 · 1,486 citations
- Safe Deep Semi-Supervised Learning for Unseen-Class Unlabeled DataLan-Zhe Guo, Zhenyu Zhang, Yuan Jiang, Yufeng Li et al.ICML 2020 · 243 citations
Related papers
- Optimizing Data Usage via Differentiable RewardsXinyi Wang, Hieu Pham, Paul Michel, Antonios Anastasopoulos et al.ICML 2020 · 73 citations
- Curation Leaks: Membership Inference Attacks against Data Curation for Machine LearningDariush Wahdany, Matthew Jagielski, Adam Dziedzic, Franziska BoenischICLR 2026
- Meta-Learning without MemorizationMingzhang Yin, George Tucker, Mingyuan Zhou, Sergey Levine et al.ICLR 2020 · 201 citations
- Capturing the Temporal Dependence of Training Data InfluenceJiachen T. Wang, Dawn Song, James Zou, Prateek Mittal et al.ICLR 2025
- Two Birds with One Stone: Efficient Deep Learning over Mislabeled Data through Subset SelectionYuhao Deng, Chengliang Chai, Kaisen Jin, Linan Zheng et al.SIGMOD 2025 · 2 citations
