Can Small Training Runs Reliably Guide Data Curation? Rethinking Proxy-Model Practice
Jiachen T. Wang, Tong Wu, Kaifeng Lyu, James Zou, Dawn Song, Ruoxi Jia, Prateek Mittal
Abstract
Data teams at frontier AI companies routinely train small proxy models to make critical decisions about pretraining data recipes for full-scale training runs. However, the community has a limited understanding of whether and when conclusions drawn from small-scale experiments reliably transfer to full-scale model training. In this work, we uncover a subtle yet critical issue in the standard experimental protocol for data recipe assessment: the use of identical small-scale model training configurations across all data recipes in the name of "fair" comparison. We show that the experiment conclusions about data quality can flip with even minor adjustments to training hyperparameters, as the optimal training configuration is inherently data-dependent. Moreover, this fixed-configuration protocol diverges from full-scale model development pipelines, where hyperparameter optimization is a standard step. Consequently, we posit that the objective of data recipe assessment should be to identify the recipe that yields the best performance under data-specific tuning. To mitigate the high cost of hyperparameter tuning, we introduce a simple patch to the evaluation protocol: using reduced learning rates for proxy model training. We show that this approach yields relative performance that strongly correlates with that of fully tuned large-scale LLM pretraining runs. Theoretically, we prove that for random-feature models, this approach preserves the ordering of datasets according to their optimal achievable loss. Empirically, we validate this approach across 23 data recipes covering four critical dimensions of data curation, demonstrating dramatic improvements in the reliability of small-scale experiments. * Work done while at UC Berkeley. † Equal contribution as senior authors.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9a5a81da-8dd1-428a-b49e-a4c06e5c1b51Builds on19
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley et al.ICML 2023 · 1,822 citations
- Selection via Proxy: Efficient Data Selection for Deep LearningCody Coleman, Christopher Yeh, Stephen Mussmann, Baharan Mirzasoleiman et al.ICLR 2020 · 462 citations
- DoReMi: Optimizing Data Mixtures Speeds Up Language Model PretrainingSang Michael Xie, Hieu Pham, Xuanyi Dong, Nan Du et al.NeurIPS 2023 · 457 citations
- Skill-it! A data-driven skills framework for understanding and training language modelsMayee F. Chen, Nicholas Roberts, Kush Bhatia, Jue Wang et al.NeurIPS 2023 · 143 citations
Related papers
- Scaling Optimal LR Across Token HorizonsJohan Bjorck, Alon Benhaim, Vishrav Chaudhary, Furu Wei et al.ICLR 2025
- Data Mixture Optimization: A Multi-fidelity Multi-scale Bayesian FrameworkThomson Yen, Andrew Siah, Haozhe Chen, C. Guetta et al.NeurIPS 2025 · 12 citations
- DataDecide: How to Predict Best Pretraining Data with Small ExperimentsIan Magnusson, Nguyen Tai, Ben Bogin, David Heineman et al.ICML 2025
- Predicting LLM Reasoning Performance with Small Proxy ModelWoosung Koh, Juyoung Suk, Sungjun Han, Se-Young Yun et al.ICLR 2026 · 5 citations
- Meta-rater: A Multi-dimensional Data Selection Method for Pre-training Language ModelsXinlin Zhuang, Jiahui Peng, Ren Ma, Yinfan Wang et al.ACL 2025 · 15 citations
