Demystifying Synthetic Data in LLM Pre-training: A Systematic Study of Scaling Laws, Benefits, and Pitfalls
Feiyang Kang, Newsha Ardalani, Michael Kuchnik, Youssef Emad, Mostafa Elhoushi, Shubhabrata Sengupta, Shang-wen Li, Ramya Raghavendra, Ruoxi Jia, Carole-Jean Wu
Abstract
Training data plays a crucial role in Large Language Models (LLM) scaling, yet high quality data is of limited supply. Synthetic data techniques offer a potential path toward sidestepping these limitations.We conduct a large-scale empirical investigation (>1000 LLMs with >100k GPU hours) using a unified protocol and scaling laws, comparing natural web data, diverse synthetic types (rephrased text, generated textbooks), and mixtures of natural and synthetic data. Specifically, we found pre-training on rephrased synthetic data alone is not faster than pre-training on natural web texts; while pre-training on 1/3 rephrased synthetic data mixed with 2/3 natural web texts can speed up 5-10x (to reach the same validation loss) at larger data budgets. Pre-training on textbookstyle synthetic data alone results in notably higher loss on many downstream domains especially at small data budgets. "Good" ratios of synthetic data in training data mixtures depend on the model size and data budget, empirically converging to ∼30% for rephrased synthetic data. Larger generator models do not necessarily yield better pre-training data than ∼8Bparam models. These results contribute mixed evidence on "model collapse" during largescale single-round (n=1) model training on synthetic data-training on rephrased synthetic data shows no degradation in performance in foreseeable scales whereas training on mixtures of textbook-style pure-generated synthetic data shows patterns predicted by "model collapse". Our work demystifies synthetic data in pretraining, validates its conditional benefits, and offers practical guidance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9284aa5a-4cb0-4f54-8901-e9c58767528eCited by top-tier papers6
- On the Generalization Gap in Self-Evolving Language Model ReasoningZhenting Qi, Susanna Maria Baby, Stefanie Baby, Kan Yuan et al.ICML 2026
- Principled Synthetic Data Enables the First Scaling Laws for LLMs in RecommendationBenyu Zhang, Qiang Zhang, Jianpeng Cheng, Hong-You Chen et al.ICML 2026
- Your Keywords Know Each Other: Breaking SSE with <1% Leaked DocumentsMingyu Bian, Jiabei Wang, Dandan Xu, Guangyu Huang et al.USENIX Security 2026
- How Text Quality Interventions Reshape Neural Scaling Laws for LLMs: Empirical StudyNewsha Ardalani, Feiyang Kang, Michael Kuchnik, Mostafa Elhoushi et al.ICLR 2026
- KoCo: Conditioning Language Model Pre-training on Knowledge CoordinatesYudong Li, Jiawei Cai, Linlin ShenACL 2026
Builds on11
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu et al.ACL 2023 · 540 citations
- DoReMi: Optimizing Data Mixtures Speeds Up Language Model PretrainingSang Michael Xie, Hieu Pham, Xuanyi Dong, Nan Du et al.NeurIPS 2023 · 457 citations
- A Tale of Tails: Model Collapse as a Change of Scaling LawsElvis Dohmatob, Yunzhen Feng, Pu Yang, François Charton et al.ICML 2024 · 123 citations
- Performance Scaling via Optimal Transport: Enabling Data Selection from Partially Revealed SourcesFeiyang Kang, Hoang Anh Just, Anit Kumar Sahu, Ruoxi JiaNeurIPS 2023 · 21 citations
- Pre-Training to Learn in ContextYuxian Gu, Li Dong, Furu Wei, Minlie HuangACL 2023 · 18 citations
Related papers
- Rephrasing the Web: A Recipe for Compute and Data-Efficient Language ModelingPratyush Maini, Skyler Seto, Richard He Bai, David Grangier et al.ACL 2024 · 13 citations
- Strong Model CollapseElvis Dohmatob, Yunzhen Feng, Arjun Subramonian, Julia KempeICLR 2025 · 3 citations
- Collapse or Thrive: Perils and Promises of Synthetic Data in a Self-Generating WorldJoshua Kazdan, Rylan Schaeffer, Apratim Dey, Matthias Gerstgrasser et al.ICML 2025
- Beyond Model Collapse: Scaling Up with Synthesized Data Requires VerificationYunzhen Feng, Elvis Dohmatob, Pu Yang, François Charton et al.ICLR 2025 · 6 citations
- Language Generation with Replay: A Learning-Theoretic View of Model CollapseGiorgio Racca, Michal Valko, Amartya SanyalICML 2026 · 4 citations
