Demystifying Synthetic Data in LLM Pre-training: A Systematic Study of Scaling Laws, Benefits, and Pitfalls
Feiyang Kang, Newsha Ardalani, Michael Kuchnik, Youssef Emad, Mostafa Elhoushi, Shubhabrata Sengupta, Shang-wen Li, Ramya Raghavendra, Ruoxi Jia, Carole-Jean Wu
摘要
Training data plays a crucial role in Large Language Models (LLM) scaling, yet high quality data is of limited supply. Synthetic data techniques offer a potential path toward sidestepping these limitations.We conduct a large-scale empirical investigation (>1000 LLMs with >100k GPU hours) using a unified protocol and scaling laws, comparing natural web data, diverse synthetic types (rephrased text, generated textbooks), and mixtures of natural and synthetic data. Specifically, we found pre-training on rephrased synthetic data alone is not faster than pre-training on natural web texts; while pre-training on 1/3 rephrased synthetic data mixed with 2/3 natural web texts can speed up 5-10x (to reach the same validation loss) at larger data budgets. Pre-training on textbookstyle synthetic data alone results in notably higher loss on many downstream domains especially at small data budgets. "Good" ratios of synthetic data in training data mixtures depend on the model size and data budget, empirically converging to ∼30% for rephrased synthetic data. Larger generator models do not necessarily yield better pre-training data than ∼8Bparam models. These results contribute mixed evidence on "model collapse" during largescale single-round (n=1) model training on synthetic data-training on rephrased synthetic data shows no degradation in performance in foreseeable scales whereas training on mixtures of textbook-style pure-generated synthetic data shows patterns predicted by "model collapse". Our work demystifies synthetic data in pretraining, validates its conditional benefits, and offers practical guidance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- On the Generalization Gap in Self-Evolving Language Model ReasoningZhenting Qi, Susanna Maria Baby, Stefanie Baby, Kan Yuan 等ICML 2026
- Principled Synthetic Data Enables the First Scaling Laws for LLMs in RecommendationBenyu Zhang, Qiang Zhang, Jianpeng Cheng, Hong-You Chen 等ICML 2026
- Your Keywords Know Each Other: Breaking SSE with <1% Leaked DocumentsMingyu Bian, Jiabei Wang, Dandan Xu, Guangyu Huang 等USENIX Security 2026
- How Text Quality Interventions Reshape Neural Scaling Laws for LLMs: Empirical StudyNewsha Ardalani, Feiyang Kang, Michael Kuchnik, Mostafa Elhoushi 等ICLR 2026
- KoCo: Conditioning Language Model Pre-training on Knowledge CoordinatesYudong Li, Jiawei Cai, Linlin ShenACL 2026
它引用的顶会 Paper11
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu 等ACL 2023 · 被引用 540 次
- DoReMi: Optimizing Data Mixtures Speeds Up Language Model PretrainingSang Michael Xie, Hieu Pham, Xuanyi Dong, Nan Du 等NeurIPS 2023 · 被引用 457 次
- A Tale of Tails: Model Collapse as a Change of Scaling LawsElvis Dohmatob, Yunzhen Feng, Pu Yang, François Charton 等ICML 2024 · 被引用 123 次
- Performance Scaling via Optimal Transport: Enabling Data Selection from Partially Revealed SourcesFeiyang Kang, Hoang Anh Just, Anit Kumar Sahu, Ruoxi JiaNeurIPS 2023 · 被引用 21 次
- Pre-Training to Learn in ContextYuxian Gu, Li Dong, Furu Wei, Minlie HuangACL 2023 · 被引用 18 次
相关 Paper
- Rephrasing the Web: A Recipe for Compute and Data-Efficient Language ModelingPratyush Maini, Skyler Seto, Richard He Bai, David Grangier 等ACL 2024 · 被引用 13 次
- Strong Model CollapseElvis Dohmatob, Yunzhen Feng, Arjun Subramonian, Julia KempeICLR 2025 · 被引用 3 次
- Collapse or Thrive: Perils and Promises of Synthetic Data in a Self-Generating WorldJoshua Kazdan, Rylan Schaeffer, Apratim Dey, Matthias Gerstgrasser 等ICML 2025
- Beyond Model Collapse: Scaling Up with Synthesized Data Requires VerificationYunzhen Feng, Elvis Dohmatob, Pu Yang, François Charton 等ICLR 2025 · 被引用 6 次
- Language Generation with Replay: A Learning-Theoretic View of Model CollapseGiorgio Racca, Michal Valko, Amartya SanyalICML 2026 · 被引用 4 次
