How to Synthesize Text Data without Model Collapse?
Xuekai Zhu, Daixuan Cheng, Hengli Li, Kaiyan Zhang, Ermo Hua, Xingtai Lv, Ning Ding, Zhouhan Lin, Zilong Zheng, Bowen Zhou
Abstract
Model collapse in synthetic data indicates that iterative training on self-generated data leads to a gradual decline in performance. With the proliferation of AI models, synthetic data will fundamentally reshape the web data ecosystem. Future GPT-n models will inevitably be trained on a blend of synthetic and human-produced data. In this paper, we focus on two questions: what is the impact of synthetic data on language model training, and how to synthesize data without model collapse? We first pre-train language models across different proportions of synthetic data, revealing a negative correlation between the proportion of synthetic data and model performance. We further conduct statistical analysis on synthetic data to uncover distributional shift phenomenon and over-concentration of n-gram features. Inspired by the above findings, we propose token editing on human-produced data to obtain semi-synthetic data. As a proof of concept, we theoretically demonstrate that token-level editing can prevent model collapse, as the test error is constrained by a finite upper bound. We conduct extensive experiments on pre-training from scratch, continual pre-training, and supervised fine-tuning. The results validate our theoretical proof that token-level editing improves model performance. How to Synthesize Text Data without Model Collapse? โ Model Collapse Setting โข ๐ฌ ๐๐๐๐ = ๐ ๐ ๐ ๐ป'๐ '๐ ร ๐ โก Token-Level Editing โข ๐ฌ ๐๐๐๐ โค ๐ ๐ ๐ ๐ป'๐ '๐ ร ๐ โข Avoiding Model Collapse ๐ท๐๐ก๐ ! Training ๐ ) Synthesizing ๐ท๐๐ก๐ " โฆ ๐ * ๐ + Training ๐ ) ๐ท๐๐ก๐ # ๐ท๐๐ก๐ ! Training ๐ + Training Editing ๐ * ๐ท๐๐ก๐ * + (1 -๐ * )๐ท๐๐ก๐ ) โฆ Source Real Data: ๐ท๐๐ก๐ ! Test Error ๐ธ "#$" Editing Operation Matrix M % Iterations ๐ โ 1, โฆ , ๐
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b9c79014-7486-468b-b92e-449a3253676dCited by top-tier papers4
- Escaping Model Collapse via Synthetic Data Verification: Near-term Improvements and Long-term ConvergenceBingji Yi, Qiyuan Liu, Yuwei Cheng, Haifeng XuICLR 2026 ยท 5 citations
- When Hard Negatives Hurt: Bridging the Generative Discriminative Gap in Hard Negative Synthesis for RetrievalZhicheng Zhang, Jiwei Tang, Kuicai Dong, Xiaopeng Li et al.KDD 2026 ยท 1 citation
- Data Pollination: An Emergent Ecological Process Driving AI Population EvolutionShufang Xie, Qizhi Pei, Ang Lv, Jingyang Hu et al.ACL 2026
- Machine-generated text detection prevents language model collapseGeorge Drayson, Emine Yilmaz, Vasileios LamposEMNLP 2025
Builds on26
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 ยท 22,562 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 ยท 13,123 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 ยท 1,016 citations
- The Flan Collection: Designing Data and Methods for Effective Instruction TuningShayne Longpre, Le Hou, Tu Vu, Albert Webson et al.ICML 2023 ยท 908 citations
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu et al.ACL 2023 ยท 540 citations
Related papers
- A Tale of Tails: Model Collapse as a Change of Scaling LawsElvis Dohmatob, Yunzhen Feng, Pu Yang, Franรงois Charton et al.ICML 2024 ยท 123 citations
- Demystifying Synthetic Data in LLM Pre-training: A Systematic Study of Scaling Laws, Benefits, and PitfallsFeiyang Kang, Newsha Ardalani, Michael Kuchnik, Youssef Emad et al.EMNLP 2025
- Collapse or Thrive: Perils and Promises of Synthetic Data in a Self-Generating WorldJoshua Kazdan, Rylan Schaeffer, Apratim Dey, Matthias Gerstgrasser et al.ICML 2025
- Strong Model CollapseElvis Dohmatob, Yunzhen Feng, Arjun Subramonian, Julia KempeICLR 2025 ยท 3 citations
- Beyond Model Collapse: Scaling Up with Synthesized Data Requires VerificationYunzhen Feng, Elvis Dohmatob, Pu Yang, Franรงois Charton et al.ICLR 2025 ยท 6 citations
