How to Synthesize Text Data without Model Collapse?
Xuekai Zhu, Daixuan Cheng, Hengli Li, Kaiyan Zhang, Ermo Hua, Xingtai Lv, Ning Ding, Zhouhan Lin, Zilong Zheng, Bowen Zhou
摘要
Model collapse in synthetic data indicates that iterative training on self-generated data leads to a gradual decline in performance. With the proliferation of AI models, synthetic data will fundamentally reshape the web data ecosystem. Future GPT-n models will inevitably be trained on a blend of synthetic and human-produced data. In this paper, we focus on two questions: what is the impact of synthetic data on language model training, and how to synthesize data without model collapse? We first pre-train language models across different proportions of synthetic data, revealing a negative correlation between the proportion of synthetic data and model performance. We further conduct statistical analysis on synthetic data to uncover distributional shift phenomenon and over-concentration of n-gram features. Inspired by the above findings, we propose token editing on human-produced data to obtain semi-synthetic data. As a proof of concept, we theoretically demonstrate that token-level editing can prevent model collapse, as the test error is constrained by a finite upper bound. We conduct extensive experiments on pre-training from scratch, continual pre-training, and supervised fine-tuning. The results validate our theoretical proof that token-level editing improves model performance. How to Synthesize Text Data without Model Collapse? ① Model Collapse Setting ⇢ 𝑬 𝒕𝒆𝒔𝒕 = 𝝈 𝟐 𝒅 𝑻'𝒅'𝟏 × 𝒏 ② Token-Level Editing ⇢ 𝑬 𝒕𝒆𝒔𝒕 ≤ 𝝈 𝟐 𝒅 𝑻'𝒅'𝟏 × 𝟐 ⇢ Avoiding Model Collapse 𝐷𝑎𝑡𝑎 ! Training 𝑓 ) Synthesizing 𝐷𝑎𝑡𝑎 " … 𝑓 * 𝑓 + Training 𝑓 ) 𝐷𝑎𝑡𝑎 # 𝐷𝑎𝑡𝑎 ! Training 𝑓 + Training Editing 𝑀 * 𝐷𝑎𝑡𝑎 * + (1 -𝑀 * )𝐷𝑎𝑡𝑎 ) … Source Real Data: 𝐷𝑎𝑡𝑎 ! Test Error 𝐸 "#$" Editing Operation Matrix M % Iterations 𝑖 ∈ 1, … , 𝑛
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Escaping Model Collapse via Synthetic Data Verification: Near-term Improvements and Long-term ConvergenceBingji Yi, Qiyuan Liu, Yuwei Cheng, Haifeng XuICLR 2026 · 被引用 5 次
- When Hard Negatives Hurt: Bridging the Generative Discriminative Gap in Hard Negative Synthesis for RetrievalZhicheng Zhang, Jiwei Tang, Kuicai Dong, Xiaopeng Li 等KDD 2026 · 被引用 1 次
- Data Pollination: An Emergent Ecological Process Driving AI Population EvolutionShufang Xie, Qizhi Pei, Ang Lv, Jingyang Hu 等ACL 2026
- Machine-generated text detection prevents language model collapseGeorge Drayson, Emine Yilmaz, Vasileios LamposEMNLP 2025
它引用的顶会 Paper26
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- The Flan Collection: Designing Data and Methods for Effective Instruction TuningShayne Longpre, Le Hou, Tu Vu, Albert Webson 等ICML 2023 · 被引用 908 次
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu 等ACL 2023 · 被引用 540 次
相关 Paper
- A Tale of Tails: Model Collapse as a Change of Scaling LawsElvis Dohmatob, Yunzhen Feng, Pu Yang, François Charton 等ICML 2024 · 被引用 123 次
- Demystifying Synthetic Data in LLM Pre-training: A Systematic Study of Scaling Laws, Benefits, and PitfallsFeiyang Kang, Newsha Ardalani, Michael Kuchnik, Youssef Emad 等EMNLP 2025
- Collapse or Thrive: Perils and Promises of Synthetic Data in a Self-Generating WorldJoshua Kazdan, Rylan Schaeffer, Apratim Dey, Matthias Gerstgrasser 等ICML 2025
- Strong Model CollapseElvis Dohmatob, Yunzhen Feng, Arjun Subramonian, Julia KempeICLR 2025 · 被引用 3 次
- Beyond Model Collapse: Scaling Up with Synthesized Data Requires VerificationYunzhen Feng, Elvis Dohmatob, Pu Yang, François Charton 等ICLR 2025 · 被引用 6 次
