A Theoretical Perspective: How to Prevent Model Collapse in Self-consuming Training Loops
Shi Fu, Yingjie Wang, Yuzhu Chen, Xinmei Tian, Dacheng Tao
摘要
High-quality data is essential for training large generative models, yet the vast reservoir of real data available online has become nearly depleted. Consequently, models increasingly generate their own data for further training, forming Self-consuming Training Loops (STLs). However, the empirical results have been strikingly inconsistent: some models degrade or even collapse, while others successfully avoid these failures, leaving a significant gap in theoretical understanding to explain this discrepancy. This paper introduces the intriguing notion of recursive stability and presents the first theoretical generalization analysis, revealing how both model architecture and the proportion between real and synthetic data influence the success of STLs. We further extend this analysis to transformers in in-context learning, showing that even a constant-sized proportion of real data ensures convergence, while also providing insights into optimal synthetic data sizing.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper9
- A Closer Look at Model Collapse: From a Generalization-to-Memorization PerspectiveLianghe Shi, Meng Wu, Huijie Zhang, Zekai Zhang 等NeurIPS 2025 · 被引用 22 次
- When Models Don't Collapse: On the Consistency of Iterative MLEDaniel Barzilai, Ohad ShamirNeurIPS 2025 · 被引用 10 次
- Self-Verification Provably Prevents Model Collapse in Recursive Synthetic TrainingShi Fu, Yingjie Wang, Yuzhu Chen, Li Shen 等NeurIPS 2025 · 被引用 5 次
- Escaping Model Collapse via Synthetic Data Verification: Near-term Improvements and Long-term ConvergenceBingji Yi, Qiyuan Liu, Yuwei Cheng, Haifeng XuICLR 2026 · 被引用 5 次
- Language Generation with Replay: A Learning-Theoretic View of Model CollapseGiorgio Racca, Michal Valko, Amartya SanyalICML 2026 · 被引用 4 次
它引用的顶会 Paper17
- Self-Consuming Generative Models Go MADSina Alemohammad, Josue Casco-Rodriguez, Lorenzo Luzi, Ahmed Imtiaz Humayun 等ICLR 2024 · 被引用 279 次
- Transformers as Algorithms: Generalization and Stability in In-context LearningYingcong Li, Muhammed Emrullah Ildiz, Dimitris Papailiopoulos, Samet OymakICML 2023 · 被引用 242 次
- Large Language Models Can Self-ImproveJiaxin Huang, Shixiang Gu, Le Hou, Yuexin Wu 等EMNLP 2023 · 被引用 184 次
- Fine-Grained Analysis of Stability and Generalization for Stochastic Gradient DescentYunwen Lei, Yiming YingICML 2020 · 被引用 165 次
- A Tale of Tails: Model Collapse as a Change of Scaling LawsElvis Dohmatob, Yunzhen Feng, Pu Yang, François Charton 等ICML 2024 · 被引用 123 次
相关 Paper
- On the Stability of Iterative Retraining of Generative Models on their own DataQuentin Bertrand, Avishek Joey Bose, Alexandre Duplessis, Marco Jiralerspong 等ICLR 2024 · 被引用 93 次
- Self-Correcting Self-Consuming Loops for Generative Model TrainingNate Gillman, Michael Freeman, Daksh Aggarwal, Chia-Hong Hsu 等ICML 2024 · 被引用 28 次
- Stabilizing Self-Consuming Diffusion Models with Latent Space FilteringZhongteng Cai, Yaxuan Wang, Yang Liu, Xueru ZhangAAAI 2026 · 被引用 2 次
- Towards Theoretical Understandings of Self-Consuming Generative ModelsShi Fu, Sen Zhang, Yingjie Wang, Xinmei Tian 等ICML 2024 · 被引用 26 次
- Observations and Remedies for Large Language Model Bias in Self-Consuming Performative LoopYaxuan Wang, Zhongteng Cai, Yujia Bao, Xueru Zhang 等ACL 2026 · 被引用 1 次
