When Models Don't Collapse: On the Consistency of Iterative MLE
Daniel Barzilai, Ohad Shamir
Abstract
The widespread use of generative models has created a feedback loop, in which each generation of models is trained on data partially produced by its predecessors. This process has raised concerns about model collapse: A critical degradation in performance caused by repeated training on synthetic data. However, different analyses in the literature have reached different conclusions as to the severity of model collapse. As such, it remains unclear how concerning this phenomenon is, and under which assumptions it can be avoided. To address this, we theoretically study model collapse for maximum likelihood estimation (MLE), in a natural setting where synthetic data is gradually added to the original data set. Under standard assumptions (similar to those long used for proving asymptotic consistency and normality of MLE), we establish non-asymptotic bounds showing that collapse can be avoided even as the fraction of real data vanishes. On the other hand, we prove that some assumptions (beyond MLE consistency) are indeed necessary: Without them, model collapse can occur arbitrarily quickly, even when the original data is still present in the training set. To the best of our knowledge, these are the first rigorous examples of iterative generative modeling with accumulating data that rapidly leads to model collapse.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8b3d1b4a-2d3e-4830-98d4-e01cd47a2532Cited by top-tier papers2
- Language Generation with Replay: A Learning-Theoretic View of Model CollapseGiorgio Racca, Michal Valko, Amartya SanyalICML 2026 · 4 citations
- Quantifying Error Propagation and Model Collapse in Diffusion ModelsNaïl B. Khelifa, Richard E Turner, Ramji VenkataramananICML 2026 · 3 citations
Builds on14
- Self-Consuming Generative Models Go MADSina Alemohammad, Josue Casco-Rodriguez, Lorenzo Luzi, Ahmed Imtiaz Humayun et al.ICLR 2024 · 279 citations
- A Tale of Tails: Model Collapse as a Change of Scaling LawsElvis Dohmatob, Yunzhen Feng, Pu Yang, François Charton et al.ICML 2024 · 123 citations
- Model Collapse Demystified: The Case of RegressionElvis Dohmatob, Yunzhen Feng, Julia KempeNeurIPS 2024 · 96 citations
- On the Stability of Iterative Retraining of Generative Models on their own DataQuentin Bertrand, Avishek Joey Bose, Alexandre Duplessis, Marco Jiralerspong et al.ICLR 2024 · 93 citations
- Will Large-scale Generative Models Corrupt Future Datasets?Ryuichiro Hataya, Han Bao, Hiromi AraiICCV 2023 · 76 citations
Related papers
- Collapse or Thrive: Perils and Promises of Synthetic Data in a Self-Generating WorldJoshua Kazdan, Rylan Schaeffer, Apratim Dey, Matthias Gerstgrasser et al.ICML 2025
- Self-Verification Provably Prevents Model Collapse in Recursive Synthetic TrainingShi Fu, Yingjie Wang, Yuzhu Chen, Li Shen et al.NeurIPS 2025 · 5 citations
- A Theoretical Perspective: How to Prevent Model Collapse in Self-consuming Training LoopsShi Fu, Yingjie Wang, Yuzhu Chen, Xinmei Tian et al.ICLR 2025
- Stabilizing Self-Consuming Diffusion Models with Latent Space FilteringZhongteng Cai, Yaxuan Wang, Yang Liu, Xueru ZhangAAAI 2026 · 2 citations
- Preventing Model Collapse Under Overparametrization: Optimal Mixing Ratios for Interpolation Learning and Ridge RegressionAnvit Garg, Sohom Bhattacharya, Pragya SurICLR 2026 · 9 citations
