A Closer Look at Model Collapse: From a Generalization-to-Memorization Perspective
Lianghe Shi, Meng Wu, Huijie Zhang, Zekai Zhang, Molei Tao, Qing Qu
摘要
The widespread use of diffusion models has led to an abundance of AI-generated data, raising concerns about model collapse-a phenomenon in which recursive iterations of training on synthetic data lead to performance degradation. Prior work primarily characterizes this collapse via variance shrinkage or distribution shift, but these perspectives miss practical manifestations of model collapse. This paper identifies a transition from generalization to memorization during model collapse in diffusion models, where models increasingly replicate training data instead of generating novel content during iterative training on synthetic samples. This transition is directly driven by the declining entropy of the synthetic training data produced in each training cycle, which serves as a clear indicator of model degradation. Motivated by this insight, we propose an entropy-based data selection strategy to mitigate the transition from generalization to memorization and alleviate model collapse. Empirical results show that our approach significantly enhances visual quality and diversity in recursive generation, effectively preventing collapse. The source code is available at https://github.com/shilianghe007/Model_Collapse.git
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Generalization of Diffusion Models Arises with a Balanced Representation SpaceZekai Zhang, Xiao Li, Xiang Li, Lianghe Shi 等ICLR 2026 · 被引用 14 次
- Quantifying Error Propagation and Model Collapse in Diffusion ModelsNaïl B. Khelifa, Richard E Turner, Ramji VenkataramananICML 2026 · 被引用 3 次
- Evaluating the Representation Space of Diffusion Models via Self-Supervised PrinciplesXiao Li, Yixuan Jia, Zekai Zhang, Xiang Li 等ICML 2026
- When Sample Selection Bias Precipitates Model CollapseXinbao Qiao, Xianglong Du, Wei Liu, Jingqi Zhang 等ICML 2026
- Diffusion Models Preferentially Memorize Prototypical Examples or: Why Does My Diffusion Model Love Slop?Marta Aparicio Rodriguez, Anastasia Borovykh, Grigorios A Pavliotis, Daniel KorchinskiICML 2026
它引用的顶会 Paper34
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Beyond neural scaling laws: beating power law scaling via data pruningBen Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli 等NeurIPS 2022 · 被引用 720 次
- Discrete Diffusion Modeling by Estimating the Ratios of the Data DistributionAaron Lou, Chenlin Meng, Stefano ErmonICML 2024 · 被引用 473 次
- How Faithful is your Synthetic Data? Sample-level Metrics for Evaluating and Auditing Generative ModelsAhmed M. Alaa, Boris van Breugel, Evgeny S. Saveliev, Mihaela van der SchaarICML 2022 · 被引用 287 次
- Self-Consuming Generative Models Go MADSina Alemohammad, Josue Casco-Rodriguez, Lorenzo Luzi, Ahmed Imtiaz Humayun 等ICLR 2024 · 被引用 279 次
相关 Paper
- Stabilizing Self-Consuming Diffusion Models with Latent Space FilteringZhongteng Cai, Yaxuan Wang, Yang Liu, Xueru ZhangAAAI 2026 · 被引用 2 次
- Model Collapse Demystified: The Case of RegressionElvis Dohmatob, Yunzhen Feng, Julia KempeNeurIPS 2024 · 被引用 96 次
- Understanding Hallucinations in Diffusion Models through Mode InterpolationSumukh K. Aithal, Pratyush Maini, Zachary C. Lipton, J. Zico KolterNeurIPS 2024 · 被引用 121 次
- On the Edge of Memorization in Diffusion ModelsSam Buchanan, Druv Pai, Yi Ma, Valentin De BortoliNeurIPS 2025 · 被引用 25 次
- When Pretty Isn't Useful: Investigating Why Modern Text-to-Image Models Fail as Reliable Training Data GeneratorsKrzysztof Adamkiewicz, Brian B. Moser, Stanislav Frolov, Tobias Christian Nauen 等CVPR 2026 · 被引用 8 次
