Stabilizing Self-Consuming Diffusion Models with Latent Space Filtering
Zhongteng Cai, Yaxuan Wang, Yang Liu, Xueru Zhang
Abstract
As synthetic data proliferates across the Internet, it is often reused to train successive generations of generative models. This creates a "self-consuming loop" that can lead to training instability or model collapse. Common strategies to address the issue---such as accumulating historical training data or injecting fresh real data---either increase computational cost or require expensive human annotation. In this paper, we empirically analyze the latent space dynamics of self-consuming diffusion models and observe that the low-dimensional structure of latent representations extracted from synthetic data degrade over generations. Based on this insight, we propose Latent Space Filtering (LSF), a novel approach that mitigates model collapse by filtering out less realistic synthetic data from mixed datasets. Theoretically, we present a framework that connects latent space degradation to empirical observations. Experimentally, we show that LSF consistently outperforms existing baselines across multiple real-world datasets, effectively mitigating model collapse without increasing training cost or relying on human annotation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 760e8364-97be-463d-8310-4e75b4c9cf70Cited by top-tier papers2
- Observations and Remedies for Large Language Model Bias in Self-Consuming Performative LoopYaxuan Wang, Zhongteng Cai, Yujia Bao, Xueru Zhang et al.ACL 2026 · 1 citation
- When and How Human Curation Backfires: Preference Alignment under Multi-Model Self-Consuming LoopYang Zhang, Xiukun Wei, Xueru ZhangICML 2026
Builds on15
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- Understanding Hallucinations in Diffusion Models through Mode InterpolationSumukh K. Aithal, Pratyush Maini, Zachary C. Lipton, J. Zico KolterNeurIPS 2024 · 121 citations
- On the Stability of Iterative Retraining of Generative Models on their own DataQuentin Bertrand, Avishek Joey Bose, Alexandre Duplessis, Marco Jiralerspong et al.ICLR 2024 · 93 citations
- Data Feedback Loops: Model-driven Amplification of Dataset BiasesRohan Taori, Tatsunori HashimotoICML 2023 · 67 citations
- Self-Correcting Self-Consuming Loops for Generative Model TrainingNate Gillman, Michael Freeman, Daksh Aggarwal, Chia-Hong Hsu et al.ICML 2024 · 28 citations
Related papers
- Self-Verification Provably Prevents Model Collapse in Recursive Synthetic TrainingShi Fu, Yingjie Wang, Yuzhu Chen, Li Shen et al.NeurIPS 2025 · 5 citations
- A Closer Look at Model Collapse: From a Generalization-to-Memorization PerspectiveLianghe Shi, Meng Wu, Huijie Zhang, Zekai Zhang et al.NeurIPS 2025 · 22 citations
- A Theoretical Perspective: How to Prevent Model Collapse in Self-consuming Training LoopsShi Fu, Yingjie Wang, Yuzhu Chen, Xinmei Tian et al.ICLR 2025
- Towards Theoretical Understandings of Self-Consuming Generative ModelsShi Fu, Sen Zhang, Yingjie Wang, Xinmei Tian et al.ICML 2024 · 26 citations
- Collapse or Thrive: Perils and Promises of Synthetic Data in a Self-Generating WorldJoshua Kazdan, Rylan Schaeffer, Apratim Dey, Matthias Gerstgrasser et al.ICML 2025
