When Sample Selection Bias Precipitates Model Collapse
Xinbao Qiao, Xianglong Du, Wei Liu, Jingqi Zhang, Peihua Mai, Meng Zhang, Yan (James) Pang
Abstract
The proliferation of recursive training on synthetic data can alleviate data scarcity but risks model collapse, where repeated training erodes distributional tails and homogenizes outputs. Data selection is widely viewed as a remedy, yet its reliability depends critically on the reference distribution used by the verifier. We show that in low-resource verification regimes, where each verifier observes only a small, fragmented, and biased slice of the target manifold, selection itself becomes biased. This situation naturally arises in low-resource data silos such as healthcare consortia or proprietary financial institutions, where raw data cannot be pooled and local references are inherently incomplete. As a result, selection preferentially retains samples aligned with the local manifold while pruning globally relevant tail modes, turning from a safeguard against collapse into a mechanism that precipitates it. We theoretically prove that such siloed selection accelerates collapse and induces power-law diversity decay. As an initial mitigation, we construct Wasserstein proxy references from multiple silos without sharing raw data. Empirical results confirm that local-reference selection fails on skewed distributions, whereas collaborative proxy references mitigate diversity degradation, suggesting that recursive synthetic-data pipelines require particular caution when real-data coverage is fragmented or scarce.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9b474935-e5b4-403c-bd01-6bee770d34f9Builds on36
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Beyond neural scaling laws: beating power law scaling via data pruningBen Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli et al.NeurIPS 2022 · 720 citations
- Self-Distillation Amplifies Regularization in Hilbert SpaceHossein Mobahi, Mehrdad Farajtabar, Peter L. BartlettNeurIPS 2020 · 298 citations
- Self-Consuming Generative Models Go MADSina Alemohammad, Josue Casco-Rodriguez, Lorenzo Luzi, Ahmed Imtiaz Humayun et al.ICLR 2024 · 279 citations
Related papers
- Self-Verification Provably Prevents Model Collapse in Recursive Synthetic TrainingShi Fu, Yingjie Wang, Yuzhu Chen, Li Shen et al.NeurIPS 2025 · 5 citations
- Recursive Training Loops in LLMs: How training data properties modulate distribution shift in generated data?Grgur Kovac, Jérémy Perez, Rémy Portelas, Peter Ford Dominey et al.EMNLP 2025
- Unravelling in Collaborative LearningAymeric Capitaine, Etienne Boursier, Antoine Scheid, Eric Moulines et al.NeurIPS 2024 · 8 citations
- Beyond Model Collapse: Scaling Up with Synthesized Data Requires VerificationYunzhen Feng, Elvis Dohmatob, Pu Yang, François Charton et al.ICLR 2025 · 6 citations
- Escaping Model Collapse via Synthetic Data Verification: Near-term Improvements and Long-term ConvergenceBingji Yi, Qiyuan Liu, Yuwei Cheng, Haifeng XuICLR 2026 · 5 citations
