Preventing Model Collapse Under Overparametrization: Optimal Mixing Ratios for Interpolation Learning and Ridge Regression
Anvit Garg, Sohom Bhattacharya, Pragya Sur
Abstract
Model collapse occurs when generative models degrade after repeatedly training on their own synthetic outputs. We study this effect in overparameterized linear regression in a setting where each iteration mixes fresh real labels with synthetic labels drawn from the model fitted in the previous iteration. We derive precise generalization error formulae for minimum--norm interpolation and ridge regression under this iterative scheme. Our analysis reveals intriguing properties of the optimal mixing weight that minimizes long-term prediction error and provably prevents model collapse. For instance, in the case of min--norm interpolation, we establish that the optimal real-data proportion converges to the reciprocal of the golden ratio for fairly general classes of covariate distributions. Previously, this property was known only for ordinary least squares, and additionally in low dimensions. For ridge regression, we further analyze two popular model classes -- the random-effects model and the spiked covariance model -- demonstrating how spectral geometry governs optimal weighting. In both cases, as well as for isotropic features, we uncover that the optimal mixing ratio should be at least one-half, reflecting the necessity of favoring real-data over synthetic. We study three additional settings: (i) where real data is fixed and fresh labels are not obtained at each iteration, (ii) where covariates vary across iterations but fresh real labels are available each time, and (iii) where covariates vary with time but only a fraction of them receive fresh real labels at each iteration. Across these diverse settings, we characterize when model collapse is inevitable and when synthetic data improves learning. We validate our theoretical results with extensive simulations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ee3105d7-5155-425e-8bf0-dcf54122fa01Cited by top-tier papers4
- Escaping Model Collapse via Synthetic Data Verification: Near-term Improvements and Long-term ConvergenceBingji Yi, Qiyuan Liu, Yuwei Cheng, Haifeng XuICLR 2026 · 5 citations
- Optimal Unconstrained Self-Distillation in Ridge Regression: Strict Improvements, Precise Asymptotics, and One-Shot TuningHien Dang, Pratik Patil, Alessandro RinaldoICML 2026 · 1 citation
- Why Self-Distillation Helps and Hurts: Denoising vs. Signal ForgettingMingqi Wu, Archer Yang, Qiang SunICML 2026
- Asymptotic Theory of Iterated Empirical Risk Minimization, with Applications to Active LearningHugo Cui, Yue LuICML 2026
Builds on16
- Self-Consuming Generative Models Go MADSina Alemohammad, Josue Casco-Rodriguez, Lorenzo Luzi, Ahmed Imtiaz Humayun et al.ICLR 2024 · 279 citations
- A Tale of Tails: Model Collapse as a Change of Scaling LawsElvis Dohmatob, Yunzhen Feng, Pu Yang, François Charton et al.ICML 2024 · 123 citations
- Understanding Double Descent Requires A Fine-Grained Bias-Variance DecompositionBen Adlam, Jeffrey PenningtonNeurIPS 2020 · 111 citations
- Model Collapse Demystified: The Case of RegressionElvis Dohmatob, Yunzhen Feng, Julia KempeNeurIPS 2024 · 96 citations
- On the Stability of Iterative Retraining of Generative Models on their own DataQuentin Bertrand, Avishek Joey Bose, Alexandre Duplessis, Marco Jiralerspong et al.ICLR 2024 · 93 citations
Related papers
- On Optimal Interpolation in Linear RegressionEduard Oravkin, Patrick RebeschiniNeurIPS 2021 · 6 citations
- Risk Phase Transitions in Spiked Regression: Alignment Driven Benign and Catastrophic OverfittingJiping Li, Rishi SonthaliaICLR 2026 · 2 citations
- When Models Don't Collapse: On the Consistency of Iterative MLEDaniel Barzilai, Ohad ShamirNeurIPS 2025 · 10 citations
- On the Optimal Weighted Regularization in Overparameterized Linear RegressionDenny Wu, Ji XuNeurIPS 2020 · 151 citations
- Implicit Regularization Leads to Benign Overfitting for Sparse Linear RegressionMo Zhou, Rong GeICML 2023 · 4 citations
