Preventing Model Collapse Under Overparametrization: Optimal Mixing Ratios for Interpolation Learning and Ridge Regression
Anvit Garg, Sohom Bhattacharya, Pragya Sur
摘要
Model collapse occurs when generative models degrade after repeatedly training on their own synthetic outputs. We study this effect in overparameterized linear regression in a setting where each iteration mixes fresh real labels with synthetic labels drawn from the model fitted in the previous iteration. We derive precise generalization error formulae for minimum--norm interpolation and ridge regression under this iterative scheme. Our analysis reveals intriguing properties of the optimal mixing weight that minimizes long-term prediction error and provably prevents model collapse. For instance, in the case of min--norm interpolation, we establish that the optimal real-data proportion converges to the reciprocal of the golden ratio for fairly general classes of covariate distributions. Previously, this property was known only for ordinary least squares, and additionally in low dimensions. For ridge regression, we further analyze two popular model classes -- the random-effects model and the spiked covariance model -- demonstrating how spectral geometry governs optimal weighting. In both cases, as well as for isotropic features, we uncover that the optimal mixing ratio should be at least one-half, reflecting the necessity of favoring real-data over synthetic. We study three additional settings: (i) where real data is fixed and fresh labels are not obtained at each iteration, (ii) where covariates vary across iterations but fresh real labels are available each time, and (iii) where covariates vary with time but only a fraction of them receive fresh real labels at each iteration. Across these diverse settings, we characterize when model collapse is inevitable and when synthetic data improves learning. We validate our theoretical results with extensive simulations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Escaping Model Collapse via Synthetic Data Verification: Near-term Improvements and Long-term ConvergenceBingji Yi, Qiyuan Liu, Yuwei Cheng, Haifeng XuICLR 2026 · 被引用 5 次
- Optimal Unconstrained Self-Distillation in Ridge Regression: Strict Improvements, Precise Asymptotics, and One-Shot TuningHien Dang, Pratik Patil, Alessandro RinaldoICML 2026 · 被引用 1 次
- Why Self-Distillation Helps and Hurts: Denoising vs. Signal ForgettingMingqi Wu, Archer Yang, Qiang SunICML 2026
- Asymptotic Theory of Iterated Empirical Risk Minimization, with Applications to Active LearningHugo Cui, Yue LuICML 2026
它引用的顶会 Paper16
- Self-Consuming Generative Models Go MADSina Alemohammad, Josue Casco-Rodriguez, Lorenzo Luzi, Ahmed Imtiaz Humayun 等ICLR 2024 · 被引用 279 次
- A Tale of Tails: Model Collapse as a Change of Scaling LawsElvis Dohmatob, Yunzhen Feng, Pu Yang, François Charton 等ICML 2024 · 被引用 123 次
- Understanding Double Descent Requires A Fine-Grained Bias-Variance DecompositionBen Adlam, Jeffrey PenningtonNeurIPS 2020 · 被引用 111 次
- Model Collapse Demystified: The Case of RegressionElvis Dohmatob, Yunzhen Feng, Julia KempeNeurIPS 2024 · 被引用 96 次
- On the Stability of Iterative Retraining of Generative Models on their own DataQuentin Bertrand, Avishek Joey Bose, Alexandre Duplessis, Marco Jiralerspong 等ICLR 2024 · 被引用 93 次
相关 Paper
- On Optimal Interpolation in Linear RegressionEduard Oravkin, Patrick RebeschiniNeurIPS 2021 · 被引用 6 次
- Risk Phase Transitions in Spiked Regression: Alignment Driven Benign and Catastrophic OverfittingJiping Li, Rishi SonthaliaICLR 2026 · 被引用 2 次
- When Models Don't Collapse: On the Consistency of Iterative MLEDaniel Barzilai, Ohad ShamirNeurIPS 2025 · 被引用 10 次
- On the Optimal Weighted Regularization in Overparameterized Linear RegressionDenny Wu, Ji XuNeurIPS 2020 · 被引用 151 次
- Implicit Regularization Leads to Benign Overfitting for Sparse Linear RegressionMo Zhou, Rong GeICML 2023 · 被引用 4 次
