What happens when generative AI models train recursively on each others' outputs?
Hung Anh Vu, Galen Reeves, Emily Wenger
摘要
The internet serves as a common source of training data for generative AI (genAI) models but is increasingly populated with AI-generated content. This duality raises the possibility that future genAI models may be trained on other models' generated outputs. Prior work has studied consequences of models training on their own generated outputs, but limited work has considered what happens if models ingest content produced by other models. Given society's increasing dependence on genAI tools, understanding such data-mediated model interactions is critical. This work provides empirical evidence for how data-mediated interactions might unfold in practice, develops a theoretical model for this interactive training process, and experimentally validates the theory. We find that data-mediated interactions can benefit models by exposing them to novel concepts perhaps missed in original training data, but also can homogenize their performance on shared tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong 等ICML 2022 · 被引用 1,173 次
- Self-Consuming Generative Models Go MADSina Alemohammad, Josue Casco-Rodriguez, Lorenzo Luzi, Ahmed Imtiaz Humayun 等ICLR 2024 · 被引用 279 次
- Model Collapse Demystified: The Case of RegressionElvis Dohmatob, Yunzhen Feng, Julia KempeNeurIPS 2024 · 被引用 96 次
- Will Large-scale Generative Models Corrupt Future Datasets?Ryuichiro Hataya, Han Bao, Hiromi AraiICCV 2023 · 被引用 76 次
相关 Paper
- Would Deep Generative Models Amplify Bias in Future Models?Tianwei Chen, Yusuke Hirota, Mayu Otani, Noa Garcia 等CVPR 2024
- On the Stability of Iterative Retraining of Generative Models on their own DataQuentin Bertrand, Avishek Joey Bose, Alexandre Duplessis, Marco Jiralerspong 等ICLR 2024 · 被引用 93 次
- Exploring the Escalation of Source Bias in User, Data, and Recommender System Feedback LoopYuqi Zhou, Sunhao Dai, Liang Pang, Gang Wang 等SIGIR 2025 · 被引用 2 次
- Human vs. Generative AI in Content Creation Competition: Symbiosis or Conflict?Fan Yao, Chuanhao Li, Denis Nekipelov, Hongning Wang 等ICML 2024 · 被引用 31 次
- A Tale of Tails: Model Collapse as a Change of Scaling LawsElvis Dohmatob, Yunzhen Feng, Pu Yang, François Charton 等ICML 2024 · 被引用 123 次
