Escaping Collapse: The Strength of Weak Data for Large Language Model Training
Kareem Amin, Sara Babakniya, Alex Bie, Weiwei Kong, Umar Syed, Sergei Vassilvitskii
摘要
Synthetically-generated data plays an increasingly larger role in training large language models. However, while synthetic data has been found to be useful, studies have also shown that without proper curation it can cause LLM performance to plateau, or even"collapse", after many training iterations. In this paper, we formalize this question and develop a theoretical framework to investigate how much curation is needed in order to ensure that LLM performance continually improves. Our analysis is inspired by boosting, a classic machine learning technique that leverages a very weak learning algorithm to produce an arbitrarily good classifier. The approach we analyze subsumes many recently proposed methods for training LLMs on synthetic data, and thus our analysis sheds light on why they are successful, and also suggests opportunities for future improvement. We present experiments that validate our theory, and show that dynamically focusing labeling resources on the most challenging examples -- in much the same way that boosting focuses the efforts of the weak learner -- leads to improved performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- When Models Don't Collapse: On the Consistency of Iterative MLEDaniel Barzilai, Ohad ShamirNeurIPS 2025 · 被引用 10 次
- Escaping Model Collapse via Synthetic Data Verification: Near-term Improvements and Long-term ConvergenceBingji Yi, Qiyuan Liu, Yuwei Cheng, Haifeng XuICLR 2026 · 被引用 5 次
- From Selection to Refinement: Iterative Optimization for Instruction DataHang Hu, Ziyan Liu, Rujie Wen, Ruihui Hou 等ACL 2026
它引用的顶会 Paper19
- STaR: Bootstrapping Reasoning With ReasoningEric Zelikman, Yuhuai Wu, Jesse Mu, Noah D. GoodmanNeurIPS 2022 · 被引用 1,126 次
- Self-Rewarding Language ModelsWeizhe Yuan, Richard Yuanzhe Pang, Kyunghyun Cho, Xian Li 等ICML 2024 · 被引用 569 次
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu 等ACL 2023 · 被引用 540 次
- Self-Play Fine-Tuning Converts Weak Language Models to Strong Language ModelsZixiang Chen, Yihe Deng, Huizhuo Yuan, Kaixuan Ji 等ICML 2024 · 被引用 527 次
- Self-Consuming Generative Models Go MADSina Alemohammad, Josue Casco-Rodriguez, Lorenzo Luzi, Ahmed Imtiaz Humayun 等ICLR 2024 · 被引用 279 次
相关 Paper
- Curated LLM: Synergy of LLMs and Data Curation for tabular augmentation in low-data regimesNabeel Seedat, Nicolas Huynh, Boris van Breugel, Mihaela van der SchaarICML 2024 · 被引用 61 次
- Strong Model CollapseElvis Dohmatob, Yunzhen Feng, Arjun Subramonian, Julia KempeICLR 2025 · 被引用 3 次
- Theoretical Modeling of Large Language Model Self-Improvement Training Dynamics Through Solver-Verifier GapYifan Sun, Yushan Liang, Zhen Zhang, Xin Liu 等ICLR 2026 · 被引用 2 次
- Learning from Evolving Training Dynamics: An Entropy-Maximizing Data Curation Strategy for LLM Supervised Post-TrainingMengxiang Zhang, Lingyuan LiuACL 2026
- Language Models are Weak LearnersHariharan Manikandan, Yiding Jiang, J. Zico KolterNeurIPS 2023 · 被引用 32 次
