Larger Datasets Can Be Repeated More: A Theoretical Analysis of Multi-Epoch Scaling in Linear Regression
Tingkai Yan, Haodong Wen, Binghui Li, Kairong Luo, Wenguang Chen, Kaifeng Lyu
摘要
While data scaling laws of large language models (LLMs) have been widely examined in the one-pass regime with massive corpora, their form under limited data and repeated epochs remains largely unexplored. This paper presents a theoretical analysis of how a common workaround, training for multiple epochs on the same dataset, reshapes the data scaling laws in linear regression. Concretely, we ask: to match the performance of training on a dataset of size for epochs, how much larger must a dataset be if the model is trained for only one pass? We quantify this using the of the data, , which we define as the multiplicative factor by which the dataset must grow under one-pass training to achieve the same test loss as -epoch training. Our analysis precisely characterizes the scaling behavior of for SGD in linear regression under either strong convexity or Zipf-distributed data: (1) When is small, we prove that , indicating that every new epoch yields a linear gain; (2) As increases, plateaus at a problem-dependent value that grows with ( for the strongly-convex case), implying that larger datasets can be repeated more times before the marginal benefit vanishes. These theoretical findings point out a neglected factor in a recent empirical study by Muennighoff et al. (2023), which claimed that training LLMs for up to epochs results in negligible loss differences compared to using fresh data at each step, , for in our notation. Supported by further empirical validation with LLMs, our results reveal that the maximum value for which in fact depends on the data size and distribution, and underscore the need to explicitly model both factors in future studies of scaling laws with data reuse.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Muon in Associative Memory Learning: Training Dynamics and Scaling LawsKaifei Wang, Binghui Li, Han Zhong, Pinyan Lu 等ICML 2026 · 被引用 7 次
- Fast Catch-Up, Late Switching: Optimal Batch Size Scheduling via Functional Scaling LawsJinbo Wang, Binghui Li, Zhanpeng Zhou, Mingze Wang 等ICLR 2026 · 被引用 6 次
- Full-Batch Gradient Descent Outperforms One-Pass SGD: Sample Complexity Separation in Single-Index LearningFilip Kovačević, Hong Chang Ji, Denny Wu, Mahdi Soltanolkotabi 等ICML 2026 · 被引用 2 次
- Scaling Laws for Precision in High-Dimensional Linear RegressionDechen Zhang, Xuan Tang, Yingyu Liang, Difan ZouICML 2026 · 被引用 2 次
- Less Data, Faster Training: repeating smaller datasets speeds up learning via sampling biasesJingwen Liu, Ezra Edelman, Surbhi Goel, Bingbin LiuICML 2026
它引用的顶会 Paper29
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Deduplicating Training Data Makes Language Models BetterKatherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang 等ACL 2022 · 被引用 844 次
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 被引用 767 次
- Scaling Data-Constrained Language ModelsNiklas Muennighoff, Alexander M. Rush, Boaz Barak, Teven Le Scao 等NeurIPS 2023 · 被引用 475 次
- The Quantization Model of Neural ScalingEric J. Michaud, Ziming Liu, Uzay Girit, Max TegmarkNeurIPS 2023 · 被引用 179 次
相关 Paper
- Improved Scaling Laws in Linear Regression via Data ReuseLicong Lin, Jingfeng Wu, Peter L. BartlettNeurIPS 2025 · 被引用 7 次
- To Repeat or Not To Repeat: Insights from Scaling LLM under Token-CrisisFuzhao Xue, Yao Fu, Wangchunshu Zhou, Zangwei Zheng 等NeurIPS 2023 · 被引用 149 次
- Data Efficient Neural Scaling Law via Model ReusingPeihao Wang, Rameswar Panda, Zhangyang WangICML 2023 · 被引用 18 次
- When Data Is Scarce: Scaling Sparse Language Models with Repeated TrainingBoqian Wu, Qiao Xiao, Patrik Okanovic, Tomasz Sternal 等ICML 2026
- Functional Scaling Laws in Kernel Regression: Loss Dynamics and Learning Rate SchedulesBinghui Li, Fengling Chen, Zixun Huang, Lean Wang 等NeurIPS 2025 · 被引用 15 次
