Larger Datasets Can Be Repeated More: A Theoretical Analysis of Multi-Epoch Scaling in Linear Regression
Tingkai Yan, Haodong Wen, Binghui Li, Kairong Luo, Wenguang Chen, Kaifeng Lyu
Abstract
While data scaling laws of large language models (LLMs) have been widely examined in the one-pass regime with massive corpora, their form under limited data and repeated epochs remains largely unexplored. This paper presents a theoretical analysis of how a common workaround, training for multiple epochs on the same dataset, reshapes the data scaling laws in linear regression. Concretely, we ask: to match the performance of training on a dataset of size for epochs, how much larger must a dataset be if the model is trained for only one pass? We quantify this using the of the data, , which we define as the multiplicative factor by which the dataset must grow under one-pass training to achieve the same test loss as -epoch training. Our analysis precisely characterizes the scaling behavior of for SGD in linear regression under either strong convexity or Zipf-distributed data: (1) When is small, we prove that , indicating that every new epoch yields a linear gain; (2) As increases, plateaus at a problem-dependent value that grows with ( for the strongly-convex case), implying that larger datasets can be repeated more times before the marginal benefit vanishes. These theoretical findings point out a neglected factor in a recent empirical study by Muennighoff et al. (2023), which claimed that training LLMs for up to epochs results in negligible loss differences compared to using fresh data at each step, , for in our notation. Supported by further empirical validation with LLMs, our results reveal that the maximum value for which in fact depends on the data size and distribution, and underscore the need to explicitly model both factors in future studies of scaling laws with data reuse.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f942dcba-05de-44c9-9ab4-bf70fd67e77dCited by top-tier papers5
- Muon in Associative Memory Learning: Training Dynamics and Scaling LawsKaifei Wang, Binghui Li, Han Zhong, Pinyan Lu et al.ICML 2026 · 7 citations
- Fast Catch-Up, Late Switching: Optimal Batch Size Scheduling via Functional Scaling LawsJinbo Wang, Binghui Li, Zhanpeng Zhou, Mingze Wang et al.ICLR 2026 · 6 citations
- Full-Batch Gradient Descent Outperforms One-Pass SGD: Sample Complexity Separation in Single-Index LearningFilip Kovačević, Hong Chang Ji, Denny Wu, Mahdi Soltanolkotabi et al.ICML 2026 · 2 citations
- Scaling Laws for Precision in High-Dimensional Linear RegressionDechen Zhang, Xuan Tang, Yingyu Liang, Difan ZouICML 2026 · 2 citations
- Less Data, Faster Training: repeating smaller datasets speeds up learning via sampling biasesJingwen Liu, Ezra Edelman, Surbhi Goel, Bingbin LiuICML 2026
Builds on29
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Deduplicating Training Data Makes Language Models BetterKatherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang et al.ACL 2022 · 844 citations
- Scaling Vision TransformersXiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, Lucas BeyerCVPR 2022 · 767 citations
- Scaling Data-Constrained Language ModelsNiklas Muennighoff, Alexander M. Rush, Boaz Barak, Teven Le Scao et al.NeurIPS 2023 · 475 citations
- The Quantization Model of Neural ScalingEric J. Michaud, Ziming Liu, Uzay Girit, Max TegmarkNeurIPS 2023 · 179 citations
Related papers
- Improved Scaling Laws in Linear Regression via Data ReuseLicong Lin, Jingfeng Wu, Peter L. BartlettNeurIPS 2025 · 7 citations
- To Repeat or Not To Repeat: Insights from Scaling LLM under Token-CrisisFuzhao Xue, Yao Fu, Wangchunshu Zhou, Zangwei Zheng et al.NeurIPS 2023 · 149 citations
- Data Efficient Neural Scaling Law via Model ReusingPeihao Wang, Rameswar Panda, Zhangyang WangICML 2023 · 18 citations
- When Data Is Scarce: Scaling Sparse Language Models with Repeated TrainingBoqian Wu, Qiao Xiao, Patrik Okanovic, Tomasz Sternal et al.ICML 2026
- Functional Scaling Laws in Kernel Regression: Loss Dynamics and Learning Rate SchedulesBinghui Li, Fengling Chen, Zixun Huang, Lean Wang et al.NeurIPS 2025 · 15 citations
