Pre-training under infinite compute
Konwoo Kim, Suhas Kotha, Percy Liang, Tatsunori Hashimoto
摘要
Since compute grows much faster than web text available for language model pre-training, we ask how one should approach pre-training under fixed data and no compute constraints. We first show that existing data-constrained approaches of increasing epoch count and parameter count eventually overfit, and we significantly improve upon such recipes by properly tuning regularization, finding that the optimal weight decay is larger than standard practice. Since our regularized recipe monotonically decreases loss following a simple power law in parameter count, we estimate its best possible performance via the asymptote of its scaling law rather than the performance at a fixed compute budget. We then identify that ensembling independently trained models achieves a significantly lower loss asymptote than the regularized recipe. Our best intervention combining epoching, regularization, parameter scaling, and ensemble scaling achieves an asymptote at 200M tokens using less data than our baseline, and our data scaling laws predict that this improvement persists at higher token budgets. We find that our data efficiency gains can be realized at much smaller parameter counts as we can distill an ensemble into a student model that is 8 smaller and retains of the ensembling benefit. Finally, our interventions designed for validation loss generalize to downstream benchmarks, achieving a improvement for pre-training evals and a data efficiency improvement over continued pre-training on math mid-training data. Our results show that simple algorithmic improvements can enable significantly more data-efficient pre-training in a compute-rich future.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Maximum Likelihood Reinforcement LearningFahim Tajwar, Guanning Zeng, Yueer Zhou, Yuda Song 等ICML 2026 · 被引用 18 次
- Beyond Multi-Token Prediction: Pretraining LLMs with Future SummariesDivyat Mahajan, Sachin Goyal, Badr Youbi Idrissi, Mohammad Pezeshki 等ICLR 2026 · 被引用 15 次
- How Memory in Optimization Algorithms Implicitly Modifies the LossMatias D. Cattaneo, Boris ShigidaNeurIPS 2025 · 被引用 6 次
- Weight Decay Improves Language Model PlasticityTessa Han, Sebastian Bordt, Hanlin Zhang, Sham KakadeICML 2026 · 被引用 3 次
- Inner-layer Token Self-modulation as Another Scaling Axis for LLMsYebin Yang, Huaijin Wu, Jingtao Han, Yu Wang 等ICML 2026 · 被引用 1 次
它引用的顶会 Paper39
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs 等ICML 2022 · 被引用 1,464 次
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang 等ICLR 2020 · 被引用 1,108 次
- Be Your Own Teacher: Improve the Performance of Convolutional Neural Networks via Self DistillationLinfeng Zhang, Jiebo Song, Anni Gao, Jingwei Chen 等ICCV 2019 · 被引用 1,069 次
相关 Paper
- InfoLaw: Information Scaling Laws for Large Language Models with Quality-Weighted Mixture Data and RepetitionWeidong Zhou, Fengze Liu, LIU, Ping Guo 等ICML 2026 · 被引用 1 次
- Training compute-optimal transformer encoder modelsMegi Dervishi, Alexandre Allauzen, Gabriel Synnaeve, Yann LeCunEMNLP 2025 · 被引用 1 次
- Algorithmic progress in language modelsAnson Ho, Tamay Besiroglu, Ege Erdil, Zifan Carl Guo 等NeurIPS 2024 · 被引用 51 次
- To Repeat or Not To Repeat: Insights from Scaling LLM under Token-CrisisFuzhao Xue, Yao Fu, Wangchunshu Zhou, Zangwei Zheng 等NeurIPS 2023 · 被引用 149 次
- When Data Is Scarce: Scaling Sparse Language Models with Repeated TrainingBoqian Wu, Qiao Xiao, Patrik Okanovic, Tomasz Sternal 等ICML 2026
