Scaling Optimal LR Across Token Horizons
Johan Bjorck, Alon Benhaim, Vishrav Chaudhary, Furu Wei, Xia Song
摘要
State-of-the-art LLMs are powered by scaling -- scaling model size, dataset size and cluster size. It is economically infeasible to extensively tune hyperparameter for the largest runs. Instead, approximately optimal hyperparameters must be inferred or transferred from smaller experiments. Hyperparameter transfer across model sizes has been studied in Yang et al. However, hyperparameter transfer across dataset size -- or token horizon -- has not been studied yet. To remedy this we conduct a large scale empirical study on how optimal learning rate (LR) depends on token horizon in LLM training. We first demonstrate that the optimal LR changes significantly with token horizon -- longer training necessitates smaller LR. Secondly we demonstrate the the optimal LR follows a scaling law, and that the optimal LR for longer horizons can be accurately estimated from shorter horizons via such scaling laws. We also provide a rule-of-thumb for transferring LR across token horizons with zero overhead over current practices. Lastly we provide evidence that LLama-1 used too high LR, and estimate the performance hit from this. We thus argue that hyperparameter transfer across data size is an important and overlooked component of LLM training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Power Lines: Scaling laws for weight decay and batch size in LLM pre-trainingShane Bergsma, Nolan Dey, Gurpreet Gosal, Gavia Gray 等NeurIPS 2025 · 被引用 44 次
- Completed Hyperparameter Transfer across Modules, Width, Depth, Batch and DurationBruno Mlodozeniec, Pierre Ablin, Louis Béthune, Dan Busbridge 等ICLR 2026 · 被引用 24 次
- Gemstones: A Model Suite for Multi-Faceted Scaling LawsSean McLeish, John Kirchenbauer, David Yu Miller, Siddharth Singh 等NeurIPS 2025 · 被引用 20 次
- Functional Scaling Laws in Kernel Regression: Loss Dynamics and Learning Rate SchedulesBinghui Li, Fengling Chen, Zixun Huang, Lean Wang 等NeurIPS 2025 · 被引用 15 次
- Decoder-Hybrid-Decoder Architecture for Efficient Reasoning with Long GenerationLiliang Ren, Congcong Chen, Haoran Xu, Young Jin Kim 等NeurIPS 2025 · 被引用 12 次
它引用的顶会 Paper24
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Scaling Laws for Reward Model OveroptimizationLeo Gao, John Schulman, Jacob HiltonICML 2023 · 被引用 963 次
- Scaling Vision Transformers to 22 Billion ParametersMostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski 等ICML 2023 · 被引用 848 次
- Language Is Not All You Need: Aligning Perception with Language ModelsShaohan Huang, Li Dong, Wenhui Wang, Yaru Hao 等NeurIPS 2023 · 被引用 810 次
- Scaling Data-Constrained Language ModelsNiklas Muennighoff, Alexander M. Rush, Boaz Barak, Teven Le Scao 等NeurIPS 2023 · 被引用 475 次
相关 Paper
- LLMs on the Line: Data Determines Loss-to-Loss Scaling LawsPrasanna Mayilvahanan, Thaddäus Wiedemer, Sayak Mallick, Matthias Bethge 等ICML 2025
- Temporal Scaling Law for Large Language ModelsYizhe Xiong, Xiansheng Chen, Xin Ye, Hui Chen 等EMNLP 2025
- Weight Decay may matter more than µP for Learning Rate Transfer in PracticeAtli Kosson, Jeremy Welborn, Yang Liu, Martin Jaggi 等ICLR 2026 · 被引用 11 次
- Language models scale reliably with over-training and on downstream tasksSamir Yitzhak Gadre, Georgios Smyrnis, Vaishaal Shankar, Suchin Gururangan 等ICLR 2025 · 被引用 3 次
- Small Batch Size Training for Language Models: When Vanilla SGD Works, and Why Gradient Accumulation is WastefulMartin Marek, Sanae Lotfi, Aditya Somasundaram, Andrew Gordon Wilson 等NeurIPS 2025 · 被引用 46 次
