Scaling Law with Learning Rate Annealing
Howe Tissue, Venus Wang, Lu Wang
Abstract
We find that the cross-entropy loss curves of neural language models empirically adhere to a scaling law with learning rate (LR) annealing over training steps: where is the validation loss at step , is the area under the LR curve, is the LR annealing area, and , , , are constant parameters. This formulation takes into account two factors: (1) power-law scaling over data size, and (2) the additional loss reduction during LR annealing. Therefore, this formulation can describe the full loss curve at each step, rather than the single loss point at the end of training. Applying the scaling law with LR annealing and fitting only one or two training curves, we can accurately predict the loss at any given step across any learning rate scheduler (LRS). This approach significantly reduces computational cost in formulating scaling laws while providing more accuracy and expressiveness for training dynamics. Extensive experiments demonstrate that our findings hold across a range of hyper-parameters and model architectures, and our equation can extend to scaling effect of model sizes. Moreover, our formulation provides accurate theoretical verification and explanation for empirical results observed in numerous previous studies, particularly those focusing on LR schedule and annealing. We believe that this work is promising to enhance the understanding of LLM training dynamics while greatly democratizing scaling laws, and it can guide researchers in refining training strategies (e.g. critical LRS) for further LLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9ccbe53a-edf5-4a3a-9a58-8d5b85b98de3Cited by top-tier papers19
- Functional Scaling Laws in Kernel Regression: Loss Dynamics and Learning Rate SchedulesBinghui Li, Fengling Chen, Zixun Huang, Lean Wang et al.NeurIPS 2025 · 15 citations
- Pre-training LLM without Learning Rate Decay Enhances Supervised Fine-TuningKazuki Yano, Shun Kiyono, Sosuke Kobayashi, Sho Takase et al.ICLR 2026 · 13 citations
- Training Dynamics Impact Post-Training Quantization RobustnessAlbert Catalan-Tatjer, Niccolò Ajroldi, Jonas GeipingICLR 2026 · 13 citations
- How Learning Rate Decay Wastes Your Best Data in Curriculum-Based LLM PretrainingKairong Luo, Zhenbo Sun, Haodong Wen, Xinyu Shi et al.ICLR 2026 · 12 citations
- Larger Datasets Can Be Repeated More: A Theoretical Analysis of Multi-Epoch Scaling in Linear RegressionTingkai Yan, Haodong Wen, Binghui Li, Kairong Luo et al.ICLR 2026 · 12 citations
Builds on7
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- On the Variance of the Adaptive Learning Rate and BeyondLiyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen et al.ICLR 2020 · 2,210 citations
- Scaling Data-Constrained Language ModelsNiklas Muennighoff, Alexander M. Rush, Boaz Barak, Teven Le Scao et al.NeurIPS 2023 · 475 citations
- The Quantization Model of Neural ScalingEric J. Michaud, Ziming Liu, Uzay Girit, Max TegmarkNeurIPS 2023 · 179 citations
- Scaling Laws and Compute-Optimal Training Beyond Fixed Training DurationsAlexander Hägele, Elie Bakouch, Atli Kosson, Loubna Ben Allal et al.NeurIPS 2024 · 168 citations
Related papers
- A Multi-Power Law for Loss Curve Prediction Across Learning Rate SchedulesKairong Luo, Haodong Wen, Shengding Hu, Zhenbo Sun et al.ICLR 2025
- Learning Dynamics in Continual Pre-Training for Large Language ModelsXingjin Wang, Howe Tissue, Lu Wang, Linjing Li et al.ICML 2025
- What Scales in Cross-Entropy Scaling Law?Junxi Yan, Zixi Wei, Qingyao Ai, Yiqun Liu et al.ICLR 2026 · 1 citation
- Scaling with Collapse: Efficient and Predictable Training of LLM FamiliesShane Bergsma, Bin Claire Zhang, Nolan Simran Dey, Shaheer Muhammad et al.ICLR 2026 · 10 citations
- Temporal Scaling Law for Large Language ModelsYizhe Xiong, Xiansheng Chen, Xin Ye, Hui Chen et al.EMNLP 2025
