A Learning Rate Path Switching Training Paradigm for Version Updates of Large Language Models
Zhihao Wang, Shiyu Liu, Jianheng Huang, Wang Zheng, Yixuan Liao, Xiaoxin Chen, Junfeng Yao, Jinsong Su
Abstract
Due to the continuous emergence of new data, version updates have become an indispensable requirement for Large Language Models (LLMs). The training paradigms for version updates of LLMs include pre-training from scratch (PTFS) and continual pre-training (CPT). Preliminary experiments demonstrate that PTFS achieves better pre-training performance, while CPT has lower training cost. Moreover, their performance and training cost gaps widen progressively with version updates. To investigate the underlying reasons for this phenomenon, we analyze the effect of learning rate adjustments during the two stages of CPT: preparing an initialization checkpoint and continual pre-training based on this checkpoint. We find that a large learning rate in the first stage and a complete learning rate decay process in the second stage are crucial for version updates of LLMs. Hence, we propose a learning rate path switching training paradigm. Our paradigm comprises one main path, where we pre-train a LLM with the maximal learning rate, and multiple branching paths, each of which corresponds to an update of the LLM with newly-added training data. Extensive experiments demonstrate the effectiveness and generalization of our paradigm. Particularly, when training four versions of LLMs, our paradigm reduces the total training cost to 58% compared to PTFS, while maintaining comparable pretraining performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 40912d1b-5205-47dd-b6a3-2d0f57b563f0Cited by top-tier papers3
- Pre-training LLM without Learning Rate Decay Enhances Supervised Fine-TuningKazuki Yano, Shun Kiyono, Sosuke Kobayashi, Sho Takase et al.ICLR 2026 · 13 citations
- Advancing SMoE for Continuous Domain Adaptation of MLLMs: Adaptive Router and Domain-Specific LossLiang Zhang, Ziyao Lu, Fandong Meng, Hui Li et al.ACL 2025 · 3 citations
- Learning Dynamics in Continual Pre-Training for Large Language ModelsXingjin Wang, Howe Tissue, Lu Wang, Linjing Li et al.ICML 2025
Builds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- GLM-130B: An Open Bilingual Pre-trained ModelAohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang et al.ICLR 2023 · 295 citations
- Scalable Language Model with Generalized Continual LearningBohao Peng, Zhuotao Tian, Shu Liu, Ming-Chang Yang et al.ICLR 2024 · 36 citations
- Confidence Based Bidirectional Global Context Aware Training Framework for Neural Machine TranslationChulun Zhou, Fandong Meng, Jie Zhou, Min Zhang et al.ACL 2022 · 20 citations
Related papers
- Breaking Language Barriers: Cross-Lingual Continual Pre-Training at ScaleWenzhen Zheng, Wenbo Pan, Xu Xu, Libo Qin et al.EMNLP 2024 · 3 citations
- Pretrained Language Model in Continual Learning: A Comparative StudyTongtong Wu, Massimo Caccia, Zhuang Li, Yuan-Fang Li et al.ICLR 2022 · 76 citations
- TiC-LM: A Web-Scale Benchmark for Time-Continual LLM PretrainingJeffrey Li, Mohammadreza Armandpour, Iman Mirzadeh, Sachin Mehta et al.ACL 2025
- ADEPT: Continual Pretraining via Adaptive Expansion and Dynamic Decoupled TuningJinyang Zhang, Yue Fang, Hongxin Ding, Weibin Liao et al.ICLR 2026 · 5 citations
- Reinforcement Learning on Pre-Training DataSiheng Li, Kejiao Li, Zenan Xu, Guanhua Huang et al.ACL 2026 · 11 citations
