Velocitune: A Velocity-based Dynamic Domain Reweighting Method for Continual Pre-training
Zheheng Luo, Xin Zhang, Xiao Liu, Haoling Li, Yeyun Gong, Qi Chen, Peng Cheng
摘要
It is well-known that a diverse corpus is critical for training large language models, which are typically constructed from a mixture of various domains. In general, previous efforts resort to either sampling training data from different domains with static proportions or dynamically adjusting these proportions during training to optimise pretraining performance. However, few methods addressed the complexity of domain-adaptive continual pre-training. To fill this gap, we propose Velocitune, a novel framework that dynamically assesses learning velocity and adjusts data proportions accordingly, favouring slower learning domains while de-emphasising faster learning ones, which is guided by a scaling law to estimate the desired learning goal for each domain with a less associated cost. To evaluate the effectiveness of Velocitune, we conduct experiments on a dataset focused on reasoning tasks with CodeLlama, as well as on a corpus of system commands using Llama3 and Mistral. Velocitune achieves performance gains in both math and code reasoning tasks and command-line generation benchmarks. Further analysis reveals that key factors driving the effectiveness of Velocitune include target estimation and data ordering.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- RuleReasoner: Reinforced Rule-based Reasoning via Domain-aware Dynamic SamplingYang Liu, Jiaqi Li, Zilong ZhengICLR 2026 · 被引用 8 次
- Data Mixing Agent: Learning to Re-weight Domains for Continual Pre-trainingKailai Yang, Xiao Liu, Lei Ji, Hao Li 等ACL 2026 · 被引用 3 次
- DIDS: Domain Impact-aware Data Sampling for Large Language Model TrainingWeijie Shi, Jipeng Zhang, Yaguang Wu, Jingzhi Fang 等EMNLP 2025 · 被引用 1 次
- Enhancing Large Language Model Performance with Gradient-Based Parameter SelectionHaoling Li, Xin Zhang, Xiao Liu, Yeyun Gong 等AAAI 2025
- HanjaBridge: Resolving Semantic Ambiguity in Korean LLMs via Hanja-Augmented Pre-TrainingSeungho Choi, Sihyun Park, Minsang Kim, Chansol Park 等AAAI 2026
它引用的顶会 Paper9
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer 等NeurIPS 2022 · 被引用 2,039 次
- GLaM: Efficient Scaling of Language Models with Mixture-of-ExpertsNan Du, Yanping Huang, Andrew M. Dai, Simon Tong 等ICML 2022 · 被引用 1,173 次
- DoReMi: Optimizing Data Mixtures Speeds Up Language Model PretrainingSang Michael Xie, Hieu Pham, Xuanyi Dong, Nan Du 等NeurIPS 2023 · 被引用 457 次
- Sheared LLaMA: Accelerating Language Model Pre-training via Structured PruningMengzhou Xia, Tianyu Gao, Zhiyuan Zeng, Danqi ChenICLR 2024 · 被引用 453 次
相关 Paper
- D-CPT Law: Domain-specific Continual Pre-Training Scaling Law for Large Language ModelsHaoran Que, Jiaheng Liu, Ge Zhang, Chenchen Zhang 等NeurIPS 2024 · 被引用 47 次
- Programming Every Example: Lifting Pre-training Data Quality Like Experts at ScaleFan Zhou, Zengzhi Wang, Qian Liu, Junlong Li 等ICML 2025
- Diversity as a Reward: Fine-Tuning LLMs on a Mixture of Domain-Undetermined DataZhenqing Ling, Daoyuan Chen, Liuyi Yao, Qianli Shen 等NeurIPS 2025 · 被引用 14 次
- Efficient Domain Continual pretraining by Mitigating the Stability GapYiduo Guo, Jie Fu, Huishuai Zhang, Dongyan ZhaoACL 2025
- DRPruning: Efficient Large Language Model Pruning through Distributionally Robust OptimizationHexuan Deng, Wenxiang Jiao, Xuebo Liu, Jing Li 等ACL 2025
