Lune

NeurIPS2024顶会

Stacking Your Transformers: A Closer Look at Model Growth for Efficient LLM Pre-Training

Wenyu Du, Tongxu Luo, Zihan Qiu, Zeyu Huang, Yikang Shen, Reynold Cheng, Yike Guo, Jie Fu

2024年份
52被引次数
13顶会引用

摘要

LLMs are computationally expensive to pre-train due to their large scale. Model growth emerges as a promising approach by leveraging smaller models to accelerate the training of larger ones. However, the viability of these model growth methods in efficient LLM pre-training remains underexplored. This work identifies three critical O‾\underline{\textit{O}}bstacles: (O\textit{O}1) lack of comprehensive evaluation, (O\textit{O}2) untested viability for scaling, and (O\textit{O}3) lack of empirical guidelines. To tackle O\textit{O}1, we summarize existing approaches into four atomic growth operators and systematically evaluate them in a standardized LLM pre-training setting. Our findings reveal that a depthwise stacking operator, called GstackG_{\text{stack}}, exhibits remarkable acceleration in training, leading to decreased loss and improved overall performance on eight standard NLP benchmarks compared to strong baselines. Motivated by these promising results, we conduct extensive experiments to delve deeper into GstackG_{\text{stack}} to address O\textit{O}2 and O\textit{O}3. For O\textit{O}2 (untested scalability), our study shows that GstackG_{\text{stack}} is scalable and consistently performs well, with experiments up to 7B LLMs after growth and pre-training LLMs with 750B tokens. For example, compared to a conventionally trained 7B model using 300B tokens, our GstackG_{\text{stack}} model converges to the same loss with 194B tokens, resulting in a 54.6% speedup. We further address O\textit{O}3 (lack of empirical guidelines) by formalizing guidelines to determine growth timing and growth factor for GstackG_{\text{stack}}, making it practical in general LLM pre-training. We also provide in-depth discussions and comprehensive ablation studies of GstackG_{\text{stack}}. Our code and pre-trained model are available at https://llm-stacking.github.io.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper13

问问它们各自怎么用它

它引用的顶会 Paper20

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖