LLaMA Pro: Progressive LLaMA with Block Expansion
Chengyue Wu, Yukang Gan, Yixiao Ge, Zeyu Lu, Jiahao Wang, Ye Feng, Ying Shan, Ping Luo
Abstract
Humans generally acquire new skills without compromising the old; however, the opposite holds for Large Language Models (LLMs), e.g., from LLaMA to CodeLLaMA. To this end, we propose a new post-pretraining method for LLMs with an expansion of Transformer blocks. We tune the expanded blocks using only new corpus, efficiently and effectively improving the model's knowledge while mitigating forgetting. In this paper, we experiment on the corpus of code and math, yielding LLAMA PRO-8.3B, a versatile foundation model initialized from LLaMA2-7B, excelling in general tasks, programming, and mathematics. LLAMA PRO and its instruction-following counterpart (LLAMA PRO -INSTRUCT) achieve advanced performance among various benchmarks, demonstrating superiority over existing open models in the LLaMA family and the immense potential of reasoning and addressing diverse tasks as an intelligent agent. Our findings provide valuable insights into integrating natural and programming languages, laying a solid foundation for developing advanced language agents that operate effectively in various environments. of data, which poses a challenge to the democra-042 tization of LLM research. Consequently, another 043 line of research, known as domain-adaptive pre-044 training, focuses on post-pretraining with domain-045 specific corpora (Gururangan et al., 2020). These
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers30
- Stacking Your Transformers: A Closer Look at Model Growth for Efficient LLM Pre-TrainingWenyu Du, Tongxu Luo, Zihan Qiu, Zeyu Huang et al.NeurIPS 2024 · 52 citations
- Understand What LLM Needs: Dual Preference Alignment for Retrieval-Augmented GenerationGuanting Dong, Yutao Zhu, Chenghao Zhang, Zechen Wang et al.WWW 2025 · 44 citations
- Transformer Copilot: Learning from The Mistake Log in LLM Fine-tuningJiaru Zou, Yikun Ban, Zihao Li, Yunzhe Qi et al.NeurIPS 2025 · 29 citations
- From Seeing to Experiencing: Scaling Navigation Foundation Models with Reinforcement LearningHonglin He, Yukai Ma, Brad Squicciarini, Wayne Wu et al.ICLR 2026 · 19 citations
- Language Models as Science TutorsAlexis Chevalier, Jiayi Geng, Alexander Wettig, Howard Chen et al.ICML 2024 · 17 citations
Builds on9
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu et al.ICLR 2022 · 4,966 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- WizardCoder: Empowering Code Large Language Models with Evol-InstructZiyang Luo, Can Xu, Pu Zhao, Qingfeng Sun et al.ICLR 2024 · 945 citations
Related papers
- Programming Every Example: Lifting Pre-training Data Quality Like Experts at ScaleFan Zhou, Zengzhi Wang, Qian Liu, Junlong Li et al.ICML 2025
- Efficient Domain Continual pretraining by Mitigating the Stability GapYiduo Guo, Jie Fu, Huishuai Zhang, Dongyan ZhaoACL 2025
- ParamΔ for Direct Mixing: Post-Train Large Language Model At Zero CostSheng Cao, Mingrui Wu, Karthik Prasad, Yuandong Tian et al.ICLR 2025
- Rewriting Pre-Training Data Boosts LLM Performance in Math and CodeKazuki Fujii, Yukito Tajima, Sakae Mizuki, Masaki Kawamura et al.ICLR 2026 · 21 citations
- Adapting a Language Model While Preserving its General KnowledgeZixuan Ke, Yijia Shao, Haowei Lin, Hu Xu et al.EMNLP 2022 · 6 citations
