ParamΔ for Direct Mixing: Post-Train Large Language Model At Zero Cost
Sheng Cao, Mingrui Wu, Karthik Prasad, Yuandong Tian, Zechun Liu
摘要
The post-training phase of large language models is essential for enhancing capabilities such as instruction-following, reasoning, and alignment with human preferences. However, it demands extensive high-quality data and poses risks like overfitting, alongside significant computational costs due to repeated posttraining and evaluation after each base model update. This paper introduces Param∆, a novel method that streamlines post-training by transferring knowledge from an existing post-trained model to a newly updated base model with zero additional training. By computing the difference between post-trained model weights (Θ post ) and base model weights (Θ base ), and adding this to the updated base model (Θ ′ base ), we define Param∆ Model as: Θ Param∆ = Θ post -Θ base +Θ ′ base . This approach surprisingly equips the new base model with post-trained capabilities, achieving performance comparable to direct post-training. We did analysis on LLama3, Llama3.1, Qwen, and DeepSeek-distilled models. Results indicate Param∆ Model effectively replicates traditional post-training. For example, the Param∆ Model obtained from 70B Llama3-inst, Llama3-base, Llama3.1-base models attains approximately 95% of Llama3.1-inst model's performance on average. Param∆ brings a new perspective on how to fully leverage models in the open-weight community, where checkpoints for base and instruct models are readily available and frequently updated, by providing a cost-free framework to accelerate the iterative cycle of model development.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Knowledge is Not Enough: Injecting RL Skills for Continual AdaptationPingzhi Tang, Yiding Wang, Muhan ZhangACL 2026 · 被引用 2 次
- Fine-Tune Once, Reuse Across Models: Bayesian Task-Update Factors and ApproximationsSiyang Guo, Junbo Wang, Zibin ZhengICML 2026
它引用的顶会 Paper17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs 等ICML 2022 · 被引用 1,464 次
相关 Paper
- LLaMA Pro: Progressive LLaMA with Block ExpansionChengyue Wu, Yukang Gan, Yixiao Ge, Zeyu Lu 等ACL 2024
- Efficient Model Development through Fine-tuning TransferPin-Jie Lin, Rishab Balasubramanian, Fengyuan Liu, Nikhil Kandpal 等EMNLP 2025
- Steering Information Utility in Key-Value Memory for Language Model Post-TrainingChunyuan Deng, Ruidi Chang, Hanjie ChenNeurIPS 2025 · 被引用 2 次
- Composing Parameter-Efficient Modules with Arithmetic OperationJinghan Zhang, Shiqi Chen, Junteng Liu, Junxian HeNeurIPS 2023 · 被引用 164 次
- Model Extrapolation Expedites AlignmentChujie Zheng, Ziqi Wang, Heng Ji, Minlie Huang 等ACL 2025 · 被引用 34 次
