ParamΔ for Direct Mixing: Post-Train Large Language Model At Zero Cost
Sheng Cao, Mingrui Wu, Karthik Prasad, Yuandong Tian, Zechun Liu
Abstract
The post-training phase of large language models is essential for enhancing capabilities such as instruction-following, reasoning, and alignment with human preferences. However, it demands extensive high-quality data and poses risks like overfitting, alongside significant computational costs due to repeated posttraining and evaluation after each base model update. This paper introduces Param∆, a novel method that streamlines post-training by transferring knowledge from an existing post-trained model to a newly updated base model with zero additional training. By computing the difference between post-trained model weights (Θ post ) and base model weights (Θ base ), and adding this to the updated base model (Θ ′ base ), we define Param∆ Model as: Θ Param∆ = Θ post -Θ base +Θ ′ base . This approach surprisingly equips the new base model with post-trained capabilities, achieving performance comparable to direct post-training. We did analysis on LLama3, Llama3.1, Qwen, and DeepSeek-distilled models. Results indicate Param∆ Model effectively replicates traditional post-training. For example, the Param∆ Model obtained from 70B Llama3-inst, Llama3-base, Llama3.1-base models attains approximately 95% of Llama3.1-inst model's performance on average. Param∆ brings a new perspective on how to fully leverage models in the open-weight community, where checkpoints for base and instruct models are readily available and frequently updated, by providing a cost-free framework to accelerate the iterative cycle of model development.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3fdc7fae-c69e-4050-b3b3-2bb56fb5d3ffCited by top-tier papers2
- Knowledge is Not Enough: Injecting RL Skills for Continual AdaptationPingzhi Tang, Yiding Wang, Muhan ZhangACL 2026 · 2 citations
- Fine-Tune Once, Reuse Across Models: Bayesian Task-Update Factors and ApproximationsSiyang Guo, Junbo Wang, Zibin ZhengICML 2026
Builds on17
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
Related papers
- LLaMA Pro: Progressive LLaMA with Block ExpansionChengyue Wu, Yukang Gan, Yixiao Ge, Zeyu Lu et al.ACL 2024
- Efficient Model Development through Fine-tuning TransferPin-Jie Lin, Rishab Balasubramanian, Fengyuan Liu, Nikhil Kandpal et al.EMNLP 2025
- Steering Information Utility in Key-Value Memory for Language Model Post-TrainingChunyuan Deng, Ruidi Chang, Hanjie ChenNeurIPS 2025 · 2 citations
- Composing Parameter-Efficient Modules with Arithmetic OperationJinghan Zhang, Shiqi Chen, Junteng Liu, Junxian HeNeurIPS 2023 · 164 citations
- Model Extrapolation Expedites AlignmentChujie Zheng, Ziqi Wang, Heng Ji, Minlie Huang et al.ACL 2025 · 34 citations
