Scalable Model Merging with Progressive Layer-wise Distillation
Jing Xu, Jiazheng Li, Jingzhao Zhang
Abstract
Model merging offers an effective way to integrate the capabilities of multiple fine-tuned models. However, the performance degradation of the merged model remains a challenge, particularly when none or few data are available. This paper first highlights the necessity of domain-specific data for model merging by proving that data-agnostic algorithms can have arbitrarily bad worst-case performance. Building on this theoretical insight, we explore the relationship between model merging and distillation, introducing a novel few-shot merging algorithm, ProDistill (Progressive Layer-wise Distillation). Unlike common belief that layerwise training hurts performance, we show that layer-wise teacher-student distillation not only enhances the scalability but also improves model merging performance. We conduct extensive experiments to show that compared to existing fewshot merging methods, ProDistill achieves state-of-the-art performance, with up to 6.14% and 6.61% improvements in vision and NLU tasks. Furthermore, we extend the experiments to models with over 10B parameters, showcasing the exceptional scalability of ProDistill. Scalable Model Merging with Progressive Layer-wise Distillation Preliminaries We consider model merging in a pretrain-to-finetune setup. Let θ 0 denote the weights of a pre-trained model. Consider a set of T tasks, each with a model θ i fine-tuned from θ 0 . Model merging aims to combine the knowledge learned by task-specific models θ i into a unified model θ, which preserves the generalization ability of the pre-trained model and incorporates the specialized knowledge from each task.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- Hawaii: Hierarchical Visual Knowledge Transfer for Efficient Vision-Language ModelsYimu Wang, Mozhgan Nasr Azadani, Sean Sedwards, Krzysztof CzarneckiNeurIPS 2025 · 6 citations
- Depth-Progressive Monotonic Learning without Global BackpropagationChenhao Ye, Rongguang Ye, Yuchao Zhang, Ming TangICML 2026 · 1 citation
- SyMerge: From Non-Interference to Synergistic Merging via Single-Layer AdaptationAecheon Jung, Seunghwan Lee, Dongyoon Han, Sungeun HongICML 2026 · 1 citation
- A Study on PAVE Specification for LearnwareHao-Yu Shi, Zhi-Hao Tan, Zi-Chen Zhao, Yang Yu et al.ICLR 2026
- Composable Cross-prompt Essay Scoring by Merging ModelsSanwoo Lee, Kun Liang, Yunfang WuEMNLP 2025
Builds on43
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
Related papers
- Leveraging Normalization Layer in Adapters with Progressive Learning and Adaptive Distillation for Cross-Domain Few-Shot LearningYongjin Yang, Taehyeon Kim, Se-Young YunAAAI 2024 · 12 citations
- Few Shot Network Compression via Cross DistillationHaoli Bai, Jiaxiang Wu, Irwin King, Michael R. LyuAAAI 2020 · 66 citations
- Twin-Merging: Dynamic Integration of Modular Expertise in Model MergingZhenyi Lu, Chenghao Fan, Wei Wei, Xiaoye Qu et al.NeurIPS 2024 · 139 citations
- Dataless Knowledge Fusion by Merging Weights of Language ModelsXisen Jin, Xiang Ren, Daniel Preotiuc-Pietro, Pengxiang ChengICLR 2023 · 8 citations
- MergOPT: A Merge-Aware Optimizer for Robust Model MergingEnneng Yang, Qun Yang, Peng Wang, Anke Tang et al.ICLR 2026
