Scalable Model Merging with Progressive Layer-wise Distillation
Jing Xu, Jiazheng Li, Jingzhao Zhang
摘要
Model merging offers an effective way to integrate the capabilities of multiple fine-tuned models. However, the performance degradation of the merged model remains a challenge, particularly when none or few data are available. This paper first highlights the necessity of domain-specific data for model merging by proving that data-agnostic algorithms can have arbitrarily bad worst-case performance. Building on this theoretical insight, we explore the relationship between model merging and distillation, introducing a novel few-shot merging algorithm, ProDistill (Progressive Layer-wise Distillation). Unlike common belief that layerwise training hurts performance, we show that layer-wise teacher-student distillation not only enhances the scalability but also improves model merging performance. We conduct extensive experiments to show that compared to existing fewshot merging methods, ProDistill achieves state-of-the-art performance, with up to 6.14% and 6.61% improvements in vision and NLU tasks. Furthermore, we extend the experiments to models with over 10B parameters, showcasing the exceptional scalability of ProDistill. Scalable Model Merging with Progressive Layer-wise Distillation Preliminaries We consider model merging in a pretrain-to-finetune setup. Let θ 0 denote the weights of a pre-trained model. Consider a set of T tasks, each with a model θ i fine-tuned from θ 0 . Model merging aims to combine the knowledge learned by task-specific models θ i into a unified model θ, which preserves the generalization ability of the pre-trained model and incorporates the specialized knowledge from each task.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Hawaii: Hierarchical Visual Knowledge Transfer for Efficient Vision-Language ModelsYimu Wang, Mozhgan Nasr Azadani, Sean Sedwards, Krzysztof CzarneckiNeurIPS 2025 · 被引用 6 次
- Depth-Progressive Monotonic Learning without Global BackpropagationChenhao Ye, Rongguang Ye, Yuchao Zhang, Ming TangICML 2026 · 被引用 1 次
- SyMerge: From Non-Interference to Synergistic Merging via Single-Layer AdaptationAecheon Jung, Seunghwan Lee, Dongyoon Han, Sungeun HongICML 2026 · 被引用 1 次
- A Study on PAVE Specification for LearnwareHao-Yu Shi, Zhi-Hao Tan, Zi-Chen Zhao, Yang Yu 等ICLR 2026
- Composable Cross-prompt Essay Scoring by Merging ModelsSanwoo Lee, Kun Liang, Yunfang WuEMNLP 2025
它引用的顶会 Paper43
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs 等ICML 2022 · 被引用 1,464 次
相关 Paper
- Leveraging Normalization Layer in Adapters with Progressive Learning and Adaptive Distillation for Cross-Domain Few-Shot LearningYongjin Yang, Taehyeon Kim, Se-Young YunAAAI 2024 · 被引用 12 次
- Few Shot Network Compression via Cross DistillationHaoli Bai, Jiaxiang Wu, Irwin King, Michael R. LyuAAAI 2020 · 被引用 66 次
- Twin-Merging: Dynamic Integration of Modular Expertise in Model MergingZhenyi Lu, Chenghao Fan, Wei Wei, Xiaoye Qu 等NeurIPS 2024 · 被引用 139 次
- Dataless Knowledge Fusion by Merging Weights of Language ModelsXisen Jin, Xiang Ren, Daniel Preotiuc-Pietro, Pengxiang ChengICLR 2023 · 被引用 8 次
- MergOPT: A Merge-Aware Optimizer for Robust Model MergingEnneng Yang, Qun Yang, Peng Wang, Anke Tang 等ICLR 2026
