Why Do More Experts Fail? A Theoretical Analysis of Model Merging
Zijing Wang, Xingle Xu, Yongkang Liu, Yiqun Zhang, Peiqin Lin, Shi Feng, Daling Wang, Xiaocui Yang, Hinrich Schütze
摘要
Model merging dramatically reduces storage and computational resources by combining multiple expert models into a single multi-task model. Although recent model merging methods have shown promising results, they struggle to maintain performance gains as the number of merged models increases. In this paper, we investigate the key obstacles that limit the scalability of model merging when integrating a large number of expert models. First, we prove that there is an upper bound on model merging. Further theoretical analysis reveals that the limited effective parameter space imposes a strict constraint on the number of models that can be successfully merged. Gaussian Width shows that the marginal benefit of merging additional models diminishes according to a strictly concave function. This implies that the effective parameter space becomes rapidly saturated as the number of merged models increases. Furthermore, using Approximate Kinematics Theory, we prove the existence of a unique optimal threshold beyond which adding more models does not yield significant performance improvements. At the same time, we introduce a straightforward Reparameterized Heavy-Tailed method (RHT) to extend the coverage of the merged model, thereby enhancing its performance. Empirical results on 12 benchmarks, including both knowledgeintensive and general-purpose tasks, validate our theoretical analysis. We believe that these results spark further research beyond the current scope of model merging. The source code is in the Github repository: https://github.com/wzj1718/ ModelMergingAnalysis .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Model Merging Scaling Laws in Large Language ModelsYuanyi Wang, Yanggan Gu, Yiming Zhang, Qi ZHOU 等ICML 2026 · 被引用 8 次
- Sharpness-aware Model Merging with Salience Recovery for LLM-based Cross-Domain Sequential RecommendationHuwei Ji, Jiajie Su, Yuyuan Li, Xiaohua Feng 等KDD 2026 · 被引用 1 次
- Curriculum Model Merging: Harmonizing Chemical LLMs for Enhanced Cross-Task GeneralizationBaoyi He, Luotian Yuan, Ying Wei, Fei WuNeurIPS 2025
- Saliency-Aware Model MergingJungin Park, Jiyoung Lee, Kwanghoon SohnICML 2026
它引用的顶会 Paper15
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- TIES-Merging: Resolving Interference When Merging ModelsPrateek Yadav, Derek Tam, Leshem Choshen, Colin A. Raffel 等NeurIPS 2023 · 被引用 999 次
- Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free LunchLe Yu, Bowen Yu, Haiyang Yu, Fei Huang 等ICML 2024 · 被引用 605 次
- Rewarded soups: towards Pareto-optimal alignment by interpolating weights fine-tuned on diverse rewardsAlexandre Ramé, Guillaume Couairon, Corentin Dancette, Jean-Baptiste Gaya 等NeurIPS 2023 · 被引用 295 次
相关 Paper
- Merging Multi-Task Models via Weight-Ensembling Mixture of ExpertsAnke Tang, Li Shen, Yong Luo, Nan Yin 等ICML 2024 · 被引用 96 次
- Free-Merging: Fourier Transform for Efficient Model MergingShenghe Zheng, Hongzhi WangICCV 2025 · 被引用 12 次
- HM3: Hierarchical Multi-Objective Model Merging for Pretrained ModelsYu Zhou, Xingyu Wu, Jibin Wu, Liang Feng 等NeurIPS 2025 · 被引用 14 次
- Pareto Merging: Multi-Objective Optimization for Preference-Aware Model MergingWeiyu Chen, James T. KwokICML 2025
- From Memorization to Parameter Interference: How Overtraining Experts Harms Model MergingStefan Horoi, Guy Wolf, Eugene Belilovsky, Gintare Karolina DziugaiteICML 2026
