Why Do More Experts Fail? A Theoretical Analysis of Model Merging
Zijing Wang, Xingle Xu, Yongkang Liu, Yiqun Zhang, Peiqin Lin, Shi Feng, Daling Wang, Xiaocui Yang, Hinrich Schütze
Abstract
Model merging dramatically reduces storage and computational resources by combining multiple expert models into a single multi-task model. Although recent model merging methods have shown promising results, they struggle to maintain performance gains as the number of merged models increases. In this paper, we investigate the key obstacles that limit the scalability of model merging when integrating a large number of expert models. First, we prove that there is an upper bound on model merging. Further theoretical analysis reveals that the limited effective parameter space imposes a strict constraint on the number of models that can be successfully merged. Gaussian Width shows that the marginal benefit of merging additional models diminishes according to a strictly concave function. This implies that the effective parameter space becomes rapidly saturated as the number of merged models increases. Furthermore, using Approximate Kinematics Theory, we prove the existence of a unique optimal threshold beyond which adding more models does not yield significant performance improvements. At the same time, we introduce a straightforward Reparameterized Heavy-Tailed method (RHT) to extend the coverage of the merged model, thereby enhancing its performance. Empirical results on 12 benchmarks, including both knowledgeintensive and general-purpose tasks, validate our theoretical analysis. We believe that these results spark further research beyond the current scope of model merging. The source code is in the Github repository: https://github.com/wzj1718/ ModelMergingAnalysis .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 57983fa6-e040-491d-97ff-96248487ee7eCited by top-tier papers4
- Model Merging Scaling Laws in Large Language ModelsYuanyi Wang, Yanggan Gu, Yiming Zhang, Qi ZHOU et al.ICML 2026 · 8 citations
- Sharpness-aware Model Merging with Salience Recovery for LLM-based Cross-Domain Sequential RecommendationHuwei Ji, Jiajie Su, Yuyuan Li, Xiaohua Feng et al.KDD 2026 · 1 citation
- Curriculum Model Merging: Harmonizing Chemical LLMs for Enhanced Cross-Task GeneralizationBaoyi He, Luotian Yuan, Ying Wei, Fei WuNeurIPS 2025
- Saliency-Aware Model MergingJungin Park, Jiyoung Lee, Kwanghoon SohnICML 2026
Builds on15
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- TIES-Merging: Resolving Interference When Merging ModelsPrateek Yadav, Derek Tam, Leshem Choshen, Colin A. Raffel et al.NeurIPS 2023 · 999 citations
- Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free LunchLe Yu, Bowen Yu, Haiyang Yu, Fei Huang et al.ICML 2024 · 605 citations
- Rewarded soups: towards Pareto-optimal alignment by interpolating weights fine-tuned on diverse rewardsAlexandre Ramé, Guillaume Couairon, Corentin Dancette, Jean-Baptiste Gaya et al.NeurIPS 2023 · 295 citations
Related papers
- Merging Multi-Task Models via Weight-Ensembling Mixture of ExpertsAnke Tang, Li Shen, Yong Luo, Nan Yin et al.ICML 2024 · 96 citations
- Free-Merging: Fourier Transform for Efficient Model MergingShenghe Zheng, Hongzhi WangICCV 2025 · 12 citations
- HM3: Hierarchical Multi-Objective Model Merging for Pretrained ModelsYu Zhou, Xingyu Wu, Jibin Wu, Liang Feng et al.NeurIPS 2025 · 14 citations
- Pareto Merging: Multi-Objective Optimization for Preference-Aware Model MergingWeiyu Chen, James T. KwokICML 2025
- From Memorization to Parameter Interference: How Overtraining Experts Harms Model MergingStefan Horoi, Guy Wolf, Eugene Belilovsky, Gintare Karolina DziugaiteICML 2026
