AdaMMS: Model Merging for Heterogeneous Multimodal Large Language Models with Unsupervised Coefficient Optimization
Yiyang Du, Xiaochen Wang, Chi Chen, Jiabo Ye, Yiru Wang, Peng Li, Ming Yan, Ji Zhang, Fei Huang, Zhifang Sui, Maosong Sun, Yang Liu
Abstract
Recently, model merging methods have demonstrated powerful strengths in combining abilities on various tasks from multiple Large Language Models (LLMs). While previous model merging methods mainly focus on merging homogeneous models with identical architecture, they meet challenges when dealing with Multimodal Large Language Models (MLLMs) with inherent heterogeneous property, including differences in model architecture and the asymmetry in the parameter space. In this work, we propose AdaMMS 1 , a novel model merging method tailored for heterogeneous MLLMs. Our method tackles the challenges in three steps: mapping, merging and searching. Specifically, we first design mapping function between models to apply model merging on MLLMs with different architecture. Then we apply linear interpolation on model weights to actively adapt the asymmetry in the heterogeneous MLLMs. Finally in the hyper-parameter searching step, we propose an unsupervised hyper-parameter selection method for model merging. As the first model merging method capable of merging heterogeneous MLLMs without labeled data, extensive experiments on various model combinations demonstrated that AdaMMS outperforms previous model merging methods on various vision-language benchmarks. 2 * Equal contribution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2b36fb93-37ba-4a03-b3f6-7c958b9106b5Cited by top-tier papers8
- MergeVLA: Cross-Skill Model Merging Toward a Generalist Vision-Language-Action AgentYuxia Fu, Zhizhen Zhang, Yuqi Zhang, Zijian Wang et al.CVPR 2026 · 21 citations
- OptMerge: Unifying Multimodal LLM Capabilities and Modalities via Model MergingYongxian Wei, Runxi Cheng, Weike Jin, Enneng Yang et al.ICLR 2026 · 10 citations
- Label-Free Cross-Task LoRA Merging with Null-Space CompressionWonyoung Lee, Wooseong Jeong, Kuk-Jin YoonCVPR 2026 · 3 citations
- FedMerge: Federated Model Merging for PersonalizationShutong Chen, Tianyi Zhou, Guodong Long, Jing Jiang et al.AAAI 2026 · 2 citations
- Preference-Aligned LoRA Merging: Preserving Subspace Coverage and Addressing Directional AnisotropyWooseong Jeong, Wonyoung Lee, Kuk-Jin YoonCVPR 2026 · 1 citation
Builds on10
- TIES-Merging: Resolving Interference When Merging ModelsPrateek Yadav, Derek Tam, Leshem Choshen, Colin A. Raffel et al.NeurIPS 2023 · 999 citations
- CogVLM: Visual Expert for Pretrained Language ModelsWeihan Wang, Qingsong Lv, Wenmeng Yu, Wenyi Hong et al.NeurIPS 2024 · 858 citations
- Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free LunchLe Yu, Bowen Yu, Haiyang Yu, Fei Huang et al.ICML 2024 · 605 citations
- MMMU: A Massive Multi-Discipline Multimodal Understanding and Reasoning Benchmark for Expert AGIXiang Yue, Yuansheng Ni, Tianyu Zheng, Kai Zhang et al.CVPR 2024 · 213 citations
- mPLUG-OwI2: Revolutionizing Multi-modal Large Language Model with Modality CollaborationQinghao Ye, Haiyang Xu, Jiabo Ye, Ming Yan et al.CVPR 2024 · 144 citations
Related papers
- REMEDY: Recipe Merging Dynamics in Large Vision-Language ModelsDidi Zhu, Yibing Song, Tao Shen, Ziyu Zhao et al.ICLR 2025
- Bring Reason to Vision: Understanding Perception and Reasoning through Model MergingShiqi Chen, Jinghan Zhang, Tongyao Zhu, Wei Liu et al.ICML 2025
- Twin-Merging: Dynamic Integration of Modular Expertise in Model MergingZhenyi Lu, Chenghao Fan, Wei Wei, Xiaoye Qu et al.NeurIPS 2024 · 139 citations
- Investigating Cross-Modal Skill Injection: Scenarios, Methods, and HyperparametersZhiyu Xu, Lean Wang, Yuanxin Liu, Lei Li et al.ACL 2026
- Training-free LLM Merging for Multi-task LearningZichuan Fu, Xian Wu, Yejing Wang, Wanyu Wang et al.ACL 2025
