OptMerge: Unifying Multimodal LLM Capabilities and Modalities via Model Merging
Yongxian Wei, Runxi Cheng, Weike Jin, Enneng Yang, Li Shen, Lu Hou, Sinan Du, Chun Yuan, Xiaochun Cao, Dacheng Tao
摘要
Foundation models update slowly due to resource-intensive training, whereas domain-specific models evolve rapidly between releases. Model merging seeks to combine multiple expert models into a single, more capable model, reducing storage and serving costs while supporting decentralized development. Despite its potential, previous studies have primarily focused on merging visual classification models or Large Language Models (LLMs) for code and math tasks. Recently, Multimodal LLMs (MLLMs) that extend LLMs through large-scale multimodal training have gained traction. However, no benchmark exists for model merging research that clearly divides the tasks of MLLM training and evaluation. In this paper, (i) we introduce a model merging benchmark for MLLMs, which includes multiple tasks such as VQA, Geometry, Chart, OCR, and Grounding, studying both LoRA and full fine-tuning models. Moreover, we explore how model merging can combine different modalities (e.g., vision-language, audio-language, and videolanguage models), moving toward the Omni-language model. (ii) We implement 10 model merging algorithms on the benchmark. Furthermore, we propose a novel method that removes noise from task vectors and robustly optimizes the merged vector based on a loss defined over task vector interactions, achieving an average performance gain of 2.48%. (iii) We find that model merging offers a promising way for building improved MLLMs without requiring training data. Our results also demonstrate that the complementarity among multiple modalities outperforms individual modalities. All code and checkpoints are publicly available here.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Adaptive Defense against Harmful Fine-Tuning for Large Language Models via Bayesian Data SchedulerZixuan Hu, Li Shen, Zhenyi Wang, Yongxian Wei 等NeurIPS 2025 · 被引用 16 次
- VQRAE: Representation Quantization Autoencoders for Multimodal Understanding, Generation and ReconstructionSinan Du, Jiahao Guo, Bo Li, Shuhao Cui 等CVPR 2026 · 被引用 11 次
- SCOUT: Teaching Pre-trained Language Models to Enhance Reasoning via Flow Chain-of-ThoughtGuanghao Li, Wenhao Jiang, Mingfeng Chen, Yan Li 等NeurIPS 2025 · 被引用 8 次
- Label-Free Cross-Task LoRA Merging with Null-Space CompressionWonyoung Lee, Wooseong Jeong, Kuk-Jin YoonCVPR 2026 · 被引用 3 次
- Text-Guided Visual Prompt DINO for Generic SegmentationYuchen Guan, Chong Sun, Canmiao Fu, Zhipeng Huang 等ICCV 2025 · 被引用 3 次
它引用的顶会 Paper47
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs 等ICML 2022 · 被引用 1,464 次
- Are We on the Right Way for Evaluating Large Vision-Language Models?Lin Chen, Jinsong Li, Xiaoyi Dong, Pan Zhang 等NeurIPS 2024 · 被引用 1,029 次
相关 Paper
- Unraveling LoRA Interference: Orthogonal Subspaces for Robust Model MergingHaobo Zhang, Jiayu ZhouACL 2025
- Model merging with SVD to tie the KnotsGeorge Stoica, Pratik Ramesh, Boglarka Ecsedi, Leshem Choshen 等ICLR 2025
- Bring Reason to Vision: Understanding Perception and Reasoning through Model MergingShiqi Chen, Jinghan Zhang, Tongyao Zhu, Wei Liu 等ICML 2025
- VisualAgentBench: Towards Large Multimodal Models as Visual Foundation AgentsXiao Liu, Tianjie Zhang, Yu Gu, Iat Long Iong 等ICLR 2025
- VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding ModelsLingjie Jiang, Shaohan Huang, Xun Wu, Yixia Li 等ICLR 2026 · 被引用 15 次
