REMEDY: Recipe Merging Dynamics in Large Vision-Language Models
Didi Zhu, Yibing Song, Tao Shen, Ziyu Zhao, Jinluan Yang, Min Zhang, Chao Wu
Abstract
Model merging has emerged as a powerful technique for combining task-specific vision models into a unified and multi-functional model. Previous methods represented by task arithmetic, have demonstrated effectiveness and scalability in this domain. When large vision-language models (LVLMs) arise with model size scaling up, this design becomes challenging to fuse different instruction-tuned LVLMs for generalization enhancement. The large scale and multi-modal nature of LVLMs present unique obstacles, including constructing reusable and modular components to accommodate the multi-component architecture of LVLMs and the requirement for dynamic fusion based on multi-modal input tokens. To address these challenges, we propose the REcipe MErging DYnamics (REMEDY) method, a scalable and flexible paradigm for model merging in LVLMs. We first define reusable modules termed recipes including the projector and shallow LLM layers, enhancing visual-language understanding. Then, we introduce a modalityaware allocator dynamically generates weights in a one-shot manner based on input relevance to existing recipes, enabling efficient cross-modal knowledge integration. REMEDY thus offers an adaptive solution for LVLMs to tackle both seen (i.e., multi-task learning) and unseen (i.e., zero-shot generalization) tasks. Experimental results demonstrate that our method consistently improves performance on both seen and unseen tasks, underscoring the effectiveness of REMEDY in diverse multi-modal scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 96d2b005-c42a-4cab-9f91-ea9ca1f2df12Cited by top-tier papers10
- VisCodex: Unified Multimodal Code Generation via Merging Vision and Coding ModelsLingjie Jiang, Shaohan Huang, Xun Wu, Yixia Li et al.ICLR 2026 · 15 citations
- Label-Free Cross-Task LoRA Merging with Null-Space CompressionWonyoung Lee, Wooseong Jeong, Kuk-Jin YoonCVPR 2026 · 3 citations
- Boomerang Distillation Enables Zero-Shot Model Size InterpolationSara Kangaslahti, Nihal V. Nayak, Jonathan Geuter, Marco Fumero et al.ICLR 2026 · 3 citations
- Preference-Aligned LoRA Merging: Preserving Subspace Coverage and Addressing Directional AnisotropyWooseong Jeong, Wonyoung Lee, Kuk-Jin YoonCVPR 2026 · 1 citation
- Latent Score-Based Reweighting for Robust Classification on Imbalanced Tabular DataYunze Tong, Fengda Zhang, Zihao Tang, Kaifeng Gao et al.ICML 2025
Builds on35
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question AnsweringPan Lu, Swaroop Mishra, Tanglin Xia, Liang Qiu et al.NeurIPS 2022 · 2,727 citations
- MM-Vet: Evaluating Large Multimodal Models for Integrated CapabilitiesWeihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang et al.ICML 2024 · 1,191 citations
Related papers
- Improving Cross-Modal Recipe Retrieval with Component-Aware Prompted CLIP EmbeddingXu Huang, Jin Liu, Zhizhong Zhang, Yuan XieACM MM 2023 · 11 citations
- AdaMMS: Model Merging for Heterogeneous Multimodal Large Language Models with Unsupervised Coefficient OptimizationYiyang Du, Xiaochen Wang, Chi Chen, Jiabo Ye et al.CVPR 2025
- Investigating Cross-Modal Skill Injection: Scenarios, Methods, and HyperparametersZhiyu Xu, Lean Wang, Yuanxin Liu, Lei Li et al.ACL 2026
- Beyond Layer-Wise Merging: Chain-of-Merging for Vision-Language ModelsXinyu Zhang, Yuxuan Dong, Lingling Zhang, Chengyou Jia et al.CVPR 2026
- Bring Reason to Vision: Understanding Perception and Reasoning through Model MergingShiqi Chen, Jinghan Zhang, Tongyao Zhu, Wei Liu et al.ICML 2025
