Beyond Layer-Wise Merging: Chain-of-Merging for Vision-Language Models
Xinyu Zhang, Yuxuan Dong, Lingling Zhang, Chengyou Jia, Zhuohang Dang, Yixing Yao, Yaqiang Wu, Basura Fernando, Jun Liu
Abstract
While model merging has demonstrated remarkable success across diverse domains for large language models (LLMs), its application to vision-language models (VLMs) remains largely underexplored. Recent methods attempt to enhance VLM reasoning capabilities by integrating specialized LLM parameters through layer-wise merging. However, existing paradigms suffer from two critical limitations: (1) strict positional correspondence, which enforces rigid one-to-one layer alignment, and (2) uniform merging weights applied indiscriminately across all layers. These constraints fail to account for substantial functional disparities between corresponding layers in VLMs and LLMs, potentially misaligning incompatible layers and leading to detrimental parameter combinations.To address these, we propose Chain-of-Merging (CoM) framework that adaptively adjusts merging plans for different images and questions, including two key stages: (1) Adaptive Layer Matching, which identifies optimal layer pairings based on structural and semantic matching scores while filtering incompatible pairings, and (2) Dynamic Weight Merging, which determines layer-specific merging weights based on matching scores and employs spherical linear interpolation to minimize memory overhead.Extensive experiments demonstrate that CoM achieves substantial performance improvements, with Qwen2.5-VL-7B + Qwen2.5-Math-7B attaining a 4.4% average improvement on mathematical reasoning benchmarks while enhancing general visual understanding, significantly outperforming existing training-free methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cf59b1f8-983d-4fcc-b9e2-10e51648379eRelated papers
- Bring Reason to Vision: Understanding Perception and Reasoning through Model MergingShiqi Chen, Jinghan Zhang, Tongyao Zhu, Wei Liu et al.ICML 2025
- Can MLLMs Absorb Math Reasoning Abilities from LLMs as Free Lunch?Yijie Hu, Zihao Zhou, Kaizhu Huang, Xiaowei Huang et al.NeurIPS 2025 · 3 citations
- Activation-Guided Consensus Merging for Large Language ModelsYuxuan Yao, Shuqi Liu, Zehua Liu, Qintong Li et al.NeurIPS 2025 · 14 citations
- FRISM: Fine-Grained Reasoning Injection via Subspace-Level Model Merging for Vision–Language ModelsChenyu Huang, Peng Ye, Xudong Tan, Jinhan Mu et al.ICML 2026
- Outlier Matters: Efficient Long-to-Short Reasoning via Outlier-Guided Model MergingQiyuan Zhu, Dezhi Li, Lujun Li, Xiaoyu Qin et al.AAAI 2026
