Twin-Merging: Dynamic Integration of Modular Expertise in Model Merging
Zhenyi Lu, Chenghao Fan, Wei Wei, Xiaoye Qu, Dangyang Chen, Yu Cheng
摘要
In the era of large language models, model merging is a promising way to combine multiple task-specific models into a single multitask model without extra training. However, two challenges remain: (a) interference between different models and (b) heterogeneous data during testing. Traditional model merging methods often show significant performance gaps compared to fine-tuned models due to these issues. Additionally, a one-size-fits-all model lacks flexibility for diverse test data, leading to performance degradation. We show that both shared and exclusive task-specific knowledge are crucial for merging performance, but directly merging exclusive knowledge hinders overall performance. In view of this, we propose Twin-Merging, a method that encompasses two principal stages: (1) modularizing knowledge into shared and exclusive components, with compression to reduce redundancy and enhance efficiency; (2) dynamically merging shared and task-specific knowledge based on the input. This approach narrows the performance gap between merged and fine-tuned models and improves adaptability to heterogeneous data. Extensive experiments on datasets for both language and vision tasks demonstrate the effectiveness of our method, showing an average improvement of in absolute normalized score for discriminative tasks and even surpassing the fine-tuned upper bound on the generative tasks. Our implementation is available in https://github.com/LZY-the-boys/Twin-Merging
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper68
- MergeVLA: Cross-Skill Model Merging Toward a Generalist Vision-Language-Action AgentYuxia Fu, Zhizhen Zhang, Yuqi Zhang, Zijian Wang 等CVPR 2026 · 被引用 21 次
- Towards Minimizing Feature Drift in Model Merging: Layer-wise Task Vector Fusion for Adaptive Knowledge IntegrationWenju Sun, Qingyong Li, Wen Wang, Yang Liu 等NeurIPS 2025 · 被引用 20 次
- On Giant's Shoulders: Effortless Weak to Strong by Dynamic Logits FusionChenghao Fan, Zhenyi Lu, Wei Wei, Jie Tian 等NeurIPS 2024 · 被引用 17 次
- MINGLE: Mixture of Null-Space Gated Low-Rank Experts for Test-Time Continual Model MergingZihuan Qiu, Yi Xu, Chiyuan He, Fanman Meng 等NeurIPS 2025 · 被引用 16 次
- AdaRank: Adaptive Rank Pruning for Enhanced Model MergingChanhyuk Lee, Jiho Choi, Chanryeol Lee, Donggyun Kim 等ICLR 2026 · 被引用 14 次
它引用的顶会 Paper36
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 被引用 3,228 次
相关 Paper
- Unraveling LoRA Interference: Orthogonal Subspaces for Robust Model MergingHaobo Zhang, Jiayu ZhouACL 2025
- Enough is as good as a feast: A Comprehensive Analysis of How Reinforcement Learning Mitigates Task Conflicts in LLMsZixuan Ren, Jinliang Lu, Junhong Wu, Yang Zhao 等ICLR 2026 · 被引用 2 次
- AdaMMS: Model Merging for Heterogeneous Multimodal Large Language Models with Unsupervised Coefficient OptimizationYiyang Du, Xiaochen Wang, Chi Chen, Jiabo Ye 等CVPR 2025
- Dataless Knowledge Fusion by Merging Weights of Language ModelsXisen Jin, Xiang Ren, Daniel Preotiuc-Pietro, Pengxiang ChengICLR 2023 · 被引用 8 次
- Training-free LLM Merging for Multi-task LearningZichuan Fu, Xian Wu, Yejing Wang, Wanyu Wang 等ACL 2025
