MetaGPT: Merging Large Language Models Using Model Exclusive Task Arithmetic
Yuyan Zhou, Liang Song, Bingning Wang, Weipeng Chen
摘要
The advent of large language models (LLMs) like GPT-4 has catalyzed the exploration of multi-task learning (MTL), in which a single model demonstrates proficiency across diverse tasks.Task arithmetic has emerged as a costeffective approach for MTL.It enables performance enhancement across multiple tasks by adding their corresponding task vectors to a pre-trained model.However, the current lack of a method that can achieve optimal performance with low computational cost and protecting the data privacy, which limits their application to LLMs.In this paper, we propose Model Exclusive Task Arithmetic for merging GPT-scale models (MetaGPT), which formalizes the objective of model merging into a multi-task learning framework, aiming to minimize the average loss difference between the merged model and each individual task model.Since data privacy limits the use of multi-task training data, we leverage LLMs' local linearity and task vectors' orthogonality to separate the data term and scaling coefficients term and derive a model-exclusive task arithmetic method.Our proposed MetaGPT is dataagnostic and bypasses the heavy search process, making it cost-effective and easy to implement for LLMs.Extensive experiments demonstrate that MetaGPT leads to improvements in task arithmetic and achieves state-of-the-art performance on multiple tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Model Merging in Pre-training of Large Language ModelsYunshui Li, Yiyuan Ma, Shen Yan, Chaoyi Zhang 等NeurIPS 2025 · 被引用 40 次
- Ada-R1: Hybrid-CoT via Bi-Level Adaptive Reasoning OptimizationHaotian Luo, Haiying He, Yibo Wang, Jinluan Yang 等NeurIPS 2025 · 被引用 29 次
- Structure-Adaptive Multi-View Graph Clustering for Remote Sensing DataRenxiang Guan, Wenxuan Tu, Siwei Wang, Jiyuan Liu 等AAAI 2025 · 被引用 26 次
- Continual Model Merging without Data: Dual Projections for Balancing Stability and PlasticityEnneng Yang, Anke Tang, Li Shen, Guibing Guo 等NeurIPS 2025 · 被引用 13 次
- Free-Merging: Fourier Transform for Efficient Model MergingShenghe Zheng, Hongzhi WangICCV 2025 · 被引用 12 次
它引用的顶会 Paper25
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine 等NeurIPS 2020 · 被引用 2,261 次
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs 等ICML 2022 · 被引用 1,464 次
- TIES-Merging: Resolving Interference When Merging ModelsPrateek Yadav, Derek Tam, Leshem Choshen, Colin A. Raffel 等NeurIPS 2023 · 被引用 999 次
- Merging Models with Fisher-Weighted AveragingMichael Matena, Colin RaffelNeurIPS 2022 · 被引用 741 次
相关 Paper
- AdaMerging: Adaptive Model Merging for Multi-Task LearningEnneng Yang, Zhenyi Wang, Li Shen, Shiwei Liu 等ICLR 2024 · 被引用 230 次
- Post-Hoc Merging is Not Enough: Many-Shot Model Merging with Loss-Gap BalancingKyungjin Im, Miru Kim, Chanin Eom, Minhae KwonICML 2026
- Leveraging Submodule Linearity Enhances Task Arithmetic Performance in LLMsRui Dai, Sile Hu, Xu Shen, Yonggang Zhang 等ICLR 2025
- CAT Merging: A Training-Free Approach for Resolving Conflicts in Model MergingWenju Sun, Qingyong Li, Yangliao Geng, Boyang LiICML 2025
- Task Arithmetic in Trust Region: A Training-Free Model Merging Approach to Navigate Knowledge ConflictsWenju Sun, Qingyong Li, Wen Wang, Yangliao Geng 等ACM MM 2025 · 被引用 3 次
