Representation Surgery for Multi-Task Model Merging
Enneng Yang, Li Shen, Zhenyi Wang, Guibing Guo, Xiaojun Chen, Xingwei Wang, Dacheng Tao
Abstract
Multi-task learning (MTL) compresses the information from multiple tasks into a unified backbone to improve computational efficiency and generalization. Recent work directly merges multiple independently trained models to perform MTL instead of collecting their raw data for joint training, greatly expanding the application scenarios of MTL. However, by visualizing the representation distribution of existing model merging schemes, we find that the merged model often suffers from the dilemma of representation bias. That is, there is a significant discrepancy in the representation distribution between the merged and individual models, resulting in poor performance of merged MTL. In this paper, we propose a representation surgery solution called "Surgery" to reduce representation bias in the merged model. Specifically, Surgery is a lightweight task-specific module that takes the representation of the merged model as input and attempts to output the biases contained in the representation from the merged model. We then designed an unsupervised optimization objective that updates the Surgery module by minimizing the distance between the merged model's representation and the individual model's representation. Extensive experiments demonstrate significant MTL performance improvements when our Surgery module is applied to state-of-theart (SOTA) model merging schemes. The code is available at https://github.com/ EnnengYang/RepresentationSurgery .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 186d2304-a46d-4b18-93e9-724e09319347Cited by top-tier papers50
- Twin-Merging: Dynamic Integration of Modular Expertise in Model MergingZhenyi Lu, Chenghao Fan, Wei Wei, Xiaoye Qu et al.NeurIPS 2024 · 139 citations
- Accurate and Efficient Low-Rank Model Merging in Core SpaceAniello Panariello, Daniel Marczak, Simone Magistri, Angelo Porrello et al.NeurIPS 2025 · 32 citations
- Towards Minimizing Feature Drift in Model Merging: Layer-wise Task Vector Fusion for Adaptive Knowledge IntegrationWenju Sun, Qingyong Li, Wen Wang, Yang Liu et al.NeurIPS 2025 · 20 citations
- AdaRank: Adaptive Rank Pruning for Enhanced Model MergingChanhyuk Lee, Jiho Choi, Chanryeol Lee, Donggyun Kim et al.ICLR 2026 · 14 citations
- Continual Model Merging without Data: Dual Projections for Balancing Stability and PlasticityEnneng Yang, Anke Tang, Li Shen, Guibing Guo et al.NeurIPS 2025 · 13 citations
Builds on27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
- Merging Models with Fisher-Weighted AveragingMichael Matena, Colin RaffelNeurIPS 2022 · 741 citations
Related papers
- Representation Surgery in Model Merging with Probabilistic ModelingQi Wei, Shuo He, Enneng Yang, Tingcong Liu et al.ICML 2025
- AdaMerging: Adaptive Model Merging for Multi-Task LearningEnneng Yang, Zhenyi Wang, Li Shen, Shiwei Liu et al.ICLR 2024 · 230 citations
- MOS: Model Surgery for Pre-Trained Model-Based Class-Incremental LearningHai-Long Sun, Da-Wei Zhou, Hanbin Zhao, Le Gan et al.AAAI 2025 · 31 citations
- Post-Hoc Merging is Not Enough: Many-Shot Model Merging with Loss-Gap BalancingKyungjin Im, Miru Kim, Chanin Eom, Minhae KwonICML 2026
- Swiss Army Knife: Synergizing Biases in Knowledge from Vision Foundation Models for Multi-Task LearningYuxiang Lu, Shengcao Cao, Yu-Xiong WangICLR 2025
