Representation Surgery for Multi-Task Model Merging
Enneng Yang, Li Shen, Zhenyi Wang, Guibing Guo, Xiaojun Chen, Xingwei Wang, Dacheng Tao
摘要
Multi-task learning (MTL) compresses the information from multiple tasks into a unified backbone to improve computational efficiency and generalization. Recent work directly merges multiple independently trained models to perform MTL instead of collecting their raw data for joint training, greatly expanding the application scenarios of MTL. However, by visualizing the representation distribution of existing model merging schemes, we find that the merged model often suffers from the dilemma of representation bias. That is, there is a significant discrepancy in the representation distribution between the merged and individual models, resulting in poor performance of merged MTL. In this paper, we propose a representation surgery solution called "Surgery" to reduce representation bias in the merged model. Specifically, Surgery is a lightweight task-specific module that takes the representation of the merged model as input and attempts to output the biases contained in the representation from the merged model. We then designed an unsupervised optimization objective that updates the Surgery module by minimizing the distance between the merged model's representation and the individual model's representation. Extensive experiments demonstrate significant MTL performance improvements when our Surgery module is applied to state-of-theart (SOTA) model merging schemes. The code is available at https://github.com/ EnnengYang/RepresentationSurgery .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper50
- Twin-Merging: Dynamic Integration of Modular Expertise in Model MergingZhenyi Lu, Chenghao Fan, Wei Wei, Xiaoye Qu 等NeurIPS 2024 · 被引用 139 次
- Accurate and Efficient Low-Rank Model Merging in Core SpaceAniello Panariello, Daniel Marczak, Simone Magistri, Angelo Porrello 等NeurIPS 2025 · 被引用 32 次
- Towards Minimizing Feature Drift in Model Merging: Layer-wise Task Vector Fusion for Adaptive Knowledge IntegrationWenju Sun, Qingyong Li, Wen Wang, Yang Liu 等NeurIPS 2025 · 被引用 20 次
- AdaRank: Adaptive Rank Pruning for Enhanced Model MergingChanhyuk Lee, Jiho Choi, Chanryeol Lee, Donggyun Kim 等ICLR 2026 · 被引用 14 次
- Continual Model Merging without Data: Dual Projections for Balancing Stability and PlasticityEnneng Yang, Anke Tang, Li Shen, Guibing Guo 等NeurIPS 2025 · 被引用 13 次
它引用的顶会 Paper27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine 等NeurIPS 2020 · 被引用 2,261 次
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs 等ICML 2022 · 被引用 1,464 次
- Merging Models with Fisher-Weighted AveragingMichael Matena, Colin RaffelNeurIPS 2022 · 被引用 741 次
相关 Paper
- Representation Surgery in Model Merging with Probabilistic ModelingQi Wei, Shuo He, Enneng Yang, Tingcong Liu 等ICML 2025
- AdaMerging: Adaptive Model Merging for Multi-Task LearningEnneng Yang, Zhenyi Wang, Li Shen, Shiwei Liu 等ICLR 2024 · 被引用 230 次
- MOS: Model Surgery for Pre-Trained Model-Based Class-Incremental LearningHai-Long Sun, Da-Wei Zhou, Hanbin Zhao, Le Gan 等AAAI 2025 · 被引用 31 次
- Post-Hoc Merging is Not Enough: Many-Shot Model Merging with Loss-Gap BalancingKyungjin Im, Miru Kim, Chanin Eom, Minhae KwonICML 2026
- Swiss Army Knife: Synergizing Biases in Knowledge from Vision Foundation Models for Multi-Task LearningYuxiang Lu, Shengcao Cao, Yu-Xiong WangICLR 2025
