ForkMerge: Mitigating Negative Transfer in Auxiliary-Task Learning
Junguang Jiang, Baixu Chen, Junwei Pan, Ximei Wang, Dapeng Liu, Jie Jiang, Mingsheng Long
Abstract
Auxiliary-Task Learning (ATL) aims to improve the performance of the target task by leveraging the knowledge obtained from related tasks. Occasionally, learning multiple tasks simultaneously results in lower accuracy than learning only the target task, which is known as negative transfer. This problem is often attributed to the gradient conflicts among tasks, and is frequently tackled by coordinating the task gradients in previous works. However, these optimization-based methods largely overlook the auxiliary-target generalization capability. To better understand the root cause of negative transfer, we experimentally investigate it from both optimization and generalization perspectives. Based on our findings, we introduce ForkMerge, a novel approach that periodically forks the model into multiple branches, automatically searches the varying task weights by minimizing target validation errors, and dynamically merges all branches to filter out detrimental task-parameter updates. On a series of auxiliary-task learning benchmarks, ForkMerge outperforms existing methods and effectively mitigates negative transfer.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext da022c8d-1326-4932-be65-d54ddb33b053Cited by top-tier papers14
- AdaMerging: Adaptive Model Merging for Multi-Task LearningEnneng Yang, Zhenyi Wang, Li Shen, Shiwei Liu et al.ICLR 2024 · 230 citations
- Twin-Merging: Dynamic Integration of Modular Expertise in Model MergingZhenyi Lu, Chenghao Fan, Wei Wei, Xiaoye Qu et al.NeurIPS 2024 · 139 citations
- Task Groupings Regularization: Data-Free Meta-Learning with Heterogeneous Pre-trained ModelsYongxian Wei, Zixuan Hu, Li Shen, Zhenyi Wang et al.ICML 2024 · 11 citations
- Merging on the Fly Without Retraining: A Sequential Approach to Scalable Continual Model MergingAnke Tang, Enneng Yang, Li Shen, Yong Luo et al.NeurIPS 2025 · 8 citations
- To Remember, To Adapt, To Preempt: A Stable Continual Test-Time Adaptation Framework for Remote Physiological Measurement in Dynamic Domain ShiftsShuyang Chu, Jingang Shi, Xu Cheng, Haoyu Chen et al.ACM MM 2025 · 4 citations
Builds on24
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang et al.NeurIPS 2020 · 5,129 citations
- Big Self-Supervised Models are Strong Semi-Supervised LearnersTing Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi et al.NeurIPS 2020 · 2,611 citations
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- Moment Matching for Multi-Source Domain AdaptationXingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang et al.ICCV 2019 · 2,239 citations
Related papers
- Selective Task Group Updates for Multi-Task OptimizationWooseong Jeong, Kuk-Jin YoonICLR 2025
- Learning Conflict-Noticed Architecture for Multi-Task LearningZhixiong Yue, Yu Zhang, Jie LiangAAAI 2023 · 9 citations
- Quantifying Task Priority for Multi-Task OptimizationWooseong Jeong, Kuk-Jin YoonCVPR 2024
- Auxiliary Learning as an Asymmetric Bargaining GameAviv Shamsian, Aviv Navon, Neta Glazer, Kenji Kawaguchi et al.ICML 2023 · 15 citations
- Task Arithmetic in Trust Region: A Training-Free Model Merging Approach to Navigate Knowledge ConflictsWenju Sun, Qingyong Li, Wen Wang, Yangliao Geng et al.ACM MM 2025 · 3 citations
