Mitigating the Backdoor Effect for Multi-Task Model Merging via Safety-Aware Subspace
Jinluan Yang, Anke Tang, Didi Zhu, Zhengyu Chen, Li Shen, Fei Wu
Abstract
Model merging has gained significant attention as a cost-effective approach to integrate multiple single-task fine-tuned models into a unified one that can perform well on multiple tasks. However, existing model merging techniques primarily focus on resolving conflicts between task-specific models, they often overlook potential security threats, particularly the risk of backdoor attacks in the open-source model ecosystem. In this paper, we first investigate the vulnerabilities of existing model merging methods to backdoor attacks, identifying two critical challenges: backdoor succession and backdoor transfer. To address these issues, we propose a novel Defense-Aware Merging (DAM) approach that simultaneously mitigates task interference and backdoor vulnerabilities. Specifically, DAM employs a meta-learning-based optimization method with dual masks to identify a shared and safety-aware subspace for model merging. These masks are alternately optimized: the Task-Shared mask identifies common beneficial parameters across tasks, aiming to preserve task-specific knowledge while reducing interference, while the Backdoor-Detection mask isolates potentially harmful parameters to neutralize security threats. This dual-mask design allows us to carefully balance the preservation of useful knowledge and the removal of potential vulnerabilities. Compared to existing merging methods, DAM achieves a more favorable balance between performance and security, reducing the attack success rate by 2-10 percentage points while sacrificing only about 1% in accuracy. Furthermore, DAM exhibits robust performance and broad applicability across various types of backdoor attacks and the number of compromised models involved in the merging process. Our codes and models can be accessed through DAM. * Corresponding Authors. • We reveal for the first time the vulnerabilities of current multi-task merging methods under backdoor attacks, identifying "backdoor succession" and "backdoor transfer" as core challenges. • We propose a novel Defense-Aware Merging (DAM) algorithm through a dual-mask optimization, that not only mitigates task interference but also effectively alleviates the backdoor effect for the multi-task model merging process. • Extensive experiments validate the effectiveness of DAM, demonstrating significant performance improvements across multiple benchmarks and backbone networks while maintaining minimal accuracy degradation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 474e992c-733a-4f6d-9d23-412cb5e15c3fCited by top-tier papers10
- Backdoor Cleaning without External Guidance in MLLM Fine-tuningXuankun Rong, Wenke Huang, Jian Liang, Jinhe Bi et al.NeurIPS 2025 · 39 citations
- Mix Data or Merge Models? Balancing the Helpfulness, Honesty, and Harmlessness of Large Language Model via Model MergingJinluan Yang, Dingnan Jin, Anke Tang, Li Shen et al.NeurIPS 2025 · 23 citations
- Knowledge Fusion of Large Language Models via Modular SkillPacksGuodong Du, Zhuo Li, Xuanning Zhou, Junlin Li et al.ICLR 2026 · 9 citations
- BadTV: Unveiling Backdoor Threats in Third-Party Task VectorsChia-Yi Hsu, Yu-Lin Tsai, Zhe Yu, Yan-Lun Chen et al.CCS 2026 · 2 citations
- Defending against Backdoor Attacks via Module SwitchingWeijun Li, Ansh Arora, Xuanli He, Mark Dras et al.ICLR 2026 · 2 citations
Builds on29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
- TIES-Merging: Resolving Interference When Merging ModelsPrateek Yadav, Derek Tam, Leshem Choshen, Colin A. Raffel et al.NeurIPS 2023 · 999 citations
- Linear Mode Connectivity and the Lottery Ticket HypothesisJonathan Frankle, Gintare Karolina Dziugaite, Daniel M. Roy, Michael CarbinICML 2020 · 750 citations
- Merging Models with Fisher-Weighted AveragingMichael Matena, Colin RaffelNeurIPS 2022 · 741 citations
Related papers
- BadMerging: Backdoor Attacks Against Model MergingJinghuai Zhang, Jianfeng Chi, Zheng Li, Kunlin Cai et al.CCS 2024 · 5 citations
- From Parameters to Feature Space: Task Arithmetic for Backdoor Mitigation in Model MergingZhenqian Zhu, Yamin Hu, Yiya Diao, Weixiang Li et al.ICML 2026 · 1 citation
- Merge Hijacking: Backdoor Attacks to Model Merging of Large Language ModelsZenghui Yuan, Yangming Xu, Jiawen Shi, Pan Zhou et al.ACL 2025 · 5 citations
- From Purity to Peril: Backdooring Merged Models From "Harmless" Benign ComponentsLijin Wang, Jingjing Wang, Tianshuo Cong, Xinlei He et al.USENIX Security 2025
- Anti-Backdoor Learning: Training Clean Models on Poisoned DataYige Li, Xixiang Lyu, Nodens Koren, Lingjuan Lyu et al.NeurIPS 2021 · 503 citations
