Merge Hijacking: Backdoor Attacks to Model Merging of Large Language Models
Zenghui Yuan, Yangming Xu, Jiawen Shi, Pan Zhou, Lichao Sun
摘要
Model merging for Large Language Models (LLMs) directly fuses the parameters of different models finetuned on various tasks, creating a unified model for multi-domain tasks. However, due to potential vulnerabilities in models available on open-source platforms, model merging is susceptible to backdoor attacks. In this paper, we propose Merge Hijacking, the first backdoor attack targeting model merging in LLMs. The attacker constructs a malicious upload model and releases it. Once a victim user merges it with any other models, the resulting merged model inherits the backdoor while maintaining utility across tasks. Merge Hijacking defines two main objectives-effectiveness and utility-and achieves them through four steps. Extensive experiments demonstrate the effectiveness of our attack across different models, merging algorithms, and tasks. Additionally, we show that the attack remains effective even when merging real-world models. Moreover, our attack demonstrates robustness against two inference-time defenses (Paraphrasing and CLEANGEN) and one training-time defense (Fine-pruning).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- BadTV: Unveiling Backdoor Threats in Third-Party Task VectorsChia-Yi Hsu, Yu-Lin Tsai, Zhe Yu, Yan-Lun Chen 等CCS 2026 · 被引用 2 次
- From Parameters to Feature Space: Task Arithmetic for Backdoor Mitigation in Model MergingZhenqian Zhu, Yamin Hu, Yiya Diao, Weixiang Li 等ICML 2026 · 被引用 1 次
- FairMerging: Rethinking Model Merging through the Lens of FairnessBing Liu, Xinrui Shan, Boyu Zhang, Qiankun Zhang 等ICML 2026
- Activation Decomposition and Steering for LLM Backdoor RemediationLingfeng Zhong, Qiongkai Xu, Usman NaseemACL 2026
它引用的顶会 Paper13
- Trojaning Attack on Neural NetworksYingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee 等NDSS 2018 · 被引用 1,377 次
- Language Models are Super Mario: Absorbing Abilities from Homologous Models as a Free LunchLe Yu, Bowen Yu, Haiyang Yu, Fei Huang 等ICML 2024 · 被引用 605 次
- BadEncoder: Backdoor Attacks to Pre-trained Encoders in Self-Supervised LearningJinyuan Jia, Yupei Liu, Neil Zhenqiang GongS&P 2022 · 被引用 200 次
- BadPre: Task-agnostic Backdoor Attacks to Pre-trained NLP Foundation ModelsKangjie Chen, Yuxian Meng, Xiaofei Sun, Shangwei Guo 等ICLR 2022 · 被引用 133 次
- Backdoor Pre-trained Models Can Transfer to AllLujia Shen, Shouling Ji, Xuhong Zhang, Jinfeng Li 等CCS 2021 · 被引用 72 次
相关 Paper
- BadMerging: Backdoor Attacks Against Model MergingJinghuai Zhang, Jianfeng Chi, Zheng Li, Kunlin Cai 等CCS 2024 · 被引用 5 次
- Merger-as-a-Stealer: Stealing Targeted PII from Aligned LLMs with Model MergingLin Lu, Zhigang Zuo, Ziji Sheng, Pan ZhouEMNLP 2025 · 被引用 1 次
- From Purity to Peril: Backdooring Merged Models From "Harmless" Benign ComponentsLijin Wang, Jingjing Wang, Tianshuo Cong, Xinlei He 等USENIX Security 2025
- Merging Triggers, Breaking Backdoors: Defensive Poisoning for Instruction-Tuned Language ModelsSan Kim, Gary LeeACL 2026
- PurMM: Attention-Guided Test-Time Backdoor Purification in Multimodal Large Language ModelsWenzheng Jiang, Ke Liang, Xuankun Rong, Jingxuan Zhou 等AAAI 2026
