Merger-as-a-Stealer: Stealing Targeted PII from Aligned LLMs with Model Merging
Lin Lu, Zhigang Zuo, Ziji Sheng, Pan Zhou
Abstract
Model merging has emerged as a promising approach for updating large language models (LLMs) by integrating multiple domain-specific models into a cross-domain merged model. Despite its utility and plug-and-play nature, unmonitored mergers can introduce significant security vulnerabilities, such as backdoor attacks and model merging abuse. In this paper, we identify a novel and more realistic attack surface where a malicious merger can extract targeted personally identifiable information (PII) from an aligned model with model merging. Specifically, we propose Merger-as-a-Stealer, a two-stage framework to achieve this attack: First, the attacker fine-tunes a malicious model to force it to respond to any PII-related queries. The attacker then uploads this malicious model to the model merging conductor and obtains the merged model. Second, the attacker inputs direct PII-related queries to the merged model to extract targeted PII. Extensive experiments demonstrate that Merger-as-a-Stealer successfully executes attacks against various LLMs and model merging methods across diverse settings, highlighting the effectiveness of the proposed framework. Given that this attack enables character-level extraction for targeted PII without requiring any additional knowledge from the attacker, we stress the necessity for improved model alignment and more robust defense mechanisms to mitigate such threats.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- LearnerCoMPASS: Intelligent Tutoring System with Dynamic Cognitive Diagnosis and Multi-Model Path PlanningZiji Sheng, Guiyao Tie, Weidong Wang, Pan Zhou et al.ACL 2026
- An Empirical Study on the Resilience of Partial Merging to Model Clone AttacksTiantong Wu, Yurong Hao, Wei Yang Bryan LimICML 2026
Builds on16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski et al.USENIX Security 2021 · 2,866 citations
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
Related papers
- Merge Hijacking: Backdoor Attacks to Model Merging of Large Language ModelsZenghui Yuan, Yangming Xu, Jiawen Shi, Pan Zhou et al.ACL 2025 · 5 citations
- BadMerging: Backdoor Attacks Against Model MergingJinghuai Zhang, Jianfeng Chi, Zheng Li, Kunlin Cai et al.CCS 2024 · 5 citations
- Private Investigator: Extracting Personally Identifiable Information from Large Language Models Using Optimized PromptsSeongho Keum, Dongwon Shin, Leo Marchyok, Sanghyun Hong et al.USENIX Security 2025
- Do Not Merge My Model! Safeguarding Open-Source LLMs Against Unauthorized Model MergingQinfeng Li, Miao Pan, Jintao Chen, Fu Teng et al.AAAI 2026 · 1 citation
- Effective PII Extraction from LLMs through Augmented Few-Shot LearningShuai Cheng, Shu Meng, Haitao Xu, Haoran Zhang et al.USENIX Security 2025
