PEFTGuard: Detecting Backdoor Attacks Against Parameter-Efficient Fine-Tuning
Zhen Sun, Tianshuo Cong, Yule Liu, Chenhao Lin, Xinlei He, Rongmao Chen, Xingshuo Han, Xinyi Huang
Abstract
Fine-tuning is an essential process to improve the performance of Large Language Models (LLMs) in specific domains, with Parameter-Efficient Fine-Tuning (PEFT) gaining popularity due to its capacity to reduce computational demands through the integration of low-rank adapters. These lightweight adapters, such as LoRA, can be shared and utilized on open-source platforms. However, adversaries could exploit this mechanism to inject backdoors into these adapters, resulting in malicious behaviors like incorrect or harmful outputs, which pose serious security risks to the community. Unfortunately, few current efforts concentrate on analyzing the backdoor patterns or detecting the backdoors in the adapters. To fill this gap, we first construct and release PADBench, a comprehensive benchmark that contains 13, 300 benign and backdoored adapters fine-tuned with various datasets, attack strategies, PEFT methods, and LLMs. Moreover, we propose PEFTGuard, the first backdoor detection framework against PEFT-based adapters. Extensive evaluation upon PADBench shows that PEFTGuard outperforms existing detection methods, achieving nearly perfect detection accuracy (100%) in most cases. Notably, PEFTGuard exhibits zero-shot transferability on three aspects, including different attacks, PEFT methods, and adapter ranks. In addition, we consider various adaptive attacks to demonstrate the high robustness of PEFTGuard. We further explore several possible backdoor mitigation defenses, finding fine-mixing to be the most effective method. We envision that our benchmark and method can shed light on future LLM backdoor detection research. 11Our code and dataset are available at: https://github.com/Vincent-HKUSTGZ/PEFTGuard.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers12
- CL-Attack: Textual Backdoor Attacks via Cross-Lingual TriggersJingyi Zheng, Tianyi Hu, Tianshuo Cong, Xinlei HeAAAI 2025 · 13 citations
- PromptCOS: Towards Content-Only System Prompt Copyright Auditing for LLMsYuchen Yang, Yiming Li, Hongwei Yao, Enhao Huang et al.S&P 2026 · 5 citations
- DecepChain: Inducing Deceptive Reasoning in Large Language ModelsWei Shen, Han Wang, Haoyu Li, Huan ZhangICML 2026 · 4 citations
- Causal-Guided Detoxify Backdoor Attack of Open-Weight LoRA ModelsLinzhi Chen, Yang Sun, Hongru Wei, Yuqi ChenNDSS 2026 · 4 citations
- 6DAttack: Backdoor Attacks in the 6DoF Pose EstimationJihui Guo, Zongmin Zhang, Zhen Sun, Yuhao Yang et al.AAAI 2026 · 2 citations
Builds on34
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
Related papers
- ELBA-Bench: An Efficient Learning Backdoor Attacks Benchmark for Large Language ModelsXuxu Liu, Siyuan Liang, Mengya Han, Yong Luo et al.ACL 2025 · 13 citations
- INDEXGUARD: Index-only Backdoor Vetting for Secure Federated PEFT of Large Language ModelsJavad Dogani, Devriş İşler, Nikolaos LaoutarisICML 2026
- SaLoRA: Safety-Alignment Preserved Low-Rank AdaptationMingjie Li, Wai Man Si, Michael Backes, Yang Zhang et al.ICLR 2025
- Refining Salience-Aware Sparse Fine-Tuning Strategies for Language ModelsXinxin Liu, Aaron Thomas, Cheng Zhang, Jianyi Cheng et al.ACL 2025 · 3 citations
- Trans-LoRA: towards data-free Transferable Parameter Efficient FinetuningRunqian Wang, Soumya Ghosh, David D. Cox, Diego Antognini et al.NeurIPS 2024 · 15 citations
