Defense against Backdoor Attack on Pre-trained Language Models via Head Pruning and Attention Normalization
Xingyi Zhao, Depeng Xu, Shuhan Yuan
Abstract
Pre-trained language models (PLMs) are commonly used for various downstream natural language processing tasks via fine-tuning. However, recent studies have demonstrated that PLMs are vulnerable to backdoor attacks, which can mislabel poisoned samples to target outputs even after a vanilla fine-tuning process. The key challenge for defending against the backdoored PLMs is that end users who adopt the PLMs for their downstream tasks usually do not have any knowledge about the attacking strategies, such as triggers. To tackle this challenge, in this work, we propose a backdoor mitigation approach, PURE, via head pruning and normalization of attention weights. The idea is to prune the attention heads that are potentially affected by poisoned texts with only clean texts on hand and then further normalize the weights of remaining attention heads to mitigate the backdoor impacts. We conduct experiments to defend against various backdoor attacks on the classification task. The experimental results show the effectiveness of PURE in lowering the attack success rate without sacrificing the performance on clean texts. The code is available at https: //github.com/xingyizhao/PURE . Defense against Backdoor Attack on Pre-trained Language Models via Head Pruning and Attention Normalization Backdoor model detection employs various trigger inversion techniques to reverse-engineer the injected trigger which is then utilized to ascertain whether a PLM has been poisoned. Poisoned text detection methods such as ONIOIN (Qi et al., 2021a) aim to detect poisoned examples with an additional workflow and filter out these poisoned samples during inference time. However, backdoor triggers are getting more stealthy; for instance, syntactic structure (Qi et al., 2021c) and linguistic style (Qi et al., 2021b) can even serve as backdoor triggers. Consequently, it is challenging to reverse or detect these triggers. Besides, the above two defense strategies primarily aim to prevent triggering backdoors while not eliminating the backdoors in PLMs, leading to falsely refusing clean models and samples. Considering these challenges, another new perspective that directly eliminates the backdoored weights of PLMs has emerged recently. Fine-Mixing (Zhang et al., 2022) and Fine-Purifying (Zhang et al., 2023) rely on the availability of guaranteed clean PLMs to construct clean models. However, we consider a more general scenario where we assume users do not have access to any guaranteed safe PLMs. Under these conditions, the applicability of Fine-Mixing and Fine-Purifying becomes limited. Liu et al. ( 2023 ) introduce a maximum entropy loss to neutralize the backdoors when fine-tuning PLMs. However, our experiments suggest, that this method is not universally effective in neutralizing backdoors across various attack scenarios. Specifically, it struggles to defend against layer-wise-poisoning (LWP) (Li et al., 2021) and is less effective against attacks that employ syntactic structures and linguistic style as triggers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ffa0d816-c659-40d9-930a-d2fa6f66e293Cited by top-tier papers7
- Purifying Generative LLMs from Backdoors without Prior Knowledge or Clean ReferenceJianwei Li, Jung-Eun KimICLR 2026 · 8 citations
- Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag ParadigmYan Pang, Wenlong Meng, Xiaojing Liao, Tianhao WangNDSS 2026 · 5 citations
- Unmasking Backdoors: An Explainable Defense via Gradient-Attention Anomaly Scoring for Pre-trained Language ModelsAnindya Sundar Das, Kangjie Chen, Monowar BhuyanICLR 2026 · 4 citations
- Defending against Backdoor Attacks via Module SwitchingWeijun Li, Ansh Arora, Xuanli He, Mark Dras et al.ICLR 2026 · 2 citations
- BeDKD: Backdoor Defense Based on Directional Mapping Module and Adversarial Knowledge DistillationZhengxian Wu, Juan Wen, Wanli Peng, Yinghan Zhou et al.AAAI 2026 · 2 citations
Builds on13
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li et al.S&P 2019 · 1,801 citations
- Weight Poisoning Attacks on Pretrained ModelsKeita Kurita, Paul Michel, Graham NeubigACL 2020 · 312 citations
- BadPre: Task-agnostic Backdoor Attacks to Pre-trained NLP Foundation ModelsKangjie Chen, Yuxian Meng, Xiaofei Sun, Shangwei Guo et al.ICLR 2022 · 133 citations
- Mind the Style of Text! Adversarial and Backdoor Attacks Based on Text Style TransferFanchao Qi, Yangyi Chen, Xurui Zhang, Mukai Li et al.EMNLP 2021 · 114 citations
- Backdoor Attacks on Pre-trained Models by Layerwise Weight PoisoningLinyang Li, Demin Song, Xiaonan Li, Jiehang Zeng et al.EMNLP 2021 · 93 citations
Related papers
- Moderate-fitting as a Natural Backdoor Defender for Pre-trained Language ModelsBiru Zhu, Yujia Qin, Ganqu Cui, Yangyi Chen et al.NeurIPS 2022 · 29 citations
- Purifier: Plug-and-play Backdoor Mitigation for Pre-trained Models Via Anomaly Activation SuppressionXiaoyu Zhang, Yulin Jin, Tao Wang, Jian Lou et al.ACM MM 2022 · 12 citations
- Test-Time Attention Purification for Backdoored Large Vision Language ModelsZhifang Zhang, Bojun Yang, Shuo He, Weitong Chen et al.CVPR 2026 · 7 citations
- PurMM: Attention-Guided Test-Time Backdoor Purification in Multimodal Large Language ModelsWenzheng Jiang, Ke Liang, Xuankun Rong, Jingxuan Zhou et al.AAAI 2026
- LT-Defense: Searching-free Backdoor Defense via Exploiting the Long-tailed EffectYixiao Xu, Binxing Fang, Mohan Li, Keke Tang et al.NeurIPS 2024 · 7 citations
