Defense against Backdoor Attack on Pre-trained Language Models via Head Pruning and Attention Normalization
Xingyi Zhao, Depeng Xu, Shuhan Yuan
摘要
Pre-trained language models (PLMs) are commonly used for various downstream natural language processing tasks via fine-tuning. However, recent studies have demonstrated that PLMs are vulnerable to backdoor attacks, which can mislabel poisoned samples to target outputs even after a vanilla fine-tuning process. The key challenge for defending against the backdoored PLMs is that end users who adopt the PLMs for their downstream tasks usually do not have any knowledge about the attacking strategies, such as triggers. To tackle this challenge, in this work, we propose a backdoor mitigation approach, PURE, via head pruning and normalization of attention weights. The idea is to prune the attention heads that are potentially affected by poisoned texts with only clean texts on hand and then further normalize the weights of remaining attention heads to mitigate the backdoor impacts. We conduct experiments to defend against various backdoor attacks on the classification task. The experimental results show the effectiveness of PURE in lowering the attack success rate without sacrificing the performance on clean texts. The code is available at https: //github.com/xingyizhao/PURE . Defense against Backdoor Attack on Pre-trained Language Models via Head Pruning and Attention Normalization Backdoor model detection employs various trigger inversion techniques to reverse-engineer the injected trigger which is then utilized to ascertain whether a PLM has been poisoned. Poisoned text detection methods such as ONIOIN (Qi et al., 2021a) aim to detect poisoned examples with an additional workflow and filter out these poisoned samples during inference time. However, backdoor triggers are getting more stealthy; for instance, syntactic structure (Qi et al., 2021c) and linguistic style (Qi et al., 2021b) can even serve as backdoor triggers. Consequently, it is challenging to reverse or detect these triggers. Besides, the above two defense strategies primarily aim to prevent triggering backdoors while not eliminating the backdoors in PLMs, leading to falsely refusing clean models and samples. Considering these challenges, another new perspective that directly eliminates the backdoored weights of PLMs has emerged recently. Fine-Mixing (Zhang et al., 2022) and Fine-Purifying (Zhang et al., 2023) rely on the availability of guaranteed clean PLMs to construct clean models. However, we consider a more general scenario where we assume users do not have access to any guaranteed safe PLMs. Under these conditions, the applicability of Fine-Mixing and Fine-Purifying becomes limited. Liu et al. ( 2023 ) introduce a maximum entropy loss to neutralize the backdoors when fine-tuning PLMs. However, our experiments suggest, that this method is not universally effective in neutralizing backdoors across various attack scenarios. Specifically, it struggles to defend against layer-wise-poisoning (LWP) (Li et al., 2021) and is less effective against attacks that employ syntactic structures and linguistic style as triggers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Purifying Generative LLMs from Backdoors without Prior Knowledge or Clean ReferenceJianwei Li, Jung-Eun KimICLR 2026 · 被引用 8 次
- Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag ParadigmYan Pang, Wenlong Meng, Xiaojing Liao, Tianhao WangNDSS 2026 · 被引用 5 次
- Unmasking Backdoors: An Explainable Defense via Gradient-Attention Anomaly Scoring for Pre-trained Language ModelsAnindya Sundar Das, Kangjie Chen, Monowar BhuyanICLR 2026 · 被引用 4 次
- Defending against Backdoor Attacks via Module SwitchingWeijun Li, Ansh Arora, Xuanli He, Mark Dras 等ICLR 2026 · 被引用 2 次
- BeDKD: Backdoor Defense Based on Directional Mapping Module and Adversarial Knowledge DistillationZhengxian Wu, Juan Wen, Wanli Peng, Yinghan Zhou 等AAAI 2026 · 被引用 2 次
它引用的顶会 Paper13
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li 等S&P 2019 · 被引用 1,801 次
- Weight Poisoning Attacks on Pretrained ModelsKeita Kurita, Paul Michel, Graham NeubigACL 2020 · 被引用 312 次
- BadPre: Task-agnostic Backdoor Attacks to Pre-trained NLP Foundation ModelsKangjie Chen, Yuxian Meng, Xiaofei Sun, Shangwei Guo 等ICLR 2022 · 被引用 133 次
- Mind the Style of Text! Adversarial and Backdoor Attacks Based on Text Style TransferFanchao Qi, Yangyi Chen, Xurui Zhang, Mukai Li 等EMNLP 2021 · 被引用 114 次
- Backdoor Attacks on Pre-trained Models by Layerwise Weight PoisoningLinyang Li, Demin Song, Xiaonan Li, Jiehang Zeng 等EMNLP 2021 · 被引用 93 次
相关 Paper
- Moderate-fitting as a Natural Backdoor Defender for Pre-trained Language ModelsBiru Zhu, Yujia Qin, Ganqu Cui, Yangyi Chen 等NeurIPS 2022 · 被引用 29 次
- Purifier: Plug-and-play Backdoor Mitigation for Pre-trained Models Via Anomaly Activation SuppressionXiaoyu Zhang, Yulin Jin, Tao Wang, Jian Lou 等ACM MM 2022 · 被引用 12 次
- Test-Time Attention Purification for Backdoored Large Vision Language ModelsZhifang Zhang, Bojun Yang, Shuo He, Weitong Chen 等CVPR 2026 · 被引用 7 次
- PurMM: Attention-Guided Test-Time Backdoor Purification in Multimodal Large Language ModelsWenzheng Jiang, Ke Liang, Xuankun Rong, Jingxuan Zhou 等AAAI 2026
- LT-Defense: Searching-free Backdoor Defense via Exploiting the Long-tailed EffectYixiao Xu, Binxing Fang, Mohan Li, Keke Tang 等NeurIPS 2024 · 被引用 7 次
