Unmasking Backdoors: An Explainable Defense via Gradient-Attention Anomaly Scoring for Pre-trained Language Models
Anindya Sundar Das, Kangjie Chen, Monowar Bhuyan
摘要
Pre-trained language models have achieved remarkable success across a wide range of natural language processing (NLP) tasks, particularly when fine-tuned on large, domain-relevant datasets. However, they remain vulnerable to backdoor attacks, where adversaries embed malicious behaviors using trigger patterns in the training data. These triggers remain dormant during normal usage, but, when activated, can cause targeted misclassifications. In this work, we investigate the internal behavior of backdoored pre-trained encoder-based language models, focusing on the consistent shift in attention and gradient attribution when processing poisoned inputs; where the trigger token dominates both attention and gradient signals, overriding the surrounding context. We propose an inference-time defense that constructs anomaly scores by combining token-level attention and gradient information. Extensive experiments on text classification tasks across diverse backdoor attack scenarios demonstrate that our method significantly reduces attack success rates compared to existing baselines. Furthermore, we provide an interpretability-driven analysis of the scoring mechanism, shedding light on trigger localization and the robustness of the proposed defense.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Trojaning Attack on Neural NetworksYingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee 等NDSS 2018 · 被引用 1,377 次
- Weight Poisoning Attacks on Pretrained ModelsKeita Kurita, Paul Michel, Graham NeubigACL 2020 · 被引用 312 次
- BadPre: Task-agnostic Backdoor Attacks to Pre-trained NLP Foundation ModelsKangjie Chen, Yuxian Meng, Xiaofei Sun, Shangwei Guo 等ICLR 2022 · 被引用 133 次
- Mind the Style of Text! Adversarial and Backdoor Attacks Based on Text Style TransferFanchao Qi, Yangyi Chen, Xurui Zhang, Mukai Li 等EMNLP 2021 · 被引用 114 次
- Backdoor Attacks on Pre-trained Models by Layerwise Weight PoisoningLinyang Li, Demin Song, Xiaonan Li, Jiehang Zeng 等EMNLP 2021 · 被引用 93 次
相关 Paper
- Backdoor Pre-trained Models Can Transfer to AllLujia Shen, Shouling Ji, Xuhong Zhang, Jinfeng Li 等CCS 2021 · 被引用 72 次
- When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated ExplanationsHuaizhi Ge, Yiming Li, Qifan Wang, Yongfeng Zhang 等ACL 2025
- Test-Time Attention Purification for Backdoored Large Vision Language ModelsZhifang Zhang, Bojun Yang, Shuo He, Weitong Chen 等CVPR 2026 · 被引用 7 次
- Defense against Backdoor Attack on Pre-trained Language Models via Head Pruning and Attention NormalizationXingyi Zhao, Depeng Xu, Shuhan YuanICML 2024 · 被引用 17 次
- Moderate-fitting as a Natural Backdoor Defender for Pre-trained Language ModelsBiru Zhu, Yujia Qin, Ganqu Cui, Yangyi Chen 等NeurIPS 2022 · 被引用 29 次
