When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations
Huaizhi Ge, Yiming Li, Qifan Wang, Yongfeng Zhang, Ruixiang Tang
摘要
Large Language Models (LLMs) are known to be vulnerable to backdoor attacks, where triggers embedded in poisoned samples can maliciously alter LLMs' behaviors. In this paper, we move beyond attacking LLMs and instead examine backdoor attacks through the novel lens of natural language explanations. Specifically, we leverage LLMs' generative capabilities to produce human-readable explanations for their decisions, enabling direct comparisons between explanations for clean and poisoned samples. Our results show that backdoored models produce coherent explanations for clean inputs but diverse and logically flawed explanations for poisoned data, a pattern consistent across classification and generation tasks for different backdoor attacks. Further analysis reveals key insights into the explanation generation process. At the token level, explanation tokens associated with poisoned samples only appear in the final few transformer layers. At the sentence level, attention dynamics indicate that poisoned inputs shift attention away from the original input context during explanation generation. These observations enhance our understanding of the mechanisms behind backdoor attacks in LLMs and shed light on leveraging explanations for backdoor detection.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Beyond the Surface: Enhancing LLM-as-a-Judge Alignment with Human via Internal RepresentationsPeng Lai, Jianjie Zheng, Sijie Cheng, Yun Chen 等NeurIPS 2025 · 被引用 16 次
- Echoes within the Reasoning: Stealthy and Effective Watermarking via Chain of ThoughtJiacheng Lu, Yiming Li, Tao Song, Weijian Wang 等ICML 2026 · 被引用 2 次
- Hollow-LLM Attack: Computationally Trivial Weights in Zero-Knowledge Verification of LLM InferenceChen Gong, Beijie Liu, Mengyuan LiS&P 2026 · 被引用 2 次
- MirageBackdoor: A Stealthy Attack that Induces Think-Well-Answer-Wrong ReasoningYizhe Zeng, Wei Zhang, Yunpeng Li, Juxin Xiao 等ACL 2026
- Uncovering Hidden Triggers: Backdoor Attribution in Language ModelsMiao Yu, Zhenhong Zhou, Moayad Aloqaily, Kun Wang 等ICML 2026
它引用的顶会 Paper10
- Trojaning Attack on Neural NetworksYingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee 等NDSS 2018 · 被引用 1,377 次
- Poisoning Language Models During Instruction TuningAlexander Wan, Eric Wallace, Sheng Shen, Dan KleinICML 2023 · 被引用 319 次
- Weight Poisoning Attacks on Pretrained ModelsKeita Kurita, Paul Michel, Graham NeubigACL 2020 · 被引用 312 次
- The Unreliability of Explanations in Few-shot Prompting for Textual ReasoningXi Ye, Greg DurrettNeurIPS 2022 · 被引用 272 次
- An Embarrassingly Simple Approach for Trojan Attack in Deep Neural NetworksRuixiang Tang, Mengnan Du, Ninghao Liu, Fan Yang 等KDD 2020 · 被引用 164 次
相关 Paper
- EmbedX: Embedding-Based Cross-Trigger Backdoor Attack Against Large Language ModelsNan Yan, Yuqing Li, Xiong Wang, Jing Chen 等USENIX Security 2025
- Lethe: Purifying Backdoored Large Language Models with Knowledge DilutionChen Chen, Yuchen Sun, Jiaxin Gao, Xueluan Gong 等USENIX Security 2026 · 被引用 1 次
- Unmasking Backdoors: An Explainable Defense via Gradient-Attention Anomaly Scoring for Pre-trained Language ModelsAnindya Sundar Das, Kangjie Chen, Monowar BhuyanICLR 2026 · 被引用 4 次
- Backdoor Attacks in Token Selection of Attention MechanismYunjuan Wang, Raman AroraICML 2025
- Critical-CoT: A Robust Defense Framework against Reasoning-Level Backdoor Attacks in Large Language ModelsVu Tuan Truong, Long Bao LeACL 2026
