When Backdoors Speak: Understanding LLM Backdoor Attacks Through Model-Generated Explanations
Huaizhi Ge, Yiming Li, Qifan Wang, Yongfeng Zhang, Ruixiang Tang
Abstract
Large Language Models (LLMs) are known to be vulnerable to backdoor attacks, where triggers embedded in poisoned samples can maliciously alter LLMs' behaviors. In this paper, we move beyond attacking LLMs and instead examine backdoor attacks through the novel lens of natural language explanations. Specifically, we leverage LLMs' generative capabilities to produce human-readable explanations for their decisions, enabling direct comparisons between explanations for clean and poisoned samples. Our results show that backdoored models produce coherent explanations for clean inputs but diverse and logically flawed explanations for poisoned data, a pattern consistent across classification and generation tasks for different backdoor attacks. Further analysis reveals key insights into the explanation generation process. At the token level, explanation tokens associated with poisoned samples only appear in the final few transformer layers. At the sentence level, attention dynamics indicate that poisoned inputs shift attention away from the original input context during explanation generation. These observations enhance our understanding of the mechanisms behind backdoor attacks in LLMs and shed light on leveraging explanations for backdoor detection.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c44e1032-8770-4066-8ab1-f962b2d6fd0bCited by top-tier papers5
- Beyond the Surface: Enhancing LLM-as-a-Judge Alignment with Human via Internal RepresentationsPeng Lai, Jianjie Zheng, Sijie Cheng, Yun Chen et al.NeurIPS 2025 · 16 citations
- Echoes within the Reasoning: Stealthy and Effective Watermarking via Chain of ThoughtJiacheng Lu, Yiming Li, Tao Song, Weijian Wang et al.ICML 2026 · 2 citations
- Hollow-LLM Attack: Computationally Trivial Weights in Zero-Knowledge Verification of LLM InferenceChen Gong, Beijie Liu, Mengyuan LiS&P 2026 · 2 citations
- MirageBackdoor: A Stealthy Attack that Induces Think-Well-Answer-Wrong ReasoningYizhe Zeng, Wei Zhang, Yunpeng Li, Juxin Xiao et al.ACL 2026
- Uncovering Hidden Triggers: Backdoor Attribution in Language ModelsMiao Yu, Zhenhong Zhou, Moayad Aloqaily, Kun Wang et al.ICML 2026
Builds on10
- Trojaning Attack on Neural NetworksYingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee et al.NDSS 2018 · 1,377 citations
- Poisoning Language Models During Instruction TuningAlexander Wan, Eric Wallace, Sheng Shen, Dan KleinICML 2023 · 319 citations
- Weight Poisoning Attacks on Pretrained ModelsKeita Kurita, Paul Michel, Graham NeubigACL 2020 · 312 citations
- The Unreliability of Explanations in Few-shot Prompting for Textual ReasoningXi Ye, Greg DurrettNeurIPS 2022 · 272 citations
- An Embarrassingly Simple Approach for Trojan Attack in Deep Neural NetworksRuixiang Tang, Mengnan Du, Ninghao Liu, Fan Yang et al.KDD 2020 · 164 citations
Related papers
- EmbedX: Embedding-Based Cross-Trigger Backdoor Attack Against Large Language ModelsNan Yan, Yuqing Li, Xiong Wang, Jing Chen et al.USENIX Security 2025
- Lethe: Purifying Backdoored Large Language Models with Knowledge DilutionChen Chen, Yuchen Sun, Jiaxin Gao, Xueluan Gong et al.USENIX Security 2026 · 1 citation
- Unmasking Backdoors: An Explainable Defense via Gradient-Attention Anomaly Scoring for Pre-trained Language ModelsAnindya Sundar Das, Kangjie Chen, Monowar BhuyanICLR 2026 · 4 citations
- Backdoor Attacks in Token Selection of Attention MechanismYunjuan Wang, Raman AroraICML 2025
- Critical-CoT: A Robust Defense Framework against Reasoning-Level Backdoor Attacks in Large Language ModelsVu Tuan Truong, Long Bao LeACL 2026
