Causality Based Front-door Defense Against Backdoor Attack on Language Models
Yiran Liu, Xiaoang Xu, Zhiyi Hou, Yang Yu
Abstract
We have developed a new framework based on the theory of causal inference to protect language models against backdoor attacks. Backdoor attackers can poison language models with different types of triggers, such as words, sentences, grammar, and style, enabling them to selectively modify the decision-making of the victim model. However, existing defense approaches are only effective when the backdoor attack form meets specific assumptions, making it difficult to counter diverse backdoor attacks. We propose a new defense framework Front-door Adjustment for Backdoor Elimination (FABE) based on causal reasoning that does not rely on assumptions about the form of triggers. This method effectively differentiates between spurious and legitimate associations by creating a 'front door' that maps out the actual causal relationships. The term 'front door' refers to a text that retains the semantic equivalence of the initial input, which is generated by an additional, fine-tuned language model, denoted as the defense model. Our defense experiments against various attack methods at the token, sentence, and syntactic levels reduced the attack success rate from 93.63% to 15.12%, improving the defense effect by 2.91 times compared to the best baseline result of 66.61%, achieving state-of-the-art results. Through ablation study analysis, we analyzed the effect of each module in FABE, demonstrating the importance of complying with the front-door criterion and front-door adjustment for-* Equal contribution
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- ICLScan: Detecting Backdoors in Black-Box Large Language Models via Targeted In-context IlluminationXiaoyi Pang, Xuanyi Hao, Song Guo, Qi Luo et al.NeurIPS 2025 · 7 citations
- Paladin: Defending LLM-enabled Phishing Emails with a New Trigger-Tag ParadigmYan Pang, Wenlong Meng, Xiaojing Liao, Tianhao WangNDSS 2026 · 5 citations
- Rethinking and Red-Teaming Protective Perturbation in Personalized Diffusion ModelsYixin Liu, Ruoxi Chen, Xun Chen, Lichao SunKDD 2026 · 3 citations
- Lethe: Purifying Backdoored Large Language Models with Knowledge DilutionChen Chen, Yuchen Sun, Jiaxin Gao, Xueluan Gong et al.USENIX Security 2026 · 1 citation
- BYORn: Bootstrap Your Own Responses to Defend Large Vision-Language Models Against Backdoor AttacksIvan Sabolic, Marin Oršić, Josip Šarić, Sven LoncaricICML 2026
Builds on16
- Cross-Task Generalization via Natural Language Crowdsourcing InstructionsSwaroop Mishra, Daniel Khashabi, Chitta Baral, Hannaneh HajishirziACL 2022 · 887 citations
- Investigating Gender Bias in Language Models Using Causal Mediation AnalysisJesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian et al.NeurIPS 2020 · 851 citations
- Causal Intervention for Weakly-Supervised Semantic SegmentationDong Zhang, Hanwang Zhang, Jinhui Tang, Xian-Sheng Hua et al.NeurIPS 2020 · 563 citations
- Long-Tailed Classification by Keeping the Good and Removing the Bad Momentum Causal EffectKaihua Tang, Jianqiang Huang, Hanwang ZhangNeurIPS 2020 · 533 citations
- Extractive Summarization as Text MatchingMing Zhong, Pengfei Liu, Yiran Chen, Danqing Wang et al.ACL 2020 · 410 citations
Related papers
- Unmasking Backdoors: An Explainable Defense via Gradient-Attention Anomaly Scoring for Pre-trained Language ModelsAnindya Sundar Das, Kangjie Chen, Monowar BhuyanICLR 2026 · 4 citations
- Critical-CoT: A Robust Defense Framework against Reasoning-Level Backdoor Attacks in Large Language ModelsVu Tuan Truong, Long Bao LeACL 2026
- Backdoor Collapse: Eliminating Unknown Threats Via Known Backdoor Aggregation In Language ModelsLiang Lin, Miao Yu, Moayad Aloqaily, Zhenhong Zhou et al.ACL 2026 · 4 citations
- Moderate-fitting as a Natural Backdoor Defender for Pre-trained Language ModelsBiru Zhu, Yujia Qin, Ganqu Cui, Yangyi Chen et al.NeurIPS 2022 · 29 citations
- Mitigating Backdoor Poisoning Attacks through the Lens of Spurious CorrelationXuanli He, Qiongkai Xu, Jun Wang, Benjamin I. P. Rubinstein et al.EMNLP 2023 · 9 citations
