BeDKD: Backdoor Defense Based on Directional Mapping Module and Adversarial Knowledge Distillation
Zhengxian Wu, Juan Wen, Wanli Peng, Yinghan Zhou, Changtong Dou, Yiming Xue
Abstract
Although existing backdoor defenses have gained success in mitigating backdoor attacks, they still face substantial challenges. In particular, most of them rely on large amounts of clean data to weaken the backdoor mapping but generally struggle with residual trigger effects, resulting in persistently high attack success rates (ASR). Therefore, in this paper, we propose a novel Backdoor defense method based on Directional mapping module and adversarial Knowledge Distillation (BeDKD), which balances the trade-off between defense effectiveness and model performance using a small amount of clean and poisoned data. We first introduce a directional mapping module to identify poisoned data, which destroys clean mapping while keeping backdoor mapping on a small set of flipped clean data. Then, the adversarial knowledge distillation is designed to reinforce clean mapping and suppress backdoor mapping through a cycle iteration mechanism between trust and punish distillations using clean and identified poisoned data. We conduct experiments to mitigate mainstream attacks on three datasets, and experimental results demonstrate that BeDKD surpasses the state-of-the-art defenses and reduces the ASR by 98% without significantly reducing the CACC.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 96c9f930-205b-4917-87ca-5aae406cdf4eBuilds on12
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 1,305 citations
- Attack of the Tails: Yes, You Really Can Backdoor Federated LearningHongyi Wang, Kartik Sreenivasan, Shashank Rajput, Harit Vishwakarma et al.NeurIPS 2020 · 862 citations
- Neural Attention Distillation: Erasing Backdoor Triggers from Deep Neural NetworksYige Li, Xixiang Lyu, Nodens Koren, Lingjuan Lyu et al.ICLR 2021 · 548 citations
- Defending against Backdoor Attacks in Natural Language GenerationXiaofei Sun, Xiaoya Li, Yuxian Meng, Xiang Ao et al.AAAI 2023 · 65 citations
- Defending Pre-trained Language Models as Few-shot Learners against Backdoor AttacksZhaohan Xi, Tianyu Du, Changjiang Li, Ren Pang et al.NeurIPS 2023 · 61 citations
Related papers
- Backdoor Attacks Against Dataset DistillationYugeng Liu, Zheng Li, Michael Backes, Yun Shen et al.NDSS 2023
- Beating Backdoor Attack at Its Own GameMin Liu, Alberto L. Sangiovanni-Vincentelli, Xiangyu YueICCV 2023 · 19 citations
- Redeem Myself: Purifying Backdoors in Deep Learning Models using Self Attention DistillationXueluan Gong, Yanjiao Chen, Wang Yang, Qian Wang et al.S&P 2023
- From Toxic to Trustworthy: Using Self-Distillation and Semi-supervised Methods to Refine Neural NetworksXianda Zhang, Baolin Zheng, Jianbao Hu, Chengyang Li et al.AAAI 2024 · 3 citations
- MultiKD: Backdoor Defense in Federated Graph Learning via Attention-Guided Multi-Teacher DistillationJiale Zhang, Yanan Wang, Bosen Rao, Chengcheng Zhu et al.AAAI 2026
