Redeem Myself: Purifying Backdoors in Deep Learning Models using Self Attention Distillation
Xueluan Gong, Yanjiao Chen, Wang Yang, Qian Wang, Yuzhe Gu, Huayang Huang, Chao Shen
摘要
Recent works have revealed the vulnerability of deep neural networks to backdoor attacks, where a backdoored model orchestrates targeted or untargeted misclassification when activated by a trigger. A line of purification methods (e.g., fine-pruning, neural attention transfer, MCR [69]) have been proposed to remove the backdoor in a model. However, they either fail to reduce the attack success rate of more advanced backdoor attacks or largely degrade the prediction capacity of the model for clean samples. In this paper, we put forward a new purification defense framework, dubbed SAGE, which utilizes self-attention distillation to purge models of backdoors. Unlike traditional attention transfer mechanisms that require a teacher model to supervise the distillation process, SAGE can realize self-purification with a small number of clean samples. To enhance the defense performance, we further propose a dynamic learning rate adjustment strategy that carefully tracks the prediction accuracy of clean samples to guide the learning rate adjustment. We compare the defense performance of SAGE with 6 state-of-the-art defense approaches against 8 backdoor attacks on 4 datasets. It is shown that SAGE can reduce the attack success rate by as much as 90% with less than 3% decrease in prediction accuracy for clean samples. We will open-source our codes upon publication.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper8
- UBA-Inf: Unlearning Activated Backdoor Attack with Influence-Driven CamouflageZirui Huang, Yunlong Mao, Sheng ZhongUSENIX Security 2024 · 被引用 16 次
- Watch Out! Simple Horizontal Class Backdoor Can Trivially Evade DefenseHua Ma, Shang Wang, Yansong Gao, Zhi Zhang 等CCS 2024 · 被引用 7 次
- FAMOS: Robust Privacy-Preserving Authentication on Payment Apps via Federated Multi-Modal Contrastive LearningYifeng Cai, Ziqi Zhang, Jiaping Gui, Bingyan Liu 等USENIX Security 2024 · 被引用 6 次
- InverTune: A Backdoor Defense Method for Multimodal Contrastive Learning via Backdoor-Adversarial Correlation AnalysisMengyuan Sun, Yu Li, Yunjie Ge, Yuchen Liu 等NDSS 2026 · 被引用 3 次
- Lethe: Purifying Backdoored Large Language Models with Knowledge DilutionChen Chen, Yuchen Sun, Jiaxin Gao, Xueluan Gong 等USENIX Security 2026 · 被引用 1 次
相关 Paper
- Neural Attention Distillation: Erasing Backdoor Triggers from Deep Neural NetworksYige Li, Xixiang Lyu, Nodens Koren, Lingjuan Lyu 等ICLR 2021 · 被引用 548 次
- From Toxic to Trustworthy: Using Self-Distillation and Semi-supervised Methods to Refine Neural NetworksXianda Zhang, Baolin Zheng, Jianbao Hu, Chengyang Li 等AAAI 2024 · 被引用 3 次
- Beating Backdoor Attack at Its Own GameMin Liu, Alberto L. Sangiovanni-Vincentelli, Xiangyu YueICCV 2023 · 被引用 19 次
- Towards Stable Backdoor Purification through Feature Shift TuningRui Min, Zeyu Qin, Li Shen, Minhao ChengNeurIPS 2023 · 被引用 43 次
- ATTEQ-NN: Attention-based QoE-aware Evasive Backdoor AttacksXueluan Gong, Yanjiao Chen, Jianshuo Dong, Qian WangNDSS 2022
