Shared Adversarial Unlearning: Backdoor Mitigation by Unlearning Shared Adversarial Examples
Shaokui Wei, Mingda Zhang, Hongyuan Zha, Baoyuan Wu
摘要
Backdoor attacks are serious security threats to machine learning models where an adversary can inject poisoned samples into the training set, causing a backdoored model which predicts poisoned samples with particular triggers to particular target classes, while behaving normally on benign samples. In this paper, we explore the task of purifying a backdoored model using a small clean dataset. By establishing the connection between backdoor risk and adversarial risk, we derive a novel upper bound for backdoor risk, which mainly captures the risk on the shared adversarial examples (SAEs) between the backdoored model and the purified model. This upper bound further suggests a novel bi-level optimization problem for mitigating backdoor using adversarial training techniques. To solve it, we propose Shared Adversarial Unlearning (SAU). Specifically, SAU first generates SAEs, and then, unlearns the generated SAEs such that they are either correctly classified by the purified model and/or differently classified by the two models, such that the backdoor effect in the backdoored model will be mitigated in the purified model. Experiments on various benchmark datasets and network architectures show that our proposed method achieves state-of-the-art performance for backdoor defense.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- Breaking the False Sense of Security in Backdoor Defense through Re-Activation AttackMingli Zhu, Siyuan Liang, Baoyuan WuNeurIPS 2024 · 被引用 38 次
- TERD: A Unified Framework for Safeguarding Diffusion Models Against BackdoorsYichuan Mo, Hui Huang, Mingjie Li, Ang Li 等ICML 2024 · 被引用 31 次
- Mitigating Backdoor Attack by Injecting Proactive Defensive BackdoorShaokui Wei, Hongyuan Zha, Baoyuan WuNeurIPS 2024 · 被引用 20 次
- Unveiling and Mitigating Backdoor Vulnerabilities based on Unlearning Weight Changes and Backdoor ActivenessWeilin Lin, Li Liu, Shaokui Wei, Jianze Li 等NeurIPS 2024 · 被引用 16 次
- Adversarial-Inspired Backdoor Defense via Bridging Backdoor and Adversarial AttacksJia-Li Yin, Weijian Wang, Lyhwa, Wei Lin 等AAAI 2025 · 被引用 9 次
它引用的顶会 Paper27
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li 等S&P 2019 · 被引用 1,801 次
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 被引用 1,352 次
- Improving Adversarial Robustness Requires Revisiting Misclassified ExamplesYisen Wang, Difan Zou, Jinfeng Yi, James Bailey 等ICLR 2020 · 被引用 829 次
- Invisible Backdoor Attack with Sample-Specific TriggersYuezun Li, Yiming Li, Baoyuan Wu, Longkang Li 等ICCV 2021 · 被引用 639 次
- Input-Aware Dynamic Backdoor AttackTuan Anh Nguyen, Anh Tuan TranNeurIPS 2020 · 被引用 601 次
相关 Paper
- Backdoor Defense with Machine UnlearningYang Liu, Mingyuan Fan, Cen Chen, Ximeng Liu 等INFOCOM 2022 · 被引用 89 次
- Backdoor Attacks via Machine UnlearningZihao Liu, Tianhao Wang, Mengdi Huai, Chenglin MiaoAAAI 2024 · 被引用 46 次
- Adversarial Unlearning of Backdoors via Implicit HypergradientYi Zeng, Si Chen, Won Park, Zhuoqing Mao 等ICLR 2022 · 被引用 235 次
- Revisiting the Assumption of Latent Separability for Backdoor DefensesXiangyu Qi, Tinghao Xie, Yiming Li, Saeed Mahloujifar 等ICLR 2023
- Anti-Backdoor Learning: Training Clean Models on Poisoned DataYige Li, Xixiang Lyu, Nodens Koren, Lingjuan Lyu 等NeurIPS 2021 · 被引用 503 次
