Two Sides of the Same Coin: Learning the Backdoor to Remove the Backdoor
Qi Zhao, Christian Wressnegger
摘要
The community has recently developed various training-time defenses to counter neural backdoors introduced through data poisoning. In light of the observation that a model learns poisonous samples responsible for the backdoor easier than benign samples, these approaches either use a fixed threshold of the training loss for splitting (Li et al. 2021a; Huang et al. 2022; Chen et al. 2022) or iteratively learn a reference model as an oracle for identifying benign samples (Gao et al. 2023; Zhang et al. 2023) . In particular, the latter has proven effective for anti-backdoor learning. Our method, HARVEY, leverages a similar yet crucially different technique: learning an oracle for poisonous rather than benign samples. Learning a backdoored reference model is significantly easier than learning a reference model on benign data. Consequently, we can identify poisonous samples much more accurately than related work identifies benign samples. This crucial difference enables near-perfect backdoor removal as we demonstrate in our evaluation. HARVEY substantially outperforms related approaches across attack types, datasets, and architectures, lowering the attack success rate to the very minimum at a negligible loss in natural accuracy. Supplementary material is available at https://intellisec.de/research/harvey
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- LoSplit: Loss-Guided Dynamic Split for Training-Time Defense Against Graph Backdoor AttacksDi Jin, Yuxiang Zhang, Bingdao Feng, Xiaobao Wang 等NeurIPS 2025 · 被引用 4 次
- Anti-Backdoor Coreset Selection via Cumulative EntropyQi Zhao, Christian WressneggerICML 2026
- Seal Your Backdoor with Variational DefenseIvan Sabolic, Matej Grcic, Sinisa SegvicICCV 2025
它引用的顶会 Paper18
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li 等S&P 2019 · 被引用 1,801 次
- Trojaning Attack on Neural NetworksYingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee 等NDSS 2018 · 被引用 1,377 次
- Symmetric Cross Entropy for Robust Learning With Noisy LabelsYisen Wang, Xingjun Ma, Zaiyi Chen, Yuan Luo 等ICCV 2019 · 被引用 1,125 次
- Input-Aware Dynamic Backdoor AttackTuan Anh Nguyen, Anh Tuan TranNeurIPS 2020 · 被引用 601 次
- Neural Attention Distillation: Erasing Backdoor Triggers from Deep Neural NetworksYige Li, Xixiang Lyu, Nodens Koren, Lingjuan Lyu 等ICLR 2021 · 被引用 548 次
相关 Paper
- Progressive Poisoned Data Isolation for Training-Time Backdoor DefenseYiming Chen, Haiwei Wu, Jiantao ZhouAAAI 2024 · 被引用 19 次
- Mitigating Backdoor Attack by Injecting Proactive Defensive BackdoorShaokui Wei, Hongyuan Zha, Baoyuan WuNeurIPS 2024 · 被引用 20 次
- Backdoor Defense via Adaptively Splitting Poisoned DatasetKuofeng Gao, Yang Bai, Jindong Gu, Yong Yang 等CVPR 2023
- Learning from Distinction: Mitigating Backdoors Using a Low-Capacity ModelHaosen Sun, Yiming Li, Xixiang Lyu, Jing MaACM MM 2024
- Training with More Confidence: Mitigating Injected and Natural Backdoors During TrainingZhenting Wang, Hailun Ding, Juan Zhai, Shiqing MaNeurIPS 2022 · 被引用 67 次
