Attack To Defend: Exploiting Adversarial Attacks for Detecting Poisoned Models
Samar Fares, Karthik Nandakumar
Abstract
Poisoning (trojan/backdoor) attacks enable an adversary to train and deploy a corrupted machine learning (ML) model, which typically works well and achieves good accuracy on clean input samples but behaves maliciously on poisoned samples containing specific trigger patterns. Using such poisoned ML models as the foundation to build realworld systems can compromise application safety. Hence, there is a critical need for algorithms that detect whether a given target model has been poisoned. This work proposes a novel approach for detecting poisoned models called Attack To Defend (A2D), which is based on the observation that poisoned models are more sensitive to adversarial perturbations compared to benign models. We propose a metric called sensitivity to adversarial perturbations (SAP) to measure the sensitivity of a ML model to adversarial attacks at a specific perturbation bound. We then generate strong adversarial attacks against an unrelated reference model and estimate the SAP value of the target model by transferring the generated attacks. The target model is deemed to be a trojan if its SAP value exceeds a decision threshold. The A2D framework requires only black-box access to the target model and a small clean set, while being computationally efficient. The A2D approach has been evaluated on four standard image datasets and its effectiveness under various types of poisoning attacks has been demonstrated.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7a961dbf-c063-4bcc-b592-8867caf73d03Cited by top-tier papers2
- Taught Well Learned Ill: Towards Distillation-conditional Backdoor AttackYukun Chen, Boheng Li, Yu Yuan, Leyi Qi et al.NeurIPS 2025 · 6 citations
- Vpr-Cloak: a First Look at Privacy Cloak Against Visual Place RecognitionShuting Dong, Mingzhi Chen, Feng Lu, Hao Yu et al.ICCV 2025 · 2 citations
Builds on24
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li et al.S&P 2019 · 1,801 citations
- Hidden Trigger Backdoor AttacksAniruddha Saha, Akshayvarun Subramanya, Hamed PirsiavashAAAI 2020 · 743 citations
- Invisible Backdoor Attack with Sample-Specific TriggersYuezun Li, Yiming Li, Baoyuan Wu, Longkang Li et al.ICCV 2021 · 639 citations
Related papers
- Beating Backdoor Attack at Its Own GameMin Liu, Alberto L. Sangiovanni-Vincentelli, Xiangyu YueICCV 2023 · 19 citations
- Effective Backdoor Defense by Exploiting Sensitivity of Poisoned SamplesWeixin Chen, Baoyuan Wu, Haoqian WangNeurIPS 2022 · 129 citations
- Invisible Poison: A Blackbox Clean Label Backdoor Attack to Deep Neural NetworksRui Ning, Jiang Li, Chunsheng Xin, Hongyi WuINFOCOM 2021 · 56 citations
- MDTD: A Multi-Domain Trojan Detector for Deep Neural NetworksArezoo Rajabi, Surudhi Asokraj, Fengqing Jiang, Luyao Niu et al.CCS 2023
- Towards A Proactive ML Approach for Detecting Backdoor Poison SamplesXiangyu Qi, Tinghao Xie, Jiachen T. Wang, Tong Wu et al.USENIX Security 2023
