USENIX Security2022Top-tier venue
Poison Forensics: Traceback of Data Poisoning Attacks in Neural Networks
Shawn Shan, Arjun Nitin Bhagoji, Haitao Zheng, Ben Y. Zhao
Abstract
In adversarial machine learning, new defenses against attacks on deep learning systems are routinely broken soon after their release by more powerful attacks. In this context, forensic tools can offer a valuable complement to existing defenses, by tracing back a successful attack to its root cause, and offering a path forward for mitigation to prevent similar attacks in the future. In this paper, we describe our efforts in developing a forensic traceback tool for poison attacks on deep neural networks. We propose a novel iterative clustering and pruning solution that trims"innocent"training samples, until all that remains is the set of poisoned data responsible for the attack. Our method clusters training samples based on their impact on model parameters, then uses an efficient data unlearning method to prune innocent clusters. We empirically demonstrate the efficacy of our system on three types of dirty-label (backdoor) poison attacks and three types of clean-label poison attacks, across domains of computer vision and malware classification. Our system achieves over 98.4% precision and 96.8% recall across all attacks. We also show that our system is robust against four anti-forensics measures specifically designed to attack it.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers22
- Poisoning Web-Scale Training Datasets is PracticalNicholas Carlini, Matthew Jagielski, Christopher A. Choquette-Choo, Daniel Paleka et al.S&P 2024 · 309 citations
- Nightshade: Prompt-Specific Poisoning Attacks on Text-to-Image Generative ModelsShawn Shan, Wenxin Ding, Josephine Passananti, Stanley Wu et al.S&P 2024 · 102 citations
- MM-BD: Post-Training Detection of Backdoor Attacks with Arbitrary Backdoor Pattern Types Using a Maximum Margin StatisticHang Wang, Zhen Xiang, David J. Miller, George KesidisS&P 2024 · 81 citations
- Traceback of Poisoning Attacks to Retrieval-Augmented GenerationBaolei Zhang, Haoran Xin, Minghong Fang, Zhuqing Liu et al.WWW 2025 · 20 citations
- "Get in Researchers; We're Measuring Reproducibility": A Reproducibility Study of Machine Learning Papers in Tier 1 Security ConferencesDaniel Olszewski, Allison Lu, Carson Stillman, Kevin Warren et al.CCS 2023 · 19 citations
Builds on27
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li et al.S&P 2019 · 1,801 citations
- Trojaning Attack on Neural NetworksYingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee et al.NDSS 2018 · 1,377 citations
- HopSkipJumpAttack: A Query-Efficient Decision-Based AttackJianbo Chen, Michael I. Jordan, Martin J. WainwrightS&P 2020 · 797 citations
- Hidden Trigger Backdoor AttacksAniruddha Saha, Akshayvarun Subramanya, Hamed PirsiavashAAAI 2020 · 743 citations
Related papers
- BEAGLE: Forensics of Deep Learning Backdoor Attack for Better DefenseSiyuan Cheng, Guanhong Tao, Yingqi Liu, Shengwei An et al.NDSS 2023
- Few-shot Backdoor Defense Using Shapley EstimationJiyang Guan, Zhuozhuo Tu, Ran He, Dacheng TaoCVPR 2022 · 36 citations
- Witches' Brew: Industrial Scale Data Poisoning via Gradient MatchingJonas Geiping, Liam H. Fowl, W. Ronny Huang, Wojciech Czaja et al.ICLR 2021 · 268 citations
- PoisonSpot: Precise Spotting of Clean-Label Backdoors via Fine-Grained Training Provenance TrackingPhilemon Hailemariam, Birhanu EsheteCCS 2025
- Tracing Back the Malicious Clients in Poisoning Attacks to Federated LearningYuqi Jia, Minghong Fang, Hongbin Liu, Jinghuai Zhang et al.NeurIPS 2025 · 8 citations
