Defense against backdoor attacks via robust covariance estimation
Jonathan Hayase, Weihao Kong, Raghav Somani, Sewoong Oh
Abstract
Modern machine learning increasingly requires training on a large collection of data from multiple sources, not all of which can be trusted. A particularly concerning scenario is when a small fraction of poisoned data changes the behavior of the trained model when triggered by an attackerspecified watermark. Such a compromised model will be deployed unnoticed as the model is accurate otherwise. There have been promising attempts to use the intermediate representations of such a model to separate corrupted examples from clean ones. However, these defenses work only when a certain spectral signature of the poisoned examples is large enough for detection. There is a wide range of attacks that cannot be protected against by the existing defenses. We propose a novel defense algorithm using robust covariance estimation to amplify the spectral signature of corrupted data. This defense provides a clean model, completely removing the backdoor, even in regimes where previous methods have no hope of detecting the poisoned examples. 2
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Rethinking the Reverse-engineering of Trojan TriggersZhenting Wang, Kai Mei, Hailun Ding, Juan Zhai et al.NeurIPS 2022 · 75 citations
- BagFlip: A Certified Defense Against Data PoisoningYuhao Zhang, Aws Albarghouthi, Loris D'AntoniNeurIPS 2022 · 32 citations
- Distilling Cognitive Backdoor Patterns within an ImageHanxun Huang, Xingjun Ma, Sarah Monazam Erfani, James BaileyICLR 2023 · 7 citations
- UNICORN: A Unified Backdoor Trigger Inversion FrameworkZhenting Wang, Kai Mei, Juan Zhai, Shiqing MaICLR 2023 · 7 citations
- Poison Forensics: Traceback of Data Poisoning Attacks in Neural NetworksShawn Shan, Arjun Nitin Bhagoji, Haitao Zheng, Ben Y. ZhaoUSENIX Security 2022
Builds on11
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li et al.S&P 2019 · 1,801 citations
- Attack of the Tails: Yes, You Really Can Backdoor Federated LearningHongyi Wang, Kartik Sreenivasan, Shashank Rajput, Harit Vishwakarma et al.NeurIPS 2020 · 862 citations
- Hidden Trigger Backdoor AttacksAniruddha Saha, Akshayvarun Subramanya, Hamed PirsiavashAAAI 2020 · 743 citations
- Robust and differentially private mean estimationXiyang Liu, Weihao Kong, Sham M. Kakade, Sewoong OhNeurIPS 2021 · 87 citations
- Robust Sub-Gaussian Principal Component Analysis and Width-Independent Schatten PackingArun Jambulapati, Jerry Li, Kevin TianNeurIPS 2020 · 45 citations
Related papers
- Narcissus: A Practical Clean-Label Backdoor Attack with Limited InformationYi Zeng, Minzhou Pan, Hoang Anh Just, Lingjuan Lyu et al.CCS 2023 · 170 citations
- Revisiting the Assumption of Latent Separability for Backdoor DefensesXiangyu Qi, Tinghao Xie, Yiming Li, Saeed Mahloujifar et al.ICLR 2023
- DEFEAT: Deep Hidden Feature Backdoor Attacks by Imperceptible Perturbation and Latent Representation ConstraintsZhendong Zhao, Xiaojun Chen, Yuexin Xuan, Ye Dong et al.CVPR 2022 · 72 citations
- PoisonSpot: Precise Spotting of Clean-Label Backdoors via Fine-Grained Training Provenance TrackingPhilemon Hailemariam, Birhanu EsheteCCS 2025
- Backdoor Secrets Unveiled: Identifying Backdoor Data with Optimized Scaled Prediction ConsistencySoumyadeep Pal, Yuguang Yao, Ren Wang, Bingquan Shen et al.ICLR 2024 · 15 citations
