Certified Robustness of Nearest Neighbors against Data Poisoning and Backdoor Attacks
Jinyuan Jia, Yupei Liu, Xiaoyu Cao, Neil Zhenqiang Gong
Abstract
Data poisoning attacks and backdoor attacks aim to corrupt a machine learning classifier via modifying, adding, and/or removing some carefully selected training examples, such that the corrupted classifier makes incorrect predictions as the attacker desires. The key idea of state-of-the-art certified defenses against data poisoning attacks and backdoor attacks is to create a majority vote mechanism to predict the label of a testing example. Moreover, each voter is a base classifier trained on a subset of the training dataset. Classical simple learning algorithms such as k nearest neighbors (kNN) and radius nearest neighbors (rNN) have intrinsic majority vote mechanisms. In this work, we show that the intrinsic majority vote mechanisms in kNN and rNN already provide certified robustness guarantees against data poisoning attacks and backdoor attacks. Moreover, our evaluation results on MNIST and CIFAR10 show that the intrinsic certified robustness guarantees of kNN and rNN outperform those provided by state-of-the-art certified defenses. Our results serve as standard baselines for future certified defenses against data poisoning attacks and backdoor attacks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext cf689300-803d-4dbd-8a4f-80850f7ced00Cited by top-tier papers26
- Elijah: Eliminating Backdoors Injected in Diffusion Models via Distribution ShiftShengwei An, Sheng-Yen Chou, Kaiyuan Zhang, Qiuling Xu et al.AAAI 2024 · 48 citations
- BadRL: Sparse Targeted Backdoor Attack against Reinforcement LearningJing Cui, Yufei Han, Yuzhe Ma, Jianbin Jiao et al.AAAI 2024 · 31 citations
- CBD: A Certified Backdoor Detector Based on Local Dominant ProbabilityZhen Xiang, Zidi Xiong, Bo LiNeurIPS 2023 · 29 citations
- On Collective Robustness of Bagging Against Data PoisoningRuoxin Chen, Zenan Li, Jie Li, Junchi Yan et al.ICML 2022 · 25 citations
- Django: Detecting Trojans in Object Detection Models via Gaussian Focus CalibrationGuangyu Shen, Siyuan Cheng, Guanhong Tao, Kaiyuan Zhang et al.NeurIPS 2023 · 18 citations
Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li et al.S&P 2019 · 1,801 citations
- Trojaning Attack on Neural NetworksYingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee et al.NDSS 2018 · 1,377 citations
- Manipulating Machine Learning: Poisoning Attacks and Countermeasures for Regression LearningMatthew Jagielski, Alina Oprea, Battista Biggio, Chang Liu et al.S&P 2018 · 867 citations
Related papers
- Intrinsic Certified Robustness of Bagging against Data Poisoning AttacksJinyuan Jia, Xiaoyu Cao, Neil Zhenqiang GongAAAI 2021 · 155 citations
- RAB: Provable Robustness Against Backdoor AttacksMaurice Weber, Xiaojun Xu, Bojan Karlas, Ce Zhang et al.S&P 2023
- Deterministic Certification of Graph Neural Networks against Graph Poisoning Attacks with Arbitrary PerturbationsJiate Li, Meng Pang, Yun Dong, Binghui WangCVPR 2025
- Deep Partition Aggregation: Provable Defenses against General Poisoning AttacksAlexander Levine, Soheil FeiziICLR 2021 · 22 citations
- Systematic Testing of the Data-Poisoning Robustness of KNNYannan Li, Jingbo Wang, Chao WangISSTA 2023 · 8 citations
