BagFlip: A Certified Defense Against Data Poisoning
Yuhao Zhang, Aws Albarghouthi, Loris D'Antoni
Abstract
Machine learning models are vulnerable to data-poisoning attacks, in which an attacker maliciously modifies the training set to change the prediction of a learned model. In a trigger-less attack, the attacker can modify the training set but not the test inputs, while in a backdoor attack the attacker can also modify test inputs. Existing model-agnostic defense approaches either cannot handle backdoor attacks or do not provide effective certificates (i.e., a proof of a defense). We present BagFlip, a model-agnostic certified approach that can effectively defend against both trigger-less and backdoor attacks. We evaluate BagFlip on image classification and malware detection datasets. BagFlip is equal to or more effective than the state-of-the-art approaches for trigger-less attacks and more effective than the state-of-the-art approaches for backdoor attacks. 36th Conference on Neural Information Processing Systems (NeurIPS 2022).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 65dfa566-6160-4f6f-990f-3ff8f1205c16Cited by top-tier papers11
- CBD: A Certified Backdoor Detector Based on Local Dominant ProbabilityZhen Xiang, Zidi Xiong, Bo LiNeurIPS 2023 · 29 citations
- SoK: Explainable Machine Learning in Adversarial EnvironmentsMaximilian Noppel, Christian WressneggerS&P 2024 · 28 citations
- Relational DNN Verification With Cross Executional Bound RefinementDebangshu Banerjee, Gagandeep SinghICML 2024 · 8 citations
- Relational Verification Leaps Forward with RABBitTarun Suresh, Debangshu Banerjee, Gagandeep SinghNeurIPS 2024 · 5 citations
- FCert: Certifiably Robust Few-Shot Classification in the Era of Foundation ModelsYanting Wang, Wei Zou, Jinyuan JiaS&P 2024 · 4 citations
Builds on19
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li et al.S&P 2019 · 1,801 citations
- Trojaning Attack on Neural NetworksYingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee et al.NDSS 2018 · 1,377 citations
- On Adaptive Attacks to Adversarial Example DefensesFlorian Tramèr, Nicholas Carlini, Wieland Brendel, Aleksander MadryNeurIPS 2020 · 1,026 citations
- Attack of the Tails: Yes, You Really Can Backdoor Federated LearningHongyi Wang, Kartik Sreenivasan, Shashank Rajput, Harit Vishwakarma et al.NeurIPS 2020 · 862 citations
- Hidden Trigger Backdoor AttacksAniruddha Saha, Akshayvarun Subramanya, Hamed PirsiavashAAAI 2020 · 743 citations
Related papers
- FLIP: A Provable Defense Framework for Backdoor Mitigation in Federated LearningKaiyuan Zhang, Guanhong Tao, Qiuling Xu, Siyuan Cheng et al.ICLR 2023 · 17 citations
- TextGuard: Provable Defense against Backdoor Attacks on Text ClassificationHengzhi Pei, Jinyuan Jia, Wenbo Guo, Bo Li et al.NDSS 2024
- Deep Partition Aggregation: Provable Defenses against General Poisoning AttacksAlexander Levine, Soheil FeiziICLR 2021 · 22 citations
- Mitigating Backdoor Attack by Injecting Proactive Defensive BackdoorShaokui Wei, Hongyuan Zha, Baoyuan WuNeurIPS 2024 · 20 citations
- DEFEAT: Deep Hidden Feature Backdoor Attacks by Imperceptible Perturbation and Latent Representation ConstraintsZhendong Zhao, Xiaojun Chen, Yuexin Xuan, Ye Dong et al.CVPR 2022 · 72 citations
