Oblivious Defense in ML Models: Backdoor Removal without Detection
Shafi Goldwasser, Jonathan Shafer, Neekon Vafa, Vinod Vaikuntanathan
Abstract
As society grows more reliant on machine learning, ensuring the security of machine learning systems against sophisticated attacks becomes a pressing concern. A recent result of Goldwasser, Kim, Vaikuntanathan, and Zamir (FOCS ’22) shows that an adversary can plant undetectable backdoors in machine learning models, allowing the adversary to covertly control the model’s behavior. Backdoors can be planted in such a way that the backdoored machine learning model is computationally indistinguishable from an honest model without backdoors. In this paper, we present strategies for defending against backdoors in ML models, even if they are undetectable. The key observation is that it is sometimes possible to provably mitigate or even remove backdoors without needing to detect them, using techniques inspired by the notion of random self-reducibility. This depends on properties of the ground-truth labels (chosen by nature), and not of the proposed ML model (which may be chosen by an attacker). We give formal definitions for secure backdoor mitigation, and proceed to show two types of results. First, we show a “global mitigation” technique, which removes all backdoors from a machine learning model under the assumption that the ground-truth labels are close to a Fourier-heavy function. Second, we consider distributions where the ground-truth labels are close to a linear or polynomial function in ℝn. Here, we show “local mitigation” techniques, which remove backdoors with high probability for every input of interest, and are computationally cheaper than global mitigation. All of our constructions are black-box, so our techniques work without needing access to the model’s representation (i.e., its code or parameters). Along the way we prove a simple result for robust mean estimation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on15
- Certified Robustness to Adversarial Examples with Differential PrivacyMathias Lécuyer, Vaggelis Atlidakis, Roxana Geambasu, Daniel Hsu et al.S&P 2019 · 1,022 citations
- Turning Your Weakness Into a Strength: Watermarking Deep Neural Networks by BackdooringYossi Adi, Carsten Baum, Moustapha Cissé, Benny Pinkas et al.USENIX Security 2018 · 832 citations
- Intrinsic Certified Robustness of Bagging against Data Poisoning AttacksJinyuan Jia, Xiaoyu Cao, Neil Zhenqiang GongAAAI 2021 · 155 citations
- Handcrafted Backdoors in Deep Neural NetworksSanghyun Hong, Nicholas Carlini, Alexey KurakinNeurIPS 2022 · 105 citations
- Rethinking Backdoor AttacksAlaa Khaddaj, Guillaume Leclerc, Aleksandar Makelov, Kristian Georgiev et al.ICML 2023 · 42 citations
Related papers
- Injecting Undetectable Backdoors in Obfuscated Neural Networks and Language ModelsAlkis Kalavasis, Amin Karbasi, Argyris Oikonomou, Katerina Sotiraki et al.NeurIPS 2024 · 5 citations
- Rethinking the Stealthiness of Cryptographically Undetectable Backdoors in Practical RFF LearningTianshuo Cong, Pei Li, Haojie Wu, Jinyuan Liu et al.KDD 2026
- Planting Undetectable Backdoors in Machine Learning Models : [Extended Abstract]Shafi Goldwasser, Michael P. Kim, Vinod Vaikuntanathan, Or ZamirFOCS 2022 · 40 citations
- MM-BD: Post-Training Detection of Backdoor Attacks with Arbitrary Backdoor Pattern Types Using a Maximum Margin StatisticHang Wang, Zhen Xiang, David J. Miller, George KesidisS&P 2024 · 81 citations
- Unelicitable Backdoors via Cryptographic Transformer CircuitsAndis Draguns, Andrew Gritsevskiy, Sumeet Ramesh Motwani, Christian Schröder de WittNeurIPS 2024 · 6 citations
