RAB: Provable Robustness Against Backdoor Attacks
Maurice Weber, Xiaojun Xu, Bojan Karlas, Ce Zhang, Bo Li
摘要
Recent studies have shown that deep neural net-works (DNNs) are vulnerable to adversarial attacks, including evasion and backdoor (poisoning) attacks. On the defense side, there have been intensive efforts on improving both empirical and provable robustness against evasion attacks; however, the provable robustness against backdoor attacks still remains largely unexplored. In this paper, we focus on certifying the machine learning model robustness against general threat models, especially backdoor attacks. We first provide a unified framework via randomized smoothing techniques and show how it can be instantiated to certify the robustness against both evasion and backdoor attacks. We then propose the first robust training process, RAB, to smooth the trained model and certify its robustness against backdoor attacks. We theoretically prove the robustness bound for machine learning models trained with RAB and prove that our robustness bound is tight. In addition, we theoretically show that it is possible to train the robust smoothed models efficiently for simple models such as K-nearest neighbor classifiers, and we propose an exact smooth-training algorithm that eliminates the need to sample from a noise distribution for such models. Empirically, we conduct comprehensive experiments for different machine learning (ML) models such as DNNs, support vector machines, and K-NN models on MNIST, CIFAR-10, and ImageNette datasets and provide the first benchmark for certified robustness against backdoor attacks. In addition, we evaluate K-NN models on a spambase tabular dataset to demonstrate the advantages of the proposed exact algorithm. Both the theoretic analysis and the comprehensive evaluation on diverse ML models and datasets shed light on further robust learning strategies against general training time attacks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper32
- BadEncoder: Backdoor Attacks to Pre-trained Encoders in Self-Supervised LearningJinyuan Jia, Yupei Liu, Neil Zhenqiang GongS&P 2022 · 被引用 200 次
- Improved Certified Defenses against Data Poisoning with (Deterministic) Finite AggregationWenxiao Wang, Alexander Levine, Soheil FeiziICML 2022 · 被引用 68 次
- Towards Understanding and Enhancing Robustness of Deep Learning Models against Malicious Unlearning AttacksWei Qian, Chenxu Zhao, Wei Le, Meiyi Ma 等KDD 2023 · 被引用 38 次
- BagFlip: A Certified Defense Against Data PoisoningYuhao Zhang, Aws Albarghouthi, Loris D'AntoniNeurIPS 2022 · 被引用 32 次
- SampDetox: Black-box Backdoor Defense via Perturbation-based Sample DetoxificationYanxin Yang, Chentao Jia, Dengke Yan, Ming Hu 等NeurIPS 2024 · 被引用 20 次
它引用的顶会 Paper16
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li 等S&P 2019 · 被引用 1,801 次
- Feature Squeezing: Detecting Adversarial Examples in Deep Neural NetworksWeilin Xu, David Evans, Yanjun QiNDSS 2018 · 被引用 1,633 次
- Certified Robustness to Adversarial Examples with Differential PrivacyMathias Lécuyer, Vaggelis Atlidakis, Roxana Geambasu, Daniel Hsu 等S&P 2019 · 被引用 1,022 次
- Hidden Trigger Backdoor AttacksAniruddha Saha, Akshayvarun Subramanya, Hamed PirsiavashAAAI 2020 · 被引用 743 次
- Detecting AI Trojans Using Meta Neural AnalysisXiaojun Xu, Qi Wang, Huichen Li, Nikita Borisov 等S&P 2021 · 被引用 381 次
相关 Paper
- Certified Robustness of Nearest Neighbors against Data Poisoning and Backdoor AttacksJinyuan Jia, Yupei Liu, Xiaoyu Cao, Neil Zhenqiang GongAAAI 2022 · 被引用 90 次
- CRFL: Certifiably Robust Federated Learning against Backdoor AttacksChulin Xie, Minghao Chen, Pin-Yu Chen, Bo LiICML 2021 · 被引用 218 次
- Machine Learning needs Better Randomness Standards: Randomised Smoothing and PRNG-based attacksPranav Dahiya, Ilia Shumailov, Ross AndersonUSENIX Security 2024 · 被引用 11 次
- Progressive Poisoned Data Isolation for Training-Time Backdoor DefenseYiming Chen, Haiwei Wu, Jiantao ZhouAAAI 2024 · 被引用 19 次
- Anti-Backdoor Learning: Training Clean Models on Poisoned DataYige Li, Xixiang Lyu, Nodens Koren, Lingjuan Lyu 等NeurIPS 2021 · 被引用 503 次
