Deep Partition Aggregation: Provable Defenses against General Poisoning Attacks
Alexander Levine, Soheil Feizi
摘要
Adversarial poisoning attacks distort training data in order to corrupt the test-time behavior of a classifier. A provable defense provides a certificate for each test sample, which is a lower bound on the magnitude of any adversarial distortion of the training set that can corrupt the test sample's classification. We propose two novel provable defenses against poisoning attacks: (i) Deep Partition Aggregation (DPA), a certified defense against a general poisoning threat model, defined as the insertion or deletion of a bounded number of samples to the training setby implication, this threat model also includes arbitrary distortions to a bounded number of images and/or labels; and (ii) Semi-Supervised DPA (SS-DPA), a certified defense against label-flipping poisoning attacks. DPA is an ensemble method where base models are trained on partitions of the training set determined by a hash function. DPA is related to both subset aggregation, a well-studied ensemble method in classical machine learning, as well as to randomized smoothing, a popular provable defense against evasion (inference) attacks. Our defense against label-flipping poison attacks, SS-DPA, uses a semi-supervised learning algorithm as its base classifier model: each base classifier is trained using the entire unlabeled training set in addition to the labels for a partition. SS-DPA significantly outperforms the existing certified defense for label-flipping attacks (Rosenfeld et al., 2020) on both MNIST and CIFAR-10: provably tolerating, for at least half of test images, over 600 label flips (vs. < 200 label flips) on MNIST and over 300 label flips (vs. 175 label flips) on CIFAR-10. Against general poisoning attacks where no prior certified defenses exists, DPA can certify ≥ 50% of test images against over 500 poison image insertions on MNIST, and nine insertions on CIFAR-10. These results establish new state-of-the-art provable defenses against general and label-flipping poison attacks. Code is available at https://github.com/alevine0/DPA .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper33
- Better Safe Than Sorry: Preventing Delusive Adversaries with Adversarial TrainingLue Tao, Lei Feng, Jinfeng Yi, Sheng-Jun Huang 等NeurIPS 2021 · 被引用 90 次
- Improved Certified Defenses against Data Poisoning with (Deterministic) Finite AggregationWenxiao Wang, Alexander Levine, Soheil FeiziICML 2022 · 被引用 68 次
- Prompt Certified Machine Unlearning with Randomized Gradient Smoothing and QuantizationZijie Zhang, Yang Zhou, Xin Zhao, Tianshi Che 等NeurIPS 2022 · 被引用 56 次
- Improved, Deterministic Smoothing for L1 Certified RobustnessAlexander Levine, Soheil FeiziICML 2021 · 被引用 49 次
- Not All Poisons are Created Equal: Robust Training against Data PoisoningYu Yang, Tian Yu Liu, Baharan MirzasoleimanICML 2022 · 被引用 45 次
它引用的顶会 Paper8
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna 等NeurIPS 2020 · 被引用 7,049 次
- Certified Robustness to Adversarial Examples with Differential PrivacyMathias Lécuyer, Vaggelis Atlidakis, Roxana Geambasu, Daniel Hsu 等S&P 2019 · 被引用 1,022 次
- (De)Randomized Smoothing for Certifiable Defense against Patch AttacksAlexander Levine, Soheil FeiziNeurIPS 2020 · 被引用 188 次
- Intrinsic Certified Robustness of Bagging against Data Poisoning AttacksJinyuan Jia, Xiaoyu Cao, Neil Zhenqiang GongAAAI 2021 · 被引用 155 次
相关 Paper
- Lethal Dose Conjecture on Data PoisoningWenxiao Wang, Alexander Levine, Soheil FeiziNeurIPS 2022 · 被引用 17 次
- Run-off Election: Improved Provable Defense against Data Poisoning AttacksKeivan Rezaei, Kiarash Banihashem, Atoosa Malemir Chegini, Soheil FeiziICML 2023 · 被引用 22 次
- Certifying Graph Neural Networks Against Label and Structure PoisoningLukas Gosch, Xichuan Chen, Yan Scholten, Stephan GünnemannICML 2026
- Certified Robustness to Label-Flipping Attacks via Randomized SmoothingElan Rosenfeld, Ezra Winston, Pradeep Ravikumar, J. Zico KolterICML 2020 · 被引用 182 次
- Deterministic Certification of Graph Neural Networks against Graph Poisoning Attacks with Arbitrary PerturbationsJiate Li, Meng Pang, Yun Dong, Binghui WangCVPR 2025
