Model-Targeted Poisoning Attacks with Provable Convergence
Fnu Suya, Saeed Mahloujifar, Anshuman Suri, David Evans, Yuan Tian
Abstract
In a poisoning attack, an adversary with control over a small fraction of the training data attempts to select that data in a way that induces a corrupted model that misbehaves in favor of the adversary. We consider poisoning attacks against convex machine learning models and propose an efficient poisoning attack designed to induce a specified model. Unlike previous model-targeted poisoning attacks, our attack comes with provable convergence to any attainable target classifier. The distance from the induced classifier to the target classifier is inversely proportional to the square root of the number of poisoning points. We also provide a lower bound on the minimum number of poisoning points needed to achieve a given target classifier. Our method uses online convex optimization, so finds poisoning points incrementally. This provides more flexibility than previous attacks which require a priori assumption about the number of poisoning points. Our attack is the first model-targeted poisoning attack that provides provable convergence for convex models, and in our experiments, it either exceeds or matches state-of-the-art attacks in terms of attack success rate and distance to the target model. Preprint.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2c8df6ec-3284-4845-b2df-907432519f7cCited by top-tier papers17
- Algorithmic Collective Action in Machine LearningMoritz Hardt, Eric Mazumdar, Celestine Mendler-Dünner, Tijana ZrnicICML 2023 · 36 citations
- Static and Sequential Malicious Attacks in the Context of Selective ForgettingChenxu Zhao, Wei Qian, Rex Ying, Mengdi HuaiNeurIPS 2023 · 30 citations
- An Equivalence Between Data Poisoning and Byzantine Gradient AttacksSadegh Farhadkhani, Rachid Guerraoui, Lê Nguyên Hoang, Oscar VillemaudICML 2022 · 30 citations
- Exploring the Limits of Model-Targeted Indiscriminate Data Poisoning AttacksYiwei Lu, Gautam Kamath, Yaoliang YuICML 2023 · 25 citations
- SoK: All You Need to Know About On-Device ML Model Extraction - The Gap Between Research and PracticeTushar Nayan, Qiming Guo, Mohammed Alduniawi, Marcus Botacin et al.USENIX Security 2024 · 20 citations
Builds on5
- Manipulating Machine Learning: Poisoning Attacks and Countermeasures for Regression LearningMatthew Jagielski, Alina Oprea, Battista Biggio, Chang Liu et al.S&P 2018 · 867 citations
- Why Do Adversarial Attacks Transfer? Explaining Transferability of Evasion and Poisoning AttacksAmbra Demontis, Marco Melis, Maura Pintor, Matthew Jagielski et al.USENIX Security 2019 · 466 citations
- Differentially Private Learning Needs Better Features (or Much More Data)Florian Tramèr, Dan BonehICLR 2021 · 325 citations
- Witches' Brew: Industrial Scale Data Poisoning via Gradient MatchingJonas Geiping, Liam H. Fowl, W. Ronny Huang, Wojciech Czaja et al.ICLR 2021 · 268 citations
- Subpopulation Data Poisoning AttacksMatthew Jagielski, Giorgio Severi, Niklas Pousette Harger, Alina OpreaCCS 2021 · 15 citations
Related papers
- On Robustness of Linear Classifiers to Targeted Data PoisoningNakshatra Gupta, Sumanth Prabhu S, Supratik Chakraborty, R. VenkateshAAAI 2026
- Not All Poisons are Created Equal: Robust Training against Data PoisoningYu Yang, Tian Yu Liu, Baharan MirzasoleimanICML 2022 · 45 citations
- MetaPoison: Practical General-purpose Clean-label Data PoisoningW. Ronny Huang, Jonas Geiping, Liam Fowl, Gavin Taylor et al.NeurIPS 2020 · 242 citations
- First-Order Efficient General-Purpose Clean-Label Data PoisoningTianhang Zheng, Baochun LiINFOCOM 2021 · 8 citations
- What Distributions are Robust to Indiscriminate Poisoning Attacks for Linear Learners?Fnu Suya, Xiao Zhang, Yuan Tian, David EvansNeurIPS 2023 · 3 citations
