Autoregressive Perturbations for Data Poisoning
Pedro Sandoval Segura, Vasu Singla, Jonas Geiping, Micah Goldblum, Tom Goldstein, David Jacobs
Abstract
The prevalence of data scraping from social media as a means to obtain datasets has led to growing concerns regarding unauthorized use of data. Data poisoning attacks have been proposed as a bulwark against scraping, as they make data "unlearnable" by adding small, imperceptible perturbations. Unfortunately, existing methods require knowledge of both the target architecture and the complete dataset so that a surrogate network can be trained, the parameters of which are used to generate the attack. In this work, we introduce autoregressive (AR) poisoning, a method that can generate poisoned data without access to the broader dataset. The proposed AR perturbations are generic, can be applied across different datasets, and can poison different architectures. Compared to existing unlearnable methods, our AR poisons are more resistant against common defenses such as adversarial training and strong data augmentations. Our analysis further provides insight into what makes an effective data poison.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers31
- Image Shortcut Squeezing: Countering Perturbative Availability Poisons with CompressionZhuoran Liu, Zhengyu Zhao, Martha A. LarsonICML 2023 · 51 citations
- What Can We Learn from Unlearnable Datasets?Pedro Sandoval Segura, Vasu Singla, Jonas Geiping, Micah Goldblum et al.NeurIPS 2023 · 28 citations
- Exploring the Limits of Model-Targeted Indiscriminate Data Poisoning AttacksYiwei Lu, Gautam Kamath, Yaoliang YuICML 2023 · 25 citations
- Purify Unlearnable Examples via Rate-Constrained Variational AutoencodersYi Yu, Yufei Wang, Song Xia, Wenhan Yang et al.ICML 2024 · 22 citations
- UnSeg: One Universal Unlearnable Example Generator is Enough against All Image SegmentationYe Sun, Hao Zhang, Tiehua Zhang, Xingjun Ma et al.NeurIPS 2024 · 18 citations
Builds on12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- CutMix: Regularization Strategy to Train Strong Classifiers With Localizable FeaturesSangdoo Yun, Dongyoon Han, Sanghyuk Chun, Seong Joon Oh et al.ICCV 2019 · 5,843 citations
- The Pitfalls of Simplicity Bias in Neural NetworksHarshay Shah, Kaustav Tamuly, Aditi Raghunathan, Prateek Jain et al.NeurIPS 2020 · 503 citations
Related papers
- MetaPoison: Practical General-purpose Clean-label Data PoisoningW. Ronny Huang, Jonas Geiping, Liam Fowl, Gavin Taylor et al.NeurIPS 2020 · 242 citations
- Detection and Defense of Unlearnable ExamplesYifan Zhu, Lijia Yu, Xiao-Shan GaoAAAI 2024 · 11 citations
- Universal Backdoor AttacksBenjamin Schneider, Nils Lukas, Florian KerschbaumICLR 2024 · 17 citations
- Manipulating Machine Learning: Poisoning Attacks and Countermeasures for Regression LearningMatthew Jagielski, Alina Oprea, Battista Biggio, Chang Liu et al.S&P 2018 · 867 citations
- Data Poisoning Won't Save You From Facial RecognitionEvani Radiya-Dixit, Sanghyun Hong, Nicholas Carlini, Florian TramèrICLR 2022 · 67 citations
