USENIX Security2021Top-tier venue
Poisoning the Unlabeled Dataset of Semi-Supervised Learning
Nicholas Carlini
Abstract
Semi-supervised machine learning models learn from a (small) set of labeled training examples, and a (large) set of unlabeled training examples. State-of-the-art models can reach within a few percentage points of fully-supervised training, while requiring 100× less labeled data.
We study a new class of vulnerabilities: poisoning attacks that modify the unlabeled dataset. In order to be useful, unlabeled datasets are given strictly less review than labeled datasets, and adversaries can therefore poison them easily. By inserting maliciously-crafted unlabeled examples totaling just 0.1% of the dataset size, we can manipulate a model trained on this poisoned dataset to misclassify arbitrary examples at test time (as any desired label). Our attacks are highly effective across datasets and semi-supervised learning methods.
We find that more accurate methods (thus more likely to be used) are significantly more vulnerable to poisoning attacks, and as such better training methods are unlikely to prevent this attack. To counter this we explore the space of defenses, and propose two methods that mitigate our attack.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7674c0e8-3c98-4db3-bb73-71e4fba42db3Cited by top-tier papers27
- Poisoning Language Models During Instruction TuningAlexander Wan, Eric Wallace, Sheng Shen, Dan KleinICML 2023 · 319 citations
- Poisoning Web-Scale Training Datasets is PracticalNicholas Carlini, Matthew Jagielski, Christopher A. Choquette-Choo, Daniel Paleka et al.S&P 2024 · 309 citations
- Cycle Self-Training for Domain AdaptationHong Liu, Jianmin Wang, Mingsheng LongNeurIPS 2021 · 236 citations
- Poisoning and Backdooring Contrastive LearningNicholas Carlini, Andreas TerzisICLR 2022 · 213 citations
- Better Safe Than Sorry: Preventing Delusive Adversaries with Adversarial TrainingLue Tao, Lei Feng, Jinfeng Yi, Sheng-Jun Huang et al.NeurIPS 2021 · 90 citations
Builds on12
- Deep Learning with Differential PrivacyMartín Abadi, Andy Chu, Ian J. Goodfellow, H. Brendan McMahan et al.CCS 2016 · 7,620 citations
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang et al.NeurIPS 2020 · 5,129 citations
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong et al.NeurIPS 2020 · 2,774 citations
- The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural NetworksNicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos et al.USENIX Security 2019 · 1,386 citations
Related papers
- The Perils of Learning From Unlabeled Data: Backdoor Attacks on Semi-supervised LearningVirat Shejwalkar, Lingjuan Lyu, Amir HoumansadrICCV 2023 · 15 citations
- On Robustness of Linear Classifiers to Targeted Data PoisoningNakshatra Gupta, Sumanth Prabhu S, Supratik Chakraborty, R. VenkateshAAAI 2026
- Narcissus: A Practical Clean-Label Backdoor Attack with Limited InformationYi Zeng, Minzhou Pan, Hoang Anh Just, Lingjuan Lyu et al.CCS 2023 · 170 citations
- Adversarial Examples Make Strong PoisonsLiam Fowl, Micah Goldblum, Ping-yeh Chiang, Jonas Geiping et al.NeurIPS 2021 · 185 citations
- Manipulating SGD with Data Ordering AttacksIlia Shumailov, Zakhar Shumaylov, Dmitry Kazhdan, Yiren Zhao et al.NeurIPS 2021 · 125 citations
