The Perils of Learning From Unlabeled Data: Backdoor Attacks on Semi-supervised Learning
Virat Shejwalkar, Lingjuan Lyu, Amir Houmansadr
摘要
Semi-supervised machine learning (SSL) is gaining popularity as it reduces the cost of training ML models. It does so by using very small amounts of (expensive, well-inspected) labeled data and large amounts of (cheap, non-inspected) unlabeled data. SSL has shown comparable or even superior performances compared to conventional fully-supervised ML techniques. In this paper, we show that the key feature of SSL that it can learn from (non-inspected) unlabeled data exposes SSL to strong poisoning attacks. In fact, we argue that, due to its reliance on non-inspected unlabeled data, poisoning is a much more severe problem in SSL than in conventional fully-supervised ML. Specifically, we design a backdoor poisoning attack on SSL that can be conducted by a weak adversary with no knowledge of target SSL pipeline. This is unlike prior poisoning attacks in fully-supervised settings that assume strong adversaries with practically-unrealistic capabilities. We show that by poisoning only 0.2% of the unlabeled training data, our attack can cause misclassification of more than 80% of test inputs (when they contain the adversary's backdoor trigger). Our attacks remain effective across twenty combinations of benchmark datasets and SSL algorithms, and even circumvent the state-of-the-art defenses against backdoor attacks. Our work raises significant concerns about the practical utility of existing SSL algorithms.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Improving Group Robustness on Spurious Correlation Requires Preciser Group InferenceYujin Han, Difan ZouICML 2024 · 被引用 13 次
- SABRE-FL: Selective and Accurate Backdoor Rejection for Federated Prompt LearningMomin Ahmad Khan, Yasra Chandio, Fatima M. AnwarICLR 2026 · 被引用 2 次
- Protecting Model Adaptation from Trojans in the Unlabeled DataLijun Sheng, Jian Liang, Ran He, Zilei Wang 等AAAI 2025
它引用的顶会 Paper23
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang 等NeurIPS 2020 · 被引用 5,129 次
- RandAugment: Practical Automated Data Augmentation with a Reduced Search SpaceEkin Dogus Cubuk, Barret Zoph, Jonathon Shlens, Quoc LeNeurIPS 2020 · 被引用 4,453 次
- Unsupervised Data Augmentation for Consistency TrainingQizhe Xie, Zihang Dai, Eduard H. Hovy, Thang Luong 等NeurIPS 2020 · 被引用 2,774 次
- FlexMatch: Boosting Semi-Supervised Learning with Curriculum Pseudo LabelingBowen Zhang, Yidong Wang, Wenxin Hou, Hao Wu 等NeurIPS 2021 · 被引用 1,389 次
相关 Paper
- DeHiB: Deep Hidden Backdoor Attack on Semi-supervised Learning via Adversarial PerturbationZhicong Yan, Gaolei Li, Yuan Tian, Jun Wu 等AAAI 2021 · 被引用 43 次
- Poisoning the Unlabeled Dataset of Semi-Supervised LearningNicholas CarliniUSENIX Security 2021 · 被引用 80 次
- An Embarrassingly Simple Backdoor Attack on Self-supervised LearningChangjiang Li, Ren Pang, Zhaohan Xi, Tianyu Du 等ICCV 2023 · 被引用 54 次
- Defending Against Patch-based Backdoor Attacks on Self-Supervised LearningAjinkya Tejankar, Maziar Sanjabi, Qifan Wang, Sinong Wang 等CVPR 2023
- Phantom: Untargeted Poisoning Attacks on Semi-Supervised LearningJonathan Knauer, Phillip Rieger, Hossein Fereidooni, Ahmad-Reza SadeghiCCS 2024
