An Embarrassingly Simple Backdoor Attack on Self-supervised Learning
Changjiang Li, Ren Pang, Zhaohan Xi, Tianyu Du, Shouling Ji, Yuan Yao, Ting Wang
Abstract
As a new paradigm in machine learning, self-supervised learning (SSL) is capable of learning high-quality representations of complex data without relying on labels. In addition to eliminating the need for labeled data, research has found that SSL improves the adversarial robustness over supervised learning since lacking labels makes it more challenging for adversaries to manipulate model predictions. However, the extent to which this robustness superiority generalizes to other types of attacks remains an open question.We explore this question in the context of backdoor attacks. Specifically, we design and evaluate Ctrl, an embarrassingly simple yet highly effective self-supervised backdoor attack. By only polluting a tiny fraction of training data (≤ 1%) with indistinguishable poisoning samples, Ctrl causes any trigger-embedded input to be misclassified to the adversary's designated class with a high probability (≥ 99%) at inference time. Our findings suggest that SSL and supervised learning are comparably vulnerable to backdoor attacks. More importantly, through the lens of Ctrl, we study the inherent vulnerability of SSL to backdoor attacks. With both empirical and analytical evidence, we reveal that the representation invariance property of SSL, which benefits adversarial robustness, may also be the very reason making SSL highly susceptible to backdoor attacks. Our findings also imply that the existing defenses against supervised backdoor attacks are not easily retrofitted to the unique vulnerability of SSL. Code is available at: https://github.com/meet-cjli/CTRL
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ad4c2fc2-6128-4f3a-b442-70a3e7375176Cited by top-tier papers28
- Distribution Preserving Backdoor Attack in Self-supervised LearningGuanhong Tao, Zhenting Wang, Shiwei Feng, Guangyu Shen et al.S&P 2024 · 32 citations
- Backdoor Contrastive Learning via Bi-level Trigger OptimizationWeiyu Sun, Xinyu Zhang, Hao Lu, Ying-Cong Chen et al.ICLR 2024 · 14 citations
- CLIP-Guided Backdoor Defense through Entropy-Based Poisoned Dataset SeparationBinyan Xu, Fan Yang, Xilin Dai, Di Tang et al.ACM MM 2025 · 12 citations
- On the Difficulty of Defending Contrastive Learning against Backdoor AttacksChangjiang Li, Ren Pang, Bochuan Cao, Zhaohan Xi et al.USENIX Security 2024 · 10 citations
- BadMerging: Backdoor Attacks Against Model MergingJinghuai Zhang, Jianfeng Chi, Zheng Li, Kunlin Cai et al.CCS 2024 · 5 citations
Builds on25
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 2,360 citations
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li et al.S&P 2019 · 1,801 citations
- Trojaning Attack on Neural NetworksYingqi Liu, Shiqing Ma, Yousra Aafer, Wen-Chuan Lee et al.NDSS 2018 · 1,377 citations
Related papers
- Invisible Backdoor Attack against Self-supervised LearningHanrong Zhang, Zhenting Wang, Boheng Li, Fulin Lin et al.CVPR 2025
- The Perils of Learning From Unlabeled Data: Backdoor Attacks on Semi-supervised LearningVirat Shejwalkar, Lingjuan Lyu, Amir HoumansadrICCV 2023 · 15 citations
- Backdoor Attacks on Self-Supervised LearningAniruddha Saha, Ajinkya Tejankar, Soroush Abbasi Koohpayegani, Hamed PirsiavashCVPR 2022 · 77 citations
- DeDe: Detecting Backdoor Samples for SSL Encoders via DecodersSizai Hou, Songze Li, Duanyi YaoCVPR 2025
- Defending Against Patch-based Backdoor Attacks on Self-Supervised LearningAjinkya Tejankar, Maziar Sanjabi, Qifan Wang, Sinong Wang et al.CVPR 2023
