SSLGuard: A Watermarking Scheme for Self-supervised Learning Pre-trained Encoders
Tianshuo Cong, Xinlei He, Yang Zhang
Abstract
Self-supervised learning is an emerging machine learning paradigm. Compared to supervised learning which leverages high-quality labeled datasets, self-supervised learning relies on unlabeled datasets to pre-train powerful encoders which can then be treated as feature extractors for various downstream tasks. The huge amount of data and computational resources consumption makes the encoders themselves become the valuable intellectual property of the model owner. Recent research has shown that the machine learning model's copyright is threatened by model stealing attacks, which aim to train a surrogate model to mimic the behavior of a given model. We empirically show that pre-trained encoders are highly vulnerable to model stealing attacks. However, most of the current efforts of copyright protection algorithms such as watermarking concentrate on classifiers. Meanwhile, the intrinsic challenges of pre-trained encoder's copyright protection remain largely unstudied. We fill the gap by proposing SSLGuard, the first watermarking scheme for pre-trained encoders. Given a clean pre-trained encoder, SSLGuard injects a watermark into it and outputs a watermarked version. The shadow training technique is also applied to preserve the watermark under potential model stealing attacks. Our extensive evaluation shows that SSLGuard is effective in watermark injection and verification, and it is robust against model stealing and other watermark removal attacks such as input noising, output perturbing, overwriting, model pruning, and fine-tuning. 1 Pre-trained Encoder Cloud Platform Surrogate Encoder Legitimate User Adversary ! "(!) !′ "(!′)
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6dc73391-9d8d-4cc7-b81f-4bfe2177de41Cited by top-tier papers26
- AdvCLIP: Downstream-agnostic Adversarial Examples in Multimodal Contrastive LearningZiqi Zhou, Shengshan Hu, Minghui Li, Hangtao Zhang et al.ACM MM 2023 · 62 citations
- Dataset Inference for Self-Supervised ModelsAdam Dziedzic, Haonan Duan, Muhammad Ahmad Kaleem, Nikita Dhawan et al.NeurIPS 2022 · 59 citations
- Downstream-agnostic Adversarial ExamplesZiqi Zhou, Shengshan Hu, Ruizhi Zhao, Qian Wang et al.ICCV 2023 · 45 citations
- StolenEncoder: Stealing Pre-trained Encoders in Self-supervised LearningYupei Liu, Jinyuan Jia, Hongbin Liu, Neil Zhenqiang GongCCS 2022 · 23 citations
- MEA-Defender: A Robust Watermark against Model Extraction AttackPeizhuo Lv, Hualong Ma, Kai Chen, Jiachen Zhou et al.S&P 2024 · 22 citations
Builds on26
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
Related papers
- SSL-WM: A Black-Box Watermarking Approach for Encoders Pre-trained by Self-Supervised LearningPeizhuo Lv, Pan Li, Shenchen Zhu, Shengzhi Zhang et al.NDSS 2024
- RAEncoder: A Label-Free Reversible Adversarial Examples Encoder for Dataset Intellectual Property ProtectionFan Xing, Zhuo Tian, Xuefeng Fan, Xiaoyi ZhouCVPR 2025
- PreGIP: Watermarking the Pretraining of Graph Neural Networks for Deep IP ProtectionEnyan Dai, Minhua Lin, Suhang WangKDD 2025 · 1 citation
- Can't Steal? Cont-Steal! Contrastive Stealing Attacks Against Image EncodersZeyang Sha, Xinlei He, Ning Yu, Michael Backes et al.CVPR 2023
- Task-Agnostic Language Model Watermarking via High Entropy Passthrough LayersVaden Masrani, Mohammad Akbari, David Ming Xuan Yue, Ahmad Rezaei et al.AAAI 2025 · 1 citation
