SSL-WM: A Black-Box Watermarking Approach for Encoders Pre-trained by Self-Supervised Learning
Peizhuo Lv, Pan Li, Shenchen Zhu, Shengzhi Zhang, Kai Chen, Ruigang Liang, Chang Yue, Fan Xiang, Yuling Cai, Hualong Ma, Yingjun Zhang, Guozhu Meng
Abstract
Recent years have witnessed tremendous success in Self-Supervised Learning (SSL), which has been widely utilized to facilitate various downstream tasks in Computer Vision (CV) and Natural Language Processing (NLP) domains. However, attackers may steal such SSL models and commercialize them for profit, making it crucial to verify the ownership of the SSL models. Most existing ownership protection solutions (e.g., backdoor-based watermarks) are designed for supervised learning models and cannot be used directly since they require that the models' downstream tasks and target labels be known and available during watermark embedding, which is not always possible in the domain of SSL. To address such a problem, especially when downstream tasks are diverse and unknown during watermark embedding, we propose a novel black-box watermarking solution, named SSL-WM, for verifying the ownership of SSL models. SSL-WM maps watermarked inputs of the protected encoders into an invariant representation space, which causes any downstream classifier to produce expected behavior, thus allowing the detection of embedded watermarks. We evaluate SSL-WM on numerous tasks, such as CV and NLP, using different SSL models both contrastive-based and generative-based. Experimental results demonstrate that SSL-WM can effectively verify the ownership of stolen SSL models in various downstream tasks. Furthermore, SSL-WM is robust against model fine-tuning, pruning, and input preprocessing attacks. Lastly, SSL-WM can also evade detection from evaluated watermark detection approaches, demonstrating its promising application in protecting the ownership of SSL models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 35bf1da0-9965-418b-a149-6ec254e5418bCited by top-tier papers9
- MEA-Defender: A Robust Watermark against Model Extraction AttackPeizhuo Lv, Hualong Ma, Kai Chen, Jiachen Zhou et al.S&P 2024 · 22 citations
- Chasing Shadows: Pitfalls in LLM Security ResearchJonathan Evertz, Niklas Risse, Nicolai Neuer, Andreas Müller et al.NDSS 2026 · 17 citations
- MFL-Owner: Ownership Protection for Multi-modal Federated Learning via Orthogonal Transform WatermarkKeke Gai, Dongjue Wang, Jing Yu, Mohan Wang et al.AAAI 2025 · 6 citations
- PreGIP: Watermarking the Pretraining of Graph Neural Networks for Deep IP ProtectionEnyan Dai, Minhua Lin, Suhang WangKDD 2025 · 1 citation
- Dataset Reduction and Watermark Removal via Self-supervised Learning for Model Extraction AttackHao Luan, Xue Tan, Zhiheng Li, Jun Dai et al.NDSS 2026 · 1 citation
Builds on19
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li et al.S&P 2019 · 1,801 citations
Related papers
- SSLGuard: A Watermarking Scheme for Self-supervised Learning Pre-trained EncodersTianshuo Cong, Xinlei He, Yang ZhangCCS 2022 · 28 citations
- Dataset Ownership Verification in Contrastive Pre-trained ModelsYuechen Xie, Jie Song, Mengqi Xue, Haofei Zhang et al.ICLR 2025
- Dataset Ownership Verification for Pre-Trained Masked ModelsYuechen Xie, Jie Song, Yicheng Shan, Xiaoyan Zhang et al.ICCV 2025 · 1 citation
- Margin-based Neural Network WatermarkingByungjoo Kim, Suyoung Lee, Seanie Lee, Sooel Son et al.ICML 2023 · 21 citations
- Domain Watermark: Effective and Harmless Dataset Copyright Protection is Closed at HandJunfeng Guo, Yiming Li, Lixu Wang, Shu-Tao Xia et al.NeurIPS 2023 · 93 citations
