Dataset Ownership Verification for Pre-Trained Masked Models
Yuechen Xie, Jie Song, Yicheng Shan, Xiaoyan Zhang, Yuanyu Wan, Shengxuming Zhang, Jiarui Duan, Mingli Song
Abstract
High-quality open-source datasets have emerged as a pivotal catalyst driving the swift advancement of deep learning, while facing the looming threat of potential exploitation. Protecting these datasets is of paramount importance for the interests of their owners. The verification of dataset ownership has evolved into a crucial approach in this domain; however, existing verification techniques are predominantly tailored to supervised models and contrastive pre-trained models, rendering them ill-suited for direct application to the increasingly prevalent masked models. In this work, we introduce the inaugural methodology addressing this critical, yet unresolved challenge, termed Dataset Ownership Verification for Masked Modeling (DOV4MM). The central objective is to ascertain whether a suspicious black-box model has been pre-trained on a particular unlabeled dataset, thereby assisting dataset owners in safeguarding their rights. DOV4MM is grounded in our empirical observation that when a model is pre-trained on the target dataset, the difficulty of reconstructing masked information within the embedding space exhibits a marked contrast to models not pre-trained on that dataset. We validated the efficacy of DOV4MM through ten masked image models on ImageNet-1K and four masked language models on WikiText-103. The results demonstrate that DOV4MM rejects the null hypothesis, with a -value considerably below 0.05, surpassing all prior approaches. Code is available at https://github.com/xieyc99/DOV4MM.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3781502f-957d-42a9-bf6f-75acd26b1ffbBuilds on29
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel et al.ICLR 2020 · 7,418 citations
- Membership Inference Attacks Against Machine Learning ModelsReza Shokri, Marco Stronati, Congzheng Song, Vitaly ShmatikovS&P 2017 · 5,137 citations
Related papers
- Dataset Ownership Verification in Contrastive Pre-trained ModelsYuechen Xie, Jie Song, Mengqi Xue, Haofei Zhang et al.ICLR 2025
- SSL-WM: A Black-Box Watermarking Approach for Encoders Pre-trained by Self-Supervised LearningPeizhuo Lv, Pan Li, Shenchen Zhu, Shengzhi Zhang et al.NDSS 2024
- DSSmoothing: Toward Certified Dataset Ownership Verification for Pre-trained Language Models via Dual-Space SmoothingTing Qiao, Xing Liu, Wenke Huang, Jianbin Li et al.WWW 2026 · 1 citation
- Untargeted Backdoor Watermark: Towards Harmless and Stealthy Dataset Copyright ProtectionYiming Li, Yang Bai, Yong Jiang, Yong Yang et al.NeurIPS 2022 · 161 citations
- Domain Watermark: Effective and Harmless Dataset Copyright Protection is Closed at HandJunfeng Guo, Yiming Li, Lixu Wang, Shu-Tao Xia et al.NeurIPS 2023 · 93 citations
