ZeroMark: Towards Dataset Ownership Verification without Disclosing Watermark
Junfeng Guo, Yiming Li, Ruibo Chen, Yihan Wu, Chenxi Liu, Heng Huang
Abstract
High-quality public datasets significantly prompt the prosperity of deep neural networks (DNNs). Currently, dataset ownership verification (DOV), which consists of dataset watermarking and ownership verification, is the only feasible so-lution to protect their copyright by preventing unauthorized use. In this paper, we revisit existing DOV methods and find that they all mainly focused on the first stage by designing different types of dataset watermarks and directly exploiting watermarked samples as the verification samples for ownership verification. As such, their success relies on an underlying assumption that verification is a one-time and privacy-preserving process, which does not necessarily hold in practice. To alleviate this problem, we propose ZeroMark to conduct ownership verification without disclosing dataset-specified watermarks. Our method is inspired by our empirical and theoretical findings of the intrinsic property of DNNs trained on the watermarked dataset. Specifically, ZeroMark first generates the closest boundary version of given benign samples and calculates their boundary gradients under the label-only black-box setting. After that, it examines whether the given suspicious method has been trained on the protected dataset by performing a hypothesis test, based on the cosine similarity measured on the boundary gradients and the watermark pattern. Extensive experiments on benchmark datasets verify the effectiveness of our ZeroMark and its resistance to potential adaptive attacks. The codes for reproducing our main experiments are publicly available at GitHub.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- VICTOR: Dataset Copyright Auditing in Video Recognition SystemsQuan Yuan, Zhikun Zhang, Linkang Du, Min Chen et al.NDSS 2026 · 2 citations
- DSSmoothing: Toward Certified Dataset Ownership Verification for Pre-trained Language Models via Dual-Space SmoothingTing Qiao, Xing Liu, Wenke Huang, Jianbin Li et al.WWW 2026 · 1 citation
- De-mark: Watermark Removal in Large Language ModelsRuibo Chen, Yihan Wu, Junfeng Guo, Heng HuangICML 2025
Builds on21
- Practical Secure Aggregation for Privacy-Preserving Machine LearningKallista A. Bonawitz, Vladimir Ivanov, Ben Kreuter, Antonio Marcedone et al.CCS 2017 · 3,936 citations
- Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural NetworksBolun Wang, Yuanshun Yao, Shawn Shan, Huiying Li et al.S&P 2019 · 1,801 citations
- Oblivious Neural Network Predictions via MiniONN TransformationsJian Liu, Mika Juuti, Yao Lu, N. AsokanCCS 2017 · 800 citations
- HopSkipJumpAttack: A Query-Efficient Decision-Based AttackJianbo Chen, Michael I. Jordan, Martin J. WainwrightS&P 2020 · 797 citations
- Minimally distorted Adversarial Examples with a Fast Adaptive Boundary AttackFrancesco Croce, Matthias HeinICML 2020 · 597 citations
Related papers
- Domain Watermark: Effective and Harmless Dataset Copyright Protection is Closed at HandJunfeng Guo, Yiming Li, Lixu Wang, Shu-Tao Xia et al.NeurIPS 2023 · 93 citations
- Untargeted Backdoor Watermark: Towards Harmless and Stealthy Dataset Copyright ProtectionYiming Li, Yang Bai, Yong Jiang, Yong Yang et al.NeurIPS 2022 · 161 citations
- EntropyMark: Towards More Harmless Backdoor Watermark via Entropy-based Constraint for Open-source Dataset Copyright ProtectionMing Sun, Rui Wang, Zixuan Zhu, Lihua Jing et al.CVPR 2025
- Watermarking Deep Neural Networks with Greedy ResidualsHanwen Liu, Zhenyu Weng, Yuesheng ZhuICML 2021 · 69 citations
- Towards Robust Model Watermark via Reducing Parametric VulnerabilityGuanhao Gan, Yiming Li, Dongxian Wu, Shu-Tao XiaICCV 2023 · 18 citations
