Hashed Watermark as a Filter: A Unified Defense Against Forging and Overwriting Attacks in Neural Network Watermarking
Yuan Yao, Jin Song, Jian Jin
Abstract
As valuable digital assets, deep neural networks necessitate robust ownership protection, positioning neural network watermarking (NNW) as a promising solution. Among NNW approaches, weight-based methods embed watermarks directly into model parameters; however, they remain generally susceptible to forging and overwriting attacks. To address those challenges, we propose NeuralMark, a robust method built around a hashed watermark filter. Specifically, we utilize a hash function to generate an irreversible binary watermark from a secret key, which is then used as a filter to select the model parameters for embedding. This design cleverly intertwines the embedding parameters with the hashed watermark, providing a robust defense against both forging and overwriting attacks. Average pooling is also incorporated to resist fine-tuning and pruning attacks. Furthermore, it can be seamlessly integrated into various neural network architectures, ensuring broad applicability. We theoretically analyze its security boundary and highlight the necessity of using a hashed watermark as a filter. Empirically, we demonstrate its effectiveness and robustness across 13 distinct Convolutional and Transformer architectures, covering five image classification tasks and one text generation task. The source codes are available at https://github.com/AIResearch- Group/NeuralMark.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 84ae87b7-d4a0-40d1-b886-4d32f76fe1faBuilds on10
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Swin Transformer V2: Scaling Up Capacity and ResolutionZe Liu, Han Hu, Yutong Lin, Zhuliang Yao et al.CVPR 2022 · 2,138 citations
- SoK: How Robust is Image Classification Deep Neural Network Watermarking?Nils Lukas, Edward Jiang, Xinda Li, Florian KerschbaumS&P 2022 · 124 citations
Related papers
- Watermarking Deep Neural Networks with Greedy ResidualsHanwen Liu, Zhenyu Weng, Yuesheng ZhuICML 2021 · 69 citations
- Identification for Deep Neural Network: Simply Adjusting Few Weights!Yingjie Lao, Peng Yang, Weijie Zhao, Ping LiICDE 2022 · 19 citations
- Towards the Resistance of Neural Network Fingerprinting to Fine-tuningLing Tang, Yuefeng Chen, Hui Xue', Quanshi ZhangNeurIPS 2025 · 5 citations
- Towards Robust Model Watermark via Reducing Parametric VulnerabilityGuanhao Gan, Yiming Li, Dongxian Wu, Shu-Tao XiaICCV 2023 · 18 citations
- Achieving Resolution-Agnostic DNN-based Image Watermarking: A Novel Perspective of Implicit Neural RepresentationYuchen Wang, Xingyu Zhu, Guanhui Ye, Shiyao Zhang et al.ACM MM 2024 · 6 citations
