Explanation as a Watermark: Towards Harmless and Multi-bit Model Ownership Verification via Watermarking Feature Attribution
Shuo Shao, Yiming Li, Hongwei Yao, Yiling He, Zhan Qin, Kui Ren
Abstract
Ownership verification is currently the most critical and widely adopted post-hoc method to safeguard model copyright. In general, model owners exploit it to identify whether a given suspicious third-party model is stolen from them by examining whether it has particular properties inherited' from their released models. Currently, backdoor-based model watermarks are the primary and cutting-edge methods to implant such properties in the released models. However, backdoor-based methods have two fatal drawbacks, including harmfulness and ambiguity. The former indicates that they introduce maliciously controllable misclassification behaviors ($i.e.$, backdoor) to the watermarked released models. The latter denotes that malicious users can easily pass the verification by finding other misclassified samples, leading to ownership ambiguity. In this paper, we argue that both limitations stem from the zero-bit' nature of existing watermarking schemes, where they exploit the status (, misclassified) of predictions for verification. Motivated by this understanding, we design a new watermarking paradigm, , Explanation as a Watermark (EaaW), that implants verification behaviors into the explanation of feature attribution instead of model predictions. Specifically, EaaW embeds a `multi-bit' watermark into the feature attribution explanation of specific trigger samples without changing the original prediction. We correspondingly design the watermark embedding and extraction algorithms inspired by explainable artificial intelligence. In particular, our approach can be used for different tasks (, image classification and text generation). Extensive experiments verify the effectiveness and harmlessness of our EaaW and its resistance to potential attacks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext db6227c9-dfc5-44c2-862f-249c320de46cCited by top-tier papers15
- How to Trace Latent Generative Model Generated Images without Artificial Watermark?Zhenting Wang, Vikash Sehwag, Chen Chen, Lingjuan Lyu et al.ICML 2024 · 24 citations
- Model Provenance Testing for Large Language ModelsIvica Nikolic, Teodora Baluta, Prateek SaxenaNeurIPS 2025 · 20 citations
- PromptCOS: Towards Content-Only System Prompt Copyright Auditing for LLMsYuchen Yang, Yiming Li, Hongwei Yao, Enhao Huang et al.S&P 2026 · 5 citations
- Is Difficulty Calibration All We Need? Towards More Practical Membership Inference AttacksYu He, Boheng Li, Yao Wang, Mengda Yang et al.CCS 2024 · 4 citations
- Watermarking Graph Neural Networks via Explanations for Ownership ProtectionJane Downer, Yingdan Shi, Ziyan Liu, Ren Wang et al.ICML 2026 · 3 citations
Builds on25
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Turning Your Weakness Into a Strength: Watermarking Deep Neural Networks by BackdooringYossi Adi, Carsten Baum, Moustapha Cissé, Benny Pinkas et al.USENIX Security 2018 · 832 citations
- Entangled Watermarks as a Defense against Model ExtractionHengrui Jia, Christopher A. Choquette-Choo, Varun Chandrasekaran, Nicolas PapernotUSENIX Security 2021 · 287 citations
- Untargeted Backdoor Watermark: Towards Harmless and Stealthy Dataset Copyright ProtectionYiming Li, Yang Bai, Yong Jiang, Yong Yang et al.NeurIPS 2022 · 161 citations
Related papers
- Towards Robust Model Watermark via Reducing Parametric VulnerabilityGuanhao Gan, Yiming Li, Dongxian Wu, Shu-Tao XiaICCV 2023 · 18 citations
- Neural Honeytrace: Plug&Play Watermarking Framework against Model Extraction AttacksYixiao Xu, Binxing Fang, Rui Wang, Yinghai Zhou et al.ICML 2026
- Are You Copying My Model? Protecting the Copyright of Large Language Models for EaaS via Backdoor WatermarkWenjun Peng, Jingwei Yi, Fangzhao Wu, Shangxi Wu et al.ACL 2023 · 39 citations
- Domain Watermark: Effective and Harmless Dataset Copyright Protection is Closed at HandJunfeng Guo, Yiming Li, Lixu Wang, Shu-Tao Xia et al.NeurIPS 2023 · 93 citations
- Hot-Swap MarkBoard: An Efficient Black-box Watermarking Approach for Large-scale Model DistributionZhicheng Zhang, Peizhuo Lv, Mengke Wan, Jiang Fang et al.ACM MM 2025 · 2 citations
