Certified Neural Network Watermarks with Randomized Smoothing
Arpit Bansal, Ping-Yeh Chiang, Michael J. Curry, Rajiv Jain, Curtis Wigington, Varun Manjunatha, John P. Dickerson, Tom Goldstein
Abstract
Watermarking is a commonly used strategy to protect creators' rights to digital images, videos and audio. Recently, watermarking methods have been extended to deep learning models -- in principle, the watermark should be preserved when an adversary tries to copy the model. However, in practice, watermarks can often be removed by an intelligent adversary. Several papers have proposed watermarking methods that claim to be empirically resistant to different types of removal attacks, but these new techniques often fail in the face of new or better-tuned adversaries. In this paper, we propose a certifiable watermarking method. Using the randomized smoothing technique proposed in Chiang et al., we show that our watermark is guaranteed to be unremovable unless the model parameters are changed by more than a certain l2 threshold. In addition to being certifiable, our watermark is also empirically more robust compared to previous watermarking methods. Our experiments can be reproduced with code at https://github.com/arpitbansal297/Certified_Watermarks
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 41b8ba7b-2fea-4a95-8b39-d28a29279d0dCited by top-tier papers23
- Tree-Rings Watermarks: Invisible Fingerprints for Diffusion ImagesYuxin Wen, John Kirchenbauer, Jonas Geiping, Tom GoldsteinNeurIPS 2023 · 253 citations
- Leveraging Optimization for Adaptive Attacks on Image WatermarksNils Lukas, Abdulrahman Diaa, Lucas Fenaux, Florian KerschbaumICLR 2024 · 50 citations
- Watermarks in the Sand: Impossibility of Strong Watermarking for Language ModelsHanlin Zhang, Benjamin L. Edelman, Danilo Francati, Daniele Venturi et al.ICML 2024 · 21 citations
- Margin-based Neural Network WatermarkingByungjoo Kim, Suyoung Lee, Seanie Lee, Sooel Son et al.ICML 2023 · 21 citations
- What can Discriminator do? Towards Box-free Ownership Verification of Generative Adversarial NetworksZiheng Huang, Boheng Li, Yan Cai, Run Wang et al.ICCV 2023 · 19 citations
Builds on8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Certified Robustness to Adversarial Examples with Differential PrivacyMathias Lécuyer, Vaggelis Atlidakis, Roxana Geambasu, Daniel Hsu et al.S&P 2019 · 1,022 citations
- Turning Your Weakness Into a Strength: Watermarking Deep Neural Networks by BackdooringYossi Adi, Carsten Baum, Moustapha Cissé, Benny Pinkas et al.USENIX Security 2018 · 832 citations
- Towards Stable and Efficient Training of Verifiably Robust Neural NetworksHuan Zhang, Hongge Chen, Chaowei Xiao, Sven Gowal et al.ICLR 2020 · 384 citations
- Certified Defenses for Adversarial PatchesPing-yeh Chiang, Renkun Ni, Ahmed Abdelkader, Chen Zhu et al.ICLR 2020 · 194 citations
Related papers
- Dimension-independent Certified Neural Network Watermarks via Mollifier SmoothingJiaxiang Ren, Yang Zhou, Jiayin Jin, Lingjuan Lyu et al.ICML 2023 · 10 citations
- Rethinking the Vulnerability of DNN Watermarking: Are Watermarks Robust against Naturalness-aware Perturbations?Run Wang, Haoxuan Li, Lingzhou Mu, Jixing Ren et al.ACM MM 2022 · 9 citations
- DeepEclipse: How to Break White-Box DNN-Watermarking SchemesAlessandro Pegoraro, Carlotta Segna, Kavita Kumari, Ahmad-Reza SadeghiUSENIX Security 2024 · 11 citations
- THE SELF-RE-WATERMARKING TRAP: FROM EXPLOIT TO RESILIENCEVithurabiman Senthuran, Yong Xiang, Iynkaran Natgunanathan, Uthayasanker ThayasivamICLR 2026
- Cert-LAS: Toward Certified Model Ownership Verification for Text-to-Image Diffusion Models via Layer-Adaptive SmoothingLeyi Qi, Yiming Li, Siyuan Liang, Zhengzhong Tu et al.ICML 2026 · 1 citation
