Certified Neural Network Watermarks with Randomized Smoothing
Arpit Bansal, Ping-Yeh Chiang, Michael J. Curry, Rajiv Jain, Curtis Wigington, Varun Manjunatha, John P. Dickerson, Tom Goldstein
摘要
Watermarking is a commonly used strategy to protect creators' rights to digital images, videos and audio. Recently, watermarking methods have been extended to deep learning models -- in principle, the watermark should be preserved when an adversary tries to copy the model. However, in practice, watermarks can often be removed by an intelligent adversary. Several papers have proposed watermarking methods that claim to be empirically resistant to different types of removal attacks, but these new techniques often fail in the face of new or better-tuned adversaries. In this paper, we propose a certifiable watermarking method. Using the randomized smoothing technique proposed in Chiang et al., we show that our watermark is guaranteed to be unremovable unless the model parameters are changed by more than a certain l2 threshold. In addition to being certifiable, our watermark is also empirically more robust compared to previous watermarking methods. Our experiments can be reproduced with code at https://github.com/arpitbansal297/Certified_Watermarks
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper23
- Tree-Rings Watermarks: Invisible Fingerprints for Diffusion ImagesYuxin Wen, John Kirchenbauer, Jonas Geiping, Tom GoldsteinNeurIPS 2023 · 被引用 253 次
- Leveraging Optimization for Adaptive Attacks on Image WatermarksNils Lukas, Abdulrahman Diaa, Lucas Fenaux, Florian KerschbaumICLR 2024 · 被引用 50 次
- Watermarks in the Sand: Impossibility of Strong Watermarking for Language ModelsHanlin Zhang, Benjamin L. Edelman, Danilo Francati, Daniele Venturi 等ICML 2024 · 被引用 21 次
- Margin-based Neural Network WatermarkingByungjoo Kim, Suyoung Lee, Seanie Lee, Sooel Son 等ICML 2023 · 被引用 21 次
- What can Discriminator do? Towards Box-free Ownership Verification of Generative Adversarial NetworksZiheng Huang, Boheng Li, Yan Cai, Run Wang 等ICCV 2023 · 被引用 19 次
它引用的顶会 Paper8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Certified Robustness to Adversarial Examples with Differential PrivacyMathias Lécuyer, Vaggelis Atlidakis, Roxana Geambasu, Daniel Hsu 等S&P 2019 · 被引用 1,022 次
- Turning Your Weakness Into a Strength: Watermarking Deep Neural Networks by BackdooringYossi Adi, Carsten Baum, Moustapha Cissé, Benny Pinkas 等USENIX Security 2018 · 被引用 832 次
- Towards Stable and Efficient Training of Verifiably Robust Neural NetworksHuan Zhang, Hongge Chen, Chaowei Xiao, Sven Gowal 等ICLR 2020 · 被引用 384 次
- Certified Defenses for Adversarial PatchesPing-yeh Chiang, Renkun Ni, Ahmed Abdelkader, Chen Zhu 等ICLR 2020 · 被引用 194 次
相关 Paper
- Dimension-independent Certified Neural Network Watermarks via Mollifier SmoothingJiaxiang Ren, Yang Zhou, Jiayin Jin, Lingjuan Lyu 等ICML 2023 · 被引用 10 次
- Rethinking the Vulnerability of DNN Watermarking: Are Watermarks Robust against Naturalness-aware Perturbations?Run Wang, Haoxuan Li, Lingzhou Mu, Jixing Ren 等ACM MM 2022 · 被引用 9 次
- DeepEclipse: How to Break White-Box DNN-Watermarking SchemesAlessandro Pegoraro, Carlotta Segna, Kavita Kumari, Ahmad-Reza SadeghiUSENIX Security 2024 · 被引用 11 次
- THE SELF-RE-WATERMARKING TRAP: FROM EXPLOIT TO RESILIENCEVithurabiman Senthuran, Yong Xiang, Iynkaran Natgunanathan, Uthayasanker ThayasivamICLR 2026
- Cert-LAS: Toward Certified Model Ownership Verification for Text-to-Image Diffusion Models via Layer-Adaptive SmoothingLeyi Qi, Yiming Li, Siyuan Liang, Zhengzhong Tu 等ICML 2026 · 被引用 1 次
