Leveraging Optimization for Adaptive Attacks on Image Watermarks
Nils Lukas, Abdulrahman Diaa, Lucas Fenaux, Florian Kerschbaum
摘要
Untrustworthy users can misuse image generators to synthesize high-quality deepfakes and engage in unethical activities. Watermarking deters misuse by marking generated content with a hidden message, enabling its detection using a secret watermarking key. A core security property of watermarking is robustness, which states that an attacker can only evade detection by substantially degrading image quality. Assessing robustness requires designing an adaptive attack for the specific watermarking algorithm. When evaluating watermarking algorithms and their (adaptive) attacks, it is challenging to determine whether an adaptive attack is optimal, i.e., the best possible attack. We solve this problem by defining an objective function and then approach adaptive attacks as an optimization problem. The core idea of our adaptive attacks is to replicate secret watermarking keys locally by creating surrogate keys that are differentiable and can be used to optimize the attack's parameters. We demonstrate for Stable Diffusion models that such an attacker can break all five surveyed watermarking methods at no visible degradation in image quality. Optimizing our attacks is efficient and requires less than 1 GPU hour to reduce the detection accuracy to 6.3% or less. Our findings emphasize the need for more rigorous robustness testing against adaptive, learnable attackers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Invisible Image Watermarks Are Provably Removable Using Generative AIXuandong Zhao, Kexun Zhang, Zihao Su, Saastha Vasan 等NeurIPS 2024 · 被引用 209 次
- WAVES: Benchmarking the Robustness of Image WatermarksBang An, Mucong Ding, Tahseen Rabbani, Aakriti Agrawal 等ICML 2024 · 被引用 86 次
- Can Simple Averaging Defeat Modern Watermarks?Pei Yang, Hai Ci, Yiren Song, Mike Zheng ShouNeurIPS 2024 · 被引用 48 次
- Watermarks in the Sand: Impossibility of Strong Watermarking for Language ModelsHanlin Zhang, Benjamin L. Edelman, Danilo Francati, Daniele Venturi 等ICML 2024 · 被引用 21 次
- FreqMark: Invisible Image Watermarking via Frequency Based Optimization in Latent SpaceYiyang Guo, Ruizhe Li, Mude Hui, Hanzhong Guo 等NeurIPS 2024 · 被引用 14 次
它引用的顶会 Paper12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
相关 Paper
- Optimizing Adaptive Attacks against Watermarks for Language ModelsAbdulrahman Diaa, Toluwani Aremu, Nils LukasICML 2025
- PTW: Pivotal Tuning Watermarking for Pre-Trained Image GeneratorsNils Lukas, Florian KerschbaumUSENIX Security 2023
- An Undetectable Watermark for Generative Image ModelsSam Gunn, Xuandong Zhao, Dawn SongICLR 2025
- WMAdapter: Adding WaterMark Control to Latent Diffusion ModelsHai Ci, Yiren Song, Pei Yang, Jinheng Xie 等ICML 2025
- ROBIN: Robust and Invisible Watermarks for Diffusion Models with Adversarial OptimizationHuayang Huang, Yu Wu, Qian WangNeurIPS 2024 · 被引用 73 次
