WRATH: Turning Watermark Robustness Against Itself via a Watermark-Agnostic Black-Box Invalidation Attack
Nan Jiang, Juan Hu, Bangjie Sun, Terence Sim, Jun Han
Abstract
Watermarks are increasingly embedded in AI-generated images to indicate their machine-generated origin. For reliable detection, watermarks are designed to remain robust against common image manipulations such as resizing and compression. In this paper, we uncover a previously overlooked vulnerability: the robustness property of watermarking schemes can inadvertently expose the image features that carry watermark signals. Exploiting this vulnerability, we introduce WRATH, the first watermark-agnostic black-box attack capable of both watermark removal and forgery. Our attack is practical and requires only a scheme's robustness information, which is typically public or can be obtained through simple robustness tests. When evaluated against state-of-theart watermarking schemes, including Amazon's, WRATH successfully attacks all evaluated schemes while preserving high perceptual image quality. We also discuss practical defenses to mitigate this vulnerability, and our findings call for a rethink of watermark security.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6797cf04-8de2-427d-8437-457fe13524dfBuilds on21
- The Stable Signature: Rooting Watermarks in Latent Diffusion ModelsPierre Fernandez, Guillaume Couairon, Hervé Jégou, Matthijs Douze et al.ICCV 2023 · 370 citations
- Tree-Rings Watermarks: Invisible Fingerprints for Diffusion ImagesYuxin Wen, John Kirchenbauer, Jonas Geiping, Tom GoldsteinNeurIPS 2023 · 253 citations
- Invisible Image Watermarks Are Provably Removable Using Generative AIXuandong Zhao, Kexun Zhang, Zihao Su, Saastha Vasan et al.NeurIPS 2024 · 209 citations
- Robustness of AI-Image Detectors: Fundamental Limits and Practical AttacksMehrdad Saberi, Vinu Sankar Sadasivan, Keivan Rezaei, Aounon Kumar et al.ICLR 2024 · 92 citations
- WAVES: Benchmarking the Robustness of Image WatermarksBang An, Mucong Ding, Tahseen Rabbani, Aakriti Agrawal et al.ICML 2024 · 86 citations
Related papers
- A Transfer Attack to Image WatermarksYuepeng Hu, Zhengyuan Jiang, Moyang Guo, Neil Zhenqiang GongICLR 2025
- Neural Dehydration: Effective Erasure of Black-box Watermarks from DNNs with Limited DataYifan Lu, Wenxuan Li, Mi Zhang, Xudong Pan et al.CCS 2024 · 2 citations
- Class-feature Watermark: A Resilient Black-box Watermark Against Model Extraction AttacksYaxin Xiao, Qingqing Ye, Zi Liang, Haoyang Li et al.AAAI 2026
- Can Simple Averaging Defeat Modern Watermarks?Pei Yang, Hai Ci, Yiren Song, Mike Zheng ShouNeurIPS 2024 · 48 citations
- WMCopier: Forging Invisible Watermarks on Arbitrary ImagesZiping Dong, Chao Shuai, Zhongjie Ba, Peng Cheng et al.NeurIPS 2025 · 2 citations
