WRATH: Turning Watermark Robustness Against Itself via a Watermark-Agnostic Black-Box Invalidation Attack
Nan Jiang, Juan Hu, Bangjie Sun, Terence Sim, Jun Han
摘要
Watermarks are increasingly embedded in AI-generated images to indicate their machine-generated origin. For reliable detection, watermarks are designed to remain robust against common image manipulations such as resizing and compression. In this paper, we uncover a previously overlooked vulnerability: the robustness property of watermarking schemes can inadvertently expose the image features that carry watermark signals. Exploiting this vulnerability, we introduce WRATH, the first watermark-agnostic black-box attack capable of both watermark removal and forgery. Our attack is practical and requires only a scheme's robustness information, which is typically public or can be obtained through simple robustness tests. When evaluated against state-of-theart watermarking schemes, including Amazon's, WRATH successfully attacks all evaluated schemes while preserving high perceptual image quality. We also discuss practical defenses to mitigate this vulnerability, and our findings call for a rethink of watermark security.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper21
- The Stable Signature: Rooting Watermarks in Latent Diffusion ModelsPierre Fernandez, Guillaume Couairon, Hervé Jégou, Matthijs Douze 等ICCV 2023 · 被引用 370 次
- Tree-Rings Watermarks: Invisible Fingerprints for Diffusion ImagesYuxin Wen, John Kirchenbauer, Jonas Geiping, Tom GoldsteinNeurIPS 2023 · 被引用 253 次
- Invisible Image Watermarks Are Provably Removable Using Generative AIXuandong Zhao, Kexun Zhang, Zihao Su, Saastha Vasan 等NeurIPS 2024 · 被引用 209 次
- Robustness of AI-Image Detectors: Fundamental Limits and Practical AttacksMehrdad Saberi, Vinu Sankar Sadasivan, Keivan Rezaei, Aounon Kumar 等ICLR 2024 · 被引用 92 次
- WAVES: Benchmarking the Robustness of Image WatermarksBang An, Mucong Ding, Tahseen Rabbani, Aakriti Agrawal 等ICML 2024 · 被引用 86 次
相关 Paper
- A Transfer Attack to Image WatermarksYuepeng Hu, Zhengyuan Jiang, Moyang Guo, Neil Zhenqiang GongICLR 2025
- Neural Dehydration: Effective Erasure of Black-box Watermarks from DNNs with Limited DataYifan Lu, Wenxuan Li, Mi Zhang, Xudong Pan 等CCS 2024 · 被引用 2 次
- Class-feature Watermark: A Resilient Black-box Watermark Against Model Extraction AttacksYaxin Xiao, Qingqing Ye, Zi Liang, Haoyang Li 等AAAI 2026
- Can Simple Averaging Defeat Modern Watermarks?Pei Yang, Hai Ci, Yiren Song, Mike Zheng ShouNeurIPS 2024 · 被引用 48 次
- WMCopier: Forging Invisible Watermarks on Arbitrary ImagesZiping Dong, Chao Shuai, Zhongjie Ba, Peng Cheng 等NeurIPS 2025 · 被引用 2 次
