Lune

S&P2026Top-tier venue

WRATH: Turning Watermark Robustness Against Itself via a Watermark-Agnostic Black-Box Invalidation Attack

Nan Jiang, Juan Hu, Bangjie Sun, Terence Sim, Jun Han

2026Year

Abstract

Watermarks are increasingly embedded in AI-generated images to indicate their machine-generated origin. For reliable detection, watermarks are designed to remain robust against common image manipulations such as resizing and compression. In this paper, we uncover a previously overlooked vulnerability: the robustness property of watermarking schemes can inadvertently expose the image features that carry watermark signals. Exploiting this vulnerability, we introduce WRATH, the first watermark-agnostic black-box attack capable of both watermark removal and forgery. Our attack is practical and requires only a scheme's robustness information, which is typically public or can be obtained through simple robustness tests. When evaluated against state-of-theart watermarking schemes, including Amazon's, WRATH successfully attacks all evaluated schemes while preserving high perceptual image quality. We also discuss practical defenses to mitigate this vulnerability, and our findings call for a rethink of watermark security.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 6797cf04-8de2-427d-8437-457fe13524df

Builds on21

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines