Evading Watermark based Detection of AI-Generated Content
Zhengyuan Jiang, Jinghuai Zhang, Neil Zhenqiang Gong
Abstract
A generative AI model can generate extremely realistic-looking content, posing growing challenges to the authenticity of information. To address the challenges, watermark has been leveraged to detect AI-generated content. Specifically, a watermark is embedded into an AI-generated content before it is released. A content is detected as AI-generated if a similar watermark can be decoded from it. In this work, we perform a systematic study on the robustness of such watermark-based AI-generated content detection. Our work shows that an attacker can post-process a watermarked image via adding a small, human-imperceptible perturbation to it, such that the post-processed image evades detection while maintaining its visual quality. We show the effectiveness of our attack both theoretically and empirically. Moreover, to evade detection, our adversarial post-processing method adds much smaller perturbations to AIgenerated images and thus better maintain their visual quality than existing popular post-processing methods such as JPEG compression, Gaussian blur, and Brightness/Contrast. Our work shows the insufficiency of existing watermark-based detection of AI-generated content, highlighting the urgent needs of new methods. Our code is publicly available: https://github.com/zhengyuan-jiang/WEvade . CCS CONCEPTS • Security and privacy → Security services.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4d5f772c-7355-420d-8cc0-ff325651212eCited by top-tier papers31
- Robustness of AI-Image Detectors: Fundamental Limits and Practical AttacksMehrdad Saberi, Vinu Sankar Sadasivan, Keivan Rezaei, Aounon Kumar et al.ICLR 2024 · 92 citations
- WAVES: Benchmarking the Robustness of Image WatermarksBang An, Mucong Ding, Tahseen Rabbani, Aakriti Agrawal et al.ICML 2024 · 86 citations
- EditGuard: Versatile Image Watermarking for Tamper Localization and Copyright ProtectionXuanyu Zhang, Runyi Li, Jiwen Yu, Youmin Xu et al.CVPR 2024 · 58 citations
- Leveraging Optimization for Adaptive Attacks on Image WatermarksNils Lukas, Abdulrahman Diaa, Lucas Fenaux, Florian KerschbaumICLR 2024 · 50 citations
- Governance of Generative AI in Creative Work: Consent, Credit, Compensation, and BeyondLin Kyi, Amruta Mahuli, Michael Six Silberman, Reuben Binns et al.CHI 2025 · 50 citations
Builds on17
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- DetectGPT: Zero-Shot Machine-Generated Text Detection using Probability CurvatureEric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D. Manning et al.ICML 2023 · 988 citations
- A Watermark for Large Language ModelsJohn Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz et al.ICML 2023 · 854 citations
Related papers
- Watermark-based Attribution of AI-Generated ContentZhengyuan Jiang, Moyang Guo, Yuepeng Hu, Yupu Wang et al.ICLR 2026 · 11 citations
- A Transfer Attack to Image WatermarksYuepeng Hu, Zhengyuan Jiang, Moyang Guo, Neil Zhenqiang GongICLR 2025
- Transferable Black-Box One-Shot Forging of Watermarks via Image Preference ModelsTomás Soucek, Sylvestre-Alvise Rebuffi, Pierre Fernandez, Nikola Jovanovic et al.NeurIPS 2025 · 11 citations
- Invisible Image Watermarks Are Provably Removable Using Generative AIXuandong Zhao, Kexun Zhang, Zihao Su, Saastha Vasan et al.NeurIPS 2024 · 209 citations
- PECCVAI: Overcoming the Brittleness of AI Image Watermarking Under Visual Paraphrasing AttacksShreyas Dixit, Ashhar Aziz, Shashwat Bajpai, Vasu Sharma et al.CVPR 2026
