Robustness of AI-Image Detectors: Fundamental Limits and Practical Attacks
Mehrdad Saberi, Vinu Sankar Sadasivan, Keivan Rezaei, Aounon Kumar, Atoosa Malemir Chegini, Wenxiao Wang, Soheil Feizi
Abstract
In light of recent advancements in generative AI models, it has become essential to distinguish genuine content from AI-generated one to prevent the malicious usage of fake materials as authentic ones and vice versa. Various techniques have been introduced for identifying AI-generated images, with watermarking emerging as a promising approach. In this paper, we analyze the robustness of various AI-image detectors including watermarking and classifier-based deepfake detectors. For watermarking methods that introduce subtle image perturbations (i.e., low perturbation budget methods), we reveal a fundamental trade-off between the evasion error rate (i.e., the fraction of watermarked images detected as non-watermarked ones) and the spoofing error rate (i.e., the fraction of non-watermarked images detected as watermarked ones) upon an application of diffusion purification attack. To validate our theoretical findings, we also provide empirical evidence demonstrating that diffusion purification effectively removes low perturbation budget watermarks by applying minimal changes to images. The diffusion purification attack is ineffective for high perturbation watermarking methods where notable changes are applied to images. In this case, we develop a model substitution adversarial attack that can successfully remove watermarks. Moreover, we show that watermarking methods are vulnerable to spoofing attacks where the attacker aims to have real images identified as watermarked ones, damaging the reputation of the developers. In particular, with black-box access to the watermarking method, a watermarked noise image can be generated and added to real images, causing them to be incorrectly classified as watermarked. Finally, we extend our theory to characterize a fundamental trade-off between the robustness and reliability of classifier-based deep fake detectors and demonstrate it through experiments.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 35cd9261-bfdb-4c45-873d-946addaef5e6Cited by top-tier papers26
- Invisible Image Watermarks Are Provably Removable Using Generative AIXuandong Zhao, Kexun Zhang, Zihao Su, Saastha Vasan et al.NeurIPS 2024 · 209 citations
- WAVES: Benchmarking the Robustness of Image WatermarksBang An, Mucong Ding, Tahseen Rabbani, Aakriti Agrawal et al.ICML 2024 · 86 citations
- Adversarial Paraphrasing: A Universal Attack for Humanizing AI-Generated TextYize Cheng, Vinu Sankar Sadasivan, Mehrdad Saberi, Shoumik Saha et al.NeurIPS 2025 · 28 citations
- Watermarks in the Sand: Impossibility of Strong Watermarking for Language ModelsHanlin Zhang, Benjamin L. Edelman, Danilo Francati, Daniele Venturi et al.ICML 2024 · 21 citations
- Watermarking Autoregressive Image GenerationNikola Jovanovic, Ismail Labiad, Tomás Soucek, Martin T. Vechev et al.NeurIPS 2025 · 21 citations
Builds on13
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- FaceForensics++: Learning to Detect Manipulated Facial ImagesAndreas Rössler, Davide Cozzolino, Luisa Verdoliva, Christian Riess et al.ICCV 2019 · 2,966 citations
- Diffusion Models for Adversarial PurificationWeili Nie, Brandon Guo, Yujia Huang, Chaowei Xiao et al.ICML 2022 · 663 citations
Related papers
- Hidden in the Noise: Two-Stage Robust Watermarking for ImagesKasra Arabi, Benjamin Feuer, R. Teal Witter, Chinmay Hegde et al.ICLR 2025
- A Transfer Attack to Image WatermarksYuepeng Hu, Zhengyuan Jiang, Moyang Guo, Neil Zhenqiang GongICLR 2025
- Evading Watermark based Detection of AI-Generated ContentZhengyuan Jiang, Jinghuai Zhang, Neil Zhenqiang GongCCS 2023 · 46 citations
- Shallow Diffuse: Robust and Invisible Watermarking through Low-Dim Subspaces in Diffusion ModelsWenda Li, Huijie Zhang, Qing QuNeurIPS 2025 · 8 citations
- RAW: A Robust and Agile Plug-and-Play Watermark Framework for AI-Generated Images with Provable GuaranteesXun Xian, Ganghua Wang, Xuan Bi, Jayanth Srinivasa et al.NeurIPS 2024 · 17 citations
