Robustness of AI-Image Detectors: Fundamental Limits and Practical Attacks
Mehrdad Saberi, Vinu Sankar Sadasivan, Keivan Rezaei, Aounon Kumar, Atoosa Malemir Chegini, Wenxiao Wang, Soheil Feizi
摘要
In light of recent advancements in generative AI models, it has become essential to distinguish genuine content from AI-generated one to prevent the malicious usage of fake materials as authentic ones and vice versa. Various techniques have been introduced for identifying AI-generated images, with watermarking emerging as a promising approach. In this paper, we analyze the robustness of various AI-image detectors including watermarking and classifier-based deepfake detectors. For watermarking methods that introduce subtle image perturbations (i.e., low perturbation budget methods), we reveal a fundamental trade-off between the evasion error rate (i.e., the fraction of watermarked images detected as non-watermarked ones) and the spoofing error rate (i.e., the fraction of non-watermarked images detected as watermarked ones) upon an application of diffusion purification attack. To validate our theoretical findings, we also provide empirical evidence demonstrating that diffusion purification effectively removes low perturbation budget watermarks by applying minimal changes to images. The diffusion purification attack is ineffective for high perturbation watermarking methods where notable changes are applied to images. In this case, we develop a model substitution adversarial attack that can successfully remove watermarks. Moreover, we show that watermarking methods are vulnerable to spoofing attacks where the attacker aims to have real images identified as watermarked ones, damaging the reputation of the developers. In particular, with black-box access to the watermarking method, a watermarked noise image can be generated and added to real images, causing them to be incorrectly classified as watermarked. Finally, we extend our theory to characterize a fundamental trade-off between the robustness and reliability of classifier-based deep fake detectors and demonstrate it through experiments.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper26
- Invisible Image Watermarks Are Provably Removable Using Generative AIXuandong Zhao, Kexun Zhang, Zihao Su, Saastha Vasan 等NeurIPS 2024 · 被引用 209 次
- WAVES: Benchmarking the Robustness of Image WatermarksBang An, Mucong Ding, Tahseen Rabbani, Aakriti Agrawal 等ICML 2024 · 被引用 86 次
- Adversarial Paraphrasing: A Universal Attack for Humanizing AI-Generated TextYize Cheng, Vinu Sankar Sadasivan, Mehrdad Saberi, Shoumik Saha 等NeurIPS 2025 · 被引用 28 次
- Watermarks in the Sand: Impossibility of Strong Watermarking for Language ModelsHanlin Zhang, Benjamin L. Edelman, Danilo Francati, Daniele Venturi 等ICML 2024 · 被引用 21 次
- Watermarking Autoregressive Image GenerationNikola Jovanovic, Ismail Labiad, Tomás Soucek, Martin T. Vechev 等NeurIPS 2025 · 被引用 21 次
它引用的顶会 Paper13
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- FaceForensics++: Learning to Detect Manipulated Facial ImagesAndreas Rössler, Davide Cozzolino, Luisa Verdoliva, Christian Riess 等ICCV 2019 · 被引用 2,966 次
- Diffusion Models for Adversarial PurificationWeili Nie, Brandon Guo, Yujia Huang, Chaowei Xiao 等ICML 2022 · 被引用 663 次
相关 Paper
- Hidden in the Noise: Two-Stage Robust Watermarking for ImagesKasra Arabi, Benjamin Feuer, R. Teal Witter, Chinmay Hegde 等ICLR 2025
- A Transfer Attack to Image WatermarksYuepeng Hu, Zhengyuan Jiang, Moyang Guo, Neil Zhenqiang GongICLR 2025
- Evading Watermark based Detection of AI-Generated ContentZhengyuan Jiang, Jinghuai Zhang, Neil Zhenqiang GongCCS 2023 · 被引用 46 次
- Shallow Diffuse: Robust and Invisible Watermarking through Low-Dim Subspaces in Diffusion ModelsWenda Li, Huijie Zhang, Qing QuNeurIPS 2025 · 被引用 8 次
- RAW: A Robust and Agile Plug-and-Play Watermark Framework for AI-Generated Images with Provable GuaranteesXun Xian, Ganghua Wang, Xuan Bi, Jayanth Srinivasa 等NeurIPS 2024 · 被引用 17 次
