Bridging the Gap Between Ideal and Real-World Evaluation: Benchmarking AI-Generated Image Detection in Challenging Scenarios
Chunxiao Li, Xiaoxiao Wang, Meiling Li, Boming Miao, Peng Sun, Yunjian Zhang, Xiangyang Ji, Yao Zhu
Abstract
With the rapid advancement of generative models, highly realistic image synthesis has posed new challenges to digital security and media credibility. Although AI-generated image detection methods have partially addressed these concerns, a substantial research gap remains in evaluating their performance under complex real-world conditions. This paper introduces the Real-World Robustness Dataset (RRDataset) for comprehensive evaluation of detection models across three dimensions: 1) Scenario Generalization -RRDataset encompasses high-quality images from seven major scenarios (War & Conflict, Disasters & Accidents, Political & Social Events, Medical & Public Health, Culture & Religion, Labor & Production, and everyday life), addressing existing dataset gaps from a content perspective. 2) Internet Transmission Robustnessexamining detector performance on images that have undergone multiple rounds of sharing across various social media platforms. 3) Re-digitization Robustness -assessing model effectiveness on images altered through four distinct re-digitization methods.
We benchmarked 17 detectors and 10 vision-language models (VLMs) on RRDataset and conducted a largescale human study involving 192 participants to investigate human few-shot learning capabilities in detecting AIgenerated images. The benchmarking results reveal the limitations of current AI detection methods under real-world conditions and underscore the importance of drawing on human adaptability to develop more robust detection algorithms. Our dataset is publicly available at: https: //zenodo.org/records/14963880.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0da4e098-81ae-4e80-b004-dd00a8589996Cited by top-tier papers4
- TreeTeaming: Autonomous Red-Teaming of Vision-Language Models via Hierarchical Strategy ExplorationChunxiao Li, Lijun Li, Jing ShaoCVPR 2026 · 4 citations
- TranX-Adapter: Bridging Artifacts and Semantics within MLLMs for Robust AI-generated Image DetectionWenbin Wang, Yuge Huang, Jianqing Xu, Yue Yu et al.ICML 2026 · 1 citation
- Detect Any AI-Counterfeited Text ImageChenfan Qu, Yiwu Zhong, Xuekang Zhu, Junchi Li et al.CVPR 2026
- Dissect and Prune: Enhancing Robustness in AI-Generated Image DetectionDahye Kim, Jaehyun Choi, Hyun Seok Seong, Seongho Kim et al.ICML 2026
Builds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 3,959 citations
Related papers
- Your One-Stop Solution for AI-Generated Video DetectionLong Ma, Zihao Xue, Yan Wang, Zhiyuan Yan et al.CVPR 2026 · 13 citations
- RealHD: A High-Quality Dataset for Robust Detection of State-of-the-Art AI-Generated ImagesHanzhe Yu, Yun Ye, Jintao Rong, Qi Xuan et al.ACM MM 2025 · 1 citation
- ILLUSION: Unveiling Truth with a Comprehensive Multi-Modal, Multi-Lingual Deepfake DatasetKartik Thakral, Rishabh Ranjan, Akanksha Singh, Akshat Jain et al.ICLR 2025
- WildFake: A Large-Scale and Hierarchical Dataset for AI-Generated Images DetectionYan Hong, Jianming Feng, Haoxing Chen, Jun Lan et al.AAAI 2025 · 13 citations
- AI-Face: A Million-Scale Demographically Annotated AI-Generated Face Dataset and Fairness BenchmarkLi Lin, Santosh Santosh, Mingyang Wu, Xin Wang et al.CVPR 2025
