Semantic Visual Anomaly Detection and Reasoning in AI-Generated Images
Chuangchuang Tan, Xiang Ming, Jinglu Wang, Renshuai Tao, Bin Li, Yunchao Wei, Yao Zhao, Yan Lu
摘要
The rapid advancement of AI-generated content (AIGC) has enabled the synthesis of visually convincing images; however, many such outputs exhibit subtle semantic anomalies, including unrealistic object configurations, violations of physical laws, or commonsense inconsistencies, which compromise the overall plausibility of the generated scenes. Detecting these semantic-level anomalies is essential for assessing the trustworthiness of AIGC media, especially in AIGC image analysis, explainable deepfake detection and semantic authenticity assessment.In this paper, we formalize semantic anomaly detection and reasoning for AIGC images and introduce AnomReason, a large-scale benchmark with structured annotations as quadruples (Name, Phenomenon, Reasoning, Severity). Annotations are produced by a modular multi-agent pipeline (AnomAgent) with lightweight human-in-the-loop verification, enabling scale while preserving quality. At construction time, AnomAgent processed approximately 4.17 B GPT-4o tokens, providing scale evidence for the resulting structured annotations. We further show that models fine-tuned on AnomReason achieve consistent gains over strong vision-language baselines under our proposed semantic matching metric (SemAP and SemF1). Applications to explainable deepfake detection and semantic reasonableness assessment of image generators demonstrate practical utility. In summary, AnomReason and AnomAgent serve as a foundation for measuring and improving the semantic plausibility of AI-generated images. The code is available at https://github.com/chuangchuangtan/Semantic-Visual-Anomaly-Detection-and-Reasoning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger 等ICLR 2020 · 被引用 8,443 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
相关 Paper
- FakeWorld 1.0: An Omni-modal Benchmark for Fake Media and ContentYifeng Gao, Yifan Ding, Li Wang, Feida Huang 等ICML 2026
- ReAlign: Generalizable Image Forgery Detection via Reasoning-Aligned RepresentationQing Huang, Zhipei Xu, Xuanyu Zhang, Xiangyu Yu 等CVPR 2026 · 被引用 3 次
- From Manipulation to Mistrust: Explaining Diverse Micro-Video Misinformation for Robust Debunking in the WildZhi Zeng, Yifei Yang, Jiaying Wu, Xulang Zhang 等WWW 2026 · 被引用 3 次
- AVFakeBench: A Comprehensive Audio-Video Forgery Detection Benchmark for AV-LMMsShuhan Xia, Peipei Li, Xuannan Liu, Dongsen Zhang 等CVPR 2026 · 被引用 1 次
- Towards Zero-Shot Anomaly Detection and Reasoning with Multimodal Large Language ModelsJiacong Xu, Shao-Yuan Lo, Bardia Safaei, Vishal M. Patel 等CVPR 2025
