Locate-Then-Examine: Grounded Region Reasoning Improves Detection of AI-Generated Images
Yikun Ji, Yan Hong, Bowen Deng, Jun Lan, Huijia Zhu, Weiqiang Wang, Liqing Zhang, Jianfu Zhang
Abstract
The rapid growth of AI-generated imagery has blurred the boundary between real and synthetic content, raising practical concerns for digital integrity. Vision-language models (VLMs) can provide natural language explanations, but standard one-pass classifiers often miss subtle artifacts in high-quality synthetic images and offer limited grounding in the pixels. We propose Locate-Then-Examine (LTE), a two-stage VLM-based forensic framework that first localizes suspicious regions and then re-examines these crops together with the full image to refine the real vs. AI-generated verdict and its explanation. LTE explicitly links each decision to localized visual evidence through region proposals and region-aware reasoning. To support training and evaluation, we introduce TRACE, a dataset of 20,000 real and high-quality synthetic images with region-level annotations and automatically generated forensic explanations, constructed by a VLM-based pipeline with additional consistency checks and quality control. Across TRACE and multiple external benchmarks, LTE achieves competitive accuracy and improved robustness while providing human-understandable, region-grounded explanations suitable for forensic deployment.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext edb0adcd-8fe5-41b0-88b8-be9bf894c860Builds on36
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao et al.ICCV 2023 · 13,211 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
Related papers
- FakeXplain: AI-Generated Image Detection via Human-Aligned Grounded ReasoningYikun Ji, Yan Hong, Qi Fan, Jun Lan et al.ICLR 2026 · 9 citations
- Spot the Fake: Large Multimodal Model-Based Synthetic Image Detection with Artifact ExplanationSiwei Wen, Junyan Ye, Peilin Feng, Hengrui Kang et al.NeurIPS 2025 · 82 citations
- From Prediction to Explanation: Multimodal, Explainable, and Interactive Deepfake Detection Framework for Non-Expert UsersShahroz Tariq, Simon S. Woo, Priyanka Singh, Irena Irmalasari et al.ACM MM 2025 · 13 citations
- ReAlign: Generalizable Image Forgery Detection via Reasoning-Aligned RepresentationQing Huang, Zhipei Xu, Xuanyu Zhang, Xiangyu Yu et al.CVPR 2026 · 3 citations
- Hermes: An Evidence-Driven Agentic Framework for Trustworthy and Explainable AI-Generated Video DetectionShuaibo Li, Pengfei HAO, Hongtao Wu, Jianfeng Dong et al.ICML 2026
